Performance.
Faster than vLLM and SGLang in all 52 matched tests on H100 , up to 3×.
Test setup & disclosures
Vision-language · H100 · ChartQA 512-request slice · Aug 30, 2026
Accuracy is scored on the same request stream for every runtime.
- Engines
- Photon 0.6.1 · SGLang 0.5.13 · vLLM 0.27.1
- Workload
- ChartQA 512-request slice
- Concurrency
- 1, 2, 4, 8
- Updated
- Aug 30, 2026
- One pinned provider and machine shape per run; driver and SKU recorded per cell.
- Runtime CUDA builds: Photon 12.8; vLLM 13.0; SGLang 13.0.
- Each Photon-versus-competitor comparison used the same requests, model revision, and concurrency.
- Prefix caching disabled everywhere; each runtime keeps its production graph policy.
Speech recognition · H100 ASR · Speech-to-text transcription · Aug 31, 2026
- GPU
- NVIDIA H100
- Engines
- Photon 2.1 · vLLM 0.28.0 · Qwen-ASR/vLLM · NeMo 3.0.0
- Workload
- Speech-to-text transcription
- Concurrency
- 1, 8
- Updated
- Aug 31, 2026
Launch post · full context and analysis.
Throughput by concurrency
How fast each engine serves the same workload. 60 matched pairs · 11 models.
Longer is better →
Faster than vLLM, Qwen-ASR/vLLM and NeMo in all 8 matched tests on H100 ASR , up to 3×.
425× realtime at c8
698× realtime at c8
548× realtime at c8
2006× realtime at c8
Methodology & repro
Vision-language
scripts/prime_bench.py --kind all --cells bench_cells_smoke.json --chart throughputSpeech recognition
Request completion latency
Typical (p50) and tail (p99) completion latency at concurrency 1. All percentiles are in the CSV.
Shorter is better →
Methodology & repro
scripts/prime_bench.py --kind all --cells bench_cells_smoke.json --chart latencyTime to first response
Process start to first response, lower is better.
Methodology & repro
scripts/prime_bench.py --kind all --cells bench_cells_smoke.json --chart cold-startMemory, power, and energy per token
Photon's GPU memory, power, temperature, and energy per output token.
| Model | VRAM | Power | Temp | J / output token |
|---|---|---|---|---|
| Gemma 4 E2B | 85.5 GB | 429 W | 56 °C | 0.15 |
| Gemma 4 E4B | 85.5 GB | 447 W | 70 °C | 0.26 |
| Moondream 3 | 85.5 GB | 518 W | 56 °C | 0.78 |
| Qwen 3.5 0.8B | 85.5 GB | 343 W | 45 °C | 0.05 |
| Qwen 3.5 2B | 85.5 GB | 432 W | 54 °C | 0.11 |
| Qwen 3.5 4B | 85.5 GB | 548 W | 78 °C | 0.26 |
| Qwen 3.5 9B | 85.5 GB | 552 W | 65 °C | 0.41 |
Methodology & repro
scripts/prime_bench.py --kind all --cells bench_cells_smoke.json --chart resourcesTest Photon on your stack.
Use the published command as a starting point, or talk to us about your model, hardware, and latency target.