· Performance

Performance.

Faster than vLLM and SGLang in all 52 matched tests on H100 , up to 3×.

Test setup & disclosures

Vision-language · H100 · ChartQA 512-request slice · Aug 30, 2026

Accuracy is scored on the same request stream for every runtime.

Engines
Photon 0.6.1 · SGLang 0.5.13 · vLLM 0.27.1
Workload
ChartQA 512-request slice
Concurrency
1, 2, 4, 8
Updated
Aug 30, 2026
  • One pinned provider and machine shape per run; driver and SKU recorded per cell.
  • Runtime CUDA builds: Photon 12.8; vLLM 13.0; SGLang 13.0.
  • Each Photon-versus-competitor comparison used the same requests, model revision, and concurrency.
  • Prefix caching disabled everywhere; each runtime keeps its production graph policy.

Speech recognition · H100 ASR · Speech-to-text transcription · Aug 31, 2026

Engines
Photon 2.1 · vLLM 0.28.0 · Qwen-ASR/vLLM · NeMo 3.0.0
Workload
Speech-to-text transcription
Concurrency
1, 8
Updated
Aug 31, 2026

Launch post · full context and analysis.

01 · Request throughput

Throughput by concurrency

How fast each engine serves the same workload. 60 matched pairs · 11 models.

Longer is better →

Batch size
Vision-languageH100 · ChartQA 512-request slice · Aug 30, 2026
Photon 0.6.1 vLLM 0.27.1 SGLang 0.5.13Bar length = measured throughput, requests/second
Gemma 4 E2BGemma
c1
Photon
vLLM
SGLang
Gemma 4 E4BGemma
c1
Photon
vLLM
SGLang
Moondream 3Moondream VLM
c1
Photon
vLLM
SGLang
Qwen 3.5 0.8BQwen 3
c1
Photon
vLLM
SGLang
Qwen 3.5 2BQwen 3
c1
Photon
vLLM
SGLang
Qwen 3.5 4BQwen 3
c1
Photon
vLLM
SGLang
Qwen 3.5 9BQwen 3
c1
Photon
vLLM
SGLang
Speech recognitionH100 ASR · Speech-to-text transcription · Aug 31, 2026

Faster than vLLM, Qwen-ASR/vLLM and NeMo in all 8 matched tests on H100 ASR , up to 3×.

Photon 2.1 vLLM 0.28.0 Qwen-ASR/vLLM NeMo 3.0.0Bar length = relative throughput, baseline engine = 1×
Whisper large-v3-turboWhisper
c1
Photon
vLLM

425× realtime at c8

Qwen3-ASR 0.6BQwen3-ASR
c1
Photon
Qwen-ASR/vLLM

698× realtime at c8

Qwen3-ASR 1.7BQwen3-ASR
c1
Photon
Qwen-ASR/vLLM

548× realtime at c8

Parakeet TDT 0.6B v3Parakeet
c1
Photon
NeMo

2006× realtime at c8

Methodology & repro

Vision-language

Models
Gemma 4 E2B (3e22461f65e89153144f8adb70e3b8c2cc9845a7) · Gemma 4 E4B (ee0ef6023621cff504d758262d4e04895a5af4a2) · Moondream 3 (5112966d1a723413b1c9a1e8bea272b72e647b35) · Qwen 3.5 0.8B (2fc06364715b967f1860aea9cf38778875588b17) · Qwen 3.5 2B (15852e8c16360a2fea060d615a32b45270f8a8fc) · Qwen 3.5 4B (851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a) · Qwen 3.5 9B (c202236235762e1c871ad0ccb60c8ee5ba337b9a)
Runtimes
Photon 0.6.1 · SGLang 0.5.13 · vLLM 0.27.1
Hardware
NVIDIA H100 80GB HBM3 · driver 580.95.05 · CUDA Photon 12.8 · vLLM 13.0 · SGLang 13.0
Workload
ChartQA 512-request slice
Sampling
512 closed-loop requests over the measured window
Date
Aug 30, 2026
$scripts/prime_bench.py --kind all --cells bench_cells_smoke.json --chart throughput

Speech recognition

Models
Whisper large-v3-turbo (deadbeefdeadbeefdeadbeefdeadbeefdeadbeef) · Qwen3-ASR 0.6B (deadbeefdeadbeefdeadbeefdeadbeefdeadbeef) · Qwen3-ASR 1.7B (deadbeefdeadbeefdeadbeefdeadbeefdeadbeef) · Parakeet TDT 0.6B v3 (deadbeefdeadbeefdeadbeefdeadbeefdeadbeef)
Runtimes
Photon 2.1 · vLLM 0.28.0 · Qwen-ASR/vLLM · NeMo 3.0.0
Hardware
NVIDIA H100
Workload
Speech-to-text transcription
Date
Aug 31, 2026
02 · Latency & jitter · Image + text

Request completion latency

Typical (p50) and tail (p99) completion latency at concurrency 1. All percentiles are in the CSV.

Shorter is better →

Photon 0.6.1 vLLM 0.27.1 SGLang 0.5.13Bar length = client-observed completion latency
Gemma 4 E2Bc1
typical (p50)
Photon
vLLM
SGLang
Gemma 4 E4Bc1
typical (p50)
Photon
vLLM
SGLang
Moondream 3c1
typical (p50)
Photon
vLLM
Qwen 3.5 0.8Bc1
typical (p50)
Photon
vLLM
SGLang
Qwen 3.5 2Bc1
typical (p50)
Photon
vLLM
SGLang
Qwen 3.5 4Bc1
typical (p50)
Photon
vLLM
SGLang
Qwen 3.5 9Bc1
typical (p50)
Photon
vLLM
SGLang
Methodology & repro
Models
Gemma 4 E2B (3e22461f65e89153144f8adb70e3b8c2cc9845a7) · Gemma 4 E4B (ee0ef6023621cff504d758262d4e04895a5af4a2) · Moondream 3 (5112966d1a723413b1c9a1e8bea272b72e647b35) · Qwen 3.5 0.8B (2fc06364715b967f1860aea9cf38778875588b17) · Qwen 3.5 2B (15852e8c16360a2fea060d615a32b45270f8a8fc) · Qwen 3.5 4B (851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a) · Qwen 3.5 9B (c202236235762e1c871ad0ccb60c8ee5ba337b9a)
Runtimes
Photon 0.6.1 · SGLang 0.5.13 · vLLM 0.27.1
Hardware
NVIDIA H100 80GB HBM3 · driver 580.95.05 · CUDA Photon 12.8 · vLLM 13.0 · SGLang 13.0
Workload
ChartQA 512-request slice
Sampling
client-observed request completion, semaphore admission to last token
Date
Aug 30, 2026
$scripts/prime_bench.py --kind all --cells bench_cells_smoke.json --chart latency
03 · Cold start · Image + text

Time to first response

Process start to first response, lower is better.

PhotonvLLMSGLangShared scale across models · lower is better
Model
Gemma 4 E2B
175s379s122s
Gemma 4 E4B
197s383s127s
Moondream 3
20.0s170s
Qwen 3.5 0.8B
18.0s300s138s
Qwen 3.5 2B
20.6s284s138s
Qwen 3.5 4B
43.6s308s133s
Qwen 3.5 9B
48.9s300s135s
Methodology & repro
Models
Gemma 4 E2B (3e22461f65e89153144f8adb70e3b8c2cc9845a7) · Gemma 4 E4B (ee0ef6023621cff504d758262d4e04895a5af4a2) · Moondream 3 (5112966d1a723413b1c9a1e8bea272b72e647b35) · Qwen 3.5 0.8B (2fc06364715b967f1860aea9cf38778875588b17) · Qwen 3.5 2B (15852e8c16360a2fea060d615a32b45270f8a8fc) · Qwen 3.5 4B (851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a) · Qwen 3.5 9B (c202236235762e1c871ad0ccb60c8ee5ba337b9a)
Runtimes
Photon 0.6.1 · SGLang 0.5.13 · vLLM 0.27.1
Hardware
NVIDIA H100 80GB HBM3 · driver 580.95.05 · CUDA Photon 12.8 · vLLM 13.0 · SGLang 13.0
Workload
ChartQA 512-request slice
Sampling
one-shot C1: server launch through the first completed request
Date
Aug 30, 2026
$scripts/prime_bench.py --kind all --cells bench_cells_smoke.json --chart cold-start
04 · Resource use · Image + text

Memory, power, and energy per token

Photon's GPU memory, power, temperature, and energy per output token.

ModelVRAMPowerTempJ / output token
Gemma 4 E2B85.5 GB429 W56 °C0.15
Gemma 4 E4B85.5 GB447 W70 °C0.26
Moondream 385.5 GB518 W56 °C0.78
Qwen 3.5 0.8B85.5 GB343 W45 °C0.05
Qwen 3.5 2B85.5 GB432 W54 °C0.11
Qwen 3.5 4B85.5 GB548 W78 °C0.26
Qwen 3.5 9B85.5 GB552 W65 °C0.41
Methodology & repro
Models
Gemma 4 E2B (3e22461f65e89153144f8adb70e3b8c2cc9845a7) · Gemma 4 E4B (ee0ef6023621cff504d758262d4e04895a5af4a2) · Moondream 3 (5112966d1a723413b1c9a1e8bea272b72e647b35) · Qwen 3.5 0.8B (2fc06364715b967f1860aea9cf38778875588b17) · Qwen 3.5 2B (15852e8c16360a2fea060d615a32b45270f8a8fc) · Qwen 3.5 4B (851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a) · Qwen 3.5 9B (c202236235762e1c871ad0ccb60c8ee5ba337b9a)
Runtimes
Photon 0.6.1
Hardware
NVIDIA H100 80GB HBM3 · driver 580.95.05 · CUDA Photon 12.8 · vLLM 13.0 · SGLang 13.0
Workload
ChartQA 512-request slice
Sampling
NVML at 10 Hz over the active window, idle-subtracted joules
Date
Aug 30, 2026
$scripts/prime_bench.py --kind all --cells bench_cells_smoke.json --chart resources
Your workload

Test Photon on your stack.

Use the published command as a starting point, or talk to us about your model, hardware, and latency target.

$pip install moondream