Photon 2.6 runs Qwen3.5-27B at 411 tokens per second on a single NVIDIA B200, increasing to 825 tok/s with two concurrent requests. It outperforms vLLM at every concurrency we tested, from one to eight requests.
This release adds FP8 support and DFlash speculative decoding to Photon. Under the hood, we extended our megakernel compiler to handle quantized projections and causal verification of draft tokens. Qwen3.5's recurrent layers made verification particularly interesting: when a draft is rejected, we need to recover the model state without running the expensive projections again.
Faster on one GPU
We ran Qwen3.5-27B-FP8 with the same DFlash checkpoint in Photon and vLLM 0.30.0, with thinking enabled in both. Photon is ahead at every concurrency we tested:
At two concurrent requests, Photon maintains over 400 tok/s per request. The chart divides total throughput by concurrency; the table below shows total throughput.
| Concurrent requests | Photon | vLLM | Speedup |
|---|---|---|---|
| 1 | 411 tok/s | 346 tok/s | 1.19x |
| 2 | 825 tok/s | 671 tok/s | 1.23x |
| 4 | 1,259 tok/s | 1,089 tok/s | 1.16x |
| 8 | 2,206 tok/s | 1,989 tok/s | 1.11x |
These numbers measure the full serving system, including prefill and host overhead, using the released packages from PyPI.
The cost of serving at 400 tok/s
Assume a B200 costs $6.25 per hour, and we keep two requests running at the measured throughput. At 825 tok/s total, that's 2.97 million output tokens per hour, or $2.10 in GPU rental per million output tokens, while maintaining over 400 tok/s per request.
The cost is comparable to hosted APIs, while per-request speed is much higher than the rates reported in our September 29 snapshot of OpenRouter's Qwen3.5-27B listings: 9 to 22 tok/s across its listed providers, compared with 412 tok/s per request in our two-request benchmark.
OpenRouter's speeds are P50 measurements of provider traffic; Photon's figure comes from the benchmark above. Provider prices as of September 29, 2026:
| Provider | Input / M | Output / M | Reported P50 tok/s |
|---|---|---|---|
| Alibaba | $0.195 | $1.56 | 21 |
| SiliconFlow | $0.25 | $2.00 | 9 |
| AtlasCloud | $0.27 | $2.16 | 13 |
| Phala | $0.30 | $2.40 | 10 |
| NovitaAI | $0.30 | $2.40 | 12 |
| DeepInfra | $0.26 | $2.60 | 22 |
Photon's GPU cost falls within that output-price range. The $2.10 estimate allocates the entire GPU rental cost, including time spent processing inputs, to output tokens. It assumes the benchmark workload runs continuously; idle time and operating costs are additional. API providers charge separately for input tokens.
Adding FP8 to the compiler
FP8 reduces weight storage from sixteen bits per value to eight, plus block scales. This is especially useful at low concurrency, where reading the weights can take longer than doing the arithmetic.
We added FP8 projections to the megakernel compiler, including activation quantization for tensor-core computation. The compiler can pipeline weight loading with compute and combine normalization with quantization to avoid writing an intermediate activation back to GPU memory.
The compiler chooses the tile geometry, too. A tile built for a large batch can leave most of the GPU idle when verifying a short sequence. Smaller tiles expose more parallel work, but change weight reuse, register pressure, and synchronization costs. Our cost model weighs those costs when selecting the execution plan.
Speculative verification in a megakernel
Ordinary decoding runs the target model once per new token. DFlash proposes a block of tokens with a smaller draft model, then asks the target to verify them in one causal pass. Keep the accepted prefix, discard the rejected suffix, and draft again.
One pass through the target's weights can now produce several output tokens. That is a good fit for a megakernel, but it isn't ordinary batching: sixteen positions in one sequence depend on each other. Sixteen independent requests don't.
The compiler now represents that causal sequence directly. At one concurrent request, Qwen3.5's target verification runs as a compiled megakernel, carrying work across operator boundaries without a CPU dispatch between each operation. Photon handles drafting, acceptance, and state commit around it.
Handling recurrent state
Qwen3.5 mixes full attention with Gated DeltaNet, a recurrent architecture. Rejected attention-cache positions can be hidden by shortening the visible history. A recurrent state has already incorporated their values. Changing a length counter won't undo that.
Our approach is inspired by ReplaySSM, which caches recent inputs and reconstructs recurrent state when needed. For generated verification, we use checkpoint-and-replay: keep the state from before the candidate block, retain the projected inputs, and reconstruct the state at the accepted boundary.
The compiler keeps Gated DeltaNet's working state in registers across the verification block. When a suffix is rejected, a replay kernel advances the checkpoint through only the accepted positions, using the retained inputs without repeating the large weight projections. Convolution history is committed to the same boundary. If the whole block is accepted, the engine uses the verified state directly.
Each request owns its draft context and target state. Two concurrent requests can accept different numbers of tokens in the same round; the engine keeps those histories separate while batching the work they can share.
Try it
Photon 2.6 is available through the Moondream Python package:
pip install --upgrade "moondream==2.6.1"To run Qwen3.5-27B with DFlash on a B200:
import moondream as md
from huggingface_hub import snapshot_download
draft = snapshot_download("z-lab/Qwen3.5-27B-DFlash")
with md.photon(
"Qwen/Qwen3.5-27B-FP8",
device="cuda",
draft_model_path=draft,
max_batch_size=1,
) as model:
result = model.chat(
[{"role": "user", "content": "Explain how binary search works."}],
reasoning=True,
settings={"temperature": 0, "max_tokens": 256},
)
print(result["message"])For concurrent Qwen3.5 requests, set max_batch_size to your desired concurrency and use page_size=64. Both configurations use the same DFlash checkpoint.
Both FP8 execution and causal verification are implemented in the compiler, rather than a separate Qwen-specific inference path. As we add models, they can use the same optimizations.
