All posts
Product Update

400 tok/s with Qwen3.5

400 tok/s with Qwen3.5

Photon 2.6 runs Qwen3.5-27B at 411 tokens per second on a single NVIDIA B200, increasing to 825 tok/s with two concurrent requests. It outperforms vLLM at every concurrency we tested, from one to eight requests.

This release adds FP8 support and DFlash speculative decoding to Photon. Under the hood, we extended our megakernel compiler to handle quantized projections and causal verification of draft tokens. Qwen3.5's recurrent layers made verification particularly interesting: when a draft is rejected, we need to recover the model state without running the expensive projections again.

Faster on one GPU

We ran Qwen3.5-27B-FP8 with the same DFlash checkpoint in Photon and vLLM 0.30.0, with thinking enabled in both. Photon is ahead at every concurrency we tested:

Throughput per request for Qwen3.5-27B-FP8 with DFlash on one B200. At 1, 2, 4, and 8 concurrent requests, Photon averages 411, 412, 315, and 276 tokens per second per request; vLLM averages 346, 335, 272, and 249. Values are total output throughput divided by concurrency.

At two concurrent requests, Photon maintains over 400 tok/s per request. The chart divides total throughput by concurrency; the table below shows total throughput.

Concurrent requestsPhotonvLLMSpeedup
1411 tok/s346 tok/s1.19x
2825 tok/s671 tok/s1.23x
41,259 tok/s1,089 tok/s1.16x
82,206 tok/s1,989 tok/s1.11x

These numbers measure the full serving system, including prefill and host overhead, using the released packages from PyPI.

The cost of serving at 400 tok/s

Assume a B200 costs $6.25 per hour, and we keep two requests running at the measured throughput. At 825 tok/s total, that's 2.97 million output tokens per hour, or $2.10 in GPU rental per million output tokens, while maintaining over 400 tok/s per request.

The cost is comparable to hosted APIs, while per-request speed is much higher than the rates reported in our September 29 snapshot of OpenRouter's Qwen3.5-27B listings: 9 to 22 tok/s across its listed providers, compared with 412 tok/s per request in our two-request benchmark.

Qwen3.5-27B speed and output cost. Photon: 412 tok/s per request and $2.10 GPU rental per million output tokens. OpenRouter reported P50 speeds and output prices: DeepInfra 22 tok/s and $2.60, Alibaba 21 and $1.56, AtlasCloud 13 and $2.16, NovitaAI 12 and $2.40, Phala 10 and $2.40, SiliconFlow 9 and $2.00. Provider snapshot from September 29, 2026.

OpenRouter's speeds are P50 measurements of provider traffic; Photon's figure comes from the benchmark above. Provider prices as of September 29, 2026:

ProviderInput / MOutput / MReported P50 tok/s
Alibaba$0.195$1.5621
SiliconFlow$0.25$2.009
AtlasCloud$0.27$2.1613
Phala$0.30$2.4010
NovitaAI$0.30$2.4012
DeepInfra$0.26$2.6022

Photon's GPU cost falls within that output-price range. The $2.10 estimate allocates the entire GPU rental cost, including time spent processing inputs, to output tokens. It assumes the benchmark workload runs continuously; idle time and operating costs are additional. API providers charge separately for input tokens.

Adding FP8 to the compiler

FP8 reduces weight storage from sixteen bits per value to eight, plus block scales. This is especially useful at low concurrency, where reading the weights can take longer than doing the arithmetic.

We added FP8 projections to the megakernel compiler, including activation quantization for tensor-core computation. The compiler can pipeline weight loading with compute and combine normalization with quantization to avoid writing an intermediate activation back to GPU memory.

The compiler chooses the tile geometry, too. A tile built for a large batch can leave most of the GPU idle when verifying a short sequence. Smaller tiles expose more parallel work, but change weight reuse, register pressure, and synchronization costs. Our cost model weighs those costs when selecting the execution plan.

Speculative verification in a megakernel

Ordinary decoding runs the target model once per new token. DFlash proposes a block of tokens with a smaller draft model, then asks the target to verify them in one causal pass. Keep the accepted prefix, discard the rejected suffix, and draft again.

One pass through the target's weights can now produce several output tokens. That is a good fit for a megakernel, but it isn't ordinary batching: sixteen positions in one sequence depend on each other. Sixteen independent requests don't.

The compiler now represents that causal sequence directly. At one concurrent request, Qwen3.5's target verification runs as a compiled megakernel, carrying work across operator boundaries without a CPU dispatch between each operation. Photon handles drafting, acceptance, and state commit around it.

Handling recurrent state

Qwen3.5 mixes full attention with Gated DeltaNet, a recurrent architecture. Rejected attention-cache positions can be hidden by shortening the visible history. A recurrent state has already incorporated their values. Changing a length counter won't undo that.

Our approach is inspired by ReplaySSM, which caches recent inputs and reconstructs recurrent state when needed. For generated verification, we use checkpoint-and-replay: keep the state from before the candidate block, retain the projected inputs, and reconstruct the state at the accepted boundary.

The compiler keeps Gated DeltaNet's working state in registers across the verification block. When a suffix is rejected, a replay kernel advances the checkpoint through only the accepted positions, using the retained inputs without repeating the large weight projections. Convolution history is committed to the same boundary. If the whole block is accepted, the engine uses the verified state directly.

DFlash supplies a block containing the current token and fifteen drafts. The target megakernel verifies the causal sequence and retains projected inputs. The example accepts four positions; replay advances the checkpoint through those positions to produce the next committed state, without repeating weight projections.

Each request owns its draft context and target state. Two concurrent requests can accept different numbers of tokens in the same round; the engine keeps those histories separate while batching the work they can share.

Try it

Photon 2.6 is available through the Moondream Python package:

pip install --upgrade "moondream==2.6.1"

To run Qwen3.5-27B with DFlash on a B200:

import moondream as md
from huggingface_hub import snapshot_download

draft = snapshot_download("z-lab/Qwen3.5-27B-DFlash")

with md.photon(
    "Qwen/Qwen3.5-27B-FP8",
    device="cuda",
    draft_model_path=draft,
    max_batch_size=1,
) as model:
    result = model.chat(
        [{"role": "user", "content": "Explain how binary search works."}],
        reasoning=True,
        settings={"temperature": 0, "max_tokens": 256},
    )
    print(result["message"])

For concurrent Qwen3.5 requests, set max_batch_size to your desired concurrency and use page_size=64. Both configurations use the same DFlash checkpoint.

Both FP8 execution and causal verification are implemented in the compiler, rather than a separate Qwen-specific inference path. As we add models, they can use the same optimizations.

Product Update