Photon 2.0

Photon

Inference built for physical AI.

Fast responses at small batch sizes, lower memory use, and local execution. Each model is compiled for the target GPU.

Moondream · Qwen · Gemma · NVIDIA H100

52/52throughput tests led
H100 throughput

Higher throughput than vLLM and SGLang in matched H100 tests.

28 against vLLM 0.25.1 and 24 against SGLang 0.5.3. Photon led every test, with up to 133% higher throughput.

NVIDIA H100 SXM 80GB · ChartQA 512-request slice · concurrency 1–8See benchmark data · Read the launch post
Built for physical AI

Live systems have different constraints.

Chat engines optimize large queues. Cameras, robots, and machines need quick responses from a few active streams, often on local hardware.

Live control loop Local · online
01ObserveCamera frames
02InferCompiled Photon program
03ActMachine response
Small batches

Fast responses for one or a few live streams.

Low latency

Built for systems that act on the next frame.

Less memory

Fit models into practical deployment limits.

Runs locally

Keep inference close to cameras, robots, and machines.

Models

Production-ready support.

Photon 2.0 supports Moondream, Qwen 3.5, and Gemma 4 on NVIDIA H100. More models and chips will follow.

3model families · one runtime

Open vision-language models for local perception and physical AI.

QueryCaptionDetectPointSegment
Releases
23
Modalities
Image + text → text
Verified hardware
13 targets · A10 24GB, A40 48GB +11 more

Open multimodal models in four sizes, compiled through the same path.

QueryCaption
Releases
3.53.6
Sizes
0.8B · 2B · 4B · 9B
Modalities
Image + text → text
Verified hardware
1 target · H100 SXM 80GB

Google open multimodal models in two sizes, verified on H100.

QueryCaption
Releases
4
Sizes
E2B · E4B
Modalities
Image + text → text
Verified hardware
1 target · H100 SXM 80GB
How Photon works

A megakernel for each model and GPU.

Photon compiles the model into a target-specific megakernel. The runtime handles requests, tokenization, and image decoding.

InputModel graphMoondream, Qwen, or Gemma
CompilePhoton compilerCross-operator optimization
GenerateTarget-specific megakernelOne GPU program per model × chip
ServePhoton runtimeRequests, tokenization, decode on CPU

One compiler path for each model, chip, and deployment goal.

Performance

Faster execution. Lower startup time.

Measured on NVIDIA H100 against vLLM 0.25.1 and SGLang 0.5.3.

See the benchmark
NVIDIA H100 SXM 80GBChartQA 512-request slicec1–8
52/52Throughput tests ledPhoton compared with each runtime
vLLM 0.25.11.01×2.33×
SGLang 0.5.31.14×1.57×
Cold startup to 66% lower

13 matched model/runtime comparisons.

Updates

What changed.

New support, performance work, and fixes.

Aug 4, 2026Fix
Large-image CPU decode fix

Requests with images above 4 MP no longer fail on the CPU decode path before reaching the megakernel.

Aug 3, 2026Model support
Qwen 3.5 and Gemma 4 support

Photon 2.0 adds production support for Qwen 3.5 (0.8B-9B) and Gemma 4 E2B/E4B.

Aug 3, 2026Hardware support
H100 80GB support

Photon 2.0 is tested on H100 SXM 80GB across all launch models.

Your stack

Bring Photon to your workload.

Tell us which models, chips, and latency targets matter to your deployment.