Photon
Inference built for physical AI.
Fast responses at small batch sizes, lower memory use, and local execution. Each model is compiled for the target GPU.
Moondream · Qwen · Gemma · NVIDIA H100
Higher throughput than vLLM and SGLang in matched H100 tests.
28 against vLLM 0.25.1 and 24 against SGLang 0.5.3. Photon led every test, with up to 133% higher throughput.
Live systems have different constraints.
Chat engines optimize large queues. Cameras, robots, and machines need quick responses from a few active streams, often on local hardware.
Production-ready support.
Photon 2.0 supports Moondream, Qwen 3.5, and Gemma 4 on NVIDIA H100. More models and chips will follow.
A megakernel for each model and GPU.
Photon compiles the model into a target-specific megakernel. The runtime handles requests, tokenization, and image decoding.
One compiler path for each model, chip, and deployment goal.
Faster execution. Lower startup time.
Measured on NVIDIA H100 against vLLM 0.25.1 and SGLang 0.5.3.
13 matched model/runtime comparisons.
What changed.
New support, performance work, and fixes.
Requests with images above 4 MP no longer fail on the CPU decode path before reaching the megakernel.
Photon 2.0 adds production support for Qwen 3.5 (0.8B-9B) and Gemma 4 E2B/E4B.
Photon 2.0 is tested on H100 SXM 80GB across all launch models.
Bring Photon to your workload.
Tell us which models, chips, and latency targets matter to your deployment.