Fast inference across four NVIDIA GPU generations
Run the Photon model catalog across Ampere, Ada, Hopper, and Blackwell GPUs.
- New Photon 2.2 vision and language performance numbers are not available yet.
Run the Photon model catalog across Ampere, Ada, Hopper, and Blackwell GPUs.
Local Moondream and Qwen workloads now return responses faster and process more requests.
Run accelerated local Whisper transcription on NVIDIA L4 and RTX 3090 GPUs.
Qwen3-ASR and Parakeet transcription now runs with lower latency and less host overhead, especially at concurrency 1.
Requests with images above 4 MP no longer fail on the CPU decode path before reaching the megakernel.
Photon 2.0 adds production support for Qwen 3.5 (0.8B-9B) and Gemma 4 E2B/E4B.
Photon 2.0 is tested on H100 SXM 80GB across all launch models.
Matched cold-start tests reached first response faster on Photon than vLLM and SGLang for the launch models.
Moondream 3 on Photon serves segmentation masks alongside query, caption, detect, and point.
Request handling, tokenization, and image decode stay on the CPU. Model execution runs inside one compiler-generated GPU program.
md.photon() needs moondream 0.11 or later. Earlier SDKs stay on Moondream-only local inference.
Moondream 2 on A10 and Jetson AGX Orin is no longer supported for new deployments.
Production support · feeds: RSS · JSON