· Updates

Changes by release.

New model and hardware support, performance work, fixes, and compatibility changes.

12 of 12 updates

Photon 2.2

Sep 9, 2026
Hardware supportv2.2Sep 9, 2026

Fast inference across four NVIDIA GPU generations

Run the Photon model catalog across Ampere, Ada, Hopper, and Blackwell GPUs.

Models
All models
Hardware
NVIDIA A10/A10G · NVIDIA A100 · NVIDIA RTX 3090 · NVIDIA L4 · NVIDIA H100 · NVIDIA B200 · NVIDIA RTX PRO 6000 Blackwell
Known limitations
  • New Photon 2.2 vision and language performance numbers are not available yet.
Performancev2.2Sep 9, 2026

Faster local Moondream and Qwen inference

Local Moondream and Qwen workloads now return responses faster and process more requests.

Models
Moondream · Qwen
Hardware
NVIDIA A10/A10G · NVIDIA A100 · NVIDIA RTX 3090 · NVIDIA L4 · NVIDIA H100 · NVIDIA B200 · NVIDIA RTX PRO 6000 Blackwell
Benchmarks
Response latency and throughput improved across supported GPUs.
Known limitations
  • Updated vision and language benchmark numbers are not available yet.
Capabilityv2.2Sep 9, 2026

Accelerated Whisper transcription on L4 and RTX 3090

Run accelerated local Whisper transcription on NVIDIA L4 and RTX 3090 GPUs.

Models
Whisper
Hardware
NVIDIA L4 · NVIDIA RTX 3090
Known limitations
  • Photon 2.2 Whisper performance numbers for L4 and RTX 3090 are not available yet.
Performancev2.2Sep 9, 2026

Lower ASR latency and host overhead

Qwen3-ASR and Parakeet transcription now runs with lower latency and less host overhead, especially at concurrency 1.

Models
Qwen3-ASR 0.6B · Qwen3-ASR 1.7B · Parakeet TDT 0.6B v3
Hardware
NVIDIA L4 · NVIDIA A10/A10G · NVIDIA A100 SXM 80GB
Benchmarks
Photon beat the relevant reference engine in 35 of 36 matched speech throughput configurations, tied one, and reached up to 4.83x throughput.

Photon 2.0.1

Aug 4, 2026
Fixv2.0.1Aug 4, 2026

Large-image CPU decode fix

Requests with images above 4 MP no longer fail on the CPU decode path before reaching the megakernel.

Models
All models
Hardware
All hardware
Known limitations
  • Images above 16 MP are still rejected at the API boundary.

Photon 2.0

Aug 3, 2026
Hardware supportv2.0Aug 3, 2026

H100 80GB support

Photon 2.0 is tested on H100 SXM 80GB across all launch models.

Models
All models
Hardware
H100
Benchmarks
Photon was ahead in all 52 matched throughput comparisons: 28 against vLLM and 24 against SGLang, from 1.01x to 2.33x on ChartQA at concurrency 1-8.
Performancev2.0Aug 3, 2026

Lower cold-start time

Matched cold-start tests reached first response faster on Photon than vLLM and SGLang for the launch models.

Models
All models
Hardware
H100
Benchmarks
Photon had up to 66% lower total cold-start time than vLLM across seven models and up to 55% lower time than SGLang across six supported models.
Capabilityv2.0Aug 3, 2026

Moondream 3 segmentation

Moondream 3 on Photon serves segmentation masks alongside query, caption, detect, and point.

Models
Moondream 3
Hardware
H100
Runtimev2.0Aug 3, 2026

Compiled megakernel execution

Request handling, tokenization, and image decode stay on the CPU. Model execution runs inside one compiler-generated GPU program.

Models
All models
Hardware
All hardware
Compatibilityv2.0Aug 3, 2026

md.photon() requires moondream SDK 0.11+

md.photon() needs moondream 0.11 or later. Earlier SDKs stay on Moondream-only local inference.

Models
All models
Hardware
All hardware
Deprecationv2.0Aug 3, 2026

Photon 1.x edge support ended

Moondream 2 on A10 and Jetson AGX Orin is no longer supported for new deployments.

Models
Moondream 2
Hardware
A10 · Jetson AGX Orin
Known limitations
  • Photon 1.x configurations receive critical fixes only.

Production support · feeds: RSS · JSON