Blog
Product Update
September 24, 2026Photon can now speak
Up to 70% lower speech startup latency than vLLM-Omni. Qwen3-TTS and Kokoro are now available in Photon.
Read more

Model Release
September 22, 2026Introducing Moondream Parakeet Redux and Parakeet Ultra
Meet Moondream Parakeet Redux and Parakeet Ultra: speech-to-text models for local transcription in 25 languages, on CPUs and GPUs.
Read more

Announcement
September 9, 2026Photon 2.2: Five New NVIDIA GPUs, Ampere to Blackwell
Photon 2.2 adds five NVIDIA GPUs, Ampere to Blackwell. Parakeet runs 3.1x to 4.8x NeMo's throughput. Qwen3-ASR beats vLLM in 23 of 24 benchmarked configurations.
Read more

Announcement
September 1, 2026Photon 2.1: Speech Recognition on H100 and B200, Up to 3.1× Faster
Photon 2.1 adds automatic speech recognition for Whisper, Qwen3-ASR, and Parakeet, beating each model's reference engine in every matched test, and brings the full model catalog to NVIDIA B200.
Read more

Engineering
August 24, 2026Live video is a different inference workload. Here's why.
A 13× throughput collapse, minute-long cold boots, memory that never comes back: why datacenter inference engines break on live video, with receipts.
Read more

Announcement
August 3, 2026Photon 2.0: Inference engine for Physical AI
Photon compiles models into GPU programs optimized for the chip and the job. The first release supports Moondream, Qwen, and Gemma on NVIDIA H100, and wins every matched throughput test against vLLM and SGLang.
Read more

Model Release
July 7, 2026Moondream 3.1: Beyond Benchmarks
Moondream 3.1 launches with best-in-class benchmark scores, a new fine-tuning motion that transfers to your tasks, and a partnership with Cloudflare.
Read more

Announcement
June 8, 2026Photon is now free
Photon 1.3.0 makes Moondream faster across NVIDIA, Mac, and Windows, runs finetunes on far more hardware, fixes an accuracy issue on older GPUs — and running Moondream locally is now completely free.
Read more
Engineering
June 4, 2026Popping the GPU Bubble
Photon, Moondream's inference engine, achieves near-realtime VLM inference (~33ms on NVIDIA B200). This is a peek into how it delivers up to 35% higher decode throughput by optimizing how the GPU works.
Read more