Engineering
Blog posts in the Engineering category

Engineering
August 24, 2026Live video is a different inference workload. Here's why.
A 13× throughput collapse, minute-long cold boots, memory that never comes back: why datacenter inference engines break on live video, with receipts.
Read more
Engineering
June 4, 2026Popping the GPU Bubble
Photon, Moondream's inference engine, achieves near-realtime VLM inference (~33ms on NVIDIA B200). This is a peek into how it delivers up to 35% higher decode throughput by optimizing how the GPU works.
Read more