Inception Labs releases Mercury 2.5 diffusion LLM, claims 40% quality jump
Inception Labs released Mercury 2.5, its most capable production model, claiming a 40% increase in intelligence over Mercury 2 while keeping the same low-latency, low-cost serving profile. The company says it is the largest diffusion language model ever trained and the most capable diffusion LLM on the market, with quality comparable to cost-optimized frontier models such as GPT-5.6 Luna, Gemini 3.5 Flash-Lite, and Claude Haiku 4.5.
Mercury 2.5 runs at 1,107 tokens per second on widely available NVIDIA GPUs, supports a 260K-token context, and adds tunable reasoning, parallel tool calls, and schema-aligned JSON. Pricing is $0.20 per million input tokens and $0.75 per million output, with an 80% launch discount to $0.04 and $0.15.
CEO Stefano Ermon said the model is the first result of a training loop sharpened by customer feedback and production failure cases from search, voice, and coding workloads. Inception cites production customers: voice-agent maker OpenCall reported median model response latency near 170 milliseconds, and Augment Code cut context-compaction latency by 82% and cost by 90% after switching to Mercury.
Mercury 2.5 shows diffusion-based language models maturing into production systems that compete with frontier transformer models on quality while keeping latency and cost low.