Mercury 2.5 is a diffusion large language model (dLLM) developed by Inception, the company founded by Stanford, UCLA, and Cornell researchers who first commercialized diffusion-based language generation. Rather than predicting tokens one at a time like autoregressive models, it refines a rough draft of the entire response in parallel steps, a structural difference Inception uses to justify both its throughput claims and its positioning for latency-sensitive workloads. Inception describes Mercury 2.5 as representing a 40% intelligence gain over Mercury 2 and reports comparable quality to cost-optimized frontier models such as GPT-5.6 Luna at Low reasoning, Gemini 3.5 Flash-Lite, and Claude Haiku 4.5, with provider-reported scores of 79% on GPQA Diamond and 77% on IFBench.
The model is engineered for real-time applications where waiting for sequential token generation is the bottleneck: search and RAG pipelines that chain dozens of model calls, voice agents that need conversational responsiveness, and in-editor coding workflows where developers iterate quickly. Inception reports generation speeds above 1,100 tokens per second on widely available NVIDIA GPUs (1,107 tokens/sec in its most precise figure) alongside a 260,000-token context window, tunable reasoning, native tool use with parallel tool calls, and schema-aligned JSON output. Production case studies cited by Inception describe dramatic latency and cost reductions for customers like OpenCall and Augment Code after switching inference to Mercury, suggesting the diffusion design is most attractive where inference economics, rather than absolute benchmark dominance, drive model selection.