Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Venice AI logo

Model details

Mercury 2.5

Mercury 2.5 is a diffusion large language model (dLLM) developed by Inception, the company founded by Stanford, UCLA, and Cornell researchers who first commercialized diffusion-based language generation. Rather than predicting tokens one at a time like autoregressive models, it refines a rough draft of the entire response in parallel steps, a structural difference Inception uses to justify both its throughput claims and its positioning for latency-sensitive workloads. Inception describes Mercury 2.5 as representing a 40% intelligence gain over Mercury 2 and reports comparable quality to cost-optimized frontier models such as GPT-5.6 Luna at Low reasoning, Gemini 3.5 Flash-Lite, and Claude Haiku 4.5, with provider-reported scores of 79% on GPQA Diamond and 77% on IFBench.

The model is engineered for real-time applications where waiting for sequential token generation is the bottleneck: search and RAG pipelines that chain dozens of model calls, voice agents that need conversational responsiveness, and in-editor coding workflows where developers iterate quickly. Inception reports generation speeds above 1,100 tokens per second on widely available NVIDIA GPUs (1,107 tokens/sec in its most precise figure) alongside a 260,000-token context window, tunable reasoning, native tool use with parallel tool calls, and schema-aligned JSON output. Production case studies cited by Inception describe dramatic latency and cost reductions for customers like OpenCall and Augment Code after switching inference to Mercury, suggesting the diffusion design is most attractive where inference economics, rather than absolute benchmark dominance, drive model selection.

Venice AImercury-2-5mercury

Quick Info

Powered by
Provider
Venice AI
Model key
mercury-2-5
Release date
Sep 8, 2026
Last updated
Sep 9, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.05
Output token cost
$0.1875

Limits

Output tokens
65,536 tokens
Context window
260,000 tokens

Transparent token rates

Compare Mercury 2.5 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Mercury 2.5

Venice AI

Coverage

Inception Labs launched Mercury 2.5 on 2026-09-08 as a diffusion large language model, described as the largest dLLM trained to date with roughly 40% more intelligence than Mercury 2, 1,107 tokens per second on NVIDIA GPUs, a 260K-token context window, a standard price of US$0.20 per million input tokens and US$0.75 pe The piece positions Mercury 2.5's quality as comparable to GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5, and notes two companion products shipped in preview on the same day: Mercury Voice and Mercury Router, with Mercury Router placing a vendor inside the routing layer that decides which model serves

Venice AI

Coverage

Inception Labs (founded by Stanford, UCLA, and Cornell faculty) launched Mercury 2.5 on 2026-09-08, described as the first commercially available diffusion LLM to reach production-scale performance, generating 1,107 tokens per second on NVIDIA GPUs, more than 3x faster than traditional autoregressive LLMs, while consum Mercury 2.5 is built on Inception's proprietary diffusion transformer architecture with a 260K-token context window (covering entire codebases, research papers, or legal contracts in a single pass), fine-grained control for constraining outputs to schemas or semantic requirements, and multimodal readiness (the current

Venice AI

Coverage

Testomat.io ran Mercury 2.5 hands-on as the planner model in its open-source QA agent Explorbot, motivated by the diffusion LLM's claim of 1,107 tokens per second on widely available NVIDIA GPUs. Each model call pauses the browser, so faster responses could let Explorbot complete more actions and test more scenarios in The article explains that diffusion language models begin with a noisy or masked sequence and refine many token positions across denoising passes, an iterative but non-sequential process, in contrast to autoregressive models that generate one token at a time left-to-right with a sequential floor on generation latency.

Venice AI

Coverage

Inception's Mercury 2.5 launch figures, as restated in this sponsored BenchLM piece, are 1,107 tokens/second output speed, a 260K-token context window (roughly doubling Mercury 2's 128K), standard pricing of $0.20 input / $0.75 output per million tokens, and a launch price at an 80% discount of $0.04 input / $0.15 outp Inception reports Mercury 2.5's quality is comparable to GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5, positioning it in the high-volume cost-optimized tier. The article frames Mercury 2.5 as suited to the dozens of invisible calls per request (compaction, routing, query rewriting, tool selection, cl

Venice AI

Coverage

TestingCatalog reports specific Mercury 2.5 technical details: 1,107 tokens per second on widely available NVIDIA GPUs, a 260K-token context window, and a 40% intelligence gain over Mercury 2 while retaining the same low-latency, low-cost serving profile. Inception claims quality is comparable to cost-optimized frontie Standard pricing is $0.20 per million input tokens and $0.75 per million output tokens, with an 80% launch discount reducing those to $0.04 and $0.15 per million respectively. Inception says Mercury 2.5 was shaped by customer feedback and production failures from Mercury 2's rapid adoption, and it targets latency-sensi

Venice AI

CoverageBenchmark

On 2026-09-08, AI company Inception announced Mercury 2.5, a model that generates text using a diffusion mechanism rather than autoregression: it creates and modifies multiple tokens in parallel, reportedly outputting 1,107 tokens per second. Its performance is described as comparable to GPT-5.6 Luna at the Low setting Inception's dLLM approach starts from a rough state and completes the response through multiple processing stages, correcting multiple tokens in parallel, applying the same parallel-refinement concept used by image-generation diffusion. Inception announced Mercury 2 in February 2026 as the prior "world's fastest" diffu

Venice AI

Coverage

Inception dropped Mercury 2.5 on 2026-09-08 via Business Wire, claiming over 1,100 tokens per second in live deployments with a diffusion LLM (dLLM) architecture that generates and refines chunks of text in parallel rather than predicting the next token autoregressively. The Redwood City startup frames Mercury 2.5 as t The launch landed in the same week as the Institute of Foundation Models' K2 Horizon and iFlytek's Spark X2.5, making September 2026 a dense AI release stretch. Two companion products shipped in preview alongside Mercury 2.5: Mercury Voice, a voice interface layered on the same diffusion backbone, and Mercury Router, w

Venice AI

Coverage

Morningstar syndicates the Business Wire press release announcing Inception's launch of Mercury 2.5 on September 8, 2026. The release identifies Inception as the company behind the first commercial diffusion large language models and describes Mercury 2.5 as the most capable dLLM and fastest reasoning LLM in production The announcement positions diffusion as having moved from a research bet to a production standard, with enterprise Mercury usage growing by over an order of magnitude since Mercury 2 to thousands of developers and dozens of enterprises. Stefano Ermon, CEO and co-founder of Inception, is quoted stating that the industry

Videos about Mercury 2.5

More models around Mercury 2.5