Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

Mercury 2.5

Mercury 2.5 is Inception’s production diffusion language model, designed to improve on Mercury 2 while retaining a low-latency serving profile. Rather than generating one token at a time, it can generate and refine multiple tokens in parallel, making it particularly relevant to responsive applications where interaction speed matters.

Inception positions Mercury 2.5 for latency-sensitive search, voice, and coding workloads, with capabilities aimed at tunable reasoning, parallel tool calls, and schema-aligned JSON. The creator reports a 40% intelligence increase over Mercury 2, while a third-party listing cites 1,107 tokens per second on standard GPUs, indicating a model intended to balance improved quality with practical generation speed.

Vercel AI Gatewayinception/mercury-2.5mercury

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
inception/mercury-2.5
Release date
Sep 8, 2026
Last updated
Sep 8, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.04
Output token cost
$0.15

Limits

Output tokens
65,536 tokens
Context window
260,000 tokens

Transparent token rates

Compare Mercury 2.5 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Mercury 2.5

OpenRouter

Official sourceAnnouncement

Inception announced Mercury 2.5, its most capable diffusion-based language model, in a blog post by CEO Stefano Ermon. The post describes Mercury 2.5 as a significant quality step up from Mercury 2 while preserving the same low-latency, low-cost serving profile, noting that since Mercury 2's launch enterprise usage has The post lists Mercury 2.5's specifications: 1,107 tokens per second on widely-available NVIDIA GPUs, a 260K-token context window, a reported 40% intelligence increase over Mercury 2, and capabilities including tunable reasoning, parallel tool calls, and schema-aligned JSON. Pricing is $0.20 per million input tokens an

Inception

Coverage

Nexforce published analysis on September 15, 2026, citing Inception's official announcement. It confirms Mercury 2.5 is a diffusion model, not an autoregressive transformer, and is the largest variant trained in that class. Instead of emitting tokens left-to-right, the dLLM generates a whole block of masked tokens in p The piece qualifies that autoregressive models served with continuous batching on Cerebras and Groq hardware already reach four digits of tokens/sec per request, so Mercury's differentiation is not raw speed alone but the economics of serving a diffusion architecture. Standard pricing is $0.20/M input and $0.75/M outpu

Inception

Coverage

Testomat.io published hands-on developer testing on September 11, 2026, integrating Mercury 2.5 as the planner model for Explorbot, an open-source QA agent for autonomous testing of web applications. The test team had previously used GPT-5.6 Luna as their planner after comparing four models. Mercury 2.5's 1,107 tokens/ The article explains the architectural contrast: autoregressive models generate output left-to-right one token at a time with a sequential floor under latency, while diffusion language models begin with a noisy or masked sequence and refine many token positions over denoising passes. This parallel refinement process is

Inception

Coverage

BenchLM.ai published a technical analysis on September 11, 2026, disclosing sponsorship by Inception and clarifying that all figures come from Inception's Mercury 2.5 announcement. The piece tabulates the headline specs: 1,107 tokens/sec output speed, 260K context window, standard pricing of $0.20 in / $0.75 out per mi The article frames the developer economics around 'invisible calls' — compaction, routing, query rewriting, tool selection, classification, extraction, and guardrail checks — that set the latency floor for downstream user-facing interactions. At 1,107 tokens/sec, a 2,000-token output lands in under two seconds, and at

OpenRouter

CoverageBenchmark

Gigazine reported on September 9, 2026 that Inception announced Mercury 2.5 on September 8, 2026, framing it as a diffusion-based AI model that generates text by creating and refining multiple tokens in parallel rather than sequentially. The article explains that Mercury 2.5 employs a "diffuse large-scale language mode According to the article, Mercury 2.5 outputs 1,107 tokens per second, with performance said to be on par with GPT-5.6 Luna on the Low setting and a standard price of approximately $0.75 (about 120 yen) per million output tokens. The piece highlights the practical motivation for diffusion decoding in agent pipelines wh

OpenRouter

Coverage

AI-TLDR's model directory entry confirms Mercury 2.5 was released on September 9, 2026 as Inception's diffusion language model. The entry details the architectural contrast with autoregressive models — Mercury 2.5 generates tokens in parallel rather than one at a time, which Inception credits for sub-300ms time-to-firs The directory notes Mercury 2.5 supports tool calling and structured outputs, ships with proprietary weights (API only, no public weights), and is served through the Inception API, Baseten, and OpenRouter — with OpenRouter listed as a distribution route rather than the creator. Two companion previews also shipped along

Inception

Coverage

Shattered.io provides long-form third-party analysis of the Mercury 2.5 launch dated September 8, 2026, confirming that Inception's Redwood City startup announced the model via Business Wire the same week as other major releases like K2 Horizon and Spark X2.5. The article emphasizes that Mercury 2.5's pitch centers on The piece explains the architectural distinction: Inception's dLLM generates and refines chunks of text in parallel like image diffusion models denoise pictures, rather than predicting tokens left-to-right autoregressively. Two companion products shipped in preview alongside Mercury 2.5: Mercury Voice, a voice interfac

Vercel AI Gateway

Coverage

This Rutland Herald article is a syndication of the same Business Wire press release announcing Mercury 2.5 on September 8, 2026. It repeats the core launch details: Mercury 2.5 runs over 1,100 tokens per second in production, is described as the most capable dLLM and fastest reasoning LLM in production, and builds on The article adds no independent technical detail beyond the originating press release but serves as additional confirmation of the launch date, creator attribution to Inception, and enterprise adoption claims. All performance figures remain provider-reported rather than independently verified.

Vercel AI Gateway

CoverageBenchmark

BenchLM's model card for Mercury 2.5 corroborates the September 8, 2026 release date and lists the API model ID as 'mercury-2.5' with proprietary reasoning and a 260K token context window. Launch pricing is confirmed at $0.04 per million input tokens and $0.15 per million output tokens, and the card records provider-re As an independent aggregator, BenchLM notes that only five benchmark rows are published and that several categories, including reasoning, math, multilingual, and multimodal, remain unmeasured. It also flags that independent runtime speed has not been measured, meaning all performance and benchmark figures cited on the

Vercel AI Gateway

Coverage

Inception announced Mercury 2.5 on September 8, 2026, positioning it as the most capable diffusion LLM and the fastest reasoning LLM in production. The Business Wire press release states Mercury 2.5 runs over 1,100 tokens per second in production, explaining that unlike autoregressive LLMs, dLLMs generate and refine to The release frames Mercury 2.5 as building on Mercury 2, which Inception says matched Claude Haiku and GPT Mini intelligence at roughly 10x the throughput. CEO and co-founder Stefano Ermon is quoted arguing that diffusion-based generation breaks the usual trade-off where LLMs cannot simultaneously get smarter, faster,

Videos about Mercury 2.5

More models around Mercury 2.5