Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Kilo Gateway logo

Model details

Inception: Mercury 2.5

Inception introduced Mercury 2.5 on September 8, 2026, branding it as the next tier of intelligence in its line of diffusion-based large language models. Unlike conventional autoregressive LLMs that produce one token at a time, the diffusion approach generates tokens in parallel, a design choice the company frames as the key to achieving lower latency and better cost efficiency for production AI workloads. The announcement frames this release as evidence that diffusion architectures are maturing from a research curiosity into a deployable standard, citing enterprise adoption of the prior Mercury 2 generation as the proving ground for the new model.

The model is accessed through the Inception developer platform, where new accounts receive an initial allocation of free tokens and developers obtain API credentials to send chat-completion requests using the mercury-2.5 identifier. Practical fit centers on use cases that benefit from fast, parallel token generation at scale, such as latency-sensitive agents and high-throughput enterprise assistants that pair the model's generation speed with reasoning-style control parameters exposed in the API. For teams already invested in the earlier Mercury generation, Mercury 2.5 represents an upgrade path aimed at extracting more intelligent behavior from the same diffusion backbone rather than a shift to a new paradigm.

Kilo Gatewayinception/mercury-2.5mercury

Quick Info

Powered by
Provider
Kilo Gateway
Model key
inception/mercury-2.5
Release date
Sep 8, 2026
Last updated
Sep 8, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$0.75

Limits

Output tokens
65,536 tokens
Context window
260,000 tokens

Transparent token rates

Compare Inception: Mercury 2.5 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Inception: Mercury 2.5

Kilo Gateway

Coverage

ai-tldr.dev provides a third-party technical summary of Mercury 2.5, listing it as released on September 9, 2026 with a 260K-token context window, proprietary weights (API only), and a diffusion language model architecture that generates and refines tokens in parallel. Inception reports 1,107 tokens per second on widel The model supports tool calling and structured outputs and is served through the Inception API as well as Baseten and OpenRouter, with enterprise deployments offering dedicated capacity. Pricing is listed at $0.20 per million input tokens and $0.75 per million output, with an 80% promotional rate of $0.04 input and $0.

Kilo Gateway

Coverage

An independent editorial piece covering Inception's Mercury 2.5 launch on September 8, 2026, distributed via Business Wire and picked up by outlets including the Las Vegas Sun. The article confirms Inception's claim of pushing past 1,100 tokens per second in live deployments and contextualizes the release within a dens The piece adds detail on two companion preview products that shipped with Mercury 2.5: Mercury Voice, targeting a voice interface layered on the same diffusion backbone, and Mercury Router, which classifies requests and routes them to the appropriate Mercury model. Inception's own materials describe Mercury 2.5 as the

Kilo Gateway

CoverageBenchmark

BenchLM.ai's Mercury 2.5 model card records the release as September 8, 2026, describing it as Inception's diffusion reasoning model with a 260K context window and launch rates of $0.04 per million input tokens and $0.15 per million output tokens. Inception reports 79% on GPQA Diamond and 77% on IFBench, which the page The page's value is precisely its transparency: it separates verified from provisional and unmeasured rows, rather than presenting benchmark numbers as third-party validation. Capability scores and speed are absent or unranked, so this entry should be cited as evidence of the launch date, context window, pricing, and t

Kilo Gateway

Coverage

On September 8, 2026, Inception announced the launch of Mercury 2.5, positioning it as "the most capable dLLM and the fastest reasoning LLM in production," running over 1,100 tokens per second. The Redwood City, California-based company, founded by Stanford, UCLA, and Cornell researchers behind the first commercial dif The press release details that unlike autoregressive models that generate tokens one at a time, Inception's dLLM starts with a rough draft and refines tokens in parallel, decoupling cost and latency from reasoning depth. CEO and co-founder Stefano Ermon said: "Nobody in this industry thinks LLMs can get smarter, faster

Kilo Gateway

CoverageRelease Notes

A model release aggregator lists Inception: Mercury 2.5 as an API launch on 2026-09-08 with a 260k token context window, providing an independent corroboration of the launch date and context length. The entry appears alongside other recent API launches such as DeepSeek V4.1 Flash, Gemini 3.8 Flash, and Qwen3.8 Max, sit The aggregator offers no primary technical detail beyond the timestamp and context window, serving mainly as evidence that Mercury 2.5 reached general availability through OpenRouter on schedule. Inception is named as the originating provider on the Inception API, not the Kilo gateway. The thin, timestamp-only nature o

Videos about Inception: Mercury 2.5

More models around Inception: Mercury 2.5