Sulat.com
AI models
Kilo Gateway logo

Model details

Inception: Mercury 2

Inception: Mercury 2 breaks from the tradition of autoregressive language models by using a diffusion-based architecture that generates multiple tokens simultaneously through parallel refinement. This diffusion approach eliminates the sequential bottleneck that limits traditional LLMs, allowing the model to produce and converge tokens in a small number of steps rather than one at a time. The design philosophy centers on production-grade reasoning speed, targeting applications where latency compounds across multi-step workflows—particularly agentic loops, retrieval pipelines, and interactive coding environments where every millisecond affects user experience and operational cost.

As the world's first reasoning diffusion language model, Mercury 2 brings tunable reasoning levels to a parallel generation framework, combining the depth of chain-of-thought processes with the throughput of diffusion inference. Native tool use, schema-aligned JSON output, and OpenAI API compatibility make it straightforward to integrate into existing stacks. The model demonstrates its practical value in coding agent loops, real-time search and voice applications, and anywhere high concurrency demands consistently low response times under load.

Kilo Gatewayinception/mercury-2mercury

Quick Info

Powered by
Provider
Kilo Gateway
Model key
inception/mercury-2
Release date
Mar 4, 2026
Last updated
Mar 4, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.25
Output token cost
$0.75

Limits

Output tokens
50,000 tokens
Context window
128,000 tokens

Transparent token rates

Compare Inception: Mercury 2 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Inception: Mercury 2

No articles yet. Fetch the latest news to show it here.

Videos about Inception: Mercury 2

More models around Inception: Mercury 2