Sulat.com
AI models
Inception logo

Model details

Mercury 2

Mercury 2 represents a departure from traditional autoregressive language models by utilizing a diffusion-based architecture. While standard models generate text sequentially, token-by-token, this model employs a coarse-to-fine process that iteratively refines outputs in parallel over a small number of steps. This design intent shifts the model's behavior from a one-way typewriter to an editor, allowing it to generate multiple tokens simultaneously. By moving away from sequential decoding, the model is built to overcome the latency bottlenecks inherent in standard inference, making it particularly well-suited for high-volume, multi-step agentic tasks.

The development of Mercury 2 is rooted in the foundational research of its creators, who have focused on commercializing diffusion techniques for text. By enabling reasoning-grade quality within real-time latency budgets, the model addresses the compounding delays often found in complex loops like retrieval pipelines, voice interactions, and automated coding. This approach provides a distinct speed-cost curve, allowing for high-throughput reasoning that remains responsive even as the complexity of the task increases. As production environments move toward more autonomous agent loops, this model is positioned to serve as a core component for applications where inference speed is a primary factor for successful deployment.

Inceptionmercury-2mercury

Quick Info

Powered by
Provider
Inception
Model key
mercury-2
Release date
Feb 24, 2026
Last updated
Feb 24, 2026
Knowledge cutoff
2025-01-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.25
Output token cost
$0.75

Limits

Output tokens
50,000 tokens
Context window
128,000 tokens

Transparent token rates

Compare Mercury 2 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Mercury 2

Inception

Coverage

Inception Labs introduces Mercury 2, a diffusion-based LLM designed for high-speed, multi-step reasoning tasks with a 128K context window.

Inception

Coverage

Inception's Mercury 2 diffusion LLM hits 1,196 tokens/sec at $0.25/M input, Meta signs $100B+ AMD compute deal, MatX raises $500M to challenge Nvidia's Rubin Ultra. (166 chars)

Inception

Coverage

The news blog specialized in Japanese culture, odd news, gadgets and all other funny stuffs. Updated everyday.

Inception

Coverage

Mercury 2 from Inception is the first diffusion-based reasoning model. Instead of generating text word by word, it refines entire passages in parallel, making it more than five times faster than conventional language models.

Inception

Coverage

Inception, the company behind the first commercial diffusion large language models (dLLMs), today announced the launch of Mercury 2, the fastest reasoning LL...

Videos about Mercury 2

More models around Mercury 2