Sulat.com
AI models
OpenRouter logo

Model details

Mercury 2

Mercury 2 represents a departure from traditional autoregressive language models by utilizing a diffusion-based architecture. While standard models generate text sequentially, token by token, this model functions through parallel refinement, allowing it to produce and revise multiple tokens simultaneously. This design intent focuses on solving the latency bottlenecks inherent in complex production environments, such as agent loops, real-time voice interactions, and automated coding tasks. By acting more like an editor refining a full draft rather than a typewriter, the model achieves significant throughput gains while maintaining reasoning-grade quality.

Developed by researchers from Stanford, UCLA, and Cornell, the model leverages the team's foundational work in diffusion to bring high-performance inference to real-world applications. By moving away from sequential decoding, the model achieves speeds exceeding 1,000 tokens per second on standard hardware, offering a distinct advantage for workflows where latency compounds across multiple steps. This architectural shift positions the model as a practical solution for high-value production deployments that require both rapid response times and the ability to handle complex, multi-step reasoning tasks at a lower cost.

OpenRouterinception/mercury-2mercury

Quick Info

Powered by
Provider
OpenRouter
Model key
inception/mercury-2
Release date
Mar 4, 2026
Last updated
Mar 4, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.25
Output token cost
$0.75

Limits

Output tokens
50,000 tokens
Context window
128,000 tokens

Transparent token rates

Compare Mercury 2 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Mercury 2

OpenRouter

CoverageBenchmark

Mercury 2 targets structured tasks with schema-aligned JSON output; supports OpenAI API drop-in integration, for simpler deployment.

OpenRouter

Coverage

Inception Labs introduces Mercury 2, a diffusion-based LLM designed for high-speed, multi-step reasoning tasks with a 128K context window.

OpenRouter

Coverage

Inception's Mercury 2 diffusion LLM hits 1,196 tokens/sec at $0.25/M input, Meta signs $100B+ AMD compute deal, MatX raises $500M to challenge Nvidia's Rubin Ultra. (166 chars)

OpenRouter

Coverage

The news blog specialized in Japanese culture, odd news, gadgets and all other funny stuffs. Updated everyday.

OpenRouter

Official sourceBenchmark

Benchmark scores and performance metrics for Inception: Mercury 2 - Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving >1,000 tokens/sec on standard GPUs. Mercury 2 i

Videos about Mercury 2

More models around Mercury 2