OpenRouter
Mercury 2 targets structured tasks with schema-aligned JSON output; supports OpenAI API drop-in integration, for simpler deployment.
Model details
Mercury 2 represents a departure from traditional autoregressive language models by utilizing a diffusion-based architecture. While standard models generate text sequentially, token by token, this model functions through parallel refinement, allowing it to produce and revise multiple tokens simultaneously. This design intent focuses on solving the latency bottlenecks inherent in complex production environments, such as agent loops, real-time voice interactions, and automated coding tasks. By acting more like an editor refining a full draft rather than a typewriter, the model achieves significant throughput gains while maintaining reasoning-grade quality.
Developed by researchers from Stanford, UCLA, and Cornell, the model leverages the team's foundational work in diffusion to bring high-performance inference to real-world applications. By moving away from sequential decoding, the model achieves speeds exceeding 1,000 tokens per second on standard hardware, offering a distinct advantage for workflows where latency compounds across multiple steps. This architectural shift positions the model as a practical solution for high-value production deployments that require both rapid response times and the ability to handle complex, multi-step reasoning tasks at a lower cost.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
OpenRouter
Mercury 2 targets structured tasks with schema-aligned JSON output; supports OpenAI API drop-in integration, for simpler deployment.
OpenRouter
Inception Labs introduces Mercury 2, a diffusion-based LLM designed for high-speed, multi-step reasoning tasks with a 128K context window.
OpenRouter
Inception's Mercury 2 diffusion LLM hits 1,196 tokens/sec at $0.25/M input, Meta signs $100B+ AMD compute deal, MatX raises $500M to challenge Nvidia's Rubin Ultra. (166 chars)
OpenRouter
The news blog specialized in Japanese culture, odd news, gadgets and all other funny stuffs. Updated everyday.
OpenRouter
Benchmark scores and performance metrics for Inception: Mercury 2 - Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving >1,000 tokens/sec on standard GPUs. Mercury 2 i
This exact model name is also listed by 4 other providers.