Inception
Inception Labs introduces Mercury 2, a diffusion-based LLM designed for high-speed, multi-step reasoning tasks with a 128K context window.
Model details
Mercury 2 represents a departure from traditional autoregressive language models by utilizing a diffusion-based architecture. While standard models generate text sequentially, token-by-token, this model employs a coarse-to-fine process that iteratively refines outputs in parallel over a small number of steps. This design intent shifts the model's behavior from a one-way typewriter to an editor, allowing it to generate multiple tokens simultaneously. By moving away from sequential decoding, the model is built to overcome the latency bottlenecks inherent in standard inference, making it particularly well-suited for high-volume, multi-step agentic tasks.
The development of Mercury 2 is rooted in the foundational research of its creators, who have focused on commercializing diffusion techniques for text. By enabling reasoning-grade quality within real-time latency budgets, the model addresses the compounding delays often found in complex loops like retrieval pipelines, voice interactions, and automated coding. This approach provides a distinct speed-cost curve, allowing for high-throughput reasoning that remains responsive even as the complexity of the task increases. As production environments move toward more autonomous agent loops, this model is positioned to serve as a core component for applications where inference speed is a primary factor for successful deployment.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Inception
Inception Labs introduces Mercury 2, a diffusion-based LLM designed for high-speed, multi-step reasoning tasks with a 128K context window.
Inception
Inception's Mercury 2 diffusion LLM hits 1,196 tokens/sec at $0.25/M input, Meta signs $100B+ AMD compute deal, MatX raises $500M to challenge Nvidia's Rubin Ultra. (166 chars)
Inception
The news blog specialized in Japanese culture, odd news, gadgets and all other funny stuffs. Updated everyday.
Inception
Mercury 2 from Inception is the first diffusion-based reasoning model. Instead of generating text word by word, it refines entire passages in parallel, making it more than five times faster than conventional language models.
Inception
Inception, the company behind the first commercial diffusion large language models (dLLMs), today announced the launch of Mercury 2, the fastest reasoning LL...
This exact model name is also listed by 4 other providers.