Currently listed through these providers:
Model details
Inception: Mercury 2
Inception: Mercury 2 breaks from the tradition of autoregressive language models by using a diffusion-based architecture that generates multiple tokens simultaneously through parallel refinement. This diffusion approach eliminates the sequential bottleneck that limits traditional LLMs, allowing the model to produce and converge tokens in a small number of steps rather than one at a time. The design philosophy centers on production-grade reasoning speed, targeting applications where latency compounds across multi-step workflows—particularly agentic loops, retrieval pipelines, and interactive coding environments where every millisecond affects user experience and operational cost.
As the world's first reasoning diffusion language model, Mercury 2 brings tunable reasoning levels to a parallel generation framework, combining the depth of chain-of-thought processes with the throughput of diffusion inference. Native tool use, schema-aligned JSON output, and OpenAI API compatibility make it straightforward to integrate into existing stacks. The model demonstrates its practical value in coding agent loops, real-time search and voice applications, and anywhere high concurrency demands consistently low response times under load.
Quick Info
Powered by- Provider
- Kilo Gateway
- Model key
- inception/mercury-2
- Release date
- Mar 4, 2026
- Last updated
- Mar 4, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.25
- Output token cost
- $0.75
Limits
- Output tokens
- 50,000 tokens
- Context window
- 128,000 tokens
Transparent token rates
Compare Inception: Mercury 2 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Inception: Mercury 2
No articles yet. Fetch the latest news to show it here.