Currently listed through these providers:
Model details
Mercury 2
Mercury 2 is designed for production reasoning workflows that involve search, planning, and retrieval. Its diffusion-style generation avoids the sequential token decoding of conventional autoregressive models, while streaming can display the response being refined in real time. This architecture is especially relevant to agents that need to coordinate multiple steps quickly and remain responsive during ongoing use.
Inception evaluated Mercury 2 on PinchBench, an open-source agent benchmark built on OpenClaw, where the company reports a 78% task success rate. That result matched or exceeded the figures Inception cited for GPT-5 Mini, Gemini 2.5 Flash, DeepSeek Chat, and GPT-4o, alongside the fastest execution time among models with comparable accuracy. The practical fit is strongest for latency-aware research agents, parallel search systems, and other workflows that benefit from inspectable intermediate results and source-grounded responses.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- inception/mercury-2
- Release date
- Feb 24, 2026
- Last updated
- Mar 6, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.25
- Output token cost
- $0.75
Limits
- Output tokens
- 128,000 tokens
- Context window
- 128,000 tokens
Transparent token rates
Compare Mercury 2 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Mercury 2
Videos about Mercury 2
More models around Mercury 2
This exact model name is also listed by 4 other providers.