Sulat.com
AI models
Vercel AI Gateway logo

Model details

Mercury 2

Mercury 2 is designed for production reasoning workflows that involve search, planning, and retrieval. Its diffusion-style generation avoids the sequential token decoding of conventional autoregressive models, while streaming can display the response being refined in real time. This architecture is especially relevant to agents that need to coordinate multiple steps quickly and remain responsive during ongoing use.

Inception evaluated Mercury 2 on PinchBench, an open-source agent benchmark built on OpenClaw, where the company reports a 78% task success rate. That result matched or exceeded the figures Inception cited for GPT-5 Mini, Gemini 2.5 Flash, DeepSeek Chat, and GPT-4o, alongside the fastest execution time among models with comparable accuracy. The practical fit is strongest for latency-aware research agents, parallel search systems, and other workflows that benefit from inspectable intermediate results and source-grounded responses.

Vercel AI Gatewayinception/mercury-2mercury

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
inception/mercury-2
Release date
Feb 24, 2026
Last updated
Mar 6, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.25
Output token cost
$0.75

Limits

Output tokens
128,000 tokens
Context window
128,000 tokens

Transparent token rates

Compare Mercury 2 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Mercury 2

Videos about Mercury 2

More models around Mercury 2