OpenRouter
Inception announced Mercury 2.5, its most capable diffusion-based language model, in a blog post by CEO Stefano Ermon. The post describes Mercury 2.5 as a significant quality step up from Mercury 2 while preserving the same low-latency, low-cost serving profile, noting that since Mercury 2's launch enterprise usage has The post lists Mercury 2.5's specifications: 1,107 tokens per second on widely-available NVIDIA GPUs, a 260K-token context window, a reported 40% intelligence increase over Mercury 2, and capabilities including tunable reasoning, parallel tool calls, and schema-aligned JSON. Pricing is $0.20 per million input tokens an