Inception
Inception introduced Mercury 2, described as the world's fastest reasoning language model, built on diffusion-based parallel refinement rather than autoregressive sequential decoding. The post claims 5x faster generation, 1,009 tokens/sec on NVIDIA Blackwell GPUs, a 128K context window, native tool use, tunable reasoni The launch post argues that diffusion-based reasoning enables reasoning-grade quality inside real-time latency budgets, because higher intelligence no longer requires linearly more sequential test-time compute. An included quote from NVIDIA's Shruti Koparkar highlights Mercury 2 surpassing 1,000 tokens/sec on NVIDIA GP