Currently listed through these providers:
Model details
Mercury Coder Small
Mercury Coder Small belongs to a generation of commercial-scale language models built by Inception Labs using a diffusion approach rather than conventional autoregressive decoding. The underlying architecture is a Transformer, but instead of generating one token at a time, the model is trained to predict multiple tokens in parallel, which is what enables its distinctive throughput characteristics. Mercury Coder was introduced as Inception Labs' first product line aimed specifically at coding workflows, shipping in two size variants labeled Mini and Small so users can choose a balance between latency and capacity. This diffusion-based design positions Mercury Coder as a practical option for developers who want rapid code completion, generation, and refactoring inside interactive environments where response speed materially affects the experience.
In head-to-head comparisons reported by the Inception Labs team, the Mercury Coder family has been shown to outperform speed-optimized frontier models by a wide margin on token throughput while preserving comparable code quality. Independent measurements cited in the technical report place Mercury Coder Mini and Small at the top of the speed-quality frontier, with real-world developer validation on Copilot Arena confirming strong quality rankings alongside the fastest available latency. Beyond raw performance, the models are commercially distributed through cloud marketplaces, giving teams a straightforward path to adoption without managing their own inference infrastructure. Mercury Coder Small therefore fits well into coding assistants, automated code review pipelines, and any production setting where both responsiveness and generation quality are non-negotiable.
Quick Info
Powered by- Provider
- NanoGPT
- Model key
- mercury-coder-small
- Release date
- Jan 1, 2024
- Last updated
- Jan 1, 2024
- Input modalities
- Output modalities
- Capabilities
- Base catalog fields only
Cost
- Input token cost
- $0.25
- Output token cost
- $1.00
Limits
- Input tokens
- 32,768 tokens
- Output tokens
- 16,384 tokens
- Context window
- 32,768 tokens