Currently listed through these providers:
Model details
Nvidia Nemotron Nano 9B V2
The Nemotron Nano 9B V2 is built on a hybrid Mamba-Transformer architecture that combines Mamba-2 layers with a small number of attention layers, designed specifically to handle both reasoning and non-reasoning tasks efficiently. Rather than treating these as separate capabilities, NVIDIA trained this model as a unified system that can generate an internal reasoning trace before producing its final answer, which generally yields higher quality solutions on harder problems. Users can toggle reasoning traces on or off via system prompt depending on whether they want to see the model's intermediate thinking, though disabling them may slightly reduce accuracy on complex prompts. The architecture is optimized for generating long thinking traces without the computational overhead of full attention throughout, making it well-suited for tasks requiring step-by-step problem solving.
The model traces its roots to a 12-billion parameter base model that was pre-trained on 20 trillion tokens using an FP8 training recipe, then compressed and distilled down to 9 billion parameters using NVIDIA's Minitron strategy. This compression approach preserves much of the original model's capabilities while enabling deployment on more accessible hardware with up to 128K token context windows. The instruction-tuned version was further improved using Qwen-aligned training techniques and supports six languages including English, German, Spanish, French, Italian, and Japanese. In reasoning-enabled evaluation, the model achieved 72.1% on AIME25, 97.8% on MATH500, and 64.0% on GPQA, demonstrating competitive performance for a compact model in mathematical and analytical tasks.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- nvidia/nemotron-nano-9b-v2
- Release date
- Aug 18, 2025
- Last updated
- Aug 18, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.06
- Output token cost
- $0.23
Limits
- Output tokens
- 131,072 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare Nvidia Nemotron Nano 9B V2 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Nvidia Nemotron Nano 9B V2
No articles yet. Fetch the latest news to show it here.