Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

Nvidia Nemotron Nano 9B V2

The Nemotron Nano 9B V2 is built on a hybrid Mamba-Transformer architecture that combines Mamba-2 layers with a small number of attention layers, designed specifically to handle both reasoning and non-reasoning tasks efficiently. Rather than treating these as separate capabilities, NVIDIA trained this model as a unified system that can generate an internal reasoning trace before producing its final answer, which generally yields higher quality solutions on harder problems. Users can toggle reasoning traces on or off via system prompt depending on whether they want to see the model's intermediate thinking, though disabling them may slightly reduce accuracy on complex prompts. The architecture is optimized for generating long thinking traces without the computational overhead of full attention throughout, making it well-suited for tasks requiring step-by-step problem solving.

The model traces its roots to a 12-billion parameter base model that was pre-trained on 20 trillion tokens using an FP8 training recipe, then compressed and distilled down to 9 billion parameters using NVIDIA's Minitron strategy. This compression approach preserves much of the original model's capabilities while enabling deployment on more accessible hardware with up to 128K token context windows. The instruction-tuned version was further improved using Qwen-aligned training techniques and supports six languages including English, German, Spanish, French, Italian, and Japanese. In reasoning-enabled evaluation, the model achieved 72.1% on AIME25, 97.8% on MATH500, and 64.0% on GPQA, demonstrating competitive performance for a compact model in mathematical and analytical tasks.

Vercel AI Gatewaynvidia/nemotron-nano-9b-v2nemotron

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
nvidia/nemotron-nano-9b-v2
Release date
Aug 18, 2025
Last updated
Aug 18, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.06
Output token cost
$0.23

Limits

Output tokens
131,072 tokens
Context window
131,072 tokens

Transparent token rates

Compare Nvidia Nemotron Nano 9B V2 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nvidia Nemotron Nano 9B V2

No articles yet. Fetch the latest news to show it here.

Videos about Nvidia Nemotron Nano 9B V2

More models around Nvidia Nemotron Nano 9B V2