Currently listed through these providers:
Model details
NVIDIA Nemotron Nano 9B v2
NVIDIA built Nemotron Nano 9B v2 as a unified model that handles both reasoning-heavy and standard tasks in a single architecture. The design centers on a hybrid Mamba-2 and MLP structure paired with just four attention layers, which lets it match or exceed Qwen3-8B accuracy while delivering up to six times the throughput. Users can control whether the model generates explicit reasoning traces before answering or provides direct responses, trading some accuracy on harder problems for faster output when preferred. The thinking budget is adjustable via system prompt, giving developers a dial to match response style to task demands.
NVIDIA trained this model from scratch and refined it using Qwen to create a commercially ready LLM. In reasoning-on mode it posted strong benchmark scores including 72.1% on AIME25, 97.8% on MATH500, and 64.0% on GPQA, showing solid performance across math and general reasoning tasks. The model supports English, German, Spanish, French, Italian, and Japanese, extending its utility beyond English-only deployments. This combination of efficient architecture, benchmark-validated reasoning, and multilingual support positions it well for enterprise applications needing reliable performance at scale.
Quick Info
Powered by- Provider
- Amazon Bedrock
- Model key
- nvidia.nemotron-nano-9b-v2
- Release date
- Aug 18, 2025
- Last updated
- Aug 18, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.06
- Output token cost
- $0.23
Limits
- Output tokens
- 8,192 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare NVIDIA Nemotron Nano 9B v2 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about NVIDIA Nemotron Nano 9B v2
No articles yet. Fetch the latest news to show it here.