Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Nvidia logo

Model details

nvidia-nemotron-nano-9b-v2

NVIDIA built Nemotron Nano 9B v2 from the ground up as a unified model that handles both reasoning and non-reasoning tasks in a single system. The architecture takes a hybrid approach, pairing the efficient selective state-space mechanics of Mamba-2 with just four attention layers, rather than relying on a traditional transformer-only design. This design allows the model to generate an internal reasoning trace before producing its final answer, which generally yields higher quality solutions on complex queries. Users can toggle this behavior on or off via system prompt depending on whether they want explicit reasoning steps or direct responses, though disabling reasoning traces results in slightly lower accuracy on harder problems that benefit from step-by-step thinking.

The model builds on the Nemotron-Hybrid lineage documented in NVIDIA's technical reports and draws on Qwen for post-training improvements to create the instruction-following version. Trained on data through May 2025, it supports six languages in the instruct variant and achieves strong results on reasoning benchmarks including 97.8% on MATH500 and 72.1% on AIME25, alongside competitive instruction-following scores. The architecture supports an extended context window of 128K tokens and includes runtime thinking budget control, giving developers flexibility to balance depth of reasoning against latency and token usage. With NVIDIA's open model license and NIM container deployment options, the model targets developers and researchers seeking a capable, deployable reasoning model that can adapt between analytical and conversational tasks.

Nvidianvidia/nvidia-nemotron-nano-9b-v2nemotrondeprecated

Quick Info

Powered by
Provider
Nvidia
Model key
nvidia/nvidia-nemotron-nano-9b-v2
Release date
Aug 18, 2025
Last updated
Aug 18, 2025
Knowledge cutoff
2024-09
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
131,072 tokens
Context window
131,072 tokens

Latest news about nvidia-nemotron-nano-9b-v2

No articles yet. Fetch the latest news to show it here.

Videos about nvidia-nemotron-nano-9b-v2

More models around nvidia-nemotron-nano-9b-v2