Currently listed through these providers:
Model details
Nemotron 3 Super
Nemotron 3 Super is built around a hybrid architecture that blends Mamba-2 state-space layers with mixture-of-experts routing and transformer attention, a combination designed to give the model both the speed of efficient state models and the rich pattern-matching that attention mechanisms provide. Multi-Token Prediction allows it to generate multiple tokens per forward pass rather than processing sequentially, which supports longer coherent outputs in a single pass. The design philosophy centers on agentic reasoning: the ability to sustain deep context over very long interactions, decompose multi-step problems, and maintain coherent task state across agent loops. This architecture positions the model for scenarios where a single query or task may involve hundreds of thousands of tokens of context and require multi-turn tool usage.
The model draws from the broader Nemotron family lineage, with NVIDIA making both pre-training and post-training datasets publicly available on Hugging Face, a signal of transparency that supports reproducibility and community refinement. The LatentMoE routing means that at inference time only a fraction of the total parameters activate per token, keeping compute manageable even though the model spans 120 billion parameters in total. Its 1M token context window makes it well-suited for tasks like comprehensive codebase analysis, multi-document reasoning, and complex agent pipelines where earlier turns of conversation or context directly inform later responses. The open-weights release invites developers to fine-tune or deploy it for specialized agentic workflows, while its reasoning and tool-calling capabilities suggest it was cultivated with autonomous problem-solving scenarios in mind.
Quick Info
Powered by- Provider
- Nvidia
- Model key
- nvidia/nemotron-3-super-120b-a12b
- Release date
- Mar 11, 2026
- Last updated
- Mar 11, 2026
- Knowledge cutoff
- 2024-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.20
- Output token cost
- $0.80
Limits
- Output tokens
- 262,144 tokens
- Context window
- 262,144 tokens
Transparent token rates
Compare Nemotron 3 Super pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Nemotron 3 Super
No articles yet. Fetch the latest news to show it here.