Currently listed through these providers:
Model details
Nemotron 3 Super 120B
The model uses a hybrid Mamba-Transformer architecture with Latent MoE, allowing it to selectively activate just 12 billion of its 120 billion total parameters during inference. This sparse activation pattern—routing to 4 experts at the cost of one—enables high throughput alongside strong accuracy. Multi-token prediction further boosts token generation efficiency, while the architecture combines the long-range reasoning strengths of transformers with the linear complexity benefits of Mamba. Multi-environment reinforcement learning over 10+ diverse settings built core agentic capabilities, with training recipes and datasets released openly. The model shows particular strength in multi-step reasoning, autonomous tool orchestration, and planning tasks that require sustained coherence across long contexts.
Its 256K token context window supports cross-document reasoning and multi-agent coordination where extended context is essential. Benchmark scores place it ahead of comparable non-reasoning models of similar weight class on agentic intelligence tasks, and it demonstrates leading accuracy on AIME 2025, TerminalBench, and SWE-Bench Verified—benchmarks that test real software engineering and mathematical reasoning ability. This combination of open weights, open training recipes, and strong agentic performance makes it practical for teams that want to customize, fine-tune, or deploy the model in their own infrastructure without being locked into a hosted API.
Quick Info
Powered by- Provider
- Perplexity Agent
- Model key
- nvidia/nemotron-3-super-120b-a12b
- Release date
- Mar 11, 2026
- Last updated
- Mar 11, 2026
- Knowledge cutoff
- 2026-02
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.25
- Output token cost
- $2.50
Limits
- Output tokens
- 32,000 tokens
- Context window
- 1,000,000 tokens
Latest news about Nemotron 3 Super 120B
No articles yet. Fetch the latest news to show it here.