Currently listed through these providers:
Model details
Nemotron 3 Super 120B
Nemotron 3 Super is NVIDIA's open-weight model engineered for complex multi-agent and long-horizon workflows. It blends a Mamba-2 backbone with attention and a Mixture-of-Experts design, activating only about 12B parameters per token despite a 120B total, while multi-token prediction helps it generate tokens far more efficiently than comparable open models. A latent MoE layout lets it call four experts at the cost of one, and the architecture opens up an enormous working memory for sustained reasoning, cross-document analysis, and planning across many steps. The model also ships with a configurable reasoning mode and built-in support for tool calling and structured output, which makes it especially well-suited to agentic systems, retrieval-augmented pipelines, and IT-style automation that need reliable tool use rather than free-form chat.
The model was developed by NVIDIA between late 2025 and early 2026, with pre-training data reaching into mid-2025 and post-training data extending to early 2026. After pre-training it was refined with multi-environment reinforcement learning spanning more than ten settings, a recipe that lifts accuracy on demanding benchmarks such as AIME 2025, TerminalBench, and SWE-Bench Verified. Because it is released openly under the NVIDIA Nemotron Open Model License, with weights, datasets, and training recipes available, teams can fine-tune, distill, or deploy it from a workstation up to multi-GPU servers. In practice this combination of efficient MoE inference, a million-token context, and strong agentic tuning makes Nemotron 3 Super a forward-looking choice for high-volume enterprise workloads that need long memory, tool integration, and customization without surrendering open-weight flexibility.
Quick Info
Powered by- Provider
- Cloudflare Workers AI
- Model key
- @cf/nvidia/nemotron-3-120b-a12b
- Release date
- Mar 11, 2026
- Last updated
- Mar 11, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.50
- Output token cost
- $1.50
Limits
- Output tokens
- 256,000 tokens
- Context window
- 256,000 tokens
Transparent token rates
Compare Nemotron 3 Super 120B pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Nemotron 3 Super 120B
No articles yet. Fetch the latest news to show it here.