Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Cloudflare Workers AI logo

Model details

Nemotron 3 Super 120B

Nemotron 3 Super is NVIDIA's open-weight model engineered for complex multi-agent and long-horizon workflows. It blends a Mamba-2 backbone with attention and a Mixture-of-Experts design, activating only about 12B parameters per token despite a 120B total, while multi-token prediction helps it generate tokens far more efficiently than comparable open models. A latent MoE layout lets it call four experts at the cost of one, and the architecture opens up an enormous working memory for sustained reasoning, cross-document analysis, and planning across many steps. The model also ships with a configurable reasoning mode and built-in support for tool calling and structured output, which makes it especially well-suited to agentic systems, retrieval-augmented pipelines, and IT-style automation that need reliable tool use rather than free-form chat.

The model was developed by NVIDIA between late 2025 and early 2026, with pre-training data reaching into mid-2025 and post-training data extending to early 2026. After pre-training it was refined with multi-environment reinforcement learning spanning more than ten settings, a recipe that lifts accuracy on demanding benchmarks such as AIME 2025, TerminalBench, and SWE-Bench Verified. Because it is released openly under the NVIDIA Nemotron Open Model License, with weights, datasets, and training recipes available, teams can fine-tune, distill, or deploy it from a workstation up to multi-GPU servers. In practice this combination of efficient MoE inference, a million-token context, and strong agentic tuning makes Nemotron 3 Super a forward-looking choice for high-volume enterprise workloads that need long memory, tool integration, and customization without surrendering open-weight flexibility.

Cloudflare Workers AI@cf/nvidia/nemotron-3-120b-a12bnemotron

Quick Info

Powered by
Provider
Cloudflare Workers AI
Model key
@cf/nvidia/nemotron-3-120b-a12b
Release date
Mar 11, 2026
Last updated
Mar 11, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.50
Output token cost
$1.50

Limits

Output tokens
256,000 tokens
Context window
256,000 tokens

Transparent token rates

Compare Nemotron 3 Super 120B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron 3 Super 120B

No articles yet. Fetch the latest news to show it here.

Videos about Nemotron 3 Super 120B

More models around Nemotron 3 Super 120B