Currently listed through these providers:
Model details
nvidia-nemotron-3-super-120b-a12b
NVIDIA Nemotron 3 Super 120B A12B is a large language model built around a hybrid Latent Mixture-of-Experts (LatentMoE) architecture, with roughly 120B total parameters and around 12B active at inference, which lets it deliver strong reasoning capacity while keeping per-token compute comparatively modest. NVIDIA trained it specifically for agentic workflows, sustained long-context reasoning, and tool use, and the model ships with a configurable reasoning mode so developers can tune how much deliberative processing is applied depending on the task. The wider Nemotron family is positioned as NVIDIA's open research line for capable, efficient reasoning models, and this Super variant sits toward the higher end in scale and capability while staying within a single, easy-to-deploy footprint.
In practice, Nemotron 3 Super 120B A12B is well suited to multi-step agent pipelines that combine tool calls, structured decisions, and extended context, such as research assistants, code agents, and complex planning systems that need to keep large documents or conversation histories in view. Multilingual coverage of English, French, German, Italian, Japanese, Spanish, and Chinese makes it a reasonable choice for global products where a single model must serve several locales without losing reasoning strength. The same architecture is also being explored in quantized forms like an NVFP4 build discussed in NVIDIA developer channels and in BF16 hosting from third-party platforms, signaling that the model is intended to remain flexible across different serving environments and hardware targets.
Quick Info
Powered by- Provider
- Requesty
- Model key
- nvidia-nemotron-3-super-120b-a12b
- Release date
- Mar 11, 2026
- Last updated
- Mar 11, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.10
- Output token cost
- $0.50
Limits
- Output tokens
- 262,144 tokens
- Context window
- 262,144 tokens
Transparent token rates
Compare nvidia-nemotron-3-super-120b-a12b pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about nvidia-nemotron-3-super-120b-a12b
No articles yet. Fetch the latest news to show it here.
Videos about nvidia-nemotron-3-super-120b-a12b
More models around nvidia-nemotron-3-super-120b-a12b
This exact model name is also listed by 2 other providers.
