Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Requesty logo

Model details

nvidia-nemotron-3-super-120b-a12b

NVIDIA Nemotron 3 Super 120B A12B is a large language model built around a hybrid Latent Mixture-of-Experts (LatentMoE) architecture, with roughly 120B total parameters and around 12B active at inference, which lets it deliver strong reasoning capacity while keeping per-token compute comparatively modest. NVIDIA trained it specifically for agentic workflows, sustained long-context reasoning, and tool use, and the model ships with a configurable reasoning mode so developers can tune how much deliberative processing is applied depending on the task. The wider Nemotron family is positioned as NVIDIA's open research line for capable, efficient reasoning models, and this Super variant sits toward the higher end in scale and capability while staying within a single, easy-to-deploy footprint.

In practice, Nemotron 3 Super 120B A12B is well suited to multi-step agent pipelines that combine tool calls, structured decisions, and extended context, such as research assistants, code agents, and complex planning systems that need to keep large documents or conversation histories in view. Multilingual coverage of English, French, German, Italian, Japanese, Spanish, and Chinese makes it a reasonable choice for global products where a single model must serve several locales without losing reasoning strength. The same architecture is also being explored in quantized forms like an NVFP4 build discussed in NVIDIA developer channels and in BF16 hosting from third-party platforms, signaling that the model is intended to remain flexible across different serving environments and hardware targets.

Requestynvidia-nemotron-3-super-120b-a12bnemotron

Quick Info

Powered by
Provider
Requesty
Model key
nvidia-nemotron-3-super-120b-a12b
Release date
Mar 11, 2026
Last updated
Mar 11, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.50

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Transparent token rates

Compare nvidia-nemotron-3-super-120b-a12b pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about nvidia-nemotron-3-super-120b-a12b

No articles yet. Fetch the latest news to show it here.

Videos about nvidia-nemotron-3-super-120b-a12b

More models around nvidia-nemotron-3-super-120b-a12b