Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Crusoe logo

Model details

Nemotron 3 Super 120B A12B

Nemotron 3 Super 120B A12B is NVIDIA's open-weight large language model engineered around inference throughput rather than raw parameter count. The architecture pairs 120.6 billion total parameters with roughly 12 billion active per forward pass, splitting an 88-layer stack into 40 Mamba-2 layers, 40 Latent Mixture-of-Experts layers, and only 8 grouped-query attention layers, each with 2 KV heads at a head dimension of 128. This hybrid layout is reinforced by a multi-token prediction head that serves as an internal draft mechanism for speculative decoding, and the routed expert path is compressed through Latent MoE. Together these choices shrink the attention-side key-value cache while preserving the capacity needed for long-context reasoning and tool-calling tasks typical of agentic pipelines.

NVIDIA positions the model for agentic workflows, high-volume workloads such as IT ticket automation, retrieval-augmented generation, and long-context reasoning, with a configurable reasoning mode that can be toggled on or off through the chat template. The release supports English plus several European and Asian languages and runs on infrastructure starting at 8x H100-80GB GPUs. Practically, the design trades some attention density for substantial throughput gains, making the model a good fit when serving cost, latency, or batch volume matter more than maximizing per-token reasoning depth, while still leaving room for complex multi-step tool use.

Crusoenvidia/NVIDIA-Nemotron-3-Super-120B-A12Bnemotron

Quick Info

Powered by
Provider
Crusoe
Model key
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B
Release date
Mar 11, 2026
Last updated
Mar 11, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$2.40

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Transparent token rates

Compare Nemotron 3 Super 120B A12B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron 3 Super 120B A12B

No articles yet. Fetch the latest news to show it here.

Videos about Nemotron 3 Super 120B A12B

More models around Nemotron 3 Super 120B A12B