Sulat.com
AI models
Synthetic logo

Model details

Nemotron 3 Super 120B A12B

Nemotron 3 Super 120B A12B is NVIDIA's open-weight large language model in the Nemotron family, designed as a hybrid Mamba-Transformer Mixture-of-Experts architecture with 120 billion total parameters and roughly 12 billion activated per inference call, enabling four expert calls per token. This sparse activation pattern keeps compute closer to a mid-sized model while preserving the capacity of a much larger one, making the system well suited to long-context reasoning and tool-driven workloads. NVIDIA released the model on March 11, 2026, and distributed it openly, with the open weights allowing teams to self-host, fine-tune, or deploy through managed clouds in addition to using hosted APIs.

In practical terms, the model's open-weights design and hybrid MoE backbone make it a flexible foundation for assistants, code generation, and pipeline-style reasoning tasks where both cost efficiency and long input handling matter. Independent benchmarking coverage from DeepInfra on latency and cost, and from Artificial Analysis on intelligence, performance, and price, points to active community evaluation of the model's trade-offs, while its availability on Amazon SageMaker JumpStart alongside Qwen3.5-9B and Qwen3.5-27B suggests broad enterprise interest. Developers who need a self-hostable, reasoning-capable model with sparse expert routing will find it fits naturally into retrieval-augmented, agentic, and document-heavy workflows.

Synthetichf:nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4nemotron

Quick Info

Powered by
Provider
Synthetic
Model key
hf:nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
Release date
Mar 11, 2026
Last updated
Mar 11, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$1.00

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Transparent token rates

Compare nemotron pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron 3 Super 120B A12B

Synthetic

CoverageRelease Notes

NVIDIA launches SUPER, just not the SUPER gamers wanted NVIDIA has released Nemotron 3 Super, a 120 billion parameter Mixture-of-Experts model built for

Synthetic

CoverageBenchmark

OpenRouter's listing for the NVIDIA-hosted free tier of Nemotron 3 Super (120B A12B) confirms the model is a 120B-parameter hybrid Mamba-Transformer MoE that activates 12B parameters per pass, using Latent MoE with multi-token prediction to deliver over 50% higher token generation versus leading open models. The free e The listing also notes that the model is fully open under the NVIDIA Open License with weights, datasets, and recipes, and that NVIDIA runs multi-environment RL training across 10+ environments for benchmarks such as AIME 2025, TerminalBench, and SWE-Bench Verified. OpenRouter reports P50 latency of 1.35s, throughput o

Synthetic

CoverageRelease Notes

Discover more about what's new at AWS with NVIDIA Nemotron-3-Super-120B, Qwen3.5-9B, and Qwen3.5-27B models now available on Amazon SageMaker JumpStart

Synthetic

CoverageBenchmark

Analysis of NVIDIA's NVIDIA Nemotron 3 Super 120B A12B (Reasoning) and comparison to other AI models across key metrics including quality, price, performance (tokens per second & time to first token), context window & more.

Synthetic

CoverageDiscourse

VLLM 13.1 is crashing with: bmm_fp8_internal_cublaslt failed: the library was not initialized

Videos about Nemotron 3 Super 120B A12B

More models around Nemotron 3 Super 120B A12B