Synthetic
NVIDIA launches SUPER, just not the SUPER gamers wanted NVIDIA has released Nemotron 3 Super, a 120 billion parameter Mixture-of-Experts model built for
Model details
Nemotron 3 Super 120B A12B is NVIDIA's open-weight large language model in the Nemotron family, designed as a hybrid Mamba-Transformer Mixture-of-Experts architecture with 120 billion total parameters and roughly 12 billion activated per inference call, enabling four expert calls per token. This sparse activation pattern keeps compute closer to a mid-sized model while preserving the capacity of a much larger one, making the system well suited to long-context reasoning and tool-driven workloads. NVIDIA released the model on March 11, 2026, and distributed it openly, with the open weights allowing teams to self-host, fine-tune, or deploy through managed clouds in addition to using hosted APIs.
In practical terms, the model's open-weights design and hybrid MoE backbone make it a flexible foundation for assistants, code generation, and pipeline-style reasoning tasks where both cost efficiency and long input handling matter. Independent benchmarking coverage from DeepInfra on latency and cost, and from Artificial Analysis on intelligence, performance, and price, points to active community evaluation of the model's trade-offs, while its availability on Amazon SageMaker JumpStart alongside Qwen3.5-9B and Qwen3.5-27B suggests broad enterprise interest. Developers who need a self-hostable, reasoning-capable model with sparse expert routing will find it fits naturally into retrieval-augmented, agentic, and document-heavy workflows.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Synthetic
NVIDIA launches SUPER, just not the SUPER gamers wanted NVIDIA has released Nemotron 3 Super, a 120 billion parameter Mixture-of-Experts model built for
Synthetic
OpenRouter's listing for the NVIDIA-hosted free tier of Nemotron 3 Super (120B A12B) confirms the model is a 120B-parameter hybrid Mamba-Transformer MoE that activates 12B parameters per pass, using Latent MoE with multi-token prediction to deliver over 50% higher token generation versus leading open models. The free e The listing also notes that the model is fully open under the NVIDIA Open License with weights, datasets, and recipes, and that NVIDIA runs multi-environment RL training across 10+ environments for benchmarks such as AIME 2025, TerminalBench, and SWE-Bench Verified. OpenRouter reports P50 latency of 1.35s, throughput o
Synthetic
Discover more about what's new at AWS with NVIDIA Nemotron-3-Super-120B, Qwen3.5-9B, and Qwen3.5-27B models now available on Amazon SageMaker JumpStart
Synthetic
Analysis of NVIDIA's NVIDIA Nemotron 3 Super 120B A12B (Reasoning) and comparison to other AI models across key metrics including quality, price, performance (tokens per second & time to first token), context window & more.
Synthetic
VLLM 13.1 is crashing with: bmm_fp8_internal_cublaslt failed: the library was not initialized
This exact model name is also listed by 8 other providers.