NVIDIA Nemotron 3 Super is a hybrid Mixture-of-Experts model that combines Mamba and Transformer architectures with multi-token prediction, activating just 12 billion parameters from its 120 billion total to keep inference efficient without sacrificing accuracy. Its latent MoE design calls 4 experts for the cost of one, delivering over 50% higher token generation than leading open models. Built for agent-based AI inference, it handles up to a 1 million token context window to support long-horizon multi-step task planning, cross-document reasoning, and multi-agent collaboration on a single GPU.
The model was trained with multi-environment reinforcement learning across 10 or more environments, achieving leading accuracy on benchmarks including AIME 2025, TerminalBench, and SWE-Bench Verified. Weights, datasets, and training recipes are released under the NVIDIA Open License, making it easy to customize and deploy anywhere from workstation to cloud. It excels at reasoning, tool use, and instruction following in complex multi-agent applications, with native support for Japanese and optimized for running many collaborating agents per application.