NVIDIA Nemotron 3 Super is built on a hybrid Mamba-Transformer Mixture-of-Experts architecture, combining the long-context efficiency of state-space models with the reasoning power of transformer attention. With 120 billion total parameters but only 12 billion activated per token, it achieves high quality at a fraction of the compute cost typical for models of this scale. The addition of multi-token prediction head architecture pushes token generation throughput over 50% higher than leading open alternatives, while latent MoE routing allows four expert branches to fire for the price of one. This makes the model particularly well-suited for long-horizon agent tasks where context must be maintained across thousands of turns without latency spikes.
Training involved multi-environment reinforcement learning across more than ten distinct environments, targeting benchmark accuracy on AIME 2025, TerminalBench, and SWE-Bench Verified. The model ships with full open weights, datasets, and training recipes under the NVIDIA Open License, enabling developers to reproduce results, fine-tune on domain-specific corpora, or compress the model for edge deployment. The combination of open-source transparency, million-token context support, and RL-tuned agentic behaviors positions Nemotron 3 Super as a practical foundation for anyone building multi-agent pipelines, code generation workflows, or retrieval-augmented systems that require reliable state maintenance across long execution traces.