NVIDIA Nemotron 3 Super is a hybrid Mamba-Transformer model engineered to handle the demands of complex, agentic AI systems. By utilizing a Mixture-of-Experts architecture with 120 billion total parameters and 12 billion active parameters per forward pass, the model balances high-level reasoning with significant compute efficiency. It incorporates LatentMoE for improved accuracy and MTP layers that enable native speculative decoding, allowing it to maintain high performance even as multi-agent workflows generate large volumes of tokens.
The model was developed through a rigorous training pipeline that included pre-training on 25 trillion tokens, followed by supervised fine-tuning and reinforcement learning. This process, combined with its ability to support a 1 million token context length, makes it well-suited for tasks ranging from deep literature search and molecular understanding to automated software development. By delivering high throughput and accuracy, the model provides a scalable foundation for enterprise applications that require reliable performance in data-intensive environments.