Nemotron 3 Ultra is the flagship model in the Nemotron 3 family, combining a Mixture-of-Experts Hybrid Mamba-Attention design with NVIDIA's LatentMoE technique to call four experts at the inference cost of one. The model is built around 550 billion total parameters with 55 billion active per pass, and it adds Multi-Token Prediction layers that enable native speculative decoding for faster generation on long sequences. Training is staged, beginning with an NVFP4 pre-training pass and finishing with a supervised fine-tuning, reinforcement learning, and multi-teacher on-policy distillation pipeline aimed at raising answer quality and reasoning discipline.
The model's design centers on sustained, long-running reasoning: it offers inference-time reasoning budget controls so users can trade compute against accuracy, and it supports a one-million-token context window, where it reports outperforming other leading open large language models on the RULER long-context benchmark. Throughput comparisons against other major open MoE models show substantial inference gains on long input and output workloads, making the model a practical fit for autonomous coding agents that plan, refactor, and recover across large codebases, deep research loops that synthesize across many sources, and enterprise or EDA workflows that demand extended, reliable reasoning rather than short conversational turns.