Nemotron 3 Ultra is NVIDIA's flagship entry in the Nemotron 3 family, a Mixture-of-Experts model that pairs a Hybrid Mamba-Attention backbone with NVIDIA's LatentMoE design. The model carries 550 billion total parameters with 55 billion active per forward pass, letting it draw on broad knowledge while keeping inference costs closer to a much smaller model. Multi-Token Prediction layers are baked into the architecture to accelerate generation through native speculative decoding, and pretraining in NVFP4 (NVIDIA's 4-bit floating-point format) helps make that large active footprint practical to serve.
Beyond pretraining, Nemotron 3 Ultra is post-trained with a multi-stage pipeline that combines supervised fine-tuning, reinforcement learning, and Multi-teacher On-Policy Distillation, which NVIDIA describes as designed to improve overall accuracy. The system supports inference-time reasoning budget control and is positioned as the strongest model in the Nemotron 3 lineup, with NVIDIA reporting on-par accuracy with other leading open models across a diverse benchmark set and notably higher throughput on long-output tasks. It is a natural fit for developer workflows that need open-weight flexibility, sustained long-context reasoning, and a model that can act as a drop-in engine for code generation, agentic tools, and other production pipelines.