Nemotron 3 Ultra is an open-weight frontier reasoning model from NVIDIA aimed at the most demanding agentic workloads, including complex multi-step coding agents, long-context analysis, and high-accuracy reasoning across code, math, and science. It first generates a reasoning trace and then concludes with a final response, and the reasoning behavior can be tuned through a flag in the chat template, giving developers direct control over how much deliberation the model performs per turn. The model is well suited for orchestrating long-running autonomous agents, deep research loops that must synthesize across large source sets, and enterprise workflows such as electronic design automation where reasoning must remain stable across many steps.
Under the hood, Nemotron 3 Ultra combines a hybrid Mamba-Transformer design with Latent Mixture-of-Experts layers that activate only a fraction of the total parameters per token, augmented by Multi-Token Prediction layers for faster and higher-quality long-sequence generation. This Latent MoE arrangement is described as effectively calling four experts at the inference cost of one, and the model ships fully open under the NVIDIA Open Model License with weights, training data, and recipes available for customization. A very large context window makes the model particularly attractive for tasks that require sustained reasoning across large codebases, lengthy documents, or extended agent sessions without losing track of earlier state.