Nemotron 3 Ultra is NVIDIA's open frontier-reasoning and orchestration model, designed from the ground up for long-running agentic workflows rather than short conversational exchanges. The Hugging Face model card describes a hybrid Transformer-Mamba mixture-of-experts architecture that NVIDIA calls LatentMoE, blending Mamba-2 state-space layers with MoE and attention components and adding multi-token prediction to accelerate inference and improve planning over extended sequences. With roughly 55B active parameters drawn from a 550B total pool, the design aims to keep per-token compute manageable while still unlocking frontier-quality reasoning, and the surrounding NVIDIA Nemotron 3 release also introduces complementary Nano and Super variants so teams can match model size to task complexity.
In practical terms, the model is aimed at teams building agent orchestration systems, coding agents, deep-research assistants, and other complex enterprise pipelines that benefit from sustained reasoning over very large inputs. NVIDIA publishes the full BF16 weights on Hugging Face under an OpenMDW-1.1 license and links an official technical report, pre-training and post-training v3 dataset collections, and a Nemotron developer page, so organizations that want to self-host or fine-tune have a clear path beyond the hosted free endpoint. The same release roadmap positions Nemotron 3 Ultra as a step toward more capable, agentic foundation models, with open weights and published recipes intended to let enterprises adapt it to their own long-context workflows rather than treating it as a fixed chat endpoint.