Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
DigitalOcean logo

Model details

Nemotron 3 Ultra

Nemotron 3 Ultra is a large-scale foundation model positioned for production agentic workloads rather than casual chat. It is built on a hybrid Mamba-Transformer mixture-of-experts architecture, totaling 550B parameters while activating only 55B per token, a sparsity pattern that is meant to keep long-running reasoning economical. The model is designed for autonomous agents, orchestration, complex coding, deep research, and enterprise workflows, targeting use cases where models are invoked repeatedly inside a software process rather than as a single prompt and response.

The design intent emphasizes breaking the usual trade-off between accuracy and inference speed, with reported gains of roughly five times higher throughput and up to thirty percent lower cost compared to other open models in a similar class, allowing more reasoning cycles within a fixed time budget. It accepts text input and produces text output, supports very long contexts, and runs across modern NVIDIA hardware in BF16, FP8, and NVFP4 precisions, giving deployers flexibility between maximum quality and maximum efficiency. For developer-focused teams, the model is framed as a strong fit for fast coding assistance and orchestration of multi-step agentic tasks where sustained throughput matters as much as raw benchmark scores.

DigitalOceannemotron-3-ultra-550bnemotron

Quick Info

Powered by
Provider
DigitalOcean
Model key
nemotron-3-ultra-550b
Release date
Jun 4, 2026
Last updated
Jun 12, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.90
Output token cost
$1.70

Limits

Output tokens
131,072 tokens
Context window
131,072 tokens

Transparent token rates

Compare Nemotron 3 Ultra pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron 3 Ultra

DigitalOcean

CoverageBenchmark

Divy Yadav's Towards AI article from June 13, 2026 focuses on Nemotron 3 Ultra's production-oriented design, emphasizing its 10:1 MoE sparsity ratio where only 10% of parameters activate per token. The article confirms the June 4, 2026 release date and the Computex 2026 announcement, and details the hybrid Mamba-2 stat Yadav positions Ultra as NVIDIA's reasoning engine of the Nemotron 3 family, built for sustained multi-step thinking and agentic workflows rather than fast single-turn chatbot responses. The article frames the combination of large total parameter count and low active parameter count as NVIDIA's approach to making a 550

DigitalOcean

Coverage

Navneet Guglani's June 6, 2026 Medium article details the Nemotron 3 Ultra release timeline, noting NVIDIA quietly pushed Ultra to Hugging Face on June 4, 2026, two days after Jensen Huang announced it from the Computex stage in Taipei. The article confirms Ultra's 550 billion total parameters with 55 billion active pe The piece frames Ultra within the broader three-model Nemotron 3 rollout (Nano December 2025, Super March 2026, Ultra June 2026) and emphasizes the architecture: a hybrid Mamba-Transformer MoE stack that NVIDIA released with full weights, training data, and post-training recipes. Guglani argues this distinguishes Nemot

DigitalOcean

CoverageAnalysis

Artificial Analysis reported on May 31, 2026 that NVIDIA announced Nemotron 3 Ultra at Jensen Huang's Computex keynote, describing it as the largest Nemotron 3 model at 550 billion parameters with 55 billion active per pass, and the most intelligent US open weights model based on their Intelligence Index. The article n The analysis also highlights that Ultra was benchmarked at over 300 tokens per second on a DeepInfra endpoint, significantly faster than peer models from DeepSeek and Moonshot Kimi, which the article says typically run at 50-100 tokens per second. The model is released in BF16 weights with an upcoming NVFP4 quantizatio

Videos about Nemotron 3 Ultra

More models around Nemotron 3 Ultra