Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

Nemotron 3 Ultra

Nemotron 3 Ultra is an open-weight frontier reasoning model from NVIDIA aimed at the most demanding agentic workloads, including complex multi-step coding agents, long-context analysis, and high-accuracy reasoning across code, math, and science. It first generates a reasoning trace and then concludes with a final response, and the reasoning behavior can be tuned through a flag in the chat template, giving developers direct control over how much deliberation the model performs per turn. The model is well suited for orchestrating long-running autonomous agents, deep research loops that must synthesize across large source sets, and enterprise workflows such as electronic design automation where reasoning must remain stable across many steps.

Under the hood, Nemotron 3 Ultra combines a hybrid Mamba-Transformer design with Latent Mixture-of-Experts layers that activate only a fraction of the total parameters per token, augmented by Multi-Token Prediction layers for faster and higher-quality long-sequence generation. This Latent MoE arrangement is described as effectively calling four experts at the inference cost of one, and the model ships fully open under the NVIDIA Open Model License with weights, training data, and recipes available for customization. A very large context window makes the model particularly attractive for tasks that require sustained reasoning across large codebases, lengthy documents, or extended agent sessions without losing track of earlier state.

Vercel AI Gatewaynvidia/nemotron-3-ultra-550b-a55bnemotron

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
nvidia/nemotron-3-ultra-550b-a55b
Release date
Jun 4, 2026
Last updated
Jun 4, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.60
Output token cost
$2.40

Limits

Output tokens
65,000 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare Nemotron 3 Ultra pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron 3 Ultra

Vercel AI Gateway

CoverageBenchmark

A Towards AI commentary explicitly names Nemotron 3 Ultra as a 550-billion-parameter model that activates 55 billion parameters per token (a 10:1 MoE sparsity ratio), with a hybrid Mamba-2 state-space model alternating against transformer attention layers. The article dates the release to June 4, 2026 at Jensen Huang's The piece frames Ultra's design economics as a response to the impracticality of running a 550B dense model in production, arguing that the 90% parameter inactivity enabled by MoE, combined with the Mamba-2 hybrid, makes agent workloads calling the model dozens of times per task viable. Creator attribution to NVIDIA is

Vercel AI Gateway

Coverage

A long-form community write-up on Medium traces the full Nemotron 3 family timeline: Nano released December 15, 2025; Super released March 11, 2026 at GTC; and Ultra released June 4, 2026 on Hugging Face, two days after Jensen Huang's Computex stage announcement. The article documents that all three models share a hybr The piece explicitly names the 550 billion total / 55 billion active parameter configuration for Nemotron 3 Ultra and cites the Architecture's Mamba-2 state-space layers interleaved with transformer attention as the design choice that distinguishes it from other 550B-class options. While authored by an independent comm

Vercel AI Gateway

CoverageAnalysis

NVIDIA announced the Nemotron 3 Ultra open-weights model in Jensen Huang's Computex keynote, according to Artificial Analysis. At approximately 550 billion total parameters with 55B active and ~90% sparsity, it is the largest Nemotron 3 variant released to date and the most intelligent US open-weights model Artificial The Artificial Analysis piece also reports leading throughput for Ultra's intelligence class: on a pre-release DeepInfra endpoint the model served over 300 tokens per second, compared with 50–100 tokens per second typical for peer-sized Chinese models from DeepSeek and Moonshot. gpt-oss-120b is served at similar speeds

Videos about Nemotron 3 Ultra

More models around Nemotron 3 Ultra