Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
routing.run logo

Model details

Nemotron 3 Ultra 550B A55B

Nemotron 3 Ultra is NVIDIA's open frontier-reasoning and orchestration model, designed from the ground up for long-running agentic workflows rather than short conversational exchanges. The Hugging Face model card describes a hybrid Transformer-Mamba mixture-of-experts architecture that NVIDIA calls LatentMoE, blending Mamba-2 state-space layers with MoE and attention components and adding multi-token prediction to accelerate inference and improve planning over extended sequences. With roughly 55B active parameters drawn from a 550B total pool, the design aims to keep per-token compute manageable while still unlocking frontier-quality reasoning, and the surrounding NVIDIA Nemotron 3 release also introduces complementary Nano and Super variants so teams can match model size to task complexity.

In practical terms, the model is aimed at teams building agent orchestration systems, coding agents, deep-research assistants, and other complex enterprise pipelines that benefit from sustained reasoning over very large inputs. NVIDIA publishes the full BF16 weights on Hugging Face under an OpenMDW-1.1 license and links an official technical report, pre-training and post-training v3 dataset collections, and a Nemotron developer page, so organizations that want to self-host or fine-tune have a clear path beyond the hosted free endpoint. The same release roadmap positions Nemotron 3 Ultra as a step toward more capable, agentic foundation models, with open weights and published recipes intended to let enterprises adapt it to their own long-context workflows rather than treating it as a fixed chat endpoint.

routing.runnemotron-3-ultranemotron

Quick Info

Powered by
Provider
routing.run
Model key
nemotron-3-ultra
Release date
Jun 4, 2026
Last updated
Jun 4, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.10

Limits

Output tokens
32,000 tokens
Context window
131,072 tokens

Transparent token rates

Compare Nemotron 3 Ultra 550B A55B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron 3 Ultra 550B A55B

routing.run

Coverage

A Tech Insider Canada analysis dated July 6, 2026 covers the June 4, 2026 release of NVIDIA Nemotron 3 Ultra, confirming via NVIDIA Research that the model has 550 billion total parameters with roughly 90% sparsity and about 55 billion parameters active per forward pass. The article reports a one-million-token context The piece frames the launch strategically — NVIDIA moving from chip vendor into the model layer with an open-weight, self-hostable giant — and situates it against the same week's Canadian sovereign-AI funding news. Technical specifics (hybrid Mamba-Attention MoE, 1M context, 550B/55B active) are drawn directly from NVI

routing.run

CoverageBenchmark

An NVIDIA Technical Blog post dated July 8, 2026 by Sean Lopp, Matthew Penn, and Sukrit Rao documents how to use harness engineering to improve Nemotron 3 Ultra agent accuracy without fine-tuning. It reports that tuning a LangChain Deep Agents harness profile for Nemotron 3 Ultra raised evaluation scores from 94 to 96 The post describes an automated propose-verify-keep loop implemented in the NemoClaw community repository that runs the benchmark, proposes profile changes, verifies fixes across multiple runs, and re-runs the full suite to prevent regressions. It notes the same loop can adapt Nemotron 3 Ultra to other agent frameworks

routing.run

CoverageBenchmark

NVIDIA released Nemotron 3 Ultra on June 4, 2026, announced at Jensen Huang's Computex 2026 keynote, according to a Towards AI analysis by Divy Yadav. The Nemotron 3 family comprises three models — Nano, Super, and Ultra — with Ultra positioned as the reasoning engine for sustained multi-step thinking. The article fram Architecturally, Nemotron 3 Ultra has 550 billion total parameters with 55 billion activated per token, a 10:1 Mixture-of-Experts sparsity ratio. NVIDIA embedded a Mamba-2 state space model inside the architecture, alternating it with transformer attention layers, creating a hybrid Transformer-Mamba MoE design. This co

routing.run

Coverage

A practitioner deep-dive published on Medium on June 12, 2026 documents NVIDIA's June 4, 2026 launch of Nemotron 3 Ultra at Computex. The available excerpt specifies a hybrid Mamba-Transformer mixture-of-experts architecture with 550 billion total parameters and 55 billion active per token, trained natively in NVFP4, w The article's stated scope is a practical guide for production use of the model, including calling it via cloud API, local deployment with vLLM and NVIDIA NIM, and building an agent with tool calling and explicit reasoning. Architectural details noted in the excerpt — quadratic attention cost over long contexts, sparsi

routing.run

Coverage

A Startup Fortune article dated June 5, 2026 reports NVIDIA's release of Nemotron 3 Ultra, describing it as a 550 billion parameter open-weight mixture-of-experts model with 55 billion active parameters and a one-million-token context window. Weights are available through Hugging Face under the "Nemotron 3 Ultra 550B A Citing NVIDIA's June 4 research note, the piece details the hybrid Mamba-Attention mixture-of-experts architecture, LatentMoE routing, multi-token prediction layers, and inference-time reasoning budget control as the mechanisms NVIDIA uses to make a very large model behave more efficiently on reasoning, planning, and l

routing.run

CoverageBenchmark

Nemotron 3 Ultra (550B A55B) is ranked 78th on llm-stats.com's composite LLM Stats Score with a score of 36.3, placing it in the "Average / Top half" band on the aggregated leaderboard. The model shows strong capability in legal tasks (rank 9 of 210) and healthcare (14 of 244), while scoring lower in reasoning (81 of 3 On individual benchmarks, Nemotron 3 Ultra (550B A55B) takes the top rank on RULER (0.95/1) for long-context evaluation, IMO-AnswerBench (0.92/1) for mathematical reasoning on IMO problems, and PinchBench (0.90/1) for agentic coding tasks, with scores attributed to huggingface.co sources. It also appears in LiveCodeBen

Videos about Nemotron 3 Ultra 550B A55B

More models around Nemotron 3 Ultra 550B A55B