Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Requesty logo

Model details

Nemotron 3 Ultra 550B A55B

Nemotron 3 Ultra is a large language model designed for demanding agentic, reasoning, and conversational workloads. It is positioned for complex multi-step tasks, long-context analysis, and careful work across code, mathematics, and science, making it a practical fit for systems that need to plan, inspect several stages of a problem, and produce a final answer rather than handle only isolated prompts.

The model uses a hybrid LatentMoE design with interleaved Mamba-2 and mixture-of-experts layers plus selected attention layers. It has 550B total parameters and 55B active parameters, and includes multi-token prediction layers intended to improve generation speed and output quality. Its reasoning can be configured through a chat-template flag, supporting workflows that need to control the reasoning process explicitly.

Requestynemotron-3-ultra-550b-a55bnemotron

Quick Info

Powered by
Provider
Requesty
Model key
nemotron-3-ultra-550b-a55b
Release date
Jun 4, 2026
Last updated
Jun 4, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
65,536 tokens
Context window
1,048,576 tokens

Latest news about Nemotron 3 Ultra 550B A55B

Requesty

CoverageBenchmark

A Towards AI article by Divy Yadav, published June 11, 2026, characterizes the Nemotron 3 Ultra 550B A55B as a 550 billion parameter Mixture-of-Experts model that activates only 55 billion parameters per token, giving a 10:1 sparsity ratio. The piece emphasizes NVIDIA's hybrid design that interleaves Mamba-2 state spac The article positions Nemotron 3 Ultra as an open-source model that AI developers should be aware of, presenting it as a "differently-designed" alternative to smaller models rather than a scaled-up transformer. Technical details around the 10:1 active-to-total ratio and the Mamba-2 hybrid are presented as the different

Requesty

Coverage

A third-party technical guide by Roan Brasil Monteiro, published June 12, 2026, documents NVIDIA's Nemotron 3 Ultra model that was released on June 4, 2026 at Jensen Huang's Computex 2026 keynote. The article describes the exact 550B A55B variant as a hybrid Mamba-Transformer Mixture-of-Experts architecture with 550 bi The guide further covers practical developer usage of the Nemotron 3 Ultra 550B A55B model, including cloud API invocation, local deployment via vLLM and NVIDIA NIM, and building production agents with tool calling and explicit reasoning. Stated throughput of 300+ tokens/second is highlighted alongside the architecture

Videos about Nemotron 3 Ultra 550B A55B

More models around Nemotron 3 Ultra 550B A55B