Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Fireworks AI logo

Model details

Nemotron 3 Ultra 550B A55B

NVIDIA Nemotron 3 Ultra 550B A55B is an open frontier-reasoning and orchestration model designed for complex agentic workloads. It combines a total of 550B parameters with 55B active per inference using a Mixture-of-Experts design, following a hybrid Transformer-Mamba architecture. NVIDIA has paired this release with a public technical report and openly published pre-training and post-training dataset collections, positioning the model as a foundation for transparent development in advanced reasoning systems.

The model is tuned for long-running agentic workflows such as coding agents, deep research, agent orchestration, and multi-step enterprise planning. Its hybrid MoE-Transformer-Mamba design supports text input and output with a context window reaching up to 1M tokens, making it well suited for tasks that require sustained reasoning across very long documents or chained tool interactions. The BF16 variant is distributed under the OpenMDW-1.1 license through NVIDIA's official channels, giving researchers and developers direct access to weights and supporting documentation.

Fireworks AIaccounts/fireworks/models/nemotron-3-ultra-nvfp4nemotron

Quick Info

Powered by
Provider
Fireworks AI
Model key
accounts/fireworks/models/nemotron-3-ultra-nvfp4
Release date
Jun 4, 2026
Last updated
Jun 4, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.60
Output token cost
$2.40

Limits

Output tokens
128,000 tokens
Context window
262,144 tokens

Transparent token rates

Compare Nemotron 3 Ultra 550B A55B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron 3 Ultra 550B A55B

No articles yet. Fetch the latest news to show it here.

Videos about Nemotron 3 Ultra 550B A55B

More models around Nemotron 3 Ultra 550B A55B