Currently listed through these providers:
Model details
Nemotron 3 Ultra 550B A55B
NVIDIA Nemotron 3 Ultra 550B A55B is an open frontier-reasoning and orchestration model designed for complex agentic workloads. It combines a total of 550B parameters with 55B active per inference using a Mixture-of-Experts design, following a hybrid Transformer-Mamba architecture. NVIDIA has paired this release with a public technical report and openly published pre-training and post-training dataset collections, positioning the model as a foundation for transparent development in advanced reasoning systems.
The model is tuned for long-running agentic workflows such as coding agents, deep research, agent orchestration, and multi-step enterprise planning. Its hybrid MoE-Transformer-Mamba design supports text input and output with a context window reaching up to 1M tokens, making it well suited for tasks that require sustained reasoning across very long documents or chained tool interactions. The BF16 variant is distributed under the OpenMDW-1.1 license through NVIDIA's official channels, giving researchers and developers direct access to weights and supporting documentation.
Quick Info
Powered by- Provider
- Fireworks AI
- Model key
- accounts/fireworks/models/nemotron-3-ultra-nvfp4
- Release date
- Jun 4, 2026
- Last updated
- Jun 4, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.60
- Output token cost
- $2.40
Limits
- Output tokens
- 128,000 tokens
- Context window
- 262,144 tokens
Transparent token rates
Compare Nemotron 3 Ultra 550B A55B pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Nemotron 3 Ultra 550B A55B
No articles yet. Fetch the latest news to show it here.
Videos about Nemotron 3 Ultra 550B A55B
More models around Nemotron 3 Ultra 550B A55B
This exact model name is also listed by 10 other providers.
