Currently listed through these providers:
Model details
Nemotron 3 Ultra
Nemotron 3 Ultra is a mixture-of-experts model with 550 billion total parameters but 55 billion active for each token. This design seeks to preserve the capacity of a very large model while activating only part of the network for an individual request, a useful approach for teams balancing complex work with inference efficiency. Third-party coverage positions it as a fast coding model intended to fit developer workflows rather than remain limited to conversational use.
The model is best suited to software-development and agent-oriented applications that benefit from strong reasoning, code generation, and high-throughput inference. Its reported performance includes more than 300 generated tokens per second and an Intelligence Index score of 48, while commentary from CodeRabbit highlights its relevance to coding benchmarks and real developer use. These results come from third-party sources, so teams evaluating production fit should test quality, latency, and deployment requirements in their own workloads.
Quick Info
Powered by- Provider
- Weights & Biases
- Model key
- nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B
- Release date
- Jun 4, 2026
- Last updated
- Jun 4, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.75
- Output token cost
- $2.75
Limits
- Output tokens
- 262,144 tokens
- Context window
- 262,144 tokens
Latest news about Nemotron 3 Ultra
Videos about Nemotron 3 Ultra
Recent tweets and retweets from Weights & Biases
More models around Nemotron 3 Ultra
This exact model name is also listed by 5 other providers.