Sulat.com
AI models
Weights & Biases logo

Model details

Nemotron 3 Ultra

Nemotron 3 Ultra is a mixture-of-experts model with 550 billion total parameters but 55 billion active for each token. This design seeks to preserve the capacity of a very large model while activating only part of the network for an individual request, a useful approach for teams balancing complex work with inference efficiency. Third-party coverage positions it as a fast coding model intended to fit developer workflows rather than remain limited to conversational use.

The model is best suited to software-development and agent-oriented applications that benefit from strong reasoning, code generation, and high-throughput inference. Its reported performance includes more than 300 generated tokens per second and an Intelligence Index score of 48, while commentary from CodeRabbit highlights its relevance to coding benchmarks and real developer use. These results come from third-party sources, so teams evaluating production fit should test quality, latency, and deployment requirements in their own workloads.

Weights & Biasesnvidia/NVIDIA-Nemotron-3-Ultra-550B-A55Bnemotron

Quick Info

Powered by
Provider
Weights & Biases
Model key
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B
Release date
Jun 4, 2026
Last updated
Jun 4, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.75
Output token cost
$2.75

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Latest news about Nemotron 3 Ultra

Videos about Nemotron 3 Ultra

Recent tweets and retweets from Weights & Biases

More models around Nemotron 3 Ultra