Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Baseten logo

Model details

Nemotron Ultra

Nemotron Ultra is a 550-billion-parameter open-weight large language model positioned by NVIDIA as a foundation for AI agents, multi-step reasoning, and reliable tool calling across long contexts. Coverage describes the release as a notable open-weight entry at a scale that was previously the domain of closed commercial APIs, with the specific variant identifier carrying the "3" generation and "A55B" configuration suffix, indicating a third-generation Nemotron design built around a 550B mixture-of-experts-style parameter budget. For teams building AI agents or evaluating open-weight foundations for enterprise workloads, the headline framing emphasizes that the model is engineered for agentic tasks rather than only single-turn chat: running chains of reasoning, maintaining coherence over long contexts, and invoking external tools accurately.

In practical terms, the model fits deployment scenarios where an organization needs frontier-class reasoning in a self-hostable or third-party-hosted open-weight package, particularly for agent orchestration, retrieval-augmented pipelines, and tool-mediated workflows that demand long-context retention. The combination of a very large parameter count and an explicit agent-optimized training objective suggests strong qualitative strengths in instruction following and multi-step planning, while the open-weight nature lets teams fine-tune, inspect, or self-host the model under their own governance. It is best suited to engineering teams comfortable running or integrating large foundation models who want agentic capability without depending solely on closed frontier APIs, and who can pair the model with their own evaluation harnesses to confirm fit for their specific agent workloads.

Basetennvidia/NVIDIA-Nemotron-3-Ultra-550B-A55Bnemotron

Quick Info

Powered by
Provider
Baseten
Model key
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B
Release date
Jun 4, 2026
Last updated
Jun 4, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.60
Output token cost
$2.40

Limits

Output tokens
202,800 tokens
Context window
202,800 tokens

Transparent token rates

Compare Nemotron Ultra pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron Ultra

Baseten

Official sourceAnnouncement

Baseten's engineering blog introduces NVIDIA Nemotron 3 Ultra, a mixture-of-experts language model with 550 billion total parameters and 55 billion active on any token, released June 4, 2026. The post, authored by Fred Liu, explains that NVIDIA built the model on a hybrid Transformer-Mamba architecture so that most att The blog adds that Nemotron 3 Ultra is fully open: open weights, open training data, and open recipes under the NVIDIA open model license, and that NVIDIA post-trained the model with reinforcement learning across agentic environments so it can reason, call tools, and operate inside an agent loop. The post is written by

Baseten

Official sourceRelease Notes

Baseten announced on June 4, 2026 that NVIDIA's Nemotron Ultra is available on its Model APIs platform. The model is described as a 550B-parameter mixture-of-experts architecture with 55B active parameters, accessible via a one-click deployment workflow with dedicated deployments available for larger workloads. The cha According to the Baseten changelog, the deployment exposes a 202K-token context window and supports tool calling, structured outputs, and opt-in reasoning features. Developers can invoke the model using the OpenAI or Anthropic SDKs, lowering integration friction for teams building agentic or reasoning-intensive applica

Baseten

CoverageBenchmark

A third-party AI release tracker records NVIDIA Nemotron 3 Ultra as an open-source release dated June 4, 2026, about 85 days after Nemotron 3 Super, listing it as a 550B-parameter model with weights and training code/data released under an open license. The page consolidates benchmark scores including SWE-Bench Verifie The tracker also reproduces provider pricing rows for DeepInfra ($0.50/$2.20 per 1M tokens, 256K context, fp4), BaseTen ($0.60/$2.40, 200K, fp4), and Venice ($0.625/$3.125, 256K, fp8), sourced from openrouter.ai as of September 18, 2026. Because these rows are gateway/serving-provider data rather than intrinsic model s

Baseten

CoverageBenchmark

Artificial Analysis published an independent evaluation of Nemotron 3 Ultra 550B A55B (Reasoning), placing it at 23 on the Artificial Analysis Intelligence Index against a class median of 18. The model is characterized as above average in intelligence, somewhat expensive among comparable open-weight models, notably fas The benchmark page confirms Nemotron 3 Ultra's mixture-of-experts architecture with 550B total and 55B active parameters, and flags it as the reasoning variant with a possible non-reasoning sibling. Pricing is reported at $0.60 per 1M input tokens and $2.50 per 1M output tokens, with a 73% cache discount available. NVI

Videos about Nemotron Ultra

More models around Nemotron Ultra