Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Eden AI logo

Model details

Nemotron 3 Ultra 550B A55B (Deep Infra)

Nemotron-3-Ultra-550B-A55B is a frontier-scale large language model trained by NVIDIA, positioned for demanding agentic, reasoning, and conversational workloads. According to its Microsoft Foundry listing, it is optimized for complex multi-step agents, long-context analysis, and high-accuracy reasoning over code, math, and science, reflecting NVIDIA's strategy of scaling hybrid architectures into production-ready assistants for enterprise and research users who need reliable tool-using behavior rather than simple chat.

The model introduces a hybrid latent mixture-of-experts design that interleaves Mamba-2 and MoE layers with select attention layers, and it is trained using an NVFP4 recipe to maximize compute efficiency while preserving quality. It answers questions by first generating an internal reasoning trace and then producing a final response, with reasoning behavior switchable through a flag in the chat template. NVIDIA also distributes the model through its official NIM container on the NGC Catalog, giving developers a supported inference pathway for long-horizon tasks where control over the reasoning process and sustained context handling matter more than raw chat fluency.

Eden AIdeepinfra/nemotron-3-ultra-550b-a55bnemotron

Quick Info

Powered by
Provider
Eden AI
Model key
deepinfra/nemotron-3-ultra-550b-a55b
Release date
Jun 4, 2026
Last updated
Jun 4, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.50
Output token cost
$2.20

Limits

Output tokens
128,000 tokens
Context window
262,144 tokens

Transparent token rates

Compare Nemotron 3 Ultra 550B A55B (Deep Infra) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron 3 Ultra 550B A55B (Deep Infra)

Eden AI

CoverageBenchmark

BenchmarkList's third-party benchmark page for NVIDIA Nemotron 3 Ultra 550B A55B compiles dated evaluation results (2026-06-04 to 2026-06-10) drawn from the model card and Artificial Analysis. Reported scores include PinchBench 90% (7/73, 92nd percentile), Terminal-Bench Hard 36.4% (43/326, 87th percentile), MultiChall The page is explicitly keyed to the exact Nemotron 3 Ultra 550B A55B variant and attributes the model to NVIDIA rather than to Deep Infra or Eden AI, so it qualifies as model/family-level evidence rather than gateway marketing. Readers should treat the figures as a point-in-time June 2026 snapshot from an aggregator; t

Eden AI

CoverageBenchmark

The OpenRouter model page for nvidia/nemotron-3-ultra-550b-a55b documents the model served through Eden AI's DeepInfra routing as an NVIDIA open frontier-reasoning Mixture-of-Experts model with 55B active parameters out of 550B total, built on a hybrid Transformer-Mamba architecture that supports text input and output The page also reports weighted-average effective prices of $0.3148 per 1M input tokens and $2.602 per 1M output tokens, and notes a price-history chart spanning mid-June through August. Caveats relevant to developers evaluating this route through Eden AI are that pricing and latency figures are provider-dependent and m

Videos about Nemotron 3 Ultra 550B A55B (Deep Infra)

More models around Nemotron 3 Ultra 550B A55B (Deep Infra)