Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Venice AI logo

Model details

NVIDIA Nemotron 3 Ultra

Nemotron 3 Ultra is NVIDIA's flagship entry in the Nemotron 3 family, a Mixture-of-Experts model that pairs a Hybrid Mamba-Attention backbone with NVIDIA's LatentMoE design. The model carries 550 billion total parameters with 55 billion active per forward pass, letting it draw on broad knowledge while keeping inference costs closer to a much smaller model. Multi-Token Prediction layers are baked into the architecture to accelerate generation through native speculative decoding, and pretraining in NVFP4 (NVIDIA's 4-bit floating-point format) helps make that large active footprint practical to serve.

Beyond pretraining, Nemotron 3 Ultra is post-trained with a multi-stage pipeline that combines supervised fine-tuning, reinforcement learning, and Multi-teacher On-Policy Distillation, which NVIDIA describes as designed to improve overall accuracy. The system supports inference-time reasoning budget control and is positioned as the strongest model in the Nemotron 3 lineup, with NVIDIA reporting on-par accuracy with other leading open models across a diverse benchmark set and notably higher throughput on long-output tasks. It is a natural fit for developer workflows that need open-weight flexibility, sustained long-context reasoning, and a model that can act as a drop-in engine for code generation, agentic tools, and other production pipelines.

Venice AInvidia-nemotron-3-ultra-550b-a55bnemotron

Quick Info

Powered by
Provider
Venice AI
Model key
nvidia-nemotron-3-ultra-550b-a55b
Release date
Jun 4, 2026
Last updated
Jun 11, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.625
Output token cost
$3.125

Limits

Output tokens
32,768 tokens
Context window
256,000 tokens

Transparent token rates

Compare NVIDIA Nemotron 3 Ultra pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about NVIDIA Nemotron 3 Ultra

Venice AI

CoverageBenchmark

NVIDIA published a technical blog post on July 8, 2026, detailing how tuning a LangChain Deep Agents harness profile for NVIDIA Nemotron 3 Ultra improved evaluation scores from 94 to 96 out of 127, and resolved all three failing read-file pagination tests. The optimization added a ReadFileContinuationNoticeMiddleware t The post also describes an automated improvement loop implemented in the NemoClaw community repository, which runs the benchmark, proposes profile changes, verifies fixes across multiple runs, and re-runs the full suite to prevent regressions. The same propose-verify-keep loop can adapt Nemotron 3 Ultra to other agent

Venice AI

CoverageBenchmark

NVIDIA announced on July 8, 2026 that LangChain tuned its Deep Agents harness for NVIDIA Nemotron 3 Ultra, achieving the highest accuracy among open models while completing more tasks at higher throughput and running at 10x lower inference cost per run than leading closed models. Measured against LangChain's Deep Agent LangChain's agent engineering platform has more than 200 million monthly downloads, and the collaboration demonstrates that enterprises can get strong performance from an open stack while keeping control over their agent systems. Companies including Abridge, Amdocs, and Box are embedding specialized agents directly int

Venice AI

Coverage

Kilo Code announced on June 4, 2026 that NVIDIA Nemotron 3 Ultra is available for free in its IDE/terminal product, framing it as NVIDIA's flagship open-weights model introduced by Jensen Huang at Computex 2026. The post confirms the hybrid Mamba-Transformer MoE architecture with 550B total parameters and 55B active pe Kilo positions Nemotron 3 Ultra as an upgrade over Nemotron 3 Super (120B), which it calls a daily driver for many users but limited on planning and long-horizon tasks. The post is promotional ("FREE in Kilo for a limited time") and tied to a specific IDE vendor rather than Venice, so its throughput and benchmark claim

Venice AI

CoverageAnalysis

NVIDIA announced the release of Nemotron 3 Ultra in Jensen Huang's Computex keynote on May 31, 2026. Artificial Analysis partnered with NVIDIA to evaluate the model, confirming it has 550 billion total parameters with 55 billion active per token, making it the largest Nemotron 3 model to date and the most intelligent U Nemotron 3 Ultra also leads its size class on inference speed, with Artificial Analysis measuring over 300 tokens per second on a pre-release DeepInfra endpoint, compared to 50–100 tokens per second for peer models from Chinese labs such as DeepSeek and Moonshot. The model uses BF16 weights and will also be made availa

Videos about NVIDIA Nemotron 3 Ultra

More models around NVIDIA Nemotron 3 Ultra