Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Requesty logo

Model details

nvidia-nemotron-3-ultra

Nemotron 3 Ultra is a 550-billion-parameter reasoning model with 55 billion parameters active per inference, built around a hybrid Mamba-attention Mixture-of-Experts architecture. LatentMoE and Multi-Token Prediction layers support efficient processing and faster speculative generation, while inference-time reasoning budget controls let applications balance depth against output length.

The model is aimed at long-running coding agents, deep-research systems, and enterprise workflows that require sustained synthesis, planning, and multi-step problem solving. NVIDIA reports strong inference throughput in an eight-thousand-token-input and sixty-four-thousand-token-output comparison, along with competitive accuracy across varied benchmarks and strong RULER performance at very long context lengths. Its post-training combines supervised fine-tuning, reinforcement learning, and multi-teacher on-policy distillation, reflecting a broader push toward more capable, efficient models for complex agentic tasks.

Requestynvidia-nemotron-3-ultranemotron

Quick Info

Powered by
Provider
Requesty
Model key
nvidia-nemotron-3-ultra
Release date
Jun 23, 2026
Last updated
Jun 23, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.50
Output token cost
$2.50

Limits

Output tokens
131,072 tokens
Context window
262,144 tokens

Transparent token rates

Compare nvidia-nemotron-3-ultra pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about nvidia-nemotron-3-ultra

Requesty

CoverageRelease Notes

LangChain and NVIDIA jointly announced the NemoClaw for LangChain Deep Agents blueprint on July 8, 2026, combining LangChain Deep Agents Code, NVIDIA Nemotron 3 Ultra, and NVIDIA OpenShell runtime as a reference architecture for enterprise open agent systems. On LangChain's own agent evaluation suite, Nemotron 3 Ultra LangChain CEO Harrison Chase is quoted emphasizing that improving agents means improving the system around the model, including memory, tool use, evaluation, and model behavior, positioning the blueprint as a way for enterprises to own agent memory, workflows, traces, model weights, and tuning data. For developers usin

Requesty

CoverageBenchmark

NVIDIA's developer blog (July 8, 2026) details a LangChain harness engineering effort that improved Nemotron 3 Ultra's score on the Deep Agents evaluation suite from 94 to 96 out of 127, including resolving all three failing read-file pagination tests by adding a ReadFileContinuationNoticeMiddleware to teach the model The article points to a reusable pattern: the same loop can adapt Nemotron 3 Ultra to other agent frameworks by exposing evaluation results, an editable profile file, and a frontier model to propose edits. It also links next steps including reviewing the Nemotron 3 Ultra agent harness profile published by the LangChain

Requesty

CoverageBenchmark

A Towards AI technical explainer describes NVIDIA Nemotron 3 Ultra as a 550-billion-parameter open-weight model built for production agentic workloads rather than benchmark optimization. The model activates only 55 billion parameters per token (10%), and embeds a Mamba-2 state space layer alternating with transformer a The piece emphasizes that Nemotron 3 Ultra is positioned as the reasoning engine of the Nemotron 3 family, designed for sustained multi-step thinking and long-running agentic workflows rather than fast single-turn chatbot responses. It supports text input and output with a context window of up to 1 million tokens, enab

Requesty

Coverage

vLLM announced Day-0 support for NVIDIA Nemotron 3 Ultra on June 4, 2026, confirming the model is integrated into the open-source inference engine commonly used behind routing layers like Requesty. The blog details Nemotron 3 Ultra's hybrid Transformer-Mamba MoE architecture, multi-token prediction, and NVIDIA-optimize For developers routing Nemotron 3 Ultra through providers like Requesty, this post is the most direct primary source on serving characteristics: it confirms the model targets sustained reasoning depth while keeping inference fast, and validates that vLLM was used for both training-time rollouts and post-training evalua

Requesty

CoverageAnalysis

Artificial Analysis reports that NVIDIA announced Nemotron 3 Ultra during Jensen Huang's Computex keynote on May 31, 2026, positioning it as the largest and most intelligent US open-weights model to date. The model uses a hybrid Mamba-Transformer Mixture-of-Experts architecture with approximately 550 billion total para On the Artificial Analysis Intelligence Index, Nemotron 3 Ultra scores 48, ahead of other US open-weights models such as Gemma 4 31B (39), Nemotron 3 Super (36), and gpt-oss-120b (33), though still behind the Chinese-led open-weights frontier led by Kimi K2.6 at 54. Artificial Analysis also measured over 300 tokens per

Videos about nvidia-nemotron-3-ultra

More models around nvidia-nemotron-3-ultra