Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Ollama Cloud logo

Model details

nemotron-3-ultra

Nemotron 3 Ultra is the flagship model in the Nemotron 3 family, combining a Mixture-of-Experts Hybrid Mamba-Attention design with NVIDIA's LatentMoE technique to call four experts at the inference cost of one. The model is built around 550 billion total parameters with 55 billion active per pass, and it adds Multi-Token Prediction layers that enable native speculative decoding for faster generation on long sequences. Training is staged, beginning with an NVFP4 pre-training pass and finishing with a supervised fine-tuning, reinforcement learning, and multi-teacher on-policy distillation pipeline aimed at raising answer quality and reasoning discipline.

The model's design centers on sustained, long-running reasoning: it offers inference-time reasoning budget controls so users can trade compute against accuracy, and it supports a one-million-token context window, where it reports outperforming other leading open large language models on the RULER long-context benchmark. Throughput comparisons against other major open MoE models show substantial inference gains on long input and output workloads, making the model a practical fit for autonomous coding agents that plan, refactor, and recover across large codebases, deep research loops that synthesize across many sources, and enterprise or EDA workflows that demand extended, reliable reasoning rather than short conversational turns.

Ollama Cloudnemotron-3-ultranemotron

Quick Info

Powered by
Provider
Ollama Cloud
Model key
nemotron-3-ultra
Release date
Jun 4, 2026
Last updated
Jun 4, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$3.00

Limits

Output tokens
128,000 tokens
Context window
262,144 tokens

Transparent token rates

Compare nemotron-3-ultra pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about nemotron-3-ultra

Ollama Cloud

CoverageAnalysis

NVIDIA announced Nemotron 3 Ultra in Jensen Huang's Computex keynote (May 31, 2026), releasing it as the largest model in the Nemotron 3 family at approximately 550 billion total parameters with 55 billion active (around 90% sparsity), and as the leading US open weights intelligence model according to Artificial Analys On a pre-release DeepInfra endpoint, Nemotron 3 Ultra was served at over 300 tokens per second, a leading speed for its intelligence class relative to China-based open models from DeepSeek and Moonshot, which are typically served at 50-100 tokens per second. The page notes that additional analysis and full benchmark re

Videos about nemotron-3-ultra

More models around nemotron-3-ultra