Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
OpenRouter logo

Model details

Nemotron 3 Ultra (free)

Nemotron 3 Ultra is a large open-weight reasoning and orchestration model that uses a hybrid Transformer-Mamba mixture-of-experts design, with roughly 55B active parameters drawn from a 550B total parameter pool. This sparse activation pattern lets the model keep inference costs manageable while still bringing frontier-scale capacity to bear on demanding tasks. It belongs to NVIDIA's broader Nemotron family of open models aimed specifically at agentic AI, and the model weights are publicly available on Hugging Face, reinforcing its open-weights positioning.

The model is engineered for long-running, multi-step agentic pipelines rather than single-turn chat. Its massive 1,000,000-token context window allows it to hold entire codebases, research documents, or extended interaction histories in memory, while its reasoning strength targets planning, coding agents, deep research, and complex enterprise orchestration. NVIDIA emphasizes high-throughput inference to keep up with high-volume agent traffic, and the free-tier deployment on OpenRouter makes this large MoE architecture accessible for teams that want to experiment with orchestrated AI workflows at scale.

OpenRouternvidia/nemotron-3-ultra-550b-a55b:freenemotron

Quick Info

Powered by
Provider
OpenRouter
Model key
nvidia/nemotron-3-ultra-550b-a55b:free
Release date
Jun 4, 2026
Last updated
Jun 4, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
65,536 tokens
Context window
1,000,000 tokens

Latest news about Nemotron 3 Ultra (free)

OpenRouter

Official sourceBenchmark

OpenRouter's model page for nvidia/nemotron-3-ultra-550b-a55b:free documents the exact variant: 55B active parameters out of 550B total using a hybrid Transformer-Mamba mixture-of-experts architecture, text input/output, and a 1M-token context window. The listing notes a June 4, 2026 release, a Free price tier, and pos Performance telemetry on the page reports a best-provider P50 latency of 1.57s and throughput of 48 tok/s, with 68.88% uptime, and overall latency percentiles ranging from P50 of 2.49s to P99 of 51.06s. A free-endpoint notice warns that session data is logged by NVIDIA for security and product-improvement purposes and

OpenRouter

CoverageBenchmark

Kilo Code's page for nvidia/nemotron-3-ultra-550b-a55b:free confirms the exact model identifier and its creation date of June 4, 2026, with a 1,000,000-token context window, 65,536 max completion tokens, text-only input modality, and a $0.00 per 1M tokens price for both input and output. Content moderation is listed as The page documents coding-relevant capabilities including function calling, tool choice control, structured outputs with JSON schema validation, and reasoning tokens for extended thinking. Kilo Code leaderboard rankings from the prior week show the model at rank 6 in Code mode, rank 6 in Debug, rank 9 in Orchestrator,

OpenRouter

Official sourceBenchmark

NVIDIA Nemotron 3 Ultra (free) is available through OpenRouter under the model ID nvidia/nemotron-3-ultra-550b-a55b:free. The model was released on June 4, 2026, accepts and produces text, and provides a context window of up to 1 million tokens. OpenRouter lists the free endpoint at $0 per million input tokens and $0 p The model is a 550-billion-parameter mixture-of-experts system with 55 billion parameters active per token, using a hybrid Transformer-Mamba architecture. OpenRouter positions it for agent orchestration, coding agents, deep research, and other long-running agentic workloads. The page shows an NVIDIA-hosted free provide

OpenRouter

Official sourceBenchmark

OpenRouter’s model comparison page lists NVIDIA Nemotron 3 Ultra (free) as a free variant with a 1-million-token context window and $0-per-million-token pricing for both inputs and outputs. The listing says the model was released on June 4, 2026, and displays usage-oriented category labels for Programming, Science, and The page describes Nemotron 3 Ultra as a 550-billion-parameter mixture-of-experts model with 55 billion active parameters and a hybrid Transformer-Mamba architecture. It targets long-running agentic applications, including agent orchestration, coding agents, deep research, and complex enterprise tasks, with the listing

Videos about Nemotron 3 Ultra (free)

More models around Nemotron 3 Ultra (free)