Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Kilo Gateway logo

Model details

NVIDIA: Nemotron 3 Ultra (free)

Nemotron 3 Ultra is a frontier reasoning and orchestration model built around a hybrid Transformer-Mamba mixture-of-experts design, with roughly 55B active parameters drawn from a much larger 550B total parameter pool. That combination is intended to keep inference economical on long-running agent pipelines while still drawing on the breadth of a very large model. NVIDIA positions the family as an open-weight release, and the model sits alongside other Nemotron variants aimed at different throughput and workload profiles.

In practice, the model is aimed at multi-step reasoning, planning, and agent orchestration, including coding agents and deep research workflows that need to stay coherent across very long contexts. The architecture and active-parameter count are tuned for high-throughput inference, which suits enterprise pipelines where many agent sessions run in parallel. Open weights make it adaptable for teams that want to fine-tune or self-host for specialized domains, while the hybrid Mamba-Transformer backbone helps it handle the extended reasoning chains typical of complex agentic tasks.

Kilo Gatewaynvidia/nemotron-3-ultra-550b-a55b:freenemotron

Quick Info

Powered by
Provider
Kilo Gateway
Model key
nvidia/nemotron-3-ultra-550b-a55b:free
Release date
Jun 4, 2026
Last updated
Jun 4, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
65,536 tokens
Context window
1,000,000 tokens

Latest news about NVIDIA: Nemotron 3 Ultra (free)

Kilo Gateway

Official sourceBenchmark

The Kilo Code model page for NVIDIA Nemotron 3 Ultra (free) confirms the same core specifications as OpenRouter, listing model ID nvidia/nemotron-3-ultra-550b-a55b:free, a 1,000,000-token context window, 65,536 max output tokens, text-only input modality, a $0.00 per million token price on both input and output, and a The page provides Kilo-internal community rankings placing the model at Code rank 10, Ask rank 25, Debug rank 52, and Orchestrator rank 42 on the Kilo Code leaderboard, along with weekly token-usage statistics for the model within the Kilo Code community. It is offered as part of Kilo's free access tier across 500+ mod

Kilo Gateway

CoverageBenchmark

The OpenRouter model page documents NVIDIA Nemotron 3 Ultra (free) as a 550B-total / 55B-active MoE model on a hybrid Transformer-Mamba architecture, with text input/output, a 1M-token context window, and a release date of June 4, 2026; the free variant is priced at $0 per million tokens for both input and output and i The same listing flags an important developer caveat: by using the free endpoint, users consent to NVIDIA collecting and logging session data for product improvement, and the page explicitly warns against uploading confidential information or personal data such as voices or faces, linking to NVIDIA's Privacy Policy and

Videos about NVIDIA: Nemotron 3 Ultra (free)

More models around NVIDIA: Nemotron 3 Ultra (free)