Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
OpenRouter logo

Model details

Nemotron 3 Nano Omni (free)

Nemotron 3 Nano Omni is an open multimodal system designed to consolidate vision, audio, and language processing into a single inference loop, replacing the fragmented pipelines that many enterprise agents rely on today. Built as a hybrid Mamba2 Transformer mixture-of-experts design, the 30-billion-parameter model activates only about 3 billion parameters per forward pass, delivering throughput and compute economics that resemble a much smaller dense model while still supporting complex multimodal reasoning across text, image, video, and audio inputs.

On the video side, NVIDIA pairs the MoE backbone with Conv3D video layers and its Efficient Video Sampling technique, reporting roughly two times higher throughput and two and a half times lower compute compared with separate vision and speech pipelines, and up to nine times the throughput of comparable open omnimodal models on video and document workloads. The model leads several open multimodal leaderboards, runs locally on about 25 GB of RAM, and is distributed through Hugging Face, OpenRouter, Ollama, and NVIDIA's NIM microservice, making it practical for teams building agent systems that need coherent cross-modal context without orchestrating multiple specialized models.

OpenRouternvidia/nemotron-3-nano-omni-30b-a3b-reasoning:freenemotron

Quick Info

Powered by
Provider
OpenRouter
Model key
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free
Release date
Apr 28, 2026
Last updated
Apr 28, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
65,536 tokens
Context window
256,000 tokens

Latest news about Nemotron 3 Nano Omni (free)

OpenRouter

Coverage

GIGAZINE's English coverage of NVIDIA's April 28, 2026 announcement frames Nemotron 3 Nano Omni as an omnimodal inference model that integrates visual, auditory, and linguistic processing into a single system, replacing the multi-model pipelines many AI agents rely on for video, audio, image, and text reasoning. NVIDIA For developer context, the piece explains that consolidating vision, speech, and language into one model eliminates cross-model data handoffs that introduce latency and context loss, enabling agents to respond faster and more coherently across modalities. Benchmark comparisons against the Nemotron Nano VL V2 predecesso

OpenRouter

CoverageBenchmark

BuildFastWithAI's review of NVIDIA Nemotron 3 Nano Omni provides a developer-oriented breakdown of the same 30B-A3B hybrid MoE model announced on April 28, 2026, emphasizing that it activates only ~3B parameters per token and is distributed on Hugging Face, OpenRouter (free), and build.nvidia.com as an NIM microservice The article also reports performance and deployment claims — 25GB RAM for local execution, up to 9x higher throughput than other open omni models on video and document workloads, and comparisons against Qwen3-Omni, MiMo-V2.5-Pro, and Gemini Nano — along with run instructions using Unsloth and llama.cpp. As a third-part

OpenRouter

Official sourceBenchmark

NVIDIA Nemotron 3 Nano Omni (free) is listed on OpenRouter as a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. According to the OpenRouter product page, it accepts text, image, video, and audio inputs and produces text output, enabling agents to per The same OpenRouter page describes the underlying architecture as a hybrid MoE Transformer-Mamba design with Conv3D video layers and Efficient Video Sampling (EVS), claiming approximately 2x higher throughput and 2.5x lower compute for video reasoning versus separate vision-plus-speech pipelines, and supports a 16,384

OpenRouter

Official sourceComparison

OpenRouter's compare hub for Nemotron 3 Nano Omni (free) places the model in head-to-head slots against curated cohorts of flagship, most-affordable, best-for-code, and reasoning models, including Claude Fable 5, Gemini 3.1 Pro Preview, GPT-5.5, DeepSeek V4 Flash 0423, Gemini 2.5 Flash Lite, Hy3 preview, Claude Opus 4. For developers, the page confirms that the free Nemotron 3 Nano Omni endpoint is accessible alongside hundreds of other models through a single OpenRouter API key, and that the model can be selected in the compare UI without email verification gating. It also surfaces the same NVIDIA free-endpoint data-logging and cons

Videos about Nemotron 3 Nano Omni (free)

More models around Nemotron 3 Nano Omni (free)