Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Kilo Gateway logo

Model details

NVIDIA: Nemotron 3 Nano Omni (free)

Nemotron 3 Nano Omni is a 30B-parameter model with an A3B active expert configuration, presented across listing pages as an open multimodal system that ingests text, images, video, and audio and produces text responses. The architecture is described as a hybrid mixture-of-experts Transformer, designed so the model can act as a lightweight perception and context module inside larger enterprise agent pipelines rather than as a standalone chatbot. Because the active parameter count is small relative to the total, the design favors efficient inference while still allowing broad input coverage for documents, screens, video frames, and audio clips routed through an orchestration layer.

In practice, the model fits well as a vision-and-audio preprocessor for retrieval, summarization, and grounding steps where a downstream reasoning agent needs structured observations from mixed media. Listings report a generous context window that comfortably holds long videos and audio transcripts, and the free routing makes it attractive for prototyping multimodal sub-agents without budget friction. One hosted provider reports steady throughput in the mid-fifties of tokens per second with sub-second time-to-first-token, which is serviceable for batch perception work though not benchmark-leading. No standardized benchmark scores are published yet for this checkpoint, so fit decisions are best anchored in the multimodal coverage and the sub-agent role rather than head-to-head leaderboard claims.

Kilo Gatewaynvidia/nemotron-3-nano-omni-30b-a3b-reasoning:freenemotron

Quick Info

Powered by
Provider
Kilo Gateway
Model key
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free
Release date
Apr 28, 2026
Last updated
Apr 28, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
65,536 tokens
Context window
256,000 tokens

Latest news about NVIDIA: Nemotron 3 Nano Omni (free)

Kilo Gateway

CoverageBenchmark

The OpenRouter listing for NVIDIA Nemotron 3 Nano Omni (free) documents a 30B-A3B open multimodal model released on April 28, 2026 and built on a hybrid MoE Transformer-Mamba architecture with Conv3D video layers and Efficient Video Sampling (EVS). NVIDIA positions it as a perception and context sub-agent for enterpris Operational metrics on the same OpenRouter page show a P50 latency of 0.41s and throughput of 49 tok/s for the free endpoint over the prior week, with P99 latency averaging 14.37s and end-to-end P50 of 11.80s, alongside an uptime of 88.99%. The listing also surfaces a governance caveat: users of the free endpoint conse

Videos about NVIDIA: Nemotron 3 Nano Omni (free)

More models around NVIDIA: Nemotron 3 Nano Omni (free)