Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Nvidia logo

Model details

Nemotron 3 Nano Omni

Nemotron 3 Nano Omni is positioned as a perception and context sub-agent inside enterprise agent systems, accepting text, images, video, and audio together and emitting text, so a single inference pass can reason across mixed modalities instead of stitching separate vision and speech pipelines. It is described as the first entry in the Nemotron multimodal series to natively support audio alongside the other input types, and the authors highlight real-world document understanding, long audio-video comprehension, and agentic computer use as the areas where it shows its strongest results compared to its predecessor, Nemotron Nano V2 VL.

The model is built on the Nemotron 3 Nano 30B-A3B backbone and pairs a hybrid MoE Transformer-Mamba design with Conv3D video layers and an Efficient Video Sampling (EVS) stage, using multimodal token-reduction techniques to cut inference latency and lift throughput relative to similarly sized models. Open weights are published on Hugging Face in BF16, with FP8 and FP4 variants also released alongside portions of the training data and code, making the system a practical fit for teams that want to self-host a compact multimodal reasoner or wire it into a larger agent stack where low-latency perception of documents, screens, and long audiovisual context matters.

Nvidianvidia/nemotron-3-nano-omni-30b-a3b-reasoningnemotron

Quick Info

Powered by
Provider
Nvidia
Model key
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning
Release date
Apr 28, 2026
Last updated
Apr 28, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
65,536 tokens
Context window
256,000 tokens

Latest news about Nemotron 3 Nano Omni

Nvidia

CoverageBenchmark

Artificial Analysis' third-party evaluation profiles Nemotron 3 Nano Omni 30B A3B Reasoning as an NVIDIA open-weights model released in April 2026. It reports an Intelligence Index of 10 out of 142 in its comparable class (above the median of 8), a 256K-token context window, 30B total / 3B active parameters, NVIDIA Ope Pricing for the hosted API is listed at $0.20 per 1M input tokens and $1.09 per 1M output tokens, which Artificial Analysis flags as expensive relative to comparable open-weight models (median $0.05 / $0.15). Output speed is reported as unknown. The page provides independent benchmark and positioning context but does n

Videos about Nemotron 3 Nano Omni

More models around Nemotron 3 Nano Omni