Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Deep Infra logo

Model details

Nemotron 3 Nano Omni 30B A3B Reasoning

Nemotron 3 Nano Omni 30B A3B is an open multimodal model from NVIDIA designed to perform reasoning across text, images, video, and audio within a single system. Rather than bolting together separate encoders, the model is built on a hybrid Mixture-of-Experts architecture that carries 30 billion total parameters while activating only about 3 billion per token, a sparse activation pattern that aims to keep multimodal reasoning affordable for production deployment. On the input side, the model can accept images, video frames, and audio alongside text, but its explicit chain-of-thought reasoning pathway is currently limited to text and image inputs; video and audio requests need the thinking mode disabled through the chat template options so the model can still respond without producing internal reasoning traces.

In practical use, the model behaves like a mid-sized reasoning engine that fits well into chat-completion style workflows and agentic pipelines. NVIDIA exposes it on its NIM inference microservice with the model identifier nvidia/nemotron-3-nano-omni-30b-a3b-reasoning, supporting a reasoning budget parameter so callers can control how much internal deliberation the model performs, alongside standard sampling controls such as temperature and top-p tuning. The combination of multimodal understanding, sparse MoE efficiency, and controllable reasoning budget makes it a reasonable choice for teams that want a single open model to handle mixed-media analysis, document and image question answering, and tool-assisted reasoning tasks without spinning up separate specialists for each modality.

Deep Infranvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoningnemotrondeprecated

Quick Info

Powered by
Provider
Deep Infra
Model key
nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning
Release date
Apr 28, 2026
Last updated
Apr 28, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$0.80

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Transparent token rates

Compare Nemotron 3 Nano Omni 30B A3B Reasoning pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron 3 Nano Omni 30B A3B Reasoning

Deep Infra

CoverageBenchmark

OpenRouter lists NVIDIA Nemotron 3 Nano Omni (30B-A3B) as a free endpoint released April 28, 2026, accepting text, image, video, and audio inputs with text output for use as a perception sub-agent in enterprise agent systems. The page documents a 256K context window, a 16,384-token reasoning budget, and extended-thinki The hosting page reports measured telemetry for the NVIDIA-hosted free tier: P50 latency of 0.41s, P50 throughput of 49 tokens/sec, and 88.99% uptime over the prior week. It also cites the hybrid MoE Transformer-Mamba architecture with Conv3D and Efficient Video Sampling, claiming approximately 2x higher throughput and

Videos about Nemotron 3 Nano Omni 30B A3B Reasoning

More models around Nemotron 3 Nano Omni 30B A3B Reasoning