Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vultr logo

Model details

Nemotron 3 Nano Omni 30B A3B Reasoning

Nemotron 3 Nano Omni 30B A3B Reasoning belongs to the Nemotron model family and is listed on NVIDIA's NIM model platform, indicating NVIDIA's role in publishing and hosting the model for developer access. The model's appearance in both the NVIDIA NIM catalog and the DeepInfra announcement blog on the same date suggests a coordinated launch across distribution partners, giving developers multiple entry points to evaluate the model. Its position in the Nemotron line places it within a family of models that NVIDIA has designed for practical reasoning and instruction-following workloads, where efficiency and accessibility for enterprise and research users are key priorities.

As a Nano-tier member of the Nemotron 3 generation, this model is positioned for lightweight deployment scenarios where computational resources are constrained but reasoning quality still matters. The Nano designation within Nemotron has historically targeted developers seeking a balance between capability footprint and hardware requirements, making it suitable for integration into production pipelines, on-premise deployments, and cost-sensitive inference workloads. The simultaneous availability through NVIDIA's NIM interface and third-party distribution channels like DeepInfra reflects a broader strategy of making Nemotron models broadly accessible to the developer ecosystem for building reasoning-enabled applications.

Vultrnemotron-3-nano-omni-30b-a3b-reasoningnemotron

Quick Info

Powered by
Provider
Vultr
Model key
nemotron-3-nano-omni-30b-a3b-reasoning
Release date
Apr 28, 2026
Last updated
Apr 28, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.25

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Transparent token rates

Compare Nemotron 3 Nano Omni 30B A3B Reasoning pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron 3 Nano Omni 30B A3B Reasoning

Vultr

Coverage

NVIDIA released the open Nemotron 3 Nano Omni on April 28, 2026, combining vision, audio, and language into a single model with a 30B-A3B hybrid Mixture-of-Experts architecture that activates only 3B parameters per token. NVIDIA claims up to 9x higher throughput than other open omni models at equivalent interactivity, positioning it as the perception sub-agent paired with higher-tier Nemotron 3 Super and Ultra reasoners. On a single H200 with 8K input and 16K output, the model delivers 3.3x the inference throughput of Qwen3-30B-A3B and 2.2x that of GPT-OSS-20B, per NVIDIA's developer blog. On MediaPerf video tagging it processes 9.91 hours of video per hour versus roughly 3.8 h/h for Qwen3-VL. NVIDIA published weights, datasets, and training techniques, with NVFP4 quantization, NeMo tooling, and on-prem deployment supported across over 50 early partners.

Vultr

Coverage

DeepInfra became an official launch partner for NVIDIA Nemotron 3 Nano Omni on April 28, 2026, making the first multimodal entry in the Nemotron 3 family available from day one. The open model processes images, video, audio, documents, and text in a single inference pass, with NVIDIA claiming 9x higher throughput than other open omni models at the same interactivity level. The model uses a hybrid Mixture of Experts and Mamba-Transformer backbone extended from Nemotron 3 Nano, with 3D convolution layers and Efficient Video Sampling for low-cost video reasoning. A hybrid MoE design activates 3B of 30B total parameters per token, enabling developers to build always-on multimodal sub-agents for computer use and audio-video understanding with minimal code.

Vultr

Coverage

NVIDIA announced the Nemotron 3 Nano Omni on April 28, 2026, as an open omnimodal inference model that unifies vision, speech, and language processing into a single system, eliminating the latency and context loss of chained separate models. The model targets agent workflows such as computer use, document analysis, and speech/video reasoning. NVIDIA reports the Nemotron 3 Nano Omni ranks first in six categories spanning complex document recognition, video, and audio understanding benchmarks. It significantly outperforms its predecessor, the Nemotron Nano VL V2, across the board, with particularly substantial gains on OSWorld, a benchmark for multimodal agents measuring real-world computer task completion.

Videos about Nemotron 3 Nano Omni 30B A3B Reasoning

More models around Nemotron 3 Nano Omni 30B A3B Reasoning