Vultr
NVIDIA released the open Nemotron 3 Nano Omni on April 28, 2026, combining vision, audio, and language into a single model with a 30B-A3B hybrid Mixture-of-Experts architecture that activates only 3B parameters per token. NVIDIA claims up to 9x higher throughput than other open omni models at equivalent interactivity, positioning it as the perception sub-agent paired with higher-tier Nemotron 3 Super and Ultra reasoners. On a single H200 with 8K input and 16K output, the model delivers 3.3x the inference throughput of Qwen3-30B-A3B and 2.2x that of GPT-OSS-20B, per NVIDIA's developer blog. On MediaPerf video tagging it processes 9.91 hours of video per hour versus roughly 3.8 h/h for Qwen3-VL. NVIDIA published weights, datasets, and training techniques, with NVFP4 quantization, NeMo tooling, and on-prem deployment supported across over 50 early partners.