DigitalOcean
by Dan Ferguson, Malav Shastri, and Vivek Gangasani on 28 APR 2026 in Amazon SageMaker JumpStart, Announcements, Foundational (100), Generative AI,...
Model details
Nemotron 3 Nano Omni is NVIDIA's latest entry in the Nemotron multimodal series and notably the first in the line to handle audio inputs natively, alongside text, images, and video, while emitting text outputs. It is built on the Nemotron 3 Nano 30B-A3B backbone, a highly efficient architecture that the team extended with multimodal token-reduction techniques aimed at cutting inference latency and raising throughput compared with peers of similar size. Weights are being released openly in BF16, FP8, and FP4 formats, with portions of the training data and codebase shared to support further research and downstream tuning.
The model is positioned for agentic and reasoning-heavy workloads, with NVIDIA highlighting leading real-world performance in document understanding, long audio-video comprehension, and computer-use tasks versus its predecessor, Nemotron Nano V2 VL. Reporting also frames it as a "best-in-class open omni-modal reasoning model" that can unify vision, audio, and language in a single pipeline, simplifying stacks that would otherwise stitch separate encoders and models together. Availability extends beyond the official channels into managed platforms such as Amazon SageMaker JumpStart, making it practical for teams that want a single open model to handle mixed text, visual, and audio reasoning for assistants and agents.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
DigitalOcean
by Dan Ferguson, Malav Shastri, and Vivek Gangasani on 28 APR 2026 in Amazon SageMaker JumpStart, Announcements, Foundational (100), Generative AI,...
DigitalOcean
Best-in-class open omni-modal reasoning model delivers the highest efficiency and accuracy to power agentic workflows such as computer use,...
DigitalOcean
NVIDIA's developer blog introduces Nemotron 3 Nano Omni as a unified multimodal open model designed to replace fragmented vision-language-audio stacks in agentic systems. It handles video, audio, image, and text in a single perception-to-action loop, improving convergence and reducing orchestration complexity and infer According to the blog, Nemotron 3 Nano Omni delivers best-in-class accuracy on document intelligence leaderboards (MMlongbench-Doc, OCRBenchV2) and leads on video and audio understanding benchmarks (WorldSense, DailyOmni, VoiceBench). On the open MediaPerf benchmark for real media workloads, it achieves the highest thr
DigitalOcean
NVIDIA's launch blog announces Nemotron 3 Nano Omni as an open multimodal model unifying vision, audio, and language into a single system, enabling agents to deliver faster responses with advanced reasoning across video, audio, image, and text. By combining vision and audio encoders within its 30B-A3B hybrid MoE archit Several AI and software companies have already adopted Nemotron 3 Nano Omni, including Aible, Applied Scientific Intelligence (ASI), Eka Care, Foxconn, H Company, Palantir, and Pyler, while Dell Technologies, Docusign, Infosys, K-Dense, Lila, Oracle, and Zefr are evaluating it. H Company CEO Gautier Cloix highlighted t
DigitalOcean
Artificial Analysis documents the Nemotron 3 Nano Omni 30B A3B Reasoning variant as an open weights model released in April 2026, with 30B total and 3B active parameters, a 256k token context window, and support for text, image, speech, and video input with text output under the NVIDIA Open Model License. It scores 10 Pricing is reported at $0.20 per 1M input tokens and $1.095 per 1M output tokens, described as expensive relative to the open weights class medians of $0.05 input and $0.15 output, and the page notes that a non-reasoning variant may also exist alongside the reasoning configuration shown. Model weights are hosted on Hug