Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Venice AI logo

Model details

NVIDIA Nemotron 3 Nano 30B

Nemotron 3 Nano 30B stands apart from traditional dense language models through its hybrid Mixture-of-Experts architecture layered with Mamba-2 blocks. The design splits 29 layers into 23 combined Mamba-2 and MoE layers plus 6 dedicated attention layers, allowing the model to route each token to a specialized subset of experts—128 individual experts plus one shared expert per routing layer. This sparse activation means only a fraction of the model's total 30 billion parameters fire for any given token, dramatically improving throughput while preserving the quality of a much larger dense model. The architecture is also offered in NVFP4 ultra-efficient precision, further reducing the compute footprint needed for deployment.

NVIDIA trained this model from scratch and refined it using Qwen as a foundation for improvement. The post-training data extends to late November 2025, giving it a fresher knowledge baseline than many competing open models. A distinctive feature is its configurable reasoning mode: the model can either produce explicit reasoning traces before delivering answers for complex problems, or skip directly to responses for simpler queries. This flexibility makes it adaptable as a general-purpose assistant or a more deliberate reasoning engine. The open weights, training data, and recipes are all publicly available under the Nemotron umbrella, positioning the model as a platform for the community to build on rather than a black box. Practical strengths include strong coding and reasoning performance, massive context handling up to a million tokens, and fast inference suitable for agent workflows and production applications.

Venice AInvidia-nemotron-3-nano-30b-a3bnemotron

Quick Info

Powered by
Provider
Venice AI
Model key
nvidia-nemotron-3-nano-30b-a3b
Release date
Jan 27, 2026
Last updated
Jun 11, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.075
Output token cost
$0.30

Limits

Output tokens
16,384 tokens
Context window
128,000 tokens

Transparent token rates

Compare NVIDIA Nemotron 3 Nano 30B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about NVIDIA Nemotron 3 Nano 30B

Venice AI

Coverage

HokAI's hub page profiles NVIDIA Nemotron 3 Nano 30B A3B as the entry tier of the Nemotron 3 family, released December 14, 2025 by NVIDIA, with its larger Super (120B) and Ultra (550B) siblings following in March and June 2026. It frames the model as an efficiency-focused open-weight design for high-volume agentic pipe The page details the architecture: 31.6B total parameters with only ~3.2B active per token via a Mixture-of-Experts design that routes to 6 of 128 experts plus shared experts per token, inside a hybrid stack of Mamba-2 state-space layers and grouped-query-attention Transformer layers across 52 total layers. It reports

Venice AI

CoverageBenchmark

Benchgen's model page describes NVIDIA Nemotron 3 Nano 30B A3B as an open-weight large language model built by NVIDIA on a hybrid Mamba-2/Transformer Mixture-of-Experts architecture, with 30B total parameters but only about 3.5B active per token (the "A3B" tag), enabling a 30B-class model to run on a single H100 with s Architecturally, the page specifies 52 layers split across 23 Mamba-2 layers, 23 MoE layers, and 6 grouped-query-attention layers, with each MoE layer carrying 128 routed experts plus one shared expert and routing 6 experts per token — a Mamba-heavy design the page credits with making 1M-token long-context behavior pra

Videos about NVIDIA Nemotron 3 Nano 30B

More models around NVIDIA Nemotron 3 Nano 30B