Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Kilo Gateway logo

Model details

Nemotron 3 Nano 30B A3B

Nemotron 3 Nano 30B A3B is an NVIDIA-trained large language model built as a unified system for both reasoning and non-reasoning tasks. Rather than splitting these behaviors across separate checkpoints, the model first generates an internal reasoning trace and then produces a final answer. This design lets a single deployment serve conversational workloads and analytical tasks without swapping models, and a flag in the chat template controls whether the reasoning trace is emitted or suppressed, with a slight accuracy trade-off on harder prompts when reasoning is turned off.

Under the hood, the model uses a hybrid Mixture-of-Experts architecture that interleaves Mamba-2 and MoE layers with attention layers, giving it the throughput benefits of state-space sequence modeling while retaining the long-range modeling strengths of attention. It carries 30B total parameters but only activates around 3.5B per token through its expert routing, which is the source of the A3B designation. NVIDIA has also released ultra-efficient precision variants such as the NVFP4 build, signaling a roadmap toward smaller memory footprints and faster inference on commodity and edge hardware. That combination of open weights, configurable reasoning, and a sparse-active MoE design makes the model a practical fit for teams that want strong reasoning quality without paying the full cost of a dense 30B-parameter deployment.

Kilo Gatewaynvidia/nemotron-3-nano-30b-a3bnemotron

Quick Info

Powered by
Provider
Kilo Gateway
Model key
nvidia/nemotron-3-nano-30b-a3b
Release date
Dec 15, 2025
Last updated
Dec 15, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.05
Output token cost
$0.20

Limits

Output tokens
235,929 tokens
Context window
262,144 tokens

Transparent token rates

Compare Nemotron 3 Nano 30B A3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron 3 Nano 30B A3B

OpenRouter

Coverage

HokAI's model hub entry explicitly names the NVIDIA Nemotron 3 Nano 30B-A3B and describes it as an open-weight language model built by NVIDIA, released December 14, 2025 as the entry tier of the Nemotron 3 family. It reports benchmark scores from NVIDIA's technical report (arXiv 2512.20848): 73.04% on GPQA, 68.25% on L For tool-use workloads, HokAI reports that giving Nemotron 3 Nano tools lifts AIME accuracy from 89.06% to 99.17%, positioning it as a strong fit for agentic tool-calling pipelines. The page also notes an Artificial Analysis score of 38.8% on SWE-bench Verified. The license is listed as the NVIDIA Open Model License wi

Videos about Nemotron 3 Nano 30B A3B

More models around Nemotron 3 Nano 30B A3B