Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Novita AI logo

Model details

Nemotron 3 Nano 30B A3B

We couldn't load the overview just now. Please try again in a little while.

Novita AInvidia/nemotron-3-nano-30b-a3bnemotron

Quick Info

Powered by
Provider
Novita AI
Model key
nvidia/nemotron-3-nano-30b-a3b
Release date
Dec 15, 2025
Last updated
Dec 15, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.05
Output token cost
$0.20

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Transparent token rates

Compare Nemotron 3 Nano 30B A3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron 3 Nano 30B A3B

OpenRouter

Coverage

HokAI's model hub entry explicitly names the NVIDIA Nemotron 3 Nano 30B-A3B and describes it as an open-weight language model built by NVIDIA, released December 14, 2025 as the entry tier of the Nemotron 3 family. It reports benchmark scores from NVIDIA's technical report (arXiv 2512.20848): 73.04% on GPQA, 68.25% on L For tool-use workloads, HokAI reports that giving Nemotron 3 Nano tools lifts AIME accuracy from 89.06% to 99.17%, positioning it as a strong fit for agentic tool-calling pipelines. The page also notes an Artificial Analysis score of 38.8% on SWE-bench Verified. The license is listed as the NVIDIA Open Model License wi

Kilo Gateway

Official source

NVIDIA announced the Nemotron 3 family of open models, releasing the Nano variant alongside its technical report. The family uses a hybrid Mamba-Transformer mixture-of-experts architecture designed for high-throughput agentic inference. Nemotron 3 Nano is a 3.2B active parameter (31.6B total) model that NVIDIA reports as more accurate than GPT-OSS-20B and Qwen3-30B-A3B-Thinking-2507 on popular benchmarks across categories. Nemotron 3 Nano supports context lengths up to 1M tokens and on a single H200 GPU at 8K input / 16K output it delivers 3.3x higher inference throughput than Qwen3-30B-A3B and 2.2x higher than GPT-OSS-20B. Super and Ultra variants, which add LatentMoE, Multi-Token Prediction layers, and NVFP4 training, are planned for later release. The model is positioned for cost-efficient deployment of agentic, reasoning, and conversational workloads.

Kilo Gateway

Official source

NVIDIA published a technical report for Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model with 31.6B total parameters and only 3.2B activated per forward pass. The model was pretrained on 25 trillion text tokens, including more than 3 trillion new unique tokens over Nemotron 2, followed by supervised fine-tuning and large-scale reinforcement learning on diverse environments. Both the pretrained Base and post-trained checkpoints were released on Hugging Face. According to the report, Nemotron 3 Nano achieves better or on-par accuracy compared to GPT-OSS-20B and Qwen3-30B-A3B-Thinking-2507 on popular benchmarks. The model provides up to 2.2x and 3.3x faster inference throughput than those models respectively on an 8K input / 16K output scenario. It also supports context lengths up to 1M tokens and is designed for enhanced agentic, reasoning, and chat capabilities using a granular MoE architecture with 6 of 128 experts activated per token.

Kilo Gateway

CoverageBenchmark

Nemotron 3 Nano (30B A3B) holds a composite LLM Stats Score of 13.6 and sits at rank 198 in the llm-stats.com leaderboard, placing it well behind top-tier proprietary models like Claude Opus 5.5 (60.3) but within the cost-efficient cluster alongside Gemma 4 E4B (13.8) and GPT OSS 120B (28.7). The page reports a blended price of $0.057 per million tokens, framing the model as a budget-tier option among tracked mixture-of-experts systems. On individual benchmarks, the model scores 0.99/1 on AIME 2025 with tools, ranking 10th, and is evaluated on the 55-language WMT24++ multilingual translation suite, giving the model a verifiable technical footprint beyond its average composite. Conversation-depth tracking shows quality holding near 14.2 through turns 11–30 before dipping to 13.0 at turn 31+, a useful signal for long-context agent use cases. These are aggregator figures sourced from build.nvidia.com, not original research.

Videos about Nemotron 3 Nano 30B A3B

More models around Nemotron 3 Nano 30B A3B