Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Merge Gateway logo

Model details

Nemotron 3 Nano 30B A3B

The model overview is temporarily unavailable.

Merge Gatewaynvidia/nemotron-3-nano-30b-a3bnemotron

Quick Info

Powered by
Provider
Merge Gateway
Model key
nvidia/nemotron-3-nano-30b-a3b
Release date
Dec 15, 2025
Last updated
Dec 15, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.05
Output token cost
$0.20

Limits

Output tokens
32,000 tokens
Context window
256,000 tokens

Transparent token rates

Compare Nemotron 3 Nano 30B A3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron 3 Nano 30B A3B

OpenRouter

Coverage

HokAI's model hub entry explicitly names the NVIDIA Nemotron 3 Nano 30B-A3B and describes it as an open-weight language model built by NVIDIA, released December 14, 2025 as the entry tier of the Nemotron 3 family. It reports benchmark scores from NVIDIA's technical report (arXiv 2512.20848): 73.04% on GPQA, 68.25% on L For tool-use workloads, HokAI reports that giving Nemotron 3 Nano tools lifts AIME accuracy from 89.06% to 99.17%, positioning it as a strong fit for agentic tool-calling pipelines. The page also notes an Artificial Analysis score of 38.8% on SWE-bench Verified. The license is listed as the NVIDIA Open Model License wi

Kilo Gateway

Official sourceOfficial

NVIDIA announced the Nemotron 3 family of open models, releasing the Nano variant alongside its technical report. The family uses a hybrid Mamba-Transformer mixture-of-experts architecture designed for high-throughput agentic inference. Nemotron 3 Nano is a 3.2B active parameter (31.6B total) model that NVIDIA reports as more accurate than GPT-OSS-20B and Qwen3-30B-A3B-Thinking-2507 on popular benchmarks across categories. Nemotron 3 Nano supports context lengths up to 1M tokens and on a single H200 GPU at 8K input / 16K output it delivers 3.3x higher inference throughput than Qwen3-30B-A3B and 2.2x higher than GPT-OSS-20B. Super and Ultra variants, which add LatentMoE, Multi-Token Prediction layers, and NVFP4 training, are planned for later release. The model is positioned for cost-efficient deployment of agentic, reasoning, and conversational workloads.

Videos about Nemotron 3 Nano 30B A3B

More models around Nemotron 3 Nano 30B A3B