Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Crusoe logo

Model details

Nemotron 3 Nano 30B A3B

Nemotron 3 Nano 30B A3B is a large language model trained from scratch by NVIDIA and presented as a unified system capable of handling both reasoning and non-reasoning tasks. When prompted, the model first produces an internal reasoning trace and then delivers a final answer, with the reasoning behavior togglable through a flag in the chat template. Disabling reasoning produces faster but slightly less accurate responses on harder problems, while leaving it on generally yields higher-quality solutions.

Under the hood, the model pairs a hybrid Mixture-of-Experts backbone with selective state-space components, alternating 23 Mamba-2 and MoE layers with 6 attention layers. Each MoE layer routes from 128 experts plus a shared expert, activating 6 per token, which keeps the active footprint at 3.5B parameters against a total of 30B. This design aims to deliver reasoning-grade quality while keeping inference efficient, making the model well suited for developers who want controllable thinking behavior, open-weight deployment, and a compact active parameter budget in a single artifact.

Crusoenvidia/NVIDIA-Nemotron-3-Nano-30B-A3Bnemotron

Quick Info

Powered by
Provider
Crusoe
Model key
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B
Release date
Dec 15, 2025
Last updated
Dec 15, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.05
Output token cost
$0.20

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Transparent token rates

Compare Nemotron 3 Nano 30B A3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron 3 Nano 30B A3B

No articles yet. Fetch the latest news to show it here.

Videos about Nemotron 3 Nano 30B A3B

More models around Nemotron 3 Nano 30B A3B