Currently listed through these providers:
Model details
Nemotron 3 Nano 30B A3B
Nemotron 3 Nano 30B A3B is a large language model trained from scratch by NVIDIA and presented as a unified system capable of handling both reasoning and non-reasoning tasks. When prompted, the model first produces an internal reasoning trace and then delivers a final answer, with the reasoning behavior togglable through a flag in the chat template. Disabling reasoning produces faster but slightly less accurate responses on harder problems, while leaving it on generally yields higher-quality solutions.
Under the hood, the model pairs a hybrid Mixture-of-Experts backbone with selective state-space components, alternating 23 Mamba-2 and MoE layers with 6 attention layers. Each MoE layer routes from 128 experts plus a shared expert, activating 6 per token, which keeps the active footprint at 3.5B parameters against a total of 30B. This design aims to deliver reasoning-grade quality while keeping inference efficient, making the model well suited for developers who want controllable thinking behavior, open-weight deployment, and a compact active parameter budget in a single artifact.
Quick Info
Powered by- Provider
- Crusoe
- Model key
- nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B
- Release date
- Dec 15, 2025
- Last updated
- Dec 15, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.05
- Output token cost
- $0.20
Limits
- Output tokens
- 262,144 tokens
- Context window
- 262,144 tokens
Transparent token rates
Compare Nemotron 3 Nano 30B A3B pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Nemotron 3 Nano 30B A3B
No articles yet. Fetch the latest news to show it here.
Videos about Nemotron 3 Nano 30B A3B
More models around Nemotron 3 Nano 30B A3B
This exact model name is also listed by 6 other providers.