Currently listed through these providers:
Model details
Nemotron 3 Nano 30B A3B
Nemotron 3 Nano 30B A3B is a large language model trained from scratch by NVIDIA and designed as a unified system for both reasoning and non-reasoning workloads. Rather than splitting these behaviors across separate models, it produces a reasoning trace by default and then concludes with its final answer, giving developers a single model that can handle chat, analysis, and step-by-step problem solving. A flag in the chat template lets callers switch the reasoning stage on or off, which trades a small amount of accuracy on harder prompts for shorter, more direct outputs when intermediate thinking is not needed.
Underneath that unified interface sits a hybrid architecture that mixes Mamba-2 sequence layers with mixture-of-experts layers alongside a small number of attention layers, letting a relatively compact active footprint deliver the quality of a much larger total parameter budget. This design points to practical use cases where a single deployable model must cover straightforward instruction following as well as multi-step reasoning, math, and tool-assisted workflows. The release was significant enough for Deep Infra to publish a dedicated latency and cost benchmark write-up, suggesting the model is positioned as a competitive option for hosted inference at scale.
Quick Info
Powered by- Provider
- Deep Infra
- Model key
- nvidia/Nemotron-3-Nano-30B-A3B
- Release date
- Dec 15, 2025
- Last updated
- Dec 15, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.05
- Output token cost
- $0.20
Limits
- Output tokens
- 262,144 tokens
- Context window
- 262,144 tokens
Latest news about Nemotron 3 Nano 30B A3B
Videos about Nemotron 3 Nano 30B A3B
Recent tweets and retweets from Deep Infra
More models around Nemotron 3 Nano 30B A3B
This exact model name is also listed by 6 other providers.