Currently listed through these providers:
Model details
Nemotron 3 Nano 30B A3B
Nemotron 3 Nano 30B A3B is a sparse hybrid Mamba-Transformer mixture-of-experts model that pairs 30B total parameters with only about 3B active per token. This design aims to deliver the reasoning capacity of a larger model while behaving in throughput terms much closer to a small dense one, which is useful for teams that want stronger reasoning without paying the latency or cost of a full 30B dense run.
The model is offered with an open-weights license and is positioned for agentic and reasoning workloads, with support for reasoning controls and tool calling alongside very long-context inputs. Practically, it is a good fit for applications that need a balance of quality and efficiency, such as chat assistants, retrieval over large documents, and tool-using agents where the active-parameter footprint matters more than raw parameter count.
Quick Info
Powered by- Provider
- Pioneer
- Model key
- nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- Release date
- Dec 15, 2025
- Last updated
- Dec 15, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.05
- Output token cost
- $0.20
Limits
- Output tokens
- 131,072 tokens
- Context window
- 262,144 tokens
Latest news about Nemotron 3 Nano 30B A3B
Videos about Nemotron 3 Nano 30B A3B
More models around Nemotron 3 Nano 30B A3B
This exact model name is also listed by 6 other providers.