Currently listed through these providers:
Model details
Nemotron 3 Super 120B A12B
NVIDIA Nemotron 3 Super 120B A12B is an open-weight large language model built around a hybrid Latent MoE design that interleaves Mamba-2 and MoE layers, with Multi-Token Prediction (MTP) used to accelerate generation. The model totals 120B parameters while activating only 12B per forward pass, an arrangement that lets it carry the capacity of a much larger system while keeping inference cost closer to a mid-sized model. It is targeted at agentic, reasoning, and conversational workloads, and it is offered in an FP8 quantization variant for efficient serving.
In practical deployment, the model is positioned for long-context and multilingual use, with stated support for English, French, German, Italian, Japanese, Spanish, and Chinese. It is available through both on-demand and serverless API hosting, including fine-tuning-disabled base-model endpoints and dedicated GPU deployments without rate limits, making it a practical fit for teams that want NVIDIA-grade reasoning behavior without the overhead of training and operating their own mixture-of-experts stack. The combination of MoE sparsity, hybrid attention/SSM layers, and FP8 serving makes it well suited for cost-sensitive agentic pipelines and multilingual assistants where responsiveness matters as much as raw capability.
Quick Info
Powered by- Provider
- Pioneer
- Model key
- nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8
- Release date
- Mar 11, 2026
- Last updated
- Mar 11, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.09
- Output token cost
- $0.45
Limits
- Output tokens
- 32,000 tokens
- Context window
- 256,000 tokens
Latest news about Nemotron 3 Super 120B A12B
Videos about Nemotron 3 Super 120B A12B
More models around Nemotron 3 Super 120B A12B
This exact model name is also listed by 8 other providers.