Currently listed through these providers:
Model details
Nemotron 3 Super 120B A12B
Nemotron 3 Super 120B A12B belongs to NVIDIA's third-generation Nemotron family of large language models, a line that emphasizes sparse expert architectures for efficient inference at very large parameter counts. The naming pattern itself signals its design philosophy: a 120-billion-parameter backbone with only a fraction of parameters active per token, allowing the model to scale capability without scaling compute proportionally. It sits within a broader NVIDIA effort to ship production-grade open-weight models that researchers and enterprises can fine-tune, distill, and deploy on their own hardware stacks.
Practical interest in the model is already visible beyond NVIDIA's own channels. A dedicated developer-forum thread on the DGX Spark and GB10 platform focuses on fine-tuning an NVFP4-quantized variant of the model to add new programming language support, suggesting a community workflow around adapting the weights for domain-specific coding tasks. Separately, an independent inference provider has published latency and cost benchmarks for the API form of the model roughly three weeks after release, indicating early infrastructure-level characterization that can guide deployment decisions. Together, these signals point to a model intended for teams that want a capable, sparsely activated open-weight foundation they can customize and serve on accelerated hardware.
Quick Info
Powered by- Provider
- Eden AI
- Model key
- flexai/Nemotron-3-Super-120B-A12B
- Release date
- Mar 11, 2026
- Last updated
- Mar 11, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.085
- Output token cost
- $0.40
Limits
- Output tokens
- 262,144 tokens
- Context window
- 262,144 tokens
Latest news about Nemotron 3 Super 120B A12B
Videos about Nemotron 3 Super 120B A12B
More models around Nemotron 3 Super 120B A12B
This exact model name is also listed by 8 other providers.