Sulat.com
AI models
Nebius Token Factory logo

Model details

Nemotron-3-Super-120B-A12B

Nemotron-3-Super-120B-A12B is a large open-weight language model designed for agentic, reasoning, and conversational tasks, including coding, planning, tool calling, and long-context analysis. The model uses a hybrid architecture that interleaves Mamba-2 layers with Mixture-of-Experts layers, and it adds Multi-Token Prediction (MTP) to accelerate generation. A distinctive design choice is the use of LatentMoE, where tokens are projected into a compressed latent space before expert routing, allowing roughly four times more experts to fit within the same parameter budget compared with conventional routing approaches. The model carries 120B total parameters while activating only 12B per forward pass, which is the central efficiency story behind the variant's name and its positioning for high-throughput serving. Multilingual coverage spans English, French, German, Italian, Japanese, Spanish, and Chinese.

Independent benchmarks published shortly after release evaluate the model for API latency and cost across inference providers, indicating that the active-parameter design translates into measurable inference-economics gains relative to dense 120B-class models. The architecture's hybrid Mamba-2 attention mix is well suited to long-context workloads, and the expanded context window supports sustained reasoning and planning tasks that chain multiple tool calls. These characteristics make the model a practical fit for production agents, research workflows that need long document analysis, and developer tooling that benefits from open-weight deployment and fine-tuning freedom. Compared with peers of similar total size, the combination of sparse activation, MTP, and LatentMoE offers a forward-looking path to scaling expert counts without proportional increases in compute per token.

Nebius Token Factorynvidia/nemotron-3-super-120b-a12bnemotron

Quick Info

Powered by
Provider
Nebius Token Factory
Model key
nvidia/nemotron-3-super-120b-a12b
Release date
Mar 11, 2026
Last updated
Mar 12, 2026
Knowledge cutoff
2026-02
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$0.90

Limits

Input tokens
256,000 tokens
Output tokens
32,768 tokens
Context window
256,000 tokens

Latest news about Nemotron-3-Super-120B-A12B

Nebius Token Factory

CoverageBenchmark

OpenRouter's aggregator listing provides Nebius-specific pricing data for NVIDIA Nemotron 3 Super 120B: Nebius Token Factory charges $0.30 per 1M input tokens and $0.90 per 1M output tokens, which is roughly 3–4× higher than competing providers such as DekaLLM ($0.08/$0.45) and DeepInfra ($0.085/$0.40). The listing con The page also notes that a free-tier endpoint exists where prompts and outputs are logged for model improvement, which is a relevant caveat for Nebius users evaluating the Token Factory free tier. Cross-provider metrics show Nebius Token Factory has no published latency or throughput on OpenRouter, unlike DekaLLM (2.11

Nebius Token Factory

CoverageBenchmark

Nemotron 3 Super is a 120B total / 12B active parameter hybrid Mamba-Attention Mixture-of-Experts model optimized for agentic reasoning, coding, planning, tool calling, and long-context analysis. It introduces LatentMoE (projecting tokens into a compressed latent space for expert routing, enabling 4x more experts at th

Videos about Nemotron-3-Super-120B-A12B

More models around Nemotron-3-Super-120B-A12B