Sulat.com
AI models
OpenRouter logo

Model details

Nemotron 3 Super 120B A12B

Nemotron 3 Super 120B A12B is NVIDIA's open-weight large language model released on March 11, 2026, with an NVFP4-quantized variant published on the NVIDIA developer forums for the DGX Spark / GB10 community. The naming scheme (120B-A12B) signals a hybrid architecture where a large total parameter count is paired with a much smaller set of active parameters per token, a MoE-style efficiency pattern that lets the model leverage broad capacity without paying full inference cost on every forward pass. Its open-weight availability makes it suitable for self-hosted deployments, fine-tuning experiments, and on-device research on capable NVIDIA hardware.

From a practical standpoint, the model is positioned for long-context reasoning and agent-style tasks, combining temperature control with structured output and tool calling so it can plug into retrieval pipelines, coding assistants, and multi-step workflows. The hybrid active-parameter design aims to balance quality against latency, while open weights let teams adapt behavior to domain-specific needs rather than rely solely on a hosted endpoint. It fits teams that want a frontier-class reasoning model with the freedom to run, inspect, and customize it on their own infrastructure.

OpenRouternvidia/nemotron-3-super-120b-a12bnemotron

Quick Info

Powered by
Provider
OpenRouter
Model key
nvidia/nemotron-3-super-120b-a12b
Release date
Mar 11, 2026
Last updated
Mar 11, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.085
Output token cost
$0.40

Limits

Output tokens
16,384 tokens
Context window
1,000,000 tokens

Latest news about Nemotron 3 Super 120B A12B

OpenRouter

Official sourceComparison

OpenRouter's comparison between MiniMax M3 and Nemotron 3 Super 120B A12B shows both models sitting in the ~1M-token context tier, with MiniMax M3 offering 1,048,576 tokens at $0.23 input and $0.96 output per million tokens versus Nemotron 3 Super's 1,000,000-token window at $0.085 input and $0.40 output. Either can be The page reinforces that Nemotron 3 Super is among the cheapest 1M-context listings on OpenRouter, undercutting MiniMax M3 by roughly 2.7x on input and 2.4x on output pricing while matching context scale. As with the other auto-generated compare page, no benchmark or quality metrics are present in the excerpt, so the v

OpenRouter

Official sourceComparison

OpenRouter's head-to-head comparison of Cohere's Command A and NVIDIA's Nemotron 3 Super 120B A12B positions the two models for developers choosing a 1M-context-capable alternative at very different price points. Command A offers a 256,000-token context window at $2.50 per million input tokens and $10 per million outpu The comparison emphasizes Nemotron 3 Super's roughly 29x lower input cost and 25x lower output cost against Command A, alongside a 4x larger effective context window for long-document workloads. There are no benchmark or quality scores in the supplied excerpt, so the page is most useful as a pricing and context sizing

OpenRouter

Official sourceBenchmark

OpenRouter lists NVIDIA's Nemotron 3 Super 120B A12B as an open-weight hybrid Mamba-Transformer Mixture-of-Experts model with 120B total parameters activating 12B per pass, supporting multi-agent and long-context workloads. The hosted page documents a 1M-token context window on the paid variant (262K on the :free endpo For developers routing through OpenRouter, the :free endpoint sits behind NVIDIA's free-tier data logging terms (consented on use, with logged session data used for improvement but not linked to identity), and the listed provider telemetry is P50 latency 1.35s with 58 tokens/second throughput and 99.66% uptime. The mod

Videos about Nemotron 3 Super 120B A12B

More models around Nemotron 3 Super 120B A12B