Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Amazon Bedrock logo

Model details

NVIDIA Nemotron 3 Super 120B A12B

Nemotron 3 Super 120B A12B is a 120-billion-parameter hybrid Mamba-Transformer mixture-of-experts model that activates only about 12 billion parameters per token, a design that targets compute efficiency without giving up accuracy on complex work. It belongs to the openly distributed NVIDIA Nemotron family, with weights published on Hugging Face, so teams can audit behavior, fine-tune for their own domains, or self-host when preferred. The model also incorporates multi-token prediction, a training technique that lets it generate several tokens per forward pass and lifts throughput on generation-heavy agent loops.

The practical sweet spot is long-horizon agentic work: a reported one-million-token context window supports cross-document reasoning and multi-step planning that would overflow shorter windows, and benchmark highlights such as AIME 2025, TerminalBench, and SWE-Bench Verified indicate strength on math problem solving, terminal-style tool use, and verified software engineering tasks. Native reasoning and tool calling, together with JSON-schema structured output, make it easy to slot into multi-agent pipelines that need parseable responses and external actions. Teams building research assistants, code agents, or planning systems that have to stay coherent across very large inputs are the clearest fit, especially when open weights and predictable pricing matter for procurement.

Amazon Bedrocknvidia.nemotron-super-3-120bnemotron

Quick Info

Powered by
Provider
Amazon Bedrock
Model key
nvidia.nemotron-super-3-120b
Release date
Mar 11, 2026
Last updated
Mar 11, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.65

Limits

Output tokens
131,072 tokens
Context window
262,144 tokens

Transparent token rates

Compare NVIDIA Nemotron 3 Super 120B A12B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about NVIDIA Nemotron 3 Super 120B A12B

Amazon Bedrock

Coverage

DeepInfra published a comparison of hosted API providers for NVIDIA Nemotron 3 Super 120B on 2026-05-25, explicitly naming Amazon Bedrock as one of the available deployment targets and positioning it for "AWS enterprise integration." The article describes the model as 120B total parameters with 12B active per inference The DeepInfra piece is a competitor-published guide, so its characterization of Bedrock should be read as a third-party snapshot rather than an AWS primary source, and it does not include Bedrock-specific pricing, regional availability, or GA status for Nemotron 3 Super 120B. Its practical value for engineers is the si

Amazon Bedrock

CoverageBenchmark

NVIDIA released Nemotron 3 Super 120B A12B in March 2026 as an open-weights reasoning model, according to Artificial Analysis. It uses 120.6B total parameters with 12.7B active per token and supports a 1M token context window for long-context reasoning. Artificial Analysis rates the model highly, placing it 13th on the Intelligence Index with 163.9 output tokens per second. Pricing is $0.30 per 1M input and $0.90 per 1M output tokens, with a 50% cache discount, though it runs verbose at 170M tokens per evaluation task.

Videos about NVIDIA Nemotron 3 Super 120B A12B

More models around NVIDIA Nemotron 3 Super 120B A12B