Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

NVIDIA Nemotron 3 Super 120B A12B

NVIDIA Nemotron 3 Super is a hybrid Mixture-of-Experts model that combines Mamba and Transformer architectures with multi-token prediction, activating just 12 billion parameters from its 120 billion total to keep inference efficient without sacrificing accuracy. Its latent MoE design calls 4 experts for the cost of one, delivering over 50% higher token generation than leading open models. Built for agent-based AI inference, it handles up to a 1 million token context window to support long-horizon multi-step task planning, cross-document reasoning, and multi-agent collaboration on a single GPU.

The model was trained with multi-environment reinforcement learning across 10 or more environments, achieving leading accuracy on benchmarks including AIME 2025, TerminalBench, and SWE-Bench Verified. Weights, datasets, and training recipes are released under the NVIDIA Open License, making it easy to customize and deploy anywhere from workstation to cloud. It excels at reasoning, tool use, and instruction following in complex multi-agent applications, with native support for Japanese and optimized for running many collaborating agents per application.

Vercel AI Gatewaynvidia/nemotron-3-super-120b-a12bnemotron

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
nvidia/nemotron-3-super-120b-a12b
Release date
Mar 11, 2026
Last updated
Mar 11, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.65

Limits

Output tokens
32,000 tokens
Context window
256,000 tokens

Transparent token rates

Compare NVIDIA Nemotron 3 Super 120B A12B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about NVIDIA Nemotron 3 Super 120B A12B

Vercel AI Gateway

Coverage

NVIDIA has announced Nemotron 3 Super , a new open model specialized for agent-based AI inference. Nemotron 3 Super employs the MoE architecture, which has 120 billion (120B) total parameters and 12 billion (12B) effective parameters, achieving both high computational efficiency and accuracy. Introducing Nemotron 3 Sup

Videos about NVIDIA Nemotron 3 Super 120B A12B

More models around NVIDIA Nemotron 3 Super 120B A12B