Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Baseten logo

Model details

Nemotron Super

Nemotron Super is an open-weight hybrid mixture-of-experts model built around 120B total parameters with only 12B active per token, a design that aims to deliver frontier-style reasoning at a fraction of the inference cost of much larger dense models. Third-party reporting describes it as the first model in its family pre-trained in NVFP4, combined with a Mamba2-Transformer latent MoE architecture and Multi-Token Prediction, which together shape its balance of long-context throughput and analytical depth. Independent hands-on testing by Greptile highlighted its fluency with tool calling, repository navigation, and bug identification on multi-file refactors, positioning it as a strong fit for agentic code-review and multi-step software workflows rather than general chat.

The model is intended for developers who need open weights, a long context window, and reliable structured output, with the same hands-on evaluation noting a 1M token context for tracing imports and reasoning across large codebases. DeepInfra's cross-provider benchmark post frames it as a cost-efficient option for agentic workloads where latency and token economics matter as much as raw accuracy. Buyers evaluating it on Baseten's Model API should weigh that Baseten's own changelog lists Nemotron Super 120B among models deprecated on that endpoint, so deployment on alternative NVIDIA-aligned inference providers may be the more durable path for production agentic systems.

Basetennvidia/Nemotron-120B-A12Bnemotron

Quick Info

Powered by
Provider
Baseten
Model key
nvidia/Nemotron-120B-A12B
Release date
Mar 11, 2026
Last updated
Mar 11, 2026
Knowledge cutoff
2026-02
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$0.75

Limits

Output tokens
202,800 tokens
Context window
202,800 tokens

Latest news about Nemotron Super

Baseten

CoverageBenchmark

DeepInfra's April 3, 2026 benchmark post corroborates the technical identity of Nemotron 3 Super 120B A12B as an NVIDIA open-weight model released March 11, 2026, with a hybrid Mamba2-Transformer LatentMoE architecture, Multi-Token Prediction, NVFP4 pre-training, and up to a 1 million token context window. It also docu The post frames Nemotron 3 Super 120B as the first NVFP4-pretrained model in the Nemotron 3 family and benchmarks latency and cost across multiple inference providers, giving developers a cross-provider comparison view when evaluating Baseten versus DeepInfra hosting for agentic workloads. The article also references D

Baseten

Official sourceRelease Notes

Baseten's official changelog records a "Model API Deprecation" entry dated July 8, 2026 that explicitly lists "Nemotron Super 120B" alongside GLM 5.1, GLM 5, and Kimi K2.5. This is a first-party lifecycle event affecting the exact modelKey (nvidia/Nemotron-120B-A12B) referenced by the subject, signaling that the Basete The same changelog also surfaces adjacent product activity relevant to Baseten customers, including Inkling becoming available on Baseten Model APIs on July 15, 2026, an admin visibility feature for personal API keys on July 7, and a Series F announcement banner. Developers currently routing traffic to Nemotron Super 1

Videos about Nemotron Super

More models around Nemotron Super