Nemotron Super is an open-weight hybrid mixture-of-experts model built around 120B total parameters with only 12B active per token, a design that aims to deliver frontier-style reasoning at a fraction of the inference cost of much larger dense models. Third-party reporting describes it as the first model in its family pre-trained in NVFP4, combined with a Mamba2-Transformer latent MoE architecture and Multi-Token Prediction, which together shape its balance of long-context throughput and analytical depth. Independent hands-on testing by Greptile highlighted its fluency with tool calling, repository navigation, and bug identification on multi-file refactors, positioning it as a strong fit for agentic code-review and multi-step software workflows rather than general chat.
The model is intended for developers who need open weights, a long context window, and reliable structured output, with the same hands-on evaluation noting a 1M token context for tracing imports and reasoning across large codebases. DeepInfra's cross-provider benchmark post frames it as a cost-efficient option for agentic workloads where latency and token economics matter as much as raw accuracy. Buyers evaluating it on Baseten's Model API should weigh that Baseten's own changelog lists Nemotron Super 120B among models deprecated on that endpoint, so deployment on alternative NVIDIA-aligned inference providers may be the more durable path for production agentic systems.