NVIDIA Nemotron 3 Super 120B is a large hybrid-architecture language model from NVIDIA's Nemotron family, released as a public preview with openly published weights. The base variant documented on NVIDIA's Hugging Face repository uses a hybrid Latent Mixture-of-Experts design that interleaves Mamba-2 layers with MoE layers, with the repository identifier indicating BF16 precision weights. The model was trained from scratch by NVIDIA using next-token prediction and is positioned as a starting platform for further post-training such as instruction following and coding specialization, while the family maintains an official NVIDIA developer page and a published Nemotron technical report for further reference.
With open weights available and a parameter footprint of 120B (active around 12B at inference thanks to the MoE routing), the model is aimed at developers who want a self-hostable or cloud-deployed foundation model capable of reasoning and tool calling. Its training window spans the second half of 2025 through early 2026, with pre-training data current to December 2025 and post-training data through February 2026, so it reflects recent knowledge for practical assistant and agent workloads. Teams looking to combine a capable open model with a managed inference backend will find it fits scenarios where custom fine-tuning, long context, and structured outputs matter, while keeping the ability to host weights themselves on infrastructure they control.