Nemotron 3.5 Lightning 30B A3B is a hybrid architecture that combines Mamba-2 state-space layers, Mixture-of-Experts routing, and attention into a single model with 30 billion total parameters and roughly 3 billion active per token. The active-parameter footprint keeps inference economical while the MoE capacity supports broad knowledge, and NVIDIA designed the model specifically to act as a long-running workhorse for autonomous agents, sub-agents, and other multi-step tool-using workflows. Open weights ship under the OpenMDW License Agreement v1.1, so the same checkpoint can be self-hosted for private agent stacks or served through a managed gateway, and the release is positioned for deployment scenarios where sustained reasoning and tool calling are the main workload.
Training leaned on a large pretraining corpus of more than 20 trillion tokens, with an NVFP4 recipe and Multi-Token Prediction applied to keep generation throughput high. The model handles up to one million tokens of context, which is unusually long and suits agentic traces, retrieval-augmented reasoning, and codebase-scale work that smaller-context models struggle to hold in one pass. NVIDIA recommends a sampling configuration beginning at Temperature 1.0 for stable behavior, and the supported language set covers English and common programming languages alongside Spanish, French, German, Italian, and Japanese. Together those traits make the model a practical fit for production agent pipelines that need long context, fast per-token latency, and the ability to call tools or emit structured responses without hosting a much larger dense model.