Nemotron 3.5 Lightning 30B A3B is a hybrid architecture that combines Mamba-2 state-space layers with mixture-of-experts routing and selective attention, keeping the active footprint at roughly 3B parameters while the full model spans 30B. Pre-training was conducted on more than 20 trillion tokens under an NVFP4 recipe augmented with Multi-Token Prediction, a combination aimed at accelerating inference without sacrificing throughput quality. Open weights are released under the OpenMDW License Agreement v1.1, and the model is positioned within the broader NVIDIA Nemotron family as a lightweight tier optimized for speed.
The headline practical advantage is an extended context window reaching up to one million tokens, which makes the model suitable for long-running autonomous agents, sub-agent orchestration, and other agentic pipelines that need to retain large working memories. It supports reasoning and tool calling, with a capability set oriented around text-only inputs and outputs, and it covers English along with major European languages and Japanese for multilingual deployments. Its efficiency-first design, low active parameter count, and multi-token prediction training make it a fit when teams need a responsive open-weights model for sustained agent workloads rather than a heavyweight general-purpose chat model.