Nemotron 3.5 Lightning 30B A3B is an open-weight large language model developed by NVIDIA, released on August 11, 2026 alongside openly published training data and a permissive commercial license. Its architecture is a hybrid mixture-of-experts design that combines Mamba-2 state-space layers, MoE routing, and attention blocks, totaling 30 billion parameters with 3 billion active per token. This sparse-active configuration is intended to deliver capable reasoning at a fraction of the compute cost of dense models of similar size, while still benefiting from attention's strengths on localized context. The model is positioned for long-running autonomous agents and sub-agent workhorse deployments, making the architecture choice especially relevant for sustained multi-step workflows.
For practical fit, the model supports a context length of up to one million tokens and is multilingual, covering English and coding languages alongside Spanish, French, German, Italian, and Japanese. NVIDIA designed it to run efficiently across its own hardware stack, from GeForce RTX desktop cards up to DGX Spark systems based on the GB10 chip, giving developers flexibility across consumer and data center tiers. The open weights and training data allow teams to fine-tune, audit, and integrate the model into production pipelines without licensing friction, which is valuable for organizations building specialized agentic systems that need long-context understanding and strong multilingual coverage.