Nemotron 3.5 Lightning is NVIDIA's specialized execution model for always-on AI agents, designed to handle the high-volume steps such as tool calls, result validation, and subagent delegation that dominate long-running workloads. It is a 30B parameter mixture-of-experts model with roughly 3B active parameters per token, giving it the capacity of a larger dense model at the compute cost of a small one. The architecture is a hybrid Mamba-Transformer design carried forward from its predecessor, NVIDIA Nemotron 3 Nano, and the model supports a context length of up to one million tokens, leaving ample room for long tool histories across multi-turn workflows.
Speed and customization are central to the model's design. A dedicated pretraining stage baked multi-token prediction into the weights, and the release ships with DFlash and DSpark draft models so speculative decoding can be tuned across serving scenarios, from DGX Spark up to data-center concurrency. An NVFP4 quantized checkpoint is included alongside the BF16 weights, running on the same specialized kernels that power the rest of the Nemotron family across Blackwell, Hopper, and Ampere GPUs. Open weights, training data, and recipes are released permissively under OpenMDW-1.1, enabling LoRA or full supervised fine-tuning and reinforcement learning on modest hardware. The model is well suited to coding sub-agents, local personal assistants, security workflows, and any tier where a lightweight, fast specialist should run beside a heavier frontier planner.