Nemotron 3.5 Lightning 30B A3B is built around a hybrid architecture that combines Mamba-2 with mixture-of-experts and attention layers, totaling 30 billion parameters of which roughly 3 billion activate per inference. This design choice pairs the long-range sequence modeling characteristics of Mamba-2 with the routing flexibility of MoE and the contextual precision of attention, allowing efficient handling of extended tasks without scaling full compute for every token. NVIDIA frames the model as a workhorse targeted at long-running autonomous agents and sub-agent deployments, where many tokens are spent on planning, tool use, and multi-step reasoning rather than single-shot answers.
Open weights and native function calling make the model suitable for self-hosting and tool-augmented pipelines, while built-in reasoning support lets developers opt into chain-of-thought behavior when needed. Its advertised context envelope reaches up to one million tokens in NVIDIA's reference configuration, well beyond typical chat workloads, enabling agents to operate over very long documents, multi-turn histories, or accumulated tool outputs. The combination of a lean active parameter count, agent-focused design intent, and open-weight availability positions the model as a pragmatic choice for teams building dependable autonomous workflows that need controllable reasoning and reliable tool integration.