Nemotron 3.5 Lightning 30B A3B is a hybrid mixture-of-experts model that pairs a Mamba-2 backbone with MoE and attention layers, totaling 30B parameters while activating only about 3B at inference. NVIDIA distilled it from the frontier Nemotron 3 Ultra, targeting the kind of long-running, multi-step workloads that autonomous agents run into daily, including coding, tool calling, instruction following, and extended multi-turn conversations. The architecture is purpose-built for sub-agent and always-on deployments where a smaller active footprint has to keep up with sustained throughput rather than a single deep reasoning pass.
Because the active parameter count is small relative to total capacity, the model is positioned as a high-throughput workhorse that can sit beneath an orchestrator model and handle specialized steps without becoming the bottleneck. FriendliAI reports up to roughly four times higher throughput on agent task completion compared with comparable models in its class, and offers Day-0 support on Dedicated Endpoints, signaling that production agent harnesses can integrate it immediately. With support for very long contexts and open weights under a permissive license, it is a practical fit for teams building agent systems that need a fast, customizable sub-agent rather than a general-purpose frontier chat model.