Nemotron 3.5 Lightning is an open-weights mixture-of-experts model built around a hybrid Mamba-2 and attention architecture, with roughly thirty billion total parameters but only a few billion activated per token. This sparsity keeps inference light while still giving the model enough capacity to handle the repetitive, high-throughput steps that show up inside AI agent pipelines, and it ships alongside NeMo Switchyard, an open-source router that automatically steers each step of a workflow to the most appropriate model based on quality, cost, or speed requirements. By deliberately trading peak intelligence for efficiency, the design targets the cost-per-agent-task problem rather than the general intelligence leaderboard.
In practical terms, the model is aimed at developers running large volumes of agentic tasks who need predictable, low-cost execution, with NVIDIA reporting up to four times faster output and around thirty percent faster agentic task completion compared to similarly sized open models, though its own numbers place it behind larger contemporaries such as Qwen3.6 35B and Nemotron 3 Super on broad intelligence benchmarks. The very large context window, reported at up to the cataloged API limit, supports long-running agent sessions and tool-augmented reasoning, making the model a sensible fit when Switchyard routes simpler steps to it and reserves heavier frontier models for the genuinely difficult calls.