Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
OpenRouter logo

Model details

Nemotron 3.5 Lightning (free)

Nemotron 3.5 Lightning is an open-weights mixture-of-experts model built around a hybrid Mamba-2 and attention architecture, with roughly thirty billion total parameters but only a few billion activated per token. This sparsity keeps inference light while still giving the model enough capacity to handle the repetitive, high-throughput steps that show up inside AI agent pipelines, and it ships alongside NeMo Switchyard, an open-source router that automatically steers each step of a workflow to the most appropriate model based on quality, cost, or speed requirements. By deliberately trading peak intelligence for efficiency, the design targets the cost-per-agent-task problem rather than the general intelligence leaderboard.

In practical terms, the model is aimed at developers running large volumes of agentic tasks who need predictable, low-cost execution, with NVIDIA reporting up to four times faster output and around thirty percent faster agentic task completion compared to similarly sized open models, though its own numbers place it behind larger contemporaries such as Qwen3.6 35B and Nemotron 3 Super on broad intelligence benchmarks. The very large context window, reported at up to the cataloged API limit, supports long-running agent sessions and tool-augmented reasoning, making the model a sensible fit when Switchyard routes simpler steps to it and reserves heavier frontier models for the genuinely difficult calls.

OpenRouternvidia/nemotron-3.5-lightning:freenemotron

Quick Info

Powered by
Provider
OpenRouter
Model key
nvidia/nemotron-3.5-lightning:free
Release date
Aug 11, 2026
Last updated
Aug 11, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
65,536 tokens
Context window
1,000,000 tokens

Latest news about Nemotron 3.5 Lightning (free)

OpenRouter

CoverageBenchmark

NVIDIA's published benchmark scores for the Nemotron 3.5 Lightning BF16 checkpoint place the model at SWE-bench Verified 51.56, GPQA Diamond 75.44, MMLU Pro 81.94, and PinchBench 85.37, figures that describe a model tuned for agent work and coding rather than general chat. A 51.56 on SWE-bench Verified means the model The guide contextualizes the numbers against NVIDIA's speed and efficiency claims versus gpt-oss-120b and Qwen3.6-35B, and reminds readers that the 30B-total, 3B-active MoE design lets the model behave like a small network at inference while drawing on a large pool of experts. It also flags the August 11, 2026 open-wei

OpenRouter

Coverage

NVIDIA released Nemotron 3.5 Lightning on August 11, 2026 as an open-weight mixture-of-experts model with 30 billion total parameters and 3 billion active per token, positioning it for "always-on agents" that run narrow, repetitive tasks continuously rather than single chat sessions. According to NVIDIA's developer blo The model supports context windows up to 1 million tokens, large enough to hold a full codebase or a stack of legal contracts in a single agent run, and was designed to run on NVIDIA hardware ranging from GeForce RTX desktop cards up to DGX Spark systems built around the GB10 chip. The August 11 release landed two week

Videos about Nemotron 3.5 Lightning (free)

More models around Nemotron 3.5 Lightning (free)