Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
OpenRouter logo

Model details

Nemotron 3.5 Lightning 30B A3B

Nemotron 3.5 Lightning 30B A3B is a hybrid mixture-of-experts model that pairs a Mamba-2 backbone with MoE and attention layers, totaling 30B parameters while activating only about 3B at inference. NVIDIA distilled it from the frontier Nemotron 3 Ultra, targeting the kind of long-running, multi-step workloads that autonomous agents run into daily, including coding, tool calling, instruction following, and extended multi-turn conversations. The architecture is purpose-built for sub-agent and always-on deployments where a smaller active footprint has to keep up with sustained throughput rather than a single deep reasoning pass.

Because the active parameter count is small relative to total capacity, the model is positioned as a high-throughput workhorse that can sit beneath an orchestrator model and handle specialized steps without becoming the bottleneck. FriendliAI reports up to roughly four times higher throughput on agent task completion compared with comparable models in its class, and offers Day-0 support on Dedicated Endpoints, signaling that production agent harnesses can integrate it immediately. With support for very long contexts and open weights under a permissive license, it is a practical fit for teams building agent systems that need a fast, customizable sub-agent rather than a general-purpose frontier chat model.

OpenRouternvidia/nemotron-3.5-lightningnemotron

Quick Info

Powered by
Provider
OpenRouter
Model key
nvidia/nemotron-3.5-lightning
Release date
Aug 11, 2026
Last updated
Aug 11, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.049
Output token cost
$0.14

Limits

Output tokens
131,072 tokens
Context window
262,144 tokens

Transparent token rates

Compare Nemotron 3.5 Lightning 30B A3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron 3.5 Lightning 30B A3B

OpenRouter

Coverage

An independent technical guide published on August 27, 2026 by AI-Driven Lab on note.com provides a structured walkthrough of Nemotron 3.5 Lightning's architecture, benchmarks, licensing, and enterprise adoption considerations. The article targets readers building LLM APIs or in-house agent infrastructure and positions The page is auto-translated from Japanese with an accuracy caveat displayed on the article itself, so specific technical claims about architecture, benchmarks and license terms should be cross-checked against NVIDIA primary sources rather than relied on at face value. Despite that caveat, the guide adds independent thi

OpenRouter

Coverage

Tech Insider's August 27, 2026 piece confirms that Nemotron 3.5 Lightning was released on August 11, 2026 with open weights, open training data, and a permissive commercial license (cited as NVIDIA's OpenMDW-style terms via the NVIDIA developer blog). It restates the 30-billion-parameter MoE with 3 billion active per t The article frames the release as oriented toward NVIDIA's hardware ecosystem, from GeForce RTX desktop cards to DGX Spark systems based on the company's GB10 chip, and ties the timing to NVIDIA's fiscal second-quarter data-center revenue of $89 billion (cited via CNBC). It positions the model for "always-on agents" ru

OpenRouter

CoverageDiscourse

A Hugging Face discussion thread on NVIDIA's official `nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16` repository (opened about 22 days before retrieval) adds Terminal-Bench 2.1 evaluation results to the BF16 checkpoint's eval-results directory. The proposed YAML change records a Terminal-Bench 2.1 score of 24.58, d The evaluation entry is attributed to NVIDIA's NeMo Evaluator SDK, described as the consistent in-house harness used across NVIDIA's release suite, with the source URL pointing back to the model's Hugging Face card. This is a first-party NVIDIA artifact, but it is a narrow data point covering only Terminal-Bench 2.1 an

OpenRouter

Coverage

NVIDIA's Nemotron 3.5 Lightning 30B A3B was released on August 11, 2026 as the throughput-oriented entry in the Nemotron 3.5 line, according to the Atomic Chat model page reproducing NVIDIA's BF16 model card. It is a sparse mixture-of-experts model with 30 billion total parameters and 3 billion active per token (30B-A3 The same page reports the model ships with DSpark, MTP and DFlash speculative-decoding drafter heads published alongside the checkpoint, so it can be served with speculative decoding without sourcing a separate draft model. A benchmark table drawn from the BF16 model card compares it against Qwen3.6-35B-A3B, Gemma 4 26

OpenRouter

CoverageBenchmark

A thread on the official NVIDIA developer forums reports a benchmark for the exact Nemotron 3.5 Lightning 30B A3B variant "nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4" (posted Aug 12, 2026), achieving 116.84 tokens/sec on text generation when served with vLLM on an NVIDIA DGX Spark (GB10) system, with a link to Because the reported numbers are tied to the NVFP4-quantized artifact running on DGX Spark with vLLM, the result is a concrete, hardware-specific throughput signal for the 3.5 Lightning 30B A3B family — useful for developers evaluating single-workstation inference performance for this variant, while leaving OpenRouter-

Videos about Nemotron 3.5 Lightning 30B A3B

More models around Nemotron 3.5 Lightning 30B A3B