Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Kilo Gateway logo

Model details

Nemotron 3.5 Lightning 30B A3B

Nemotron 3.5 Lightning 30B A3B is a hybrid architecture that combines Mamba-2 state-space layers, Mixture-of-Experts routing, and attention into a single model with 30 billion total parameters and roughly 3 billion active per token. The active-parameter footprint keeps inference economical while the MoE capacity supports broad knowledge, and NVIDIA designed the model specifically to act as a long-running workhorse for autonomous agents, sub-agents, and other multi-step tool-using workflows. Open weights ship under the OpenMDW License Agreement v1.1, so the same checkpoint can be self-hosted for private agent stacks or served through a managed gateway, and the release is positioned for deployment scenarios where sustained reasoning and tool calling are the main workload.

Training leaned on a large pretraining corpus of more than 20 trillion tokens, with an NVFP4 recipe and Multi-Token Prediction applied to keep generation throughput high. The model handles up to one million tokens of context, which is unusually long and suits agentic traces, retrieval-augmented reasoning, and codebase-scale work that smaller-context models struggle to hold in one pass. NVIDIA recommends a sampling configuration beginning at Temperature 1.0 for stable behavior, and the supported language set covers English and common programming languages alongside Spanish, French, German, Italian, and Japanese. Together those traits make the model a practical fit for production agent pipelines that need long context, fast per-token latency, and the ability to call tools or emit structured responses without hosting a much larger dense model.

Kilo Gatewaynvidia/nemotron-3.5-lightningnemotron

Quick Info

Powered by
Provider
Kilo Gateway
Model key
nvidia/nemotron-3.5-lightning
Release date
Aug 11, 2026
Last updated
Aug 11, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.039
Output token cost
$0.18

Limits

Output tokens
131,072 tokens
Context window
262,144 tokens

Transparent token rates

Compare Nemotron 3.5 Lightning 30B A3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron 3.5 Lightning 30B A3B

Kilo Gateway

Coverage

NVIDIA released Nemotron 3.5 Lightning on August 11, 2026 as a 30-billion-parameter MoE model with 3B active parameters per token, published with open weights, open training data, and a permissive commercial license under OpenMDW-1.1, according to the excerpted coverage. The model supports context windows up to 1 milli Coverage frames the release as targeted at NVIDIA hardware, from GeForce RTX desktop cards up to DGX Spark systems built around the GB10 chip, positioning the open release as part of the company's broader AI infrastructure strategy. The article notes the release landed two weeks before NVIDIA's fiscal Q2, reporting dat

Kilo Gateway

Coverage

The official NVIDIA Nemotron LLM info page (last updated August 2026) describes the family as open-weight, multimodal, and built for long-running, multi-step AI agents that make thousands of model calls where cost and latency of routine calls dominate. The portfolio spans four text/agentic tiers (Lightning, Nano, Super The page positions Nemotron for agent developers, enterprise platform teams needing self-hosting, regulated industries inspecting training provenance, sovereign-AI programs, and model builders post-training on open weights. Two official one-liners appear: one emphasizing highly efficient multimodal open models for self

Kilo Gateway

CoverageBenchmark

Nemotron 3.5 Lightning is documented as a 30B-total, 3B-active hybrid MoE model with interleaved Mamba-2, MoE and attention layers, distilled from Nemotron 3 Ultra and released under the OpenMDW-1.1 license. Artificial Analysis measured median output speeds of roughly 670 tokens/second on an NVFP4 endpoint, with Intell Qubrid lists pricing at $0.069 per million input tokens, $0.29 per million output tokens, and $0.0069 per million implicit-cache tokens, positioning the model as roughly 46x cheaper per agent step than a typical frontier open model on tool-calling workloads. The article notes the model trails Qwen3.6 35B A3B on raw rea

Videos about Nemotron 3.5 Lightning 30B A3B

More models around Nemotron 3.5 Lightning 30B A3B