Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Nebius Token Factory logo

Model details

Nemotron 3.5 Lightning 30B A3B

Nemotron 3.5 Lightning 30B A3B is an open-weight large language model developed by NVIDIA, released on August 11, 2026 alongside openly published training data and a permissive commercial license. Its architecture is a hybrid mixture-of-experts design that combines Mamba-2 state-space layers, MoE routing, and attention blocks, totaling 30 billion parameters with 3 billion active per token. This sparse-active configuration is intended to deliver capable reasoning at a fraction of the compute cost of dense models of similar size, while still benefiting from attention's strengths on localized context. The model is positioned for long-running autonomous agents and sub-agent workhorse deployments, making the architecture choice especially relevant for sustained multi-step workflows.

For practical fit, the model supports a context length of up to one million tokens and is multilingual, covering English and coding languages alongside Spanish, French, German, Italian, and Japanese. NVIDIA designed it to run efficiently across its own hardware stack, from GeForce RTX desktop cards up to DGX Spark systems based on the GB10 chip, giving developers flexibility across consumer and data center tiers. The open weights and training data allow teams to fine-tune, audit, and integrate the model into production pipelines without licensing friction, which is valuable for organizations building specialized agentic systems that need long-context understanding and strong multilingual coverage.

Nebius Token Factorynvidia/Nemotron-3_5-Lightningnemotron

Quick Info

Powered by
Provider
Nebius Token Factory
Model key
nvidia/Nemotron-3_5-Lightning
Release date
Aug 11, 2026
Last updated
Aug 11, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.06
Output token cost
$0.24

Limits

Output tokens
1,048,576 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Nemotron 3.5 Lightning 30B A3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron 3.5 Lightning 30B A3B

Nebius Token Factory

Coverage

The note.com AI-Driven Lab guide, published in Japanese with a system notice flagging AI auto-translation, independently corroborates the same core facts: Nemotron 3.5 Lightning was released by NVIDIA on August 11, 2026, targeting the "execution layer" of agents rather than frontier reasoning. It walks through architec Because the article is an AI-translated third-party guide rather than a primary source, its strength lies in cross-checking release date, positioning and licensing rather than in providing new authoritative data, and the system notice explicitly warns that nuance and authorial intent may not be fully reflected. It cont

Nebius Token Factory

Coverage

Tech Insider reports that NVIDIA released Nemotron 3.5 Lightning on August 11, 2026 as a 30-billion-parameter mixture-of-experts model with 3 billion parameters active per token and context windows up to 1 million tokens. It says NVIDIA positioned the model for continuously running, task-focused agents and published a The release is described as including BF16, FP8, and NVFP4 weights, allowing teams to choose a format suited to their deployment. The article also says the model can run on NVIDIA hardware ranging from GeForce RTX systems to DGX Spark with its GB10 chip. Its hardware-sales and data-center-revenue framing is opinion and

Nebius Token Factory

Coverage

Atomic Chat reports that NVIDIA released the exact Nemotron 3.5 Lightning 30B A3B checkpoint on August 11, 2026. The model is described as a 52-layer sparse mixture-of-experts architecture with 30 billion total parameters, 3 billion active per token, and six of 128 experts selected; its stack combines Mamba-2 blocks wi The page lists OpenMDW 1.1 as the license and says NVIDIA published DSpark, MTP, and DFlash speculative-decoding drafter heads alongside the model. Its reproduced NVIDIA BF16 results include 81.94 on MMLU Pro, 75.44 on GPQA Diamond, 51.56 on SWE-bench Verified, 24.58 on Terminal-Bench 2.1, and 52.00 on AA-LCR. These ar

Nebius Token Factory

CoverageBenchmark

Qubrid AI's August 11, 2026 deep dive on the Nemotron 3.5 Lightning API adds independent benchmark and architectural detail: a 30B-total / 3B-active hybrid MoE with interleaved Mamba-2, MoE and attention layers, distilled from Nemotron 3 Ultra, released under OpenMDW-1.1, with text-only 1M-token context and a runtime r The same post discloses Qubrid-specific pricing at $0.069 / $0.29 / $0.0069 per million input / output / implicit-cache tokens, framed as roughly 46x cheaper per agent step than a frontier open model on a typical tool-calling shape, with an honest caveat that Lightning trails Qwen3.6 35B A3B on raw reasoning and repo-s

Nebius Token Factory

Coverage

NVIDIA's official Nemotron LLM Info page, last updated August 2026, anchors the model in its family context: NVIDIA Nemotron is described as a portfolio of "highly efficient, multimodal, open AI models built for long-running, self-evolving agents," with text/agentic tiers that include Lightning alongside Nano, Super an The same NVIDIA page is useful for disambiguating the subject model slug on Nebius: it clarifies that "Nemotron" is a family, not a single checkpoint, so "Nemotron 3.5 Lightning" is one tier within that family rather than a standalone brand. Because the page is family-level rather than checkpoint-specific, it does not

Videos about Nemotron 3.5 Lightning 30B A3B

More models around Nemotron 3.5 Lightning 30B A3B