Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Pioneer logo

Model details

Nemotron 3.5 Lightning 30B A3B

Nemotron 3.5 Lightning 30B A3B is a hybrid architecture that combines Mamba-2 state-space layers with mixture-of-experts routing and selective attention, keeping the active footprint at roughly 3B parameters while the full model spans 30B. Pre-training was conducted on more than 20 trillion tokens under an NVFP4 recipe augmented with Multi-Token Prediction, a combination aimed at accelerating inference without sacrificing throughput quality. Open weights are released under the OpenMDW License Agreement v1.1, and the model is positioned within the broader NVIDIA Nemotron family as a lightweight tier optimized for speed.

The headline practical advantage is an extended context window reaching up to one million tokens, which makes the model suitable for long-running autonomous agents, sub-agent orchestration, and other agentic pipelines that need to retain large working memories. It supports reasoning and tool calling, with a capability set oriented around text-only inputs and outputs, and it covers English along with major European languages and Japanese for multilingual deployments. Its efficiency-first design, low active parameter count, and multi-token prediction training make it a fit when teams need a responsive open-weights model for sustained agent workloads rather than a heavyweight general-purpose chat model.

Pioneernvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16nemotron

Quick Info

Powered by
Provider
Pioneer
Model key
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
Release date
Aug 11, 2026
Last updated
Aug 11, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.50
Output token cost
$0.50

Limits

Output tokens
4,096 tokens
Context window
8,192 tokens

Transparent token rates

Compare Nemotron 3.5 Lightning 30B A3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron 3.5 Lightning 30B A3B

Pioneer

Coverage

A long-form guide published on note.com on August 27, 2026 (with an AI-translation disclaimer) walks through NVIDIA Nemotron 3.5 Lightning's release of August 11, 2026, aimed at intermediate readers selecting models for LLM APIs or in-house agent infrastructure. It summarizes the same core architecture as a 30B MoE wit The article cross-checks official NVIDIA figures against independent measurements, covers the OpenMDW-1.1 license terms, and surveys real-world company adoption examples to help readers decide whether to route agent tool-calling traffic to this checkpoint. The page explicitly warns that nuances and authorial intent may

Pioneer

Coverage

A hands-on technical article on Zenn (AI-translated from Japanese) documents an attempt to run NVIDIA's Nemotron 3.5 Lightning 30B-A3B on Google Colab, explicitly naming the exact checkpoint "NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16" hosted on Hugging Face and citing the model's NVIDIA release date of August 11, 2026 The article also serves as a practical cautionary note for practitioners: while the model card lists a 1M-token maximum context, practical usable context on Colab is bounded by the allocated runtime's GPU memory rather than the advertised ceiling, and Colab resource availability fluctuates with time, plan, and runtime

Pioneer

Coverage

The NVIDIA-hosted LLM Info page for Nemotron, last updated August 2026, provides the authoritative first-party framing of the model family. It defines Nemotron as a portfolio of high-efficiency, multimodal, open-weight AI models spanning four text/agentic tiers (Lightning, Nano, Super, Ultra), one multimodal tier (Nano The page targets long-running, multi-step AI agents where the cost and latency of routine calls dominate total spend, and notes two official one-liners are live simultaneously on NVIDIA properties: one describing Nemotron as a family of highly efficient, multimodal, open AI models built for long-running, self-evolving

Pioneer

CoverageBenchmark

A Qubrid AI blog post dated August 11, 2026 reports that NVIDIA's Nemotron 3.5 Lightning is now available as an API on Qubrid's platform. The article frames the model as a 30B-total, 3B-active hybrid Mixture-of-Experts architecture combining interleaved Mamba-2, MoE, and attention layers, distilled from Nemotron 3 Ultr The post supplies concrete performance and pricing figures: Artificial Analysis median output of nearly 670 tokens per second on a pre-release NVFP4 endpoint, an Intelligence Index score of 24 (a +9 point jump over Nemotron 3 Nano, level with gpt-oss-120b, behind Nemotron 3 Super at 26), Terminal-Bench v2.1 at 24% vers

Videos about Nemotron 3.5 Lightning 30B A3B

More models around Nemotron 3.5 Lightning 30B A3B