Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
CoreWeave logo

Model details

Nemotron 3.5 Lightning

Nemotron 3.5 Lightning is NVIDIA's specialized execution model for always-on AI agents, designed to handle the high-volume steps such as tool calls, result validation, and subagent delegation that dominate long-running workloads. It is a 30B parameter mixture-of-experts model with roughly 3B active parameters per token, giving it the capacity of a larger dense model at the compute cost of a small one. The architecture is a hybrid Mamba-Transformer design carried forward from its predecessor, NVIDIA Nemotron 3 Nano, and the model supports a context length of up to one million tokens, leaving ample room for long tool histories across multi-turn workflows.

Speed and customization are central to the model's design. A dedicated pretraining stage baked multi-token prediction into the weights, and the release ships with DFlash and DSpark draft models so speculative decoding can be tuned across serving scenarios, from DGX Spark up to data-center concurrency. An NVFP4 quantized checkpoint is included alongside the BF16 weights, running on the same specialized kernels that power the rest of the Nemotron family across Blackwell, Hopper, and Ampere GPUs. Open weights, training data, and recipes are released permissively under OpenMDW-1.1, enabling LoRA or full supervised fine-tuning and reinforcement learning on modest hardware. The model is well suited to coding sub-agents, local personal assistants, security workflows, and any tier where a lightweight, fast specialist should run beside a heavier frontier planner.

CoreWeavenvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3Bnemotron

Quick Info

Powered by
Provider
CoreWeave
Model key
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B
Release date
Aug 11, 2026
Last updated
Aug 11, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.07
Output token cost
$0.20

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Transparent token rates

Compare Nemotron 3.5 Lightning pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron 3.5 Lightning

Weights & Biases

Coverage

NVIDIA's developer blog confirms the August 11, 2026 launch of Nemotron 3.5 Lightning, an open 30B mixture-of-experts model with 3B active parameters explicitly positioned as an execution-layer model for long-running AI agents. The post names hybrid Mamba-Transformer MoE architecture, speculative decoding via multi-tok Customization is described as out-of-the-box using LoRA or full SFT through NeMo Automodel and NeMo Megatron Bridge, with reinforcement learning via NeMo RL and NeMo Gym, and NeMo Switchyard for routing between frontier planning models and Lightning for execution work. The blog claims Lightning defines the accuracy-spe

CoreWeave

Coverage

NVIDIA Nemotron 3.5 Lightning is now available on Ollama as of August 11, 2026, running completely on-device as a 30 billion total parameter (3B active per token) open Mixture-of-Experts model from NVIDIA. The model is built for agentic workloads such as reading files, calling tools, sorting results, and retrying faile Key technical features include a 1M token context window for long tool histories across multi-turn workflows, speculative decoding using multi-token prediction (MTP), DFlash, or DSpark for up to 4x higher throughput versus comparable open models, and an open model trained on open datasets under a permissive license all

CoreWeave

CoverageAnalysis

NVIDIA released Nemotron 3.5 Lightning on August 11, 2026 as the first model in the Nemotron 3.5 family and successor to Nemotron 3 Nano 30B-A3B, with 31.6B total and 3.6B active parameters on the same hybrid Mamba-Transformer architecture. The model scores 24 on the Artificial Analysis Intelligence Index, a +9 point i The largest gains over the prior generation are on agentic evaluations: GDPval-AA v2 improved by +334 ELO, surpassing gpt-oss-120b and Nemotron 3 Super, while Terminal-Bench v2.1 rose to 24% from 7%. The model ships in NVFP4 alongside BF16 weights with near-lossless quality (Intelligence Index 24 on NVFP4), features a

Videos about Nemotron 3.5 Lightning

More models around Nemotron 3.5 Lightning