Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Nvidia logo

Model details

Nemotron 3.5 Lightning 30B A3B

Nemotron 3.5 Lightning 30B A3B is an open-weights member of NVIDIA's Nemotron family, released in two precision variants (BF16 and NVFP4) and made importable into OCI Generative AI on August 11, 2026. The model uses a hybrid architecture that blends Mamba-2 with a mixture-of-experts (MoE) layer and attention, totaling 30B parameters with 3B active per forward pass. This design positions it as a workhorse for long-running autonomous agents and sub-agent deployments, where sparse activation helps control compute costs while retaining the capacity needed for sustained, multi-step workflows.

The model supports text-to-text operation across English and several additional languages (Spanish, French, German, Italian, Japanese) along with coding languages, and it offers an extended context window of up to 1M tokens, which is well suited to agentic tasks that require large working memory. NVIDIA recommends a sampling setup with temperature 1.0 for generation. The combination of a hybrid Mamba-2/MoE/Attention backbone, multilingual support, and a very large context length makes the 30B A3B particularly fitting for complex agent pipelines that juggle long histories, tool interactions, and multilingual content within a single deployment.

Nvidianvidia/nemotron-3.5-lightning-30b-a3bnemotron

Quick Info

Powered by
Provider
Nvidia
Model key
nvidia/nemotron-3.5-lightning-30b-a3b
Release date
Aug 11, 2026
Last updated
Aug 11, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Latest news about Nemotron 3.5 Lightning 30B A3B

Merge Gateway

Official sourceBenchmark

An NVIDIA developer forum benchmark thread dated August 12, 2026 reports a measured throughput of 116.84 tokens per second for text generation using nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 on an NVIDIA DGX Spark system running vLLM. The post links to the full Spark Arena benchmark entry for the same NVFP4 we Beyond the headline throughput number, the excerpt is light on architectural or methodological detail, presenting a single data point rather than a controlled comparison across configurations or hardware targets. It nevertheless directly names the NVFP4 variant and ties it to a specific deployment stack (vLLM on DGX Sp

Videos about Nemotron 3.5 Lightning 30B A3B

More models around Nemotron 3.5 Lightning 30B A3B