Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Merge Gateway logo

Model details

Nemotron 3.5 Lightning 30B A3B

Nemotron 3.5 Lightning 30B A3B is an open-weights member of NVIDIA's Nemotron family, released in two precision variants (BF16 and NVFP4) and made importable into OCI Generative AI on August 11, 2026. The model uses a hybrid architecture that blends Mamba-2 with a mixture-of-experts (MoE) layer and attention, totaling 30B parameters with 3B active per forward pass. This design positions it as a workhorse for long-running autonomous agents and sub-agent deployments, where sparse activation helps control compute costs while retaining the capacity needed for sustained, multi-step workflows.

The model supports text-to-text operation across English and several additional languages (Spanish, French, German, Italian, Japanese) along with coding languages, and it offers an extended context window of up to 1M tokens, which is well suited to agentic tasks that require large working memory. NVIDIA recommends a sampling setup with temperature 1.0 for generation. The combination of a hybrid Mamba-2/MoE/Attention backbone, multilingual support, and a very large context length makes the 30B A3B particularly fitting for complex agent pipelines that juggle long histories, tool interactions, and multilingual content within a single deployment.

Merge Gatewaynvidia/nemotron-3.5-lightning-30b-a3bnemotron

Quick Info

Powered by
Provider
Merge Gateway
Model key
nvidia/nemotron-3.5-lightning-30b-a3b
Release date
Aug 11, 2026
Last updated
Aug 11, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
262,144 tokens
Context window
1,000,000 tokens

Latest news about Nemotron 3.5 Lightning 30B A3B

Merge Gateway

Official sourceBenchmark

An NVIDIA developer forum benchmark thread dated August 12, 2026 reports a measured throughput of 116.84 tokens per second for text generation using nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 on an NVIDIA DGX Spark system running vLLM. The post links to the full Spark Arena benchmark entry for the same NVFP4 we Beyond the headline throughput number, the excerpt is light on architectural or methodological detail, presenting a single data point rather than a controlled comparison across configurations or hardware targets. It nevertheless directly names the NVFP4 variant and ties it to a specific deployment stack (vLLM on DGX Sp

Merge Gateway

Coverage

NVIDIA's official Nemotron LLM info page (last updated August 2026) frames Nemotron as a family of high-efficiency, multimodal, open-weight AI models rather than a single checkpoint. The family spans four text/agentic tiers — Lightning, Nano, Super, and Ultra — plus a multimodal Nano Omni tier and task-specific lines f The page emphasizes open weights, training data, and recipes published under the "nvidia" Hugging Face organization and the NVIDIA-NeMo/Nemotron GitHub, with technical reports at research.nvidia.com, making Nemotron suitable for self-hosted enterprise, regulated, and sovereign-AI deployments. While the supplied excerpt

Videos about Nemotron 3.5 Lightning 30B A3B

More models around Nemotron 3.5 Lightning 30B A3B