Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
RunInfra logo

Model details

Nemotron 3.5 Lightning 30B A3B

Nemotron 3.5 Lightning 30B A3B is built around a hybrid architecture that combines Mamba-2 with mixture-of-experts and attention layers, totaling 30 billion parameters of which roughly 3 billion activate per inference. This design choice pairs the long-range sequence modeling characteristics of Mamba-2 with the routing flexibility of MoE and the contextual precision of attention, allowing efficient handling of extended tasks without scaling full compute for every token. NVIDIA frames the model as a workhorse targeted at long-running autonomous agents and sub-agent deployments, where many tokens are spent on planning, tool use, and multi-step reasoning rather than single-shot answers.

Open weights and native function calling make the model suitable for self-hosting and tool-augmented pipelines, while built-in reasoning support lets developers opt into chain-of-thought behavior when needed. Its advertised context envelope reaches up to one million tokens in NVIDIA's reference configuration, well beyond typical chat workloads, enabling agents to operate over very long documents, multi-turn histories, or accumulated tool outputs. The combination of a lean active parameter count, agent-focused design intent, and open-weight availability positions the model as a pragmatic choice for teams building dependable autonomous workflows that need controllable reasoning and reliable tool integration.

RunInfranvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16nemotron

Quick Info

Powered by
Provider
RunInfra
Model key
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
Release date
Aug 11, 2026
Last updated
Aug 11, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.05
Output token cost
$0.15

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Transparent token rates

Compare Nemotron 3.5 Lightning 30B A3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron 3.5 Lightning 30B A3B

RunInfra

Coverage

NVIDIA's Nemotron LLM info page (last updated August 2026) frames the Nemotron family as a portfolio of high-efficiency, multimodal, open-weight models for long-running, multi-step AI agents, spanning text/agentic tiers (Lightning, Nano, Super, Ultra), a multimodal tier (Nano Omni), and task-specific lines for document The page notes two slightly different official NVIDIA taglines live simultaneously across nvidia.com and developer.nvidia.com, reflecting ongoing family messaging rather than contradicting Lightning's checkpoint specs. For developers evaluating RunInfra's hosted BF16 checkpoint, this provides the authoritative vendor f

Videos about Nemotron 3.5 Lightning 30B A3B

More models around Nemotron 3.5 Lightning 30B A3B