Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Fireworks AI logo

Model details

Nemotron 3.5 Lightning 30B A3B

Nemotron 3.5 Lightning 30B A3B is part of NVIDIA's open Nemotron family and is designed to act as a workhorse for long-running autonomous agents and sub-agent workflows. NVIDIA describes it as a hybrid architecture that combines Mamba-2, Mixture-of-Experts, and attention layers, with 30 billion total parameters of which roughly 3 billion are active per inference. This sparse activation pattern is intended to keep compute and latency low while still drawing on the capacity of a much larger model, making the system well suited to agentic pipelines that need fast, repeated calls.

A standout design choice is the very large context window that NVIDIA advertises, supporting up to one million tokens, which positions the model for tasks that require sustained reasoning across long documents or extended multi-step agent traces. The model is offered with English plus several major European and Asian languages alongside common coding languages, broadening its usefulness for mixed-language software work. In head-to-head benchmark comparisons surfaced by third parties, larger dense contemporaries such as Qwen3.6-35B-A3B have posted stronger results on reasoning and coding suites like GPQA, Humanity's Last Exam, MMLU-Pro, and SWE-bench, so practitioners should weigh raw capability against Nemotron's efficiency and agent-focused design when choosing a deployment.

Fireworks AIaccounts/fireworks/models/nemotron-lightning-3p5-30b-a3bnemotron

Quick Info

Powered by
Provider
Fireworks AI
Model key
accounts/fireworks/models/nemotron-lightning-3p5-30b-a3b
Release date
Aug 11, 2026
Last updated
Aug 11, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.05
Output token cost
$0.20

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

OpenCode

Model variants

Priority

Input
$0.0625
per 1M tokens
Output
$0.25
per 1M tokens
Cache Read
$0.0125
per 1M tokens

Transparent token rates

Compare Nemotron 3.5 Lightning 30B A3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Nemotron 3.5 Lightning 30B A3B

Fireworks AI

Coverage

A third-party model directory page (atomic.chat, updated August 24, 2026) documents NVIDIA's Nemotron 3.5 Lightning 30B A3B, the exact variant referenced in the subject. It describes the model as a throughput-oriented hybrid released on August 11, 2026: a sparse mixture-of-experts with 30 billion total parameters and 3 The same page reproduces a benchmark table attributed to NVIDIA's BF16 model card comparing Nemotron 3.5 Lightning against Qwen3.6-35B-A3B, Gemma 4 26B A4B, Nemotron 3 Super, and GPT-OSS 20B across MMLU Pro (81.94), GPQA Diamond (75.44), SWE-bench Verified (51.56), Terminal-Bench 2.1 (24.58), PinchBench (85.37), IFBenc

Videos about Nemotron 3.5 Lightning 30B A3B

More models around Nemotron 3.5 Lightning 30B A3B