Sulat.com
AI models
Vercel AI Gateway logo

Model details

Nemotron 3.5 Lightning 30B

Nemotron 3.5 Lightning 30B is a hybrid mixture-of-experts model that combines Mamba-2, MoE, and attention components, totaling 30B parameters with 3B active per token. NVIDIA distilled this checkpoint from its frontier Nemotron 3 Ultra, compressing the larger model's behavior into a smaller, faster engine. A 1M-token context window on the source model lets it ingest substantial codebases or multi-document agent traces in a single pass, well beyond the Vercel-listed context cap and output limit. Sampling defaults of temperature 1.0 suit tool-calling and agentic planning where reliable randomness control matters.

The model is purpose-built for always-on agents, emphasizing strong accuracy in coding, tool calling, instruction following, and multi-turn workflows. Partner deployments report up to 4× higher throughput than peer models, which translates to faster completion of specialized sub-tasks inside larger agent harnesses. It supports English plus Spanish, French, German, Italian, Japanese, and coding languages, broadening its reach for multilingual agent deployments. Its open-weight availability under a permissive license makes it a practical production workhorse alongside a larger reasoning model, which fits typical multi-model agent architectures that dispatch step-level work to smaller specialists.

Vercel AI Gatewaynvidia/nemotron-3.5-lightningnemotron

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
nvidia/nemotron-3.5-lightning
Release date
Aug 11, 2026
Last updated
Aug 11, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.05
Output token cost
$0.20

Limits

Output tokens
131,072 tokens
Context window
262,144 tokens

Latest news about Nemotron 3.5 Lightning 30B

Videos about Nemotron 3.5 Lightning 30B

Recent tweets and retweets from Vercel AI Gateway

More models around Nemotron 3.5 Lightning 30B