Currently listed through these providers:
Model details
Nemotron 3.5 Lightning 30B
Nemotron 3.5 Lightning 30B is a hybrid mixture-of-experts model that combines Mamba-2, MoE, and attention components, totaling 30B parameters with 3B active per token. NVIDIA distilled this checkpoint from its frontier Nemotron 3 Ultra, compressing the larger model's behavior into a smaller, faster engine. A 1M-token context window on the source model lets it ingest substantial codebases or multi-document agent traces in a single pass, well beyond the Vercel-listed context cap and output limit. Sampling defaults of temperature 1.0 suit tool-calling and agentic planning where reliable randomness control matters.
The model is purpose-built for always-on agents, emphasizing strong accuracy in coding, tool calling, instruction following, and multi-turn workflows. Partner deployments report up to 4× higher throughput than peer models, which translates to faster completion of specialized sub-tasks inside larger agent harnesses. It supports English plus Spanish, French, German, Italian, Japanese, and coding languages, broadening its reach for multilingual agent deployments. Its open-weight availability under a permissive license makes it a practical production workhorse alongside a larger reasoning model, which fits typical multi-model agent architectures that dispatch step-level work to smaller specialists.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- nvidia/nemotron-3.5-lightning
- Release date
- Aug 11, 2026
- Last updated
- Aug 11, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.05
- Output token cost
- $0.20
Limits
- Output tokens
- 131,072 tokens
- Context window
- 262,144 tokens