Currently listed through these providers:
Model details
nemotron-lightning-3.5-30b-a3b
Nemotron Lightning 3.5 30B A3B is a mixture-of-experts model from Nvidia that activates roughly 3 billion parameters out of a 30 billion total, giving it the footprint of a small model while drawing on a much larger expert pool per token. The Lightning variant is positioned within the Nemotron 3.5 family as the lighter option, trading the heavier planning and reasoning depth of Nemotron 3 Super and Ultra for speed and throughput. It is a text-to-text language model, making it suitable for natural language pipelines, code generation, structured data tasks, and tool-mediated workflows rather than image or audio understanding.
The model is aimed at high-throughput, agentic workloads where latency and cost per call matter more than maximum reasoning depth, such as multi-step assistants, retrieval-augmented chatbots, and routine code or text automation. Its headline technical feature is an advertised context window of up to roughly one million tokens, which lets it ingest very large documents, long conversation histories, or sizable codebases in a single pass without aggressive truncation. Because only a small fraction of the experts fire on each token, it can serve many concurrent requests efficiently, fitting well alongside larger Nemotron siblings when a deployment needs a fast, scalable workhorse for everyday language tasks.
Quick Info
Powered by- Provider
- Requesty
- Model key
- nemotron-lightning-3.5-30b-a3b
- Release date
- Aug 15, 2026
- Last updated
- Aug 15, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.05
- Output token cost
- $0.20
Limits
- Output tokens
- 262,144 tokens
- Context window
- 262,144 tokens