Currently listed through:
Model details
Nemotron 3.5 Lightning Free
Nemotron 3.5 Lightning Free comes from NVIDIA as an open Mixture-of-Experts architecture, pairing 30 billion total parameters with only 3 billion active per token. That sparse activation pattern is the core efficiency trick: the model keeps a large knowledge reservoir on disk but only spins up a small slice of experts for each request, which is why it is described as optimized for high-throughput agentic workloads and efficient inference. The lightweight active compute footprint also makes it attractive for specialized agents and domain-specific customization, where teams can fine-tune or wrap the model without paying the latency cost of a dense 30B run.
In practical terms, the design priorities point to speed, cost efficiency, and adaptability rather than raw frontier reasoning. NVIDIA frames it for scalable production deployments where many small, fast agent calls matter more than a single huge response, and the free tier removes the price barrier for prototyping those pipelines. Teams building tool-using assistants, retrieval-heavy chatbots, or background workers that need quick completions against large contexts will find the sparse MoE shape fits well, while those needing maximum reasoning depth on a single query may want to compare against denser Nemotron variants.
Quick Info
Powered by- Provider
- OpenCode Zen
- Model key
- nemotron-3.5-lightning-free
- Release date
- Aug 11, 2026
- Last updated
- Aug 11, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 262,144 tokens
- Context window
- 262,144 tokens