Currently listed through these providers:
Model details
Ling 3.1 Flash (Free)
Ling 3.1 Flash (Free) is a model from Inclusion AI's Ling family, positioned as a hybrid-reasoning mixture-of-experts (MoE) system aimed at coding tasks and tool-using workflows. Rather than separating reasoning into a distinct mode, the model blends reasoning capabilities into a single architecture, making it suitable for agent loops where planning and execution happen in the same request. This design targets developers who need a responsive model that can chain reasoning, code generation, and external tool invocations without mode-switching overhead.
The variant carries the identifier inclusionai/ling-3.1-flash and is offered without per-token charges for input, output, or cache reads, which lowers the cost barrier for teams prototyping agent pipelines or running exploratory workloads. A reported 262K-token context window gives it room to hold large codebases, multi-file refactors, or extended conversation histories in a single request. Benchmark indices for intelligence and coding remain unscored on the third-party model index consulted, so its relative standing against larger frontier systems is still unestablished, and practitioners should evaluate it directly against their own agent and coding benchmarks before committing to production use.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- inclusionai/ling-3.1-flash-free
- Release date
- Sep 29, 2026
- Last updated
- Sep 29, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 32,768 tokens
- Context window
- 262,144 tokens