Currently listed through these providers:
Model details
Ling 3.0 Flash Fin (Free)
Ling 3.0 Flash Fin Free sits in the Ling family of mixture-of-experts language models from InclusionAI, designed specifically for real-world investment and finance tasks. It is built on top of the Ling 3.0 Flash foundation and adopts an expert-routed architecture that activates 5.1B parameters out of a much larger 124B total, trading raw scale for efficiency in narrow financial workflows such as market analysis, structured financial Q&A, and decision-support summarization. The OpenRouter catalog entry that frames this description calls out the model's finance specialization as its defining trait, framing it as a domain-aligned MoE variant rather than a general-purpose chat model.
In its Free variant, the model is exposed through Vercel's AI Gateway with no per-token input or output cost, an approximate 0.6 second latency, and a reported throughput near 151 tokens per second, which positions it well for low-friction prototyping of finance agents and high-volume retrieval-augmented summarization. The Gateway backs the listing with the novita provider, fitting the model's stance as a lightweight, always-on inference option for cost-sensitive finance applications. With the cataloged API limit context window available, the model can absorb long filings, earnings transcripts, and multi-document research packets in a single pass, giving teams a low-cost way to evaluate domain-tuned MoE behavior before committing to paid tiers or larger Ling variants.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- inclusionai/ling-3.0-flash-fin-free
- Release date
- Aug 27, 2026
- Last updated
- Aug 27, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 32,000 tokens
- Context window
- 256,000 tokens