Currently listed through these providers:
Model details
inclusionAI/Ling-flash-2.0
Ling-flash-2.0 sits inside inclusionAI's Ling 2.0 lineup as the third Mixture-of-Experts release, following the Ling-mini-2.0 language model and the Ring-mini-2.0 thinking variant. Its defining design choice is sparse activation inside a very large network: 100B total parameters with only 6.1B activated per token (4.8B of those non-embedding). The model was trained on more than 20 trillion tokens of curated data and then refined with supervised fine-tuning followed by multi-stage reinforcement learning, which the authors describe as the path to its strong results relative to similarly sized dense systems.
In practical terms, the model is pitched at teams that need reasoning depth and coding ability without paying for dense inference at the same scale. The Ling-flash-2.0 authors highlight standout performance on complex reasoning benchmarks (GPQA-Diamond, MMLU-Pro, AIME 2025, Omni-MATH, OptMATH, KOR-Bench, ARC-Prize), code generation (LiveCodeBench v6, CodeForces-Elo), domain-specific tasks such as FinanceReasoning and HealthBench, and frontend development. They also position it as state-of-the-art among dense models under 40B parameters despite activating only about 6B, while remaining competitive with larger-activation MoE rivals. The same weights are mirrored through third-party distribution channels such as MedAIBase on Ollama, giving builders an additional route to run the model locally.
Quick Info
Powered by- Provider
- SiliconFlow (China)
- Model key
- inclusionAI/Ling-flash-2.0
- Release date
- Sep 18, 2025
- Last updated
- Nov 25, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.14
- Output token cost
- $0.57
Limits
- Output tokens
- 131,000 tokens
- Context window
- 131,000 tokens
Latest news about inclusionAI/Ling-flash-2.0
No articles yet. Fetch the latest news to show it here.