Ling-3.0-flash is a next-generation hybrid reasoning model from inclusionAI, designed to deliver flagship-class capability at a fraction of the compute. It carries 124B total parameters but activates only about 5.1B per token, roughly an order of magnitude smaller than the 1T-class Ring-2.6-1T it descends from, yet inclusionAI states it matches or surpasses that predecessor on key benchmarks. The architecture is a native hybrid linear attention design built from the start of pretraining, alternating Kimi Delta Attention (KDA) and MLA layers in a 5:1 stack, with KDA fine-grained diagonal gating and a 1/64 sparse MoE for long-context efficiency and low per-token cost.
The model is aimed squarely at production agentic workloads rather than maximum scale. It was trained across more than 10,000 interactive environments for end-to-end closed-loop execution of coding, general, and deep-research agent tasks, and natively integrates the SGLang HiCache plus Mooncake hierarchical caching stack for fast, repeatable serving. With reasoning, tool calling, and a 262K-token context window already exposed, Ling-3.0-flash fits well for long-document analysis, multi-step tool use, and latency-sensitive applications where a 1T-class model would be overkill, while still benefiting from the Ling family lineage.