Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Kilo Gateway logo

Model details

inclusionAI: Ling 3.0 Flash

Ling-3.0-flash is a next-generation hybrid reasoning model from inclusionAI, designed to deliver flagship-class capability at a fraction of the compute. It carries 124B total parameters but activates only about 5.1B per token, roughly an order of magnitude smaller than the 1T-class Ring-2.6-1T it descends from, yet inclusionAI states it matches or surpasses that predecessor on key benchmarks. The architecture is a native hybrid linear attention design built from the start of pretraining, alternating Kimi Delta Attention (KDA) and MLA layers in a 5:1 stack, with KDA fine-grained diagonal gating and a 1/64 sparse MoE for long-context efficiency and low per-token cost.

The model is aimed squarely at production agentic workloads rather than maximum scale. It was trained across more than 10,000 interactive environments for end-to-end closed-loop execution of coding, general, and deep-research agent tasks, and natively integrates the SGLang HiCache plus Mooncake hierarchical caching stack for fast, repeatable serving. With reasoning, tool calling, and a 262K-token context window already exposed, Ling-3.0-flash fits well for long-document analysis, multi-step tool use, and latency-sensitive applications where a 1T-class model would be overkill, while still benefiting from the Ling family lineage.

Kilo Gatewayinclusionai/ling-3.0-flashling

Quick Info

Powered by
Provider
Kilo Gateway
Model key
inclusionai/ling-3.0-flash
Release date
Jul 23, 2026
Last updated
Jul 23, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.06
Output token cost
$0.18

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Transparent token rates

Compare inclusionAI: Ling 3.0 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about inclusionAI: Ling 3.0 Flash

Kilo Gateway

Coverage

A third-party Medium explainer dated August 8, 2026 details the architecture of InclusionAI's Ling 3.0 Flash, describing it as a 124-billion-parameter hybrid-reasoning Mixture-of-Experts model that activates only 5.1B parameters per token (1/64 expert activation), with 512 routed experts plus 1 shared expert and 8 expe The same piece reports that Ling 3.0 Flash is released under an open license on Hugging Face and claims inference speeds up to 1,000 tokens per second, while compressing the expert activation ratio from 1/32 in the prior generation to 1/64, yielding 5–10x compute efficiency versus dense or less-sparse MoE models. It fr

Kilo Gateway

Coverage

inclusionAI has released Ling-3.0-flash on Hugging Face under an MIT license, a 124B-parameter hybrid Mixture-of-Experts model that activates approximately 5.1B parameters per token — about 8% of the model firing on any given step. The architecture is described as a native hybrid-linear MoE using a 5:1 alternating stac On the model card's self-reported benchmarks, Ling-3.0-flash posts 56.6% on SWE-Bench Pro, 72.4% on SWE-Bench Multilingual, 93.2 on MathArena AIME 2026, and 22.7 on HLE, with the team claiming a 60–80% reduction in Time to First Token for long inputs via hierarchical caching. The card lists compatibility with Claude Co

Kilo Gateway

CoverageBenchmark

BenchLM's independent aggregator entry for Ling 3.0 Flash, released July 23, 2026, lists the model as open-weight with a 262K-token context window and no published first-party hosted API price. The aggregator rolls up 22 source-displayable benchmark rows and assigns a composite capability score of 47.5 out of 100, rank Category percentile breakdowns show Agentic at the 29th percentile (108 of 152 ranked), Coding at the 25th percentile (113 of 151), and Knowledge at the 38th percentile (113 of 183); Reasoning, Math, Multilingual, and Multimodal are listed as not ranked. Throughput is reported at roughly 293 tokens per second with a fi

Videos about inclusionAI: Ling 3.0 Flash

More models around inclusionAI: Ling 3.0 Flash