Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
OpenRouter logo

Model details

Ling 3.0 Flash

Ling-3.0-flash is a next-generation hybrid reasoning model released by inclusionAI under Ant Group, designed from the ground up for production-grade AI agent workflows where rapid response and cost-efficiency matter. Rather than chasing raw parameter count, it uses a native hybrid linear attention architecture that alternates Kimi Delta Attention and MLA layers in a 5:1 stack, augmented with fine-grained diagonal gating on KDA and a sparse MoE routing pattern. This design lets it activate only 5.1B of its 124B total parameters per token while still delivering strong reasoning, instruction following, and long-context behavior, reportedly matching or surpassing flagship models two to three times its scale on core benchmarks.

Beyond raw architecture, Ling-3.0-flash is tuned for real-world productivity and agentic execution. Training incorporated more than 10,000 interactive environments covering coding, general assistance, and deep research tasks, enabling end-to-end closed-loop behavior rather than single-turn answers. It also integrates the SGLang HiCache plus Mooncake hierarchical caching stack, which reduces latency and inference cost for long-context serving. Available as open weights through Hugging Face, ModelScope, and OpenRouter, it is a practical fit for teams that want frontier-class reasoning and long-context throughput without operating trillion-parameter infrastructure.

OpenRouterinclusionai/ling-3.0-flashling

Quick Info

Powered by
Provider
OpenRouter
Model key
inclusionai/ling-3.0-flash
Release date
Jul 23, 2026
Last updated
Jul 23, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.021
Output token cost
$0.063

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Transparent token rates

Compare Ling 3.0 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Ling 3.0 Flash

Vercel AI Gateway

Coverage

InclusionAI officially released the weights for Ling-3.0-flash on Hugging Face and ModelScope on August 5, 2026, with both BF16 and FP8 quantized versions available under the MIT license. This followed the initial API announcement in late July and made the 124B-total/5.1B-active MoE model practically downloadable for i Architecture analysis grounded in the official model card confirms the hybrid linear attention design combining Kimi Delta Attention and MLA in a 5:1 ratio, alongside the MoE compute configuration. The piece documents FP8 quantization retention characteristics and discusses adoption considerations for engineering teams

OpenRouter

CoverageBenchmark

Ant Group announced the release of Ling-3.0-Flash on July 27, 2026 (per Business Wire), describing it as a native hybrid-reasoning foundational model aimed at production-grade AI agent workflows and serving as a high-speed execution node balancing intelligence density and cost-efficiency. The model features 124B total Architecturally, Ling-3.0-Flash moves away from scaling parameter counts and is built from the ground up with a native hybrid-linear attention architecture that alternates KDA (Kimi Delta Attention) and MLA layers at a 5:1 ratio to balance long-context efficiency with robust state memory. Key architectural advancements

OpenRouter

Coverage

Digital Applied's blog (published July 24, 2026) provides secondary technical coverage of Ling-3.0-flash, confirming the 124B total / 5.1B active per token MoE design and noting that at publication the weights were not yet on Hugging Face — explicitly warning readers that aggregator listings often misattribute the mode The article documents a practitioner-legible explanation of the KDA + MLA hybrid attention design behind the efficiency story and compares Ant's activation ratio with Moonshot AI's, while presenting the full Ling family lineage in a single table for context. It explicitly flags that the model's OpenRouter free-tier pro

OpenRouter

Coverage

Orca Router's blog provides third-party corroboration of Ling-3.0-flash's official launch by Ant Group's InclusionAI lab, announced July 24, 2026, with model weights subsequently open-sourced under the MIT license on August 7, 2026, and published on Hugging Face and ModelScope. The post describes the model as a native On the Artificial Analysis Intelligence Index, Ling-3.0-flash scores 38, a 24-point improvement over the prior fast-tier Ling-2.6-flash generation, and the write-up notes it sits on the open-weights Pareto frontier for intelligence versus total parameters. This independent benchmark context complements OpenRouter's pro

OpenRouter

Coverage

Kilo's blog (July 23, 2026) announces that Ling 3.0 Flash from inclusionAI (Ant Group) is available on the Kilo platform, with a limited-time free access promotion for developers building multi-turn coding agents or daily agentic workflows on tight token budgets. The piece frames Ling 3.0 Flash within the "flash model The Kilo post repeats the core architectural claims — 124B total parameters with 5.1B active per token, optimization for high-throughput token-efficient inference, and a native 256K context window extendable to 1M for long-horizon task stability and deep repository reasoning — and reports that early testing found the m

OpenRouter

Official sourceBenchmark

OpenRouter's official model page confirms that inclusionai/ling-3.0-flash is listed and routable on the platform as a 124B-parameter Mixture-of-Experts model with approximately 5.1B parameters activated per token, designed for token-efficient, production-scale agentic inference. The page documents a release date of Jul Two providers host the model on OpenRouter: NovitaAI (100.00% uptime, 0.69s P50 latency, 136 tok/s throughput, weighted-avg effective price ~$0.01213/$0.06326 per 1M with a 57.6% cache hit rate and 91.4% token share) and DeepInfra (93.36% uptime, 1.25s P50 latency, 36 tok/s throughput, 6.2% cache hit, 8.6% token share)

Videos about Ling 3.0 Flash

More models around Ling 3.0 Flash