Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

Ling 3.0 Flash

Ling-3.0-Flash is a foundational model from Ant Group's Ling family that is built for production-grade agent workflows, where rapid response and high reasoning density matter more than raw scale. Rather than pushing parameter counts upward, the model uses a 124B-total-parameter design with only about 5.1B parameters active per token, framed by Ant Group as matching or surpassing industry-leading models reported at roughly two to three times its parameter footprint across foundational reasoning, instruction following, and long-context benchmarks.

Underneath, Ling-3.0-Flash relies on a native hybrid-linear attention architecture that alternates KDA (Kimi Delta Attention) and MLA layers at a 5:1 ratio, which is intended to combine long-context efficiency with stronger state memory for multi-step agent tasks. The model is designed as a high-speed execution node inside larger agent pipelines, and the same backbone has since been extended through continued training into Ling-3.0-flash-Fin, a finance-focused variant that inherits the 124B/5.1B configuration and the cataloged API limit context window while adding capabilities such as source-grounded financial search, multi-document reasoning across filings, and valuation and spreadsheet workflows.

Vercel AI Gatewayinclusionai/ling-3.0-flashling

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
inclusionai/ling-3.0-flash
Release date
Aug 6, 2026
Last updated
Aug 6, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.021
Output token cost
$0.063

Limits

Output tokens
32,000 tokens
Context window
256,000 tokens

Transparent token rates

Compare Ling 3.0 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Ling 3.0 Flash

Vercel AI Gateway

Coverage

InclusionAI officially released the weights for Ling-3.0-flash on Hugging Face and ModelScope on August 5, 2026, with both BF16 and FP8 quantized versions available under the MIT license. This followed the initial API announcement in late July and made the 124B-total/5.1B-active MoE model practically downloadable for i Architecture analysis grounded in the official model card confirms the hybrid linear attention design combining Kimi Delta Attention and MLA in a 5:1 ratio, alongside the MoE compute configuration. The piece documents FP8 quantization retention characteristics and discusses adoption considerations for engineering teams

OpenRouter

CoverageBenchmark

Ant Group announced the release of Ling-3.0-Flash on July 27, 2026 (per Business Wire), describing it as a native hybrid-reasoning foundational model aimed at production-grade AI agent workflows and serving as a high-speed execution node balancing intelligence density and cost-efficiency. The model features 124B total Architecturally, Ling-3.0-Flash moves away from scaling parameter counts and is built from the ground up with a native hybrid-linear attention architecture that alternates KDA (Kimi Delta Attention) and MLA layers at a 5:1 ratio to balance long-context efficiency with robust state memory. Key architectural advancements

OpenRouter

Coverage

Digital Applied's blog (published July 24, 2026) provides secondary technical coverage of Ling-3.0-flash, confirming the 124B total / 5.1B active per token MoE design and noting that at publication the weights were not yet on Hugging Face — explicitly warning readers that aggregator listings often misattribute the mode The article documents a practitioner-legible explanation of the KDA + MLA hybrid attention design behind the efficiency story and compares Ant's activation ratio with Moonshot AI's, while presenting the full Ling family lineage in a single table for context. It explicitly flags that the model's OpenRouter free-tier pro

OpenRouter

Coverage

Kilo's blog (July 23, 2026) announces that Ling 3.0 Flash from inclusionAI (Ant Group) is available on the Kilo platform, with a limited-time free access promotion for developers building multi-turn coding agents or daily agentic workflows on tight token budgets. The piece frames Ling 3.0 Flash within the "flash model The Kilo post repeats the core architectural claims — 124B total parameters with 5.1B active per token, optimization for high-throughput token-efficient inference, and a native 256K context window extendable to 1M for long-horizon task stability and deep repository reasoning — and reports that early testing found the m

Videos about Ling 3.0 Flash

More models around Ling 3.0 Flash