Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Novita AI logo

Model details

Ling 3.0 Flash

We haven't written an overview of this model yet. New models can take a few days to gather enough reliable coverage, so check back soon.

Novita AIinclusionai/ling-3.0-flashling

Quick Info

Powered by
Provider
Novita AI
Model key
inclusionai/ling-3.0-flash
Release date
Jul 23, 2026
Last updated
Jul 23, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.06
Output token cost
$0.18

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Transparent token rates

Compare Ling 3.0 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Ling 3.0 Flash

Vercel AI Gateway

Coverage

InclusionAI officially released the weights for Ling-3.0-flash on Hugging Face and ModelScope on August 5, 2026, with both BF16 and FP8 quantized versions available under the MIT license. This followed the initial API announcement in late July and made the 124B-total/5.1B-active MoE model practically downloadable for i Architecture analysis grounded in the official model card confirms the hybrid linear attention design combining Kimi Delta Attention and MLA in a 5:1 ratio, alongside the MoE compute configuration. The piece documents FP8 quantization retention characteristics and discusses adoption considerations for engineering teams

OpenRouter

CoverageBenchmark

Ant Group announced the release of Ling-3.0-Flash on July 27, 2026 (per Business Wire), describing it as a native hybrid-reasoning foundational model aimed at production-grade AI agent workflows and serving as a high-speed execution node balancing intelligence density and cost-efficiency. The model features 124B total Architecturally, Ling-3.0-Flash moves away from scaling parameter counts and is built from the ground up with a native hybrid-linear attention architecture that alternates KDA (Kimi Delta Attention) and MLA layers at a 5:1 ratio to balance long-context efficiency with robust state memory. Key architectural advancements

OpenRouter

Coverage

Digital Applied's blog (published July 24, 2026) provides secondary technical coverage of Ling-3.0-flash, confirming the 124B total / 5.1B active per token MoE design and noting that at publication the weights were not yet on Hugging Face — explicitly warning readers that aggregator listings often misattribute the mode The article documents a practitioner-legible explanation of the KDA + MLA hybrid attention design behind the efficiency story and compares Ant's activation ratio with Moonshot AI's, while presenting the full Ling family lineage in a single table for context. It explicitly flags that the model's OpenRouter free-tier pro

OpenRouter

Coverage

Kilo's blog (July 23, 2026) announces that Ling 3.0 Flash from inclusionAI (Ant Group) is available on the Kilo platform, with a limited-time free access promotion for developers building multi-turn coding agents or daily agentic workflows on tight token budgets. The piece frames Ling 3.0 Flash within the "flash model The Kilo post repeats the core architectural claims — 124B total parameters with 5.1B active per token, optimization for high-throughput token-efficient inference, and a native 256K context window extendable to 1M for long-horizon task stability and deep repository reasoning — and reports that early testing found the m

Videos about Ling 3.0 Flash

More models around Ling 3.0 Flash