Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
DevPass (LLM Gateway) logo

Model details

InclusionAI Ling 3.0 Flash

InclusionAI's Ling 3.0 Flash is a Mixture-of-Experts language model in the Ling family that pairs a 124B-parameter total capacity with roughly 5.1B parameters activated per token, making it a sparse design aimed at keeping inference costs low while preserving the reasoning depth of a much larger dense model. The architecture prioritizes token efficiency and production-scale agentic inference, so developers can deploy it for multi-step workflows, tool-mediated tasks, and long-running assistants without paying the full price of dense 100B+ parameter models. Its 262,144-token context window is positioned for handling extensive documents, long conversational histories, and large code or data inputs in a single pass.

Because only a small slice of experts activates for any given token, Ling 3.0 Flash is well-matched to high-throughput server use, including agent pipelines that combine reasoning with tool calling and long-context retrieval. The model is offered through multiple routing options on inference gateways, and a free variant is distributed for evaluation and prototyping. The combination of a long context, sparse activation, and a developer-oriented design makes Ling 3.0 Flash a practical choice for teams building agentic applications that need to reason over large inputs while keeping per-request compute and token costs manageable.

DevPass (LLM Gateway)ling-3.0-flashling

Quick Info

Powered by
Provider
DevPass (LLM Gateway)
Model key
ling-3.0-flash
Release date
Aug 2, 2026
Last updated
Aug 2, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.06
Output token cost
$0.18

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Transparent token rates

Compare InclusionAI Ling 3.0 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about InclusionAI Ling 3.0 Flash

DevPass (LLM Gateway)

Coverage

InclusionAI (Ant Group) released Ling 3.0 Flash as a 124-billion-parameter hybrid-reasoning Mixture-of-Experts model that activates only about 5.1B parameters per token, compressing the expert activation ratio from 1/32 in the previous generation to 1/64. The architecture pairs 512 routed experts with one shared expert For developers, the core message is that Ling 3.0 Flash combines "Flash" series latency characteristics with "Ring" series reasoning depth in a single pretrained architecture, prioritizing token efficiency and production-scale agentic inference. The hybrid-linear attention design keeps memory usage stable on long conte

DevPass (LLM Gateway)

Coverage

InclusionAI announced Ling 3.0 Flash on 2 August 2026, releasing MIT-licensed BF16, FP8, FP4, and INT4 weights alongside OpenRouter availability. The model is a 124B-total, 5.1B-active hybrid-linear MoE that alternates Kimi Delta Attention layers with gated MLA layers in a 5:1 ratio, targeting agent-style workloads wit The practical implication is that Ling 3.0 Flash is positioned less on headline leaderboard scores and more on architecture-and-cost fit for long-horizon agents: a small active footprint with sparse expert routing, MIT-licensed open weights for self-hosting or fine-tuning, and tool-use/search behavior tailored to agent

DevPass (LLM Gateway)

CoverageBenchmark

BenchLM's aggregator dashboard for Ling 3.0 Flash reports a capability score of 46.8/100 against a field median of 56.3, placing the model 140th of 230 ranked entries, with category-level ranks of 110/154 in Agentic (29th percentile, 7 verified benchmarks, score 42.3), 116/154 in Coding (25th percentile, 6 verified, 40 For practitioners, the dashboard functions as comparative-context data rather than a release announcement: the strongest published evidence is in instruction following (rank 63), while math coverage is limited (3 rows, not ranked) and several capability dimensions remain unmeasured. Methodology and first-party sourcing

Videos about InclusionAI Ling 3.0 Flash

More models around InclusionAI Ling 3.0 Flash