Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Requesty logo

Model details

ling-2.6-flash

ling-2.6-flash is an instant instruct model from inclusionAI built around a Mixture-of-Experts design that keeps 7.4B of its 104B total parameters active per token. That sparse-activation pattern is positioned for real-world agents that need fast responses and high token efficiency, trading off raw model scale against per-request compute cost. Within the family, it sits below the trillion-parameter Ling-2.6-1T flagship and predates the later Ling-3.0-flash, which expanded the architecture to roughly 124B total parameters with about 5.1B active per token.

Practically, the model is aimed at instruction-following and agent-style workloads where latency and budget per call matter more than maximum reasoning depth. The sparse activation lets it return completions quickly while keeping the active compute footprint small relative to its total parameter count, which is useful for high-throughput pipelines serving tool-using assistants and automated workflows. Teams already invested in inclusionAI's stack will find it a middleweight option between compact instruct baselines and the heavier 1T-class flagship, with the newer Ling-3.0-flash available for projects that want to move to a more recent iteration of the same recipe.

Requestyling-2.6-flashling

Quick Info

Powered by
Provider
Requesty
Model key
ling-2.6-flash
Release date
Apr 21, 2026
Last updated
Apr 21, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.30

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Transparent token rates

Compare ling-2.6-flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about ling-2.6-flash

Requesty

CoverageRelease Notes

Ant Group officially announced the release of Ling-2.6-flash, a large language model built on a sparse Mixture-of-Experts architecture with 104 billion total parameters and 7.4 billion active parameters. The press release, published from Hangzhou on April 22, 2026, positions the model as prioritizing efficiency and rea According to Ant Group's release, Ling-2.6-flash achieved an Artificial Analysis Intelligence Index of 26 while consuming only 15 million output tokens, compared to over 110 million tokens for comparable models like Nemotron-3-Super, representing an 86% reduction in inference cost. The model achieves inference speeds o

Requesty

CoverageRelease Notes

Ant Group officially announced Ling-2.6-flash on April 22, 2026 (syndicated via BusinessWire to kdhnews.com), positioning it as a sparse Mixture-of-Experts large language model with 104 billion total parameters and only 7.4 billion active parameters per token. According to Artificial Analysis data cited in the release, The release highlights deployment efficiency, claiming up to 340 tokens per second prefill on a 4-card H20 setup, 2.2x the prefill throughput of Nemotron-3-Super, and a stable output speed of 215 tokens per second. Ling-2.6-flash is described as specifically enhanced for AI agent applications, combining a large knowled

Videos about ling-2.6-flash

More models around ling-2.6-flash