Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

Ling 3.1 Flash

Ling 3.1 Flash is a large mixture-of-experts model released by Ant Group's InclusionAI, built around a 560 billion total parameter design with roughly 25 billion parameters activated per token. That sparsity profile lets it deliver substantial capacity while keeping per-token compute closer to a mid-sized dense model, which is a common pattern for agent and productivity workloads where throughput matters. InclusionAI has positioned the release for agent tasks, search, office software, and other specialist applications, suggesting a focus on tool-using, instruction-following behavior rather than open-ended creative generation.

The model is designed to handle very long contexts, with a stated target window of up to one million tokens for the underlying architecture. At launch, a two-week free trial exposed it through a 256,000-token cap, and InclusionAI has indicated it plans to unlock the full one-million-token window and publish the weights as open source once that trial concludes. For practitioners, this means Ling 3.1 Flash is best understood today as a frontier-scale MoE being staged for broad availability, with its long-context and open-weights potential as the main forward-looking advantages once the broader release lands.

Vercel AI Gatewayinclusionai/ling-3.1-flashling

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
inclusionai/ling-3.1-flash
Release date
Sep 29, 2026
Last updated
Sep 29, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Latest news about Ling 3.1 Flash

Vercel AI Gateway

Coverage

InclusionAI, part of Ant Group, released Ling-3.1-flash as a mixture-of-experts language model with 560 billion total parameters and roughly 25 billion activated per token, positioning it for agent tasks, search, and office applications. The model is designed for a one-million-token context window, but the initial two-week free trial is capped at 256,000 tokens. InclusionAI stated plans to raise the context cap to its full one-million-token design and to release the model as open source once the trial period ends. The announcement, summarized by AI Weekly from a TechNode report, did not include benchmark numbers or name a spokesperson, leaving the open-source timeline and expanded context window as unconfirmed future commitments rather than completed events.

Videos about Ling 3.1 Flash

More models around Ling 3.1 Flash