Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
OpenRouter logo

Model details

Ling 3.1 Flash

Ling 3.1 Flash is a large mixture-of-experts model released by Ant Group's InclusionAI, built around a 560 billion total parameter design with roughly 25 billion parameters activated per token. That sparsity profile lets it deliver substantial capacity while keeping per-token compute closer to a mid-sized dense model, which is a common pattern for agent and productivity workloads where throughput matters. InclusionAI has positioned the release for agent tasks, search, office software, and other specialist applications, suggesting a focus on tool-using, instruction-following behavior rather than open-ended creative generation.

The model is designed to handle very long contexts, with a stated target window of up to one million tokens for the underlying architecture. At launch, a two-week free trial exposed it through a 256,000-token cap, and InclusionAI has indicated it plans to unlock the full one-million-token window and publish the weights as open source once that trial concludes. For practitioners, this means Ling 3.1 Flash is best understood today as a frontier-scale MoE being staged for broad availability, with its long-context and open-weights potential as the main forward-looking advantages once the broader release lands.

OpenRouterinclusionai/ling-3.1-flashling

Quick Info

Powered by
Provider
OpenRouter
Model key
inclusionai/ling-3.1-flash
Release date
Oct 2, 2026
Last updated
Oct 2, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Latest news about Ling 3.1 Flash

Vercel AI Gateway

CoverageBenchmark

BuildFastWithAi's review describes Ling 3.1 Flash as InclusionAI's new hybrid reasoning model released September 29, 2026, a 560B-total-parameter MoE with about 25B active per token, positioned around coding, multi-step analysis, and tool-using agents. The review explicitly distinguishes the announced 1M-token design target from the Vercel AI Gateway-served configuration, which currently exposes 262,144 tokens of context and up to 32,768 output tokens, a difference the author flags as material for production deployment planning. The review consolidates InclusionAI's launch benchmark numbers, including 52.5% on AutomationBench, 68.7% on SkillsBench, 87.9% on CyberGym, 57.9% on Finance Agent v2, 85.5% on DRACO, 40.4% on Terminal-Bench 4, 55.9% on SWE Atlas Codebase QnA, and 65.3% on HealthBench Professional, while noting promotional Vercel access is mentioned through October 13, 2026. It frames the model as a substantial scale jump from Ling 3.0 Flash and warns that the Vercel-served context is materially smaller than InclusionAI's announced design ceiling.

Vercel AI Gateway

Coverage

ArtificialWatch's wire reports that Ant Group's Ling team announced Ling-3.1-flash on September 29, 2026, describing a Mixture-of-Experts model with roughly 560B total parameters and about 25B active per token, targeting a 1M-token context window, with weights "planned to open-source soon." The article also notes that inclusionai/ling-3.1-flash is already listed on the Vercel AI Gateway and on nano-gpt, and quotes Ant's own benchmark figures including 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, and 65.35 on HealthBench Professional. The same writeup flags a discrepancy: models.dev lists Ling 3.1 Flash with a 262,144-token context and 32,768-token max output on Vercel and nano-gpt, well below the 1M Ant describes, with the Vercel route currently free and nano-gpt charging $0.075/$0.22 per million tokens in/out. No Hugging Face repository under inclusionAI exists yet, and the model is not yet on OpenRouter, so the "open-source release" framing is premature. Ant has not published its own token pricing.

Vercel AI Gateway

CoverageRelease Notes

ZICQ's article covers the Ling-3.1-flash release by Ling Labs, citing approximately 560 billion total parameters and 25 billion active parameters per token in a MoE architecture designed for long-context workloads. It lists creator-published benchmark scores of 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, and 65.35 on HealthBench Professional, framing the model as broadly applicable across programming, healthcare, and document-processing domains. A 1-million-token context window is highlighted as the headline capability. The piece states Ling Labs plans to open-source the model after a two-week free usage period, presenting this as a forthcoming contribution to the open-source AI community, though the article does not address Vercel AI Gateway specifically or any current served deployment. Multi-domain strength, long-sequence handling, and the promised open-source release are positioned as the primary draws for developers evaluating cross-industry applications.

Vercel AI Gateway

CoverageBenchmark

The Ling 3.1 Flash tracker on BenchLM records the model as carrying a 262,000-token context window and no published first-party API price, with the field-median price shown as $1 input per million tokens. Eight published benchmark rows are available, led by an Agentic category at 81.0 (verified, 22% weight) across five benchmarks, while Coding (40% weight) and Knowledge (12%) rows are verified but still score-pending. Reasoning, Multimodal, Multilingual, Instruction-Following, and Math categories are listed as Not measured with zero benchmarks. BenchLM flags Ling 3.1 Flash as unranked, citing insufficient eligible comparative evidence and no first-party hosted token rate, and describes it as a "Radar7 confirmed" entry from the Sep 25 – Oct 1, 2026 window. The page is updated as of October 1, 2026 and tracks 637 models with 508 benchmarks overall, providing only a supplementary third-party benchmark view rather than any Vercel-specific configuration detail. Decision-snapshot framing is explicitly advisory, not an absolute quality ranking.

Vercel AI Gateway

Coverage

InclusionAI, part of Ant Group, released Ling-3.1-flash as a mixture-of-experts language model with 560 billion total parameters and roughly 25 billion activated per token, positioning it for agent tasks, search, and office applications. The model is designed for a one-million-token context window, but the initial two-week free trial is capped at 256,000 tokens. InclusionAI stated plans to raise the context cap to its full one-million-token design and to release the model as open source once the trial period ends. The announcement, summarized by AI Weekly from a TechNode report, did not include benchmark numbers or name a spokesperson, leaving the open-source timeline and expanded context window as unconfirmed future commitments rather than completed events.

Videos about Ling 3.1 Flash

More models around Ling 3.1 Flash