Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

Ling 3.0 Flash VL

Ant Group's Ling-3.0-flash-VL extends the Ling-3.0-flash family into multimodal territory by accepting image and video inputs alongside text while producing text outputs, positioning it as a general-purpose sibling to the text-only Ling-3.0-flash-Fin that was built on the same base. Independent benchmark reporting places the model on the Artificial Analysis Intelligence Index with a score of 25, suggesting solid general reasoning capability relative to peers of comparable scale, and professional-task benchmarks such as GDPval-AA v2 (1225 Elo) and AA-Briefcase (986) further indicate competent performance on agents-style work like producing deliverables and managing structured tasks.

Because Ling-3.0-flash-VL is multimodal on the input side, it fits workflows that require grounding text generation in visual evidence, such as interpreting screenshots, diagrams, or short video clips for analysis, summarization, and reporting. The same independent reporting notes that a finance-focused derivative matches Ling-3.0-flash-VL on the Finance & Accounting Index, implying the base VL model carries usable domain knowledge that downstream fine-tunes can sharpen. Practically, the model is best suited for applications that need to combine textual and visual context in a single reasoning pass rather than for fully open-weight deployment, since the supplied evidence describes the sibling Fin variant as open weights without making the same claim about VL itself.

Vercel AI Gatewayinclusionai/ling-3.0-flash-vlling

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
inclusionai/ling-3.0-flash-vl
Release date
Sep 8, 2026
Last updated
Sep 8, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.075
Output token cost
$0.22

Limits

Output tokens
32,000 tokens
Context window
256,000 tokens

Transparent token rates

Compare Ling 3.0 Flash VL pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Ling 3.0 Flash VL

Vercel AI Gateway

CoverageRelease Notes

Ant Group's inclusionAI lab released Ling-3.0-flash-VL on 10 September 2026, a 124-billion-parameter open-weights mixture-of-experts model that activates 5.5 billion parameters per token and handles text, images, and video within a 262,144-token context window. The weights ship under the permissive MIT licence, with fp The vision pipeline combines a ViT encoder with a small two-layer projector and a VideoRoPE positional scheme that encodes temporal position, enabling queries that span long videos and retrieval of specific moments. The backbone alternates Kimi Delta Attention and Gated MLA at a 5:1 ratio across 42 layers, which the la

OpenRouter

CoverageBenchmark

Artificial Analysis's dedicated model page lists Ling-3.0-flash-VL as an open-weights InclusionAI model released in September 2026 with an AA Intelligence Index v4.1.1 score of 25, placing it well above the median for its open-weights size class. It records 144.6 output tokens per second (above the 134 median), 160 mil Technical specifications confirm a 124B-total-parameter MoE architecture with 5.5B active parameters per token and reasoning enabled on this variant. The page's comparison summary characterizes the model as "amongst the leading models in intelligence" within its open-weights size class and "faster than average," while

OpenRouter

CoverageBenchmark

Crafiq's model profile names Ling-3.0-flash-VL by inclusionAI, released September 2026, as the first vision-integrated model in the Ling-3.0-flash family, supporting text, image, and video inputs with text outputs for applications such as document intelligence, medical report interpretation, and agentic workflows with Visual processing is described as powered by a ViT encoder paired with a two-layer MLP projector that aligns visual features with the text embedding space, and VideoRoPE is credited with encoding both spatial and temporal frame positions to enable event localization, temporal reasoning, and visual change detection over

OpenRouter

CoverageRelease Notes

Artificial Analysis' changelog explicitly logs "New language model evaluation results available — Ling-3.0-flash-VL — Intelligence Index: 25" on 10 September 2026, directly naming inclusionAI's multimodal model and confirming its evaluation date. The changelog also records InclusionAI performance results for Ling-3.0-f This changelog entry provides a reliable, named reference to the subject model on its evaluation date, useful as a primary timestamp and benchmark anchor. It complements other Artificial Analysis surfaces by restating the Intelligence Index 25 figure for Ling-3.0-flash-VL, helping disambiguate the model from text-only

Videos about Ling 3.0 Flash VL

More models around Ling 3.0 Flash VL