Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
DevPass (LLM Gateway) logo

Model details

Ling 3.0 Flash VL

Ling-3.0-flash-VL is Ant Group's multimodal entry in the Ling 3.0 flash family, designed to process images and video alongside text input. It was released as a sibling variant to the text-focused Ling-3.0-flash base model and precedes the finance-tuned Ling-3.0-flash-Fin, giving the family a clear split between general multimodal understanding and domain-specific text reasoning. This positioning makes the VL variant suitable for tasks that require visual grounding, such as document understanding, visual question answering, and interpreting charts or screenshots within a broader workflow.

On the Artificial Analysis Intelligence Index, Ling-3.0-flash-VL scores 25, placing it in a competitive mid-tier band for multimodal models of its generation. Independent reporting also notes that it reaches 1225 Elo on GDPval-AA v2 and 986 on AA-Briefcase, benchmarks that evaluate agents on professional knowledge work, suggesting the model handles structured, multi-step professional tasks reasonably well. Compared with the later Ling-3.0-flash-Fin variant, the VL model shows lower hallucination on business knowledge (19% versus 33%) while offering broader visual input coverage, making it a better fit when input fidelity and grounded interpretation matter more than specialized financial vocabulary.

DevPass (LLM Gateway)ling-3.0-flash-vlling

Quick Info

Powered by
Provider
DevPass (LLM Gateway)
Model key
ling-3.0-flash-vl
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.06
Output token cost
$0.18

Limits

Output tokens
131,072 tokens
Context window
131,072 tokens

Transparent token rates

Compare Ling 3.0 Flash VL pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Ling 3.0 Flash VL

DevPass (LLM Gateway)

CoverageRelease Notes

The Artificial Analysis article positions Ling-3.0-flash-VL as a comparator while covering the finance-focused Ling-3.0-flash-Fin release. Ling-3.0-flash-VL scored 25 on the Artificial Analysis Intelligence Index and shares a 256K token context window with the Fin sibling built on the same MoE base. On the Finance & Accounting Index, Ling-3.0-flash-Fin matched Ling-3.0-flash-VL at 24, with the Fin variant achieving higher business knowledge accuracy (17% vs. 11%) but also higher hallucination (33% vs. 19%). Ling-3.0-flash-VL scored 1225 Elo on GDPval-AA v2 and 986 on AA-Briefcase for professional agent tasks, and 16% on AutomationBench-AA.

DevPass (LLM Gateway)

CoverageRelease Notes

Ant Group's inclusionAI lab released Ling-3.0-flash-VL on 10 September 2026 as an open-weights vision-language model under the MIT licence. It is a mixture-of-experts design with 124 billion total parameters, activating 5.5 billion per token, and ships alongside fp8, fp4, and int4 quantizations. The model accepts text, images, and video inputs and holds a 262,144-token context window, making it the multimodal counterpart to the earlier text-only Ling-3.0-flash. The architecture combines a ViT encoder with a small two-layer projector and VideoRoPE positional encoding, enabling timestamp-aware queries over long video. Its 42-layer backbone alternates Kimi Delta Attention and Gated MLA at a five-to-one ratio, keeping long-context memory costs manageable. Thinking mode is on by default but can be disabled, and inclusionAI positions the release as strengthening language reasoning while adding visual perception and visual-agent capabilities over its text predecessor.

DevPass (LLM Gateway)

Coverage

The Baidu Baike entry corroborates the September 9, 2026 release of Ling-3.0-flash-VL by Ant Group as the Bailing series' first native multimodal model. BF16 and FP8 weights were open-sourced on Hugging Face and ModelScope, with a free trial available on Ling Studio, and a 256K token context window. The architecture uses an arbitrary-resolution visual encoder, VideoRoPE, and a 42-layer hybrid language backbone alternating KDA and Gated MLA at a 5:1 ratio. The visual feedback closed-loop mechanism continuously corrects outputs in medical, frontend, and GUI automation tasks. On Artificial Analysis Intelligence Index v4.1.1, it improved by 4 points over the text-only version.

DevPass (LLM Gateway)

Coverage

Ant Group's open-source team released and open-sourced Ling-3.0-flash-VL on September 9, 2026, as the first native multimodal large model in the Bailing series. The article from aibase confirms the model is built on the MoE architecture of Ling-3.0-flash and natively supports image, text, and video inputs with a 256K token context window. Ling-3.0-flash-VL carries 124B total parameters with 5.5B activated per inference, inheriting Ling-3.0-flash's Agent-execution efficiency. It introduces a visual feedback closed-loop mechanism ("Observe → Act → Verify → Correct") for scenarios such as medical report interpretation, frontend code generation, and GUI automation, significantly improving result reliability.

Videos about Ling 3.0 Flash VL

More models around Ling 3.0 Flash VL