Sulat.com
AI models
OpenRouter logo

Model details

Ling 3.0 Flash VL

Ling 3.0 Flash VL is a native multimodal model released by inclusionAI, the AI organization associated with Ant Group, and distributed openly through the inclusionAI spaces on Hugging Face and ModelScope. It is positioned as the next-generation vision-language follow-on to Ling 3.0 Flash, inheriting that predecessor's language, reasoning, and long-context strengths while adding native image and video understanding. The model is aimed at workflows where visual information needs to feed directly into reasoning and acting, rather than being treated as a separate perception stage.

Under the hood, Ling 3.0 Flash VL combines a ViT visual encoder with a two-layer MLP projector that aligns image and video features with the language representation space, and it uses VideoRoPE to encode both spatial positions and temporal order so it can reason over visual change across frames. Its 42-layer hybrid backbone alternates KDA and Gated MLA layers at a 5:1 ratio, supporting long multimodal contexts while keeping compute efficient. With 124B total parameters and only about 5.5B activated per token, it targets practical deployments that need strong multimodal reasoning, video question answering, event localization, and agentic tool use over very long inputs, all in an open-weights package.

OpenRouterinclusionai/ling-3.0-flash-vlling

Quick Info

Powered by
Provider
OpenRouter
Model key
inclusionai/ling-3.0-flash-vl
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.06
Output token cost
$0.18

Limits

Output tokens
32,768 tokens
Context window
131,072 tokens

Transparent token rates

Compare Ling 3.0 Flash VL pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Ling 3.0 Flash VL

NanoGPT

CoverageBenchmark

Artificial Analysis has added a dedicated provider-benchmarking page for InclusionAI's Ling-3.0-flash-VL, identifying it as an open-weights model released in September 2026. The page is positioned to help users choose the best API provider for the variant, and explicitly names "Ling-3.0-flash-VL" as the subject, with I The page reports that only one provider is currently tracked for Ling-3.0-flash-VL (InclusionAI itself), with an output speed of 140.7 tokens/second, a time-to-first-token of 16.00 seconds, and a listed blended price of $0.00 per 1M tokens; cache-hit, input, output, and blended pricing charts are marked "No data availa

Videos about Ling 3.0 Flash VL

More models around Ling 3.0 Flash VL