Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

Ling 3.0 Flash VL (DeepInfra)

Ling 3.0 Flash VL extends the Ling 3.0 Flash family into multimodal territory, pairing native text, image, and video understanding with a Mixture-of-Experts backbone that activates approximately 5.5 billion parameters per token out of a total 124 billion. This sparse activation pattern keeps per-request compute low while preserving access to a large parameter pool, making the model well suited to workloads where both capability and inference efficiency matter, such as document analysis with embedded visuals, image-grounded reasoning chains, and agentic loops that interleave tool calls with visual inspection of intermediate results.

Positioned primarily for multimodal agentic workflows, long-context understanding, and multi-step reasoning, the model targets practical applications rather than benchmark-driven one-shot queries. Its large context window, served at fp16 precision on DeepInfra infrastructure, allows it to maintain coherence across extended documents that combine narrative text with diagrams, screenshots, or frames. Teams building assistants that must read, reason about, and act on visual inputs alongside conversational context will find this combination of sparse activation and broad modality support a natural fit for production pipelines.

LLM Gatewaydeepinfra/ling-3.0-flash-vlling

Quick Info

Powered by
Provider
LLM Gateway
Model key
deepinfra/ling-3.0-flash-vl
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.06
Output token cost
$0.18

Limits

Output tokens
32,768 tokens
Context window
131,072 tokens

Transparent token rates

Compare Ling 3.0 Flash VL (DeepInfra) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Ling 3.0 Flash VL (DeepInfra)

LLM Gateway

CoverageBenchmark

InclusionAI released Ling-3.0-flash-VL in September 2026 as an open-weights multimodal model under an MIT license, supporting text, image, and video inputs with text output. It uses a Mixture-of-Experts architecture with 124B total parameters and 5.5B active parameters per token during inference. The page reports a 262k token context window and reasoning capability, with weights available on Hugging Face. On the Artificial Analysis Intelligence Index the model scores 25, placing it well above the 8 median for comparable open-weights models in its size class, while generating a somewhat verbose 160M tokens. Throughput is 142.6 output tokens per second, faster than the 133 median. Pricing is reported at $0.075 per 1M input tokens and $0.22 per 1M output tokens, with an 80% cache discount.

Videos about Ling 3.0 Flash VL (DeepInfra)

More models around Ling 3.0 Flash VL (DeepInfra)