Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

InclusionAI Ling 3.0 Flash (DeepInfra)

InclusionAI's Ling 3.0 Flash, distributed through DeepInfra, is positioned as a compact text model in the broader ling family, aimed at lightweight assistant use cases such as extraction, routine automation, and fast conversational support. The aggregator profile describes it as an efficient, low-latency option rather than a flagship reasoning system, suggesting the focus is on dependable throughput for predictable workloads rather than maximum capability. Being text-only on both input and output, it is well matched to structured pipelines, retrieval-backed assistants, and tool-mediated workflows where multimodal grounding is not required.

The model is listed as generally available with closed weights and is available through the LLM Gateway channel with a large 262,144-token context window and a 32,768-token output ceiling, which is generous for a flash-tier model and useful for long-context extraction or document-grounded tasks. Its documented capabilities include reasoning, tool use, structured output, and temperature control, making it a flexible drop-in for agent-style applications that need to balance cost, latency, and instruction-following. The pricing profile, with a notably cheaper cache-read tier relative to its input and output rates, encourages prompt caching patterns, which can be a practical cost optimization for repeated context such as system prompts, retrieved documents, or conversational history.

LLM Gatewaydeepinfra/ling-3.0-flashling

Quick Info

Powered by
Provider
LLM Gateway
Model key
deepinfra/ling-3.0-flash
Release date
Aug 2, 2026
Last updated
Aug 2, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.06
Output token cost
$0.18

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Transparent token rates

Compare InclusionAI Ling 3.0 Flash (DeepInfra) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about InclusionAI Ling 3.0 Flash (DeepInfra)

No articles yet. Fetch the latest news to show it here.

Videos about InclusionAI Ling 3.0 Flash (DeepInfra)

More models around InclusionAI Ling 3.0 Flash (DeepInfra)