Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

InclusionAI Ling 3.0 Flash (NovitaAI)

InclusionAI Ling 3.0 Flash is positioned as a fast, production-oriented model in the broader ling family, offered through an LLM Gateway tier with a 262,144-token context window and a 32,768-token output cap that suit long-running agent loops and retrieval-heavy tasks. The official framing describes it as a native hybrid-reasoning foundational model engineered for production-grade AI agent workflows, acting as a high-speed execution node that balances intelligence density and cost-efficiency. In practice, this makes it a fit for latency-sensitive assistants, structured extraction, and routine automation where quick, reliable responses matter more than maximum depth.

Ant Group's announcement emphasizes that Ling-3.0-Flash delivers strong performance across multiple core benchmarks while matching or surpassing models with two to three times its parameter scale, signaling an efficiency-first design rather than a raw scale story. The model is closed-weight and text-only, which makes it accessible primarily through hosted providers rather than local deployment, and it ships with reasoning, tool use, structured output, and temperature control capabilities that align well with agent orchestration. Teams that need a responsive, controllable model for extraction pipelines, automation backbones, or hybrid reasoning tasks will find its combination of large context, generous output budget, and agent-oriented feature set a practical middle ground between lightweight chat models and heavier frontier systems.

LLM Gatewaynovita/ling-3.0-flashling

Quick Info

Powered by
Provider
LLM Gateway
Model key
novita/ling-3.0-flash
Release date
Aug 2, 2026
Last updated
Aug 2, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.06
Output token cost
$0.18

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Transparent token rates

Compare InclusionAI Ling 3.0 Flash (NovitaAI) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about InclusionAI Ling 3.0 Flash (NovitaAI)

LLM Gateway

Coverage

The aimadetools complete guide, published the day after release, documents InclusionAI Ling 3.0 Flash as a 124B total parameter MoE model with approximately 5.1B parameters activated per token, released July 23, 2026 as the successor to Ling 2.6 Flash (104B total, 7.4B active). The headline innovation is a hybrid reaso The article provides a comparison framing useful for developers: at 5.1B active parameters, Ling 3.0 Flash sits in the same inference-cost class as Gemini 3.5 Flash-Lite but with substantially more total parameter capacity (124B vs much smaller dense competitors). Key improvements over Ling 2.6 Flash are enumerated as

LLM Gateway

Coverage

Vercel added InclusionAI's Ling 3.0 Flash to its AI Gateway on July 23, 2026 under the model slug inclusionai/ling-3.0-flash-free, offering a three-week free access window through August 3, 2026. The model is described as a Mixture-of-Experts architecture with 124B total parameters and approximately 5.1B active paramet For developers already using Vercel AI Gateway, the practical advice is to route a narrow class of agent tasks to inclusionai/ling-3.0-flash-free and benchmark it against existing low-cost models on output quality, latency, retries, and token count. Suggested workloads for evaluation include issue triage, repository su

LLM Gateway

Official sourceBenchmark

The LLM Gateway model page for Ling 3.0 Flash is the primary authoritative source for this routing entry, confirming it is marked STABLE and listing NovitaAI and DeepInfra as the two serving providers, each offering 262,144-token context, streaming, tool calling, and reasoning support (but not strict JSON schema enforc From a developer perspective, the page documents concrete capability flags relevant to production use: tool/function calling is supported, soft JSON output works but strict JSON schema is not enforced by the upstream model, and streaming and reasoning modes are available. Both NovitaAI and DeepInfra show identical pric

Videos about InclusionAI Ling 3.0 Flash (NovitaAI)

More models around InclusionAI Ling 3.0 Flash (NovitaAI)