Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
NovitaAI logo

Model details

Ling-2.6-flash

Ling-2.6-flash is an open-weight instruct model positioned as a fast, execution-focused system for real-world coding, document processing, and lightweight agent workflows. It carries 104 billion total parameters with 7.4 billion active per pass, a Mixture-of-Experts-style configuration that aims to deliver state-of-the-art-comparable performance while keeping token usage low. The model is released by inclusionAI with ties to Ant Group and is described as "instant," reflecting a design priority of low-latency responses rather than maximum depth on long-running reasoning tasks.

In practical terms, Ling-2.6-flash suits teams building agents that need to parse complex documents, write and edit code, or chain tool calls quickly without paying for unnecessary output tokens. The 262K context window supports substantial documents and multi-turn agent traces, and the model's emphasis on token efficiency makes it attractive for high-volume pipelines where per-call cost compounds. Kilo's hands-on testing reported consistent strengths across coding tasks, complex document parsing, and dynamic agentic workflows, aligning with the model's stated positioning as a responsive workhorse for production agent stacks.

NovitaAIinclusionai/ling-2.6-flashling

Quick Info

Powered by
Provider
NovitaAI
Model key
inclusionai/ling-2.6-flash
Release date
Apr 24, 2026
Last updated
Apr 24, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.30

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Transparent token rates

Compare Ling-2.6-flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Ling-2.6-flash

NovitaAI

CoverageRelease Notes

The KDH News republication of the BusinessWire press release confirms Ant Group's official announcement of Ling-2.6-flash as a sparse Mixture-of-Experts model with 104B total parameters and 7.4B active, designed for token efficiency over benchmark-padding via excessive output. Artificial Analysis figures cited in the r On a hybrid linear MoE stack the model achieves up to 340 tokens per second under 4-card H20 conditions, with Prefill throughput 2.2x that of Nemotron-3-Super, and a stable 215 tokens-per-second output speed placing it in the top tier of its size class. The release states Ling-2.6-flash has been specifically enhanced f

NovitaAI

CoverageBenchmark

AIBase news coverage from April 22, 2026 reports the official launch of Ling-2.6-flash by Ant Group's Bailing large model team, confirming the 104B total / 7.4B active parameter MoE architecture. The article cites Artificial Analysis evaluation data showing an Intelligence Index of 26 achieved using only 15M output tok The piece notes that before the official announcement the model was already deployed anonymously for a one-week stress test, during which daily token usage quickly rose to the 100B level, validating stability under real-world high concurrency. Industry analysts quoted in the article position the release as a shift from

Videos about Ling-2.6-flash

More models around Ling-2.6-flash