Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Requesty logo

Model details

ling-3.0-tiny

Ling-3.0-tiny is described by community distributions as a lightweight hybrid reasoning mixture-of-experts model, combining roughly 7.9 billion total parameters with only about 1.3 billion activated per token. That sparse-activation design aims to push the open-weights reasoning frontier further per active parameter, while keeping each forward pass small enough to run efficiently on modest hardware. The model is positioned primarily as a reasoning and agentic system, with the activation pattern suggesting an architecture tuned for step-by-step problem decomposition rather than broad general-purpose chat.

Practically, Ling-3.0-tiny is pitched at developers who want reasoning-style behavior without the cost of a dense 70B-class model, with the Ollama community tag shipping at roughly 5.3 GB and exposing a 128K-token text context for local experimentation. That footprint, paired with the low active-parameter count, makes it a reasonable fit for latency-sensitive agent loops, tool-augmented assistants, and on-device prototypes where reasoning quality matters more than raw knowledge breadth. The underlying model is attributed to Ant Group, while third-party re-uploads on platforms like Ollama broaden access beyond any single hosted endpoint.

Requestyling-3.0-tinyling

Quick Info

Powered by
Provider
Requesty
Model key
ling-3.0-tiny
Release date
Aug 5, 2026
Last updated
Aug 5, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Latest news about ling-3.0-tiny

Requesty

CoverageRelease Notes

The Opper AI InclusionAI release tracker lists "Ling 3.0 Tiny" as an InclusionAI model released on August 6, 2026, appearing in the August 2026 section alongside Ling-3.0-flash and the later September 2026 Ling-3.0-flash-Fin and Ling-3.0-flash-VL releases. This confirms InclusionAI as the releasing organization and est The tracker also documents the broader InclusionAI/Ling family timeline (Ling 3.0 Tiny, Ling 3.0 Flash, Ling-3.0-flash-VL, Ling-3.0-flash-Fin, plus prior Ling-2.6 and Ring releases), situating Ling 3.0 Tiny within an active release cadence. For the Ling-3.0-tiny subject specifically, the page supplies attribution and a

Requesty

Coverage

Ling 3.0 Tiny is a lightweight hybrid reasoning model developed by InclusionAI, Ant Group's AI initiative, according to a third-party technical explainer published on August 17, 2026. The page explicitly names the exact "Ling 3.0 Tiny" variant and describes it as a mixture-of-experts model with roughly 7.9 billion tota The explainer reports a context window of up to 256K tokens, native function calling, a configurable thinking mode, and multiple deployment formats including BF16, FP8, and INT4 weights. Citing model-card figures, it lists an Artificial Analysis Intelligence Index score of 25, an Agentic Index score of 16, output speed

Requesty

CoverageBenchmark

The Artificial Analysis page for Ling 3.0 Tiny confirms it as an InclusionAI open-weights model released in August 2026, with 7.9B total parameters and 1.3B active parameters, a 262k token context window (approximately 393 A4 pages), text-in/text-out modalities, reasoning support (with a noted possible non-reasoning va The page also benchmarks Ling 3.0 Tiny against its specific peer cohort: non-reasoning models versus reasoning models, and open-weights models within the Tiny class (≤4B parameters). Its technical specifications section is model-focused rather than gateway-focused, describing the model's reasoning capability, input/out

Requesty

Coverage

The 36Kr/Zhidx article reports that on August 11, 2026, Ant Group's Bailian large model team open-sourced Ling-3.0-tiny, a lightweight hybrid-inference MoE model with 7.9B total parameters and 1.3B activated per token, released in BF16, FP8, and INT4 versions verified on DGX Spark, MacBook, and Mac mini. It cites an Ar The article additionally describes Ling-3.0-tiny's hybrid linear architecture as a 3:1 alternating stack of Moonshot AI's KDA (Kimi Dilation Attention) and DeepSeek's MLA (Multi-head Latent Attention), paired with a sparse MoE feed-forward of 128 routing experts activating 8 routing plus 1 shared expert per token, and

Requesty

Coverage

The Baidu encyclopedia entry identifies Ling-3.0-tiny as a native hybrid reasoning model launched by Ant Ling, with a total of 7.9B parameters and only 1.3B activated per token, optimized for mathematical reasoning, instruction following, and resource-sensitive deployment. It states that on August 11, 2026 the Ling Lar The page provides concrete technical and deployment framing for Ling-3.0-tiny that is model-focused rather than gateway-related: it emphasizes the parameter efficiency of activating only 1.3B out of 7.9B per token, and highlights that the model is intended to operate without cloud dependence. Together with its August 1

Videos about ling-3.0-tiny

More models around ling-3.0-tiny