Sulat.com
AI models
Get 10-25% off
Get 10-25% off from Qwen
Alibaba logo

Model details

Qwen3.5 Flash

Qwen3.5 Flash is the cloud-hosted variant of Alibaba's Qwen3.5 Medium model series, positioned as a feature-enhanced build on top of the Qwen3.5-35B-A3B mixture-of-experts checkpoint, which carries 35 billion total parameters with 3 billion active per token. The Medium series, which includes Flash alongside the open-source Qwen3.5-35B-A3B, Qwen3.5-122B-A10B, and Qwen3.5-27B releases, was published in late February 2026 by the Qwen team, and Flash itself is delivered through Alibaba Cloud's Model Studio rather than as a downloadable weights release, matching the catalog's closed-weight status. In the broader Medium lineup, the open-source checkpoints are released under the Apache 2.0 license and are reported to deliver competitive agentic tool-calling behavior, with Flash extending that capability surface in the hosted environment for developers who prefer an API integration path.

Flash inherits the MoE design philosophy that lets the underlying 35B/3B-active checkpoint route tokens through only a small slice of its parameters, which is well suited to latency-sensitive agentic flows where tool calls and short, structured responses dominate. The Medium series was positioned by third-party reporting as offering performance competitive with leading Western proprietary systems on common third-party benchmarks, and Flash carries that same underlying architecture into a managed, always-up-to-date cloud endpoint rather than a static weight download. For practitioners, this makes Flash a practical fit when an organization wants Qwen3.5 Medium-class behavior without standing up local inference infrastructure, while teams that need on-premises control can still take the sibling Qwen3.5-35B-A3B checkpoint and self-host it under Apache 2.0.

Alibabaqwen3.5-flashqwen

Quick Info

Powered by
Provider
Alibaba
Model key
qwen3.5-flash
Release date
Feb 23, 2026
Last updated
Feb 23, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.40

Limits

Output tokens
65,536 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare Qwen3.5 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.5 Flash

Alibaba

CoverageBenchmark

Alibaba's Qwen team released the Qwen3.5 Medium series on February 25, 2026, with Qwen3.5-Flash positioned as a proprietary variant available exclusively through the Alibaba Cloud Model Studio API, distinct from the three Apache-2.0 open-source siblings (Qwen3.5-35B-A3B, Qwen3.5-122B-A10B, Qwen3.5-27B) downloadable via Hugging Face and ModelScope. Flash offers a cost advantage over comparable Western proprietary models while remaining API-only. The Qwen3.5 Medium open models reportedly match or beat proprietary peers like GPT-5-mini and Claude Sonnet 4.5 on third-party benchmarks, while remaining accurate under quantization. The flagship Qwen3.5-35B-A3B achieves over 1 million tokens of context on consumer GPUs with 32GB VRAM using near-lossless 4-bit weight and KV cache quantization, enabled by a hybrid Gated Delta Networks plus sparse MoE architecture.

Videos about Qwen3.5 Flash

More models around Qwen3.5 Flash