Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Ofox logo

Model details

Qwen Flash

Qwen Flash is one of three paid API tiers offered through Alibaba Cloud Model Studio, sitting alongside the flagship Qwen-Max and the balanced Qwen-Plus. It is explicitly positioned as the cheap high-volume option within that lineup, designed for workloads where cost per token matters more than maximum capability. The tier belongs to the broader split in how Alibaba exposes its models: a free set of open-weight downloads under Apache 2.0 versus a hosted paid API, with Qwen Flash firmly in the hosted paid camp.

Its token pricing reflects that high-volume positioning, coming in well below typical flagship rates. The tier is best understood as an entry point for teams who want Alibaba-hosted inference without paying flagship prices, making it a practical fit for routing routine queries, batch summarization, and large-scale conversational traffic where budget dominates over top-end reasoning performance.

Ofoxqwen/qwen-flashqwen

Quick Info

Powered by
Provider
Ofox
Model key
qwen/qwen-flash
Release date
Jul 28, 2025
Last updated
Jul 28, 2025
Knowledge cutoff
2024-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.022
Output token cost
$0.22

Limits

Output tokens
32,000 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare Qwen Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen Flash

No articles yet. Fetch the latest news to show it here.

Videos about Qwen Flash

More models around Qwen Flash