Currently listed through these providers:
Model details
Qwen Flash
Qwen Flash is one of three paid API tiers offered through Alibaba Cloud Model Studio, sitting alongside the flagship Qwen-Max and the balanced Qwen-Plus. It is explicitly positioned as the cheap high-volume option within that lineup, designed for workloads where cost per token matters more than maximum capability. The tier belongs to the broader split in how Alibaba exposes its models: a free set of open-weight downloads under Apache 2.0 versus a hosted paid API, with Qwen Flash firmly in the hosted paid camp.
Its token pricing reflects that high-volume positioning, coming in well below typical flagship rates. The tier is best understood as an entry point for teams who want Alibaba-hosted inference without paying flagship prices, making it a practical fit for routing routine queries, batch summarization, and large-scale conversational traffic where budget dominates over top-end reasoning performance.
Quick Info
Powered by- Provider
- Ofox
- Model key
- qwen/qwen-flash
- Release date
- Jul 28, 2025
- Last updated
- Jul 28, 2025
- Knowledge cutoff
- 2024-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.022
- Output token cost
- $0.22
Limits
- Output tokens
- 32,000 tokens
- Context window
- 1,000,000 tokens
Transparent token rates
Compare Qwen Flash pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Qwen Flash
No articles yet. Fetch the latest news to show it here.