Currently listed through these providers:
Model details
Qwen Turbo
Qwen Turbo traces its lineage to the Qwen2.5 architecture, using the original Qwen tokenizer that preceded the newer Qwen3 tokenizer found in later releases. This foundation positions the model as a straightforward, text-only system built for accessible deployment rather than frontier-level capability. Its design emphasizes speed and affordability, making it well-suited for straightforward tasks that demand quick turnaround rather than complex, multi-step reasoning. The proprietary approach keeps the model optimized for controlled environments while focusing resources on inference efficiency rather than open-weight distribution.
The model scores 12 on the Artificial Analysis Intelligence Index, placing it above the average of 10 for comparable models, though it trades some raw intelligence for speed and cost-effectiveness. Performance metrics show it averages around 75 tokens per second with roughly one-second latency to first token, numbers that fall below category averages but remain adequate for simpler workloads. Built for developers needing multilingual support with tiered pricing flexibility, Qwen Turbo serves as a practical entry point in the Qwen family—capable enough for routine tasks while reserving more advanced reasoning modes for newer iterations in the model family.
Quick Info
Powered by- Provider
- Alibaba (China)
- Model key
- qwen-turbo
- Release date
- Nov 1, 2024
- Last updated
- Jul 15, 2025
- Knowledge cutoff
- 2024-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.044
- Output token cost
- $0.087
Limits
- Output tokens
- 16,384 tokens
- Context window
- 1,000,000 tokens
Transparent token rates
Compare Qwen Turbo pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Qwen Turbo
No articles yet. Fetch the latest news to show it here.