Currently listed through these providers:
Model details
Qwen3.6 Flash
Qwen3.6 Flash is positioned as the speed- and efficiency-oriented tier of Alibaba's Qwen 3.6 family, with a 1,000,000-token context window that comfortably accommodates very long documents, extended video transcripts, or multi-turn agent workflows. The model accepts text, image, and video as inputs while producing text-only outputs, giving it genuine multimodal breadth without the heavier compute footprint of larger siblings. It is served exclusively through Alibaba's infrastructure via OpenRouter, which simply forwards requests to that single upstream provider rather than routing across alternatives, helping keep latency and throughput predictable for production deployments.
In practical terms, Qwen3.6 Flash is well suited to high-volume multimodal pipelines where responsiveness matters more than top-tier reasoning depth, such as video summarization services, image-grounded customer support, and long-context retrieval or document Q&A. Tiered pricing activates above 256K tokens, and prompt caching is fully supported with explicit cache creation and cache read pricing, which materially lowers the cost of repeated or iterative prompts against the same large context. The broader Qwen 3.6 release cycle in mid-to-late April 2026 brought both open-weight MoE variants and hosted models, and Flash fits into that lineup as the accessible, multimodal, latency-friendly option for teams that want to stay within the Alibaba ecosystem.
Quick Info
Powered by- Provider
- Alibaba (China)
- Model key
- qwen3.6-flash
- Release date
- Apr 27, 2026
- Last updated
- Apr 27, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.1875
- Output token cost
- $1.125
Limits
- Output tokens
- 65,536 tokens
- Context window
- 1,000,000 tokens
Latest news about Qwen3.6 Flash
Videos about Qwen3.6 Flash
Recent tweets and retweets from Alibaba (China)
More models around Qwen3.6 Flash
This exact model name is also listed by 20 other providers.