Currently listed through these providers:
Model details
Qwen3.6 Flash
Qwen3.6 Flash is positioned within Alibaba's Qwen 3.6 series as a fast, efficient language model optimized for responsive inference. Its combination of text, image, and video input handling within a single 1,000,000-token context window makes it well suited for tasks that mix long documents with rich visual references, such as video-grounded question answering, multimodal retrieval, and extended document analysis. The model's availability on third-party routing platforms alongside its vision-capable treatment in comparison interfaces suggests an emphasis on practical, deployable multimodal reasoning rather than purely research-oriented use.
The model is designed with production economics in mind, offering prompt caching with explicit cache read and cache creation pricing to reduce repeated-input costs, plus a tiered pricing structure that activates above 256K tokens to reflect the higher cost of very long contexts. Latency figures reported around 0.39 seconds and throughput near 85 tokens per second indicate a throughput-oriented profile that fits latency-sensitive applications such as conversational agents, batch summarization, and interactive vision tools. Practically, Qwen3.6 Flash fits workflows that need a balance between multimodal breadth, very long context capacity, and cost-efficient repeated queries, making it a pragmatic choice for teams building large-context assistants, multimodal pipelines, and caching-heavy inference workloads.
Quick Info
Powered by- Provider
- DevPass (LLM Gateway)
- Model key
- qwen3.6-flash
- Release date
- Apr 27, 2026
- Last updated
- Apr 27, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.25
- Output token cost
- $1.50
Limits
- Output tokens
- 65,536 tokens
- Context window
- 1,000,000 tokens
Latest news about Qwen3.6 Flash
Videos about Qwen3.6 Flash
More models around Qwen3.6 Flash
This exact model name is also listed by 20 other providers.