Sulat.com
AI models
Get 10-25% off from Qwen
Alibaba (China) logo

Model details

Qwen3.6 Flash

Qwen3.6 Flash is positioned as the speed- and efficiency-oriented tier of Alibaba's Qwen 3.6 family, with a 1,000,000-token context window that comfortably accommodates very long documents, extended video transcripts, or multi-turn agent workflows. The model accepts text, image, and video as inputs while producing text-only outputs, giving it genuine multimodal breadth without the heavier compute footprint of larger siblings. It is served exclusively through Alibaba's infrastructure via OpenRouter, which simply forwards requests to that single upstream provider rather than routing across alternatives, helping keep latency and throughput predictable for production deployments.

In practical terms, Qwen3.6 Flash is well suited to high-volume multimodal pipelines where responsiveness matters more than top-tier reasoning depth, such as video summarization services, image-grounded customer support, and long-context retrieval or document Q&A. Tiered pricing activates above 256K tokens, and prompt caching is fully supported with explicit cache creation and cache read pricing, which materially lowers the cost of repeated or iterative prompts against the same large context. The broader Qwen 3.6 release cycle in mid-to-late April 2026 brought both open-weight MoE variants and hosted models, and Flash fits into that lineup as the accessible, multimodal, latency-friendly option for teams that want to stay within the Alibaba ecosystem.

Alibaba (China)qwen3.6-flashqwen3.6

Quick Info

Powered by
Provider
Alibaba (China)
Model key
qwen3.6-flash
Release date
Apr 27, 2026
Last updated
Apr 27, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.1875
Output token cost
$1.125

Limits

Output tokens
65,536 tokens
Context window
1,000,000 tokens

Latest news about Qwen3.6 Flash

Videos about Qwen3.6 Flash

Recent tweets and retweets from Alibaba (China)

More models around Qwen3.6 Flash