Alibaba (China)
Alibaba has introduced the new Qwen 3.5 model series. It comprises four models: Qwen3.5-Flash, Qwen3.5-35B-A3B, Qwen3.5-122B-A10B and Qwen3.5-27B.
Model details
Qwen3.5 Flash is built on a hybrid architecture that weaves a linear attention mechanism together with a sparse mixture-of-experts design, allowing the model to route specialized computation through activated expert pathways without processing every parameter at every step. The architecture is further strengthened by Gated Delta Networks, which help the model selectively update and retain information, keeping inference lean while maintaining strong multi-modal fluency. Early fusion training on interleaved text, image, and video tokens lets the model develop a unified vision-language foundation that reaches cross-generational parity with the base Qwen3 family and surpasses earlier Qwen3-VL checkpoints on benchmarks spanning reasoning, coding, agentic tasks, and visual understanding.
The Flash variant serves as the production-optimized entry point to the broader Qwen3.5 series, corresponding directly to the open 35B-parameter Qwen3.5-35B-A3B checkpoint but enhanced with longer default context handling and a more refined tool-calling pipeline. Training combines supervised fine-tuning with scalable reinforcement learning to cultivate generalization across domains, and the model ships with official built-in tools suitable for low-latency agentic workflows in enterprise and developer settings. This positions Qwen3.5 Flash as a practical choice for applications that need fast, reliable multi-modal reasoning without the overhead of deploying or managing the raw model weights.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Alibaba (China)
Alibaba has introduced the new Qwen 3.5 model series. It comprises four models: Qwen3.5-Flash, Qwen3.5-35B-A3B, Qwen3.5-122B-A10B and Qwen3.5-27B.
Alibaba (China)
The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. $0.065 per million input tokens, $0.26 per million output tokens. 1,000,000 token context window, maximum output