AIHubMix
Alibaba Unveils Qwen3.5-Plus, Undercutting Gemini 3 Pro on Cost - Chinese tech giant says new open-source AI model rivals Google’s flagship system while charging a fraction of the API price
Model details
Qwen3.5 Plus represents Alibaba's architectural pivot toward efficient large-scale intelligence, pairing a hybrid design that fuses linear attention through Gated Delta Networks with a sparse mixture-of-experts framework. This approach delivers the representational power of a 397-billion parameter model while activating only 17 billion parameters per forward pass, dramatically improving inference speed and cost compared to dense alternatives. Engineered as a native vision-language model from inception rather than retrofitted, it accepts text, images, and video while producing text output, and arrives purpose-built for the agentic AI era with integrated tool calling and a million-token context window to handle lengthy, multi-step workflows.
The model emerges from Alibaba's continued refinement of the Qwen family, with the base weights open-sourced as Qwen3.5-397B-A17B while the Plus tier offers a managed experience via cloud hosting. Its expanded language support across 201 languages and dialects broadens applicability for global deployments. The architecture's efficiency focus makes it particularly suited for scenarios requiring sustained reasoning across long documents, complex agentic tasks, and cost-sensitive production environments, positioning it as a competitive alternative to frontier models at significantly lower operational expense while maintaining frontier-class benchmark performance.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
AIHubMix
Alibaba Unveils Qwen3.5-Plus, Undercutting Gemini 3 Pro on Cost - Chinese tech giant says new open-source AI model rivals Google’s flagship system while charging a fraction of the API price
AIHubMix
Alibaba has released Qwen3.5, its latest model series, starting with o
AIHubMix
The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. $0.26 per million input tokens, $1.56 per million output tokens. 1,000,000 token context window, maximum outp