LLM Gateway
Learn about the new Qwen3.5 series of models, covering the key features, costs, how to access, and how it compares to other similar models.
Model details
Qwen3.5-397B-A17B marks the debut release in the Qwen3.5 series, introduced by the Qwen team as an open-weight native vision-language model. Its hybrid architecture fuses linear attention through Gated Delta Networks with a sparse mixture-of-experts design, reaching roughly 397 billion total parameters while activating about 17 billion per forward pass. This balance lets the model target demanding multimodal workloads without the compute profile of a fully dense counterpart, and the same launch expanded language and dialect support to 201, broadening accessibility for global developers and enterprises.
According to the launch post, the model achieves strong results across reasoning, coding, agent, and multimodal understanding benchmarks, signaling clear intent for agent-style applications that mix text and visual inputs. Distribution is broad: weights are mirrored on Hugging Face and ModelScope under the Qwen organization, the repository is available on GitHub, and NVIDIA lists a corresponding container in its NGC catalog for NIM-based deployments. Together these channels make the model a practical fit for teams building tool-using assistants, code agents, and multimodal pipelines who want open weights alongside enterprise-ready serving paths.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
LLM Gateway
Learn about the new Qwen3.5 series of models, covering the key features, costs, how to access, and how it compares to other similar models.
DevPass (LLM Gateway)
Qwen3.5-397B-A17B is the first open-weights model released in Alibaba's Qwen3.5 series, announced as a native vision-language foundation model with 397B total parameters and 17B activated per forward pass. Weights were published on Hugging Face under Apache License 2.0, compatible with Transformers, vLLM, SGLang, and K The model unifies thinking and non-thinking behavior in a single checkpoint, operating in thinking mode by default with emitted content, and unlike Qwen3 drops the /think and /nothink soft switches. Early-fusion multimodal training achieves cross-generational parity with Qwen3 while outperforming Qwen3-VL on reasoning,
DevPass (LLM Gateway)
Qwen3.5 is a flagship large language model series launched by Alibaba on February 16, 2026, comprising two versions: Qwen3.5-Plus and Qwen3.5-397B-A17B. The series supports text and multimodal tasks, utilizing a Hybrid Attention Mechanism and a Sparse Mixture of Experts (MoE) architecture, and is optimized for logical The Qwen3.5-397B-A17B variant has 397 billion total parameters with only 17 billion activated, reducing GPU memory consumption for deployment by 60%. Alibaba announced the development plan in January 2026, with a code merge request appearing on Hugging Face on February 9 and showcases at major model events on February
LLM Gateway
LLM Gateway lists Qwen3.5 397B-A17B as a stable model offering a 262,144-token context window with native multimodal MoE architecture (397B total / 17B active parameters), supporting reasoning, vision, tool use, streaming, and JSON output. The page shows gateway-tiered pricing starting at $0.17/M input and $1.03/M outp Notably, the listed gateway prices (~$0.17/$1.03 per million tokens) are substantially cheaper than direct provider pricing ($0.60/$3.60), making LLM Gateway a cost-effective routing layer for accessing this model across multiple upstream providers with unified API access.