As the first open-weight variant of the Qwen3.6 family, this sparse mixture-of-experts language model pairs 35 billion total parameters with roughly 3 billion active parameters, producing strong performance at a modest inference footprint. It is built on a hybrid stack that layers Gated DeltaNet linear attention blocks with Gated Attention blocks wrapped around MoE layers, allowing it to handle long contexts while keeping per-token compute low. The release comes as post-trained weights in the Hugging Face Transformers format and is compatible with vLLM, SGLang, and KTransformers, making it straightforward to integrate into existing open-source serving pipelines. A vision encoder is attached, so the model can jointly process text, images, video frames, and audio alongside its causal language modeling path.
The model is shaped around agentic coding workflows, with notable gains in frontend handling and repository-level reasoning that let it perform in the same range as considerably larger dense competitors. To support iterative development, the Qwen team introduced a thinking-preservation option that retains reasoning context across turns, reducing redundant re-thinking during multi-step coding sessions. The model can operate in both multimodal thinking and non-thinking modes, giving developers a choice between deeper deliberation and faster responses depending on the task. Distributed through Hugging Face and Model Scope under the Qwen organization, it offers a practical path for teams that want a capable, multimodal, open-weights coding model without paying the cost or latency of a much larger dense system.