Alibaba
The news blog specialized in Japanese culture, odd news, gadgets and all other funny stuffs. Updated everyday.
Model details
Qwen3.6-35B-A3B is a sparse mixture-of-experts model built to balance high-level performance with operational efficiency. By utilizing 35 billion total parameters while activating only 3 billion per token, the architecture achieves a lightweight footprint that allows it to rival much larger dense models in complex environments. Its design centers on agentic coding, featuring specialized capabilities for repository-level reasoning, frontend workflows, and multi-step tool calling. This makes it a versatile choice for developers who require deep analytical power without the resource demands typically associated with dense, large-scale systems.
The model benefits from a post-training lineage that emphasizes stability and real-world utility, incorporating direct community feedback to refine its responsiveness. It features a sophisticated hidden layout that integrates gated DeltaNet and gated attention mechanisms, supporting both multimodal perception and a new thinking preservation option that retains reasoning context across historical messages. These advancements streamline iterative development and improve precision in coding tasks. With broad compatibility across standard frameworks and hardware, the model is positioned as a robust tool for building next-generation agentic workflows that demand both speed and logical depth.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Alibaba
The news blog specialized in Japanese culture, odd news, gadgets and all other funny stuffs. Updated everyday.
Alibaba
On April 2, 2026, Alibaba's Qwen team released Qwen3.6-35B-A3B as an open-weight, Apache 2.0 licensed variant of the Qwen3.6 generation on Hugging Face, paired with the closed Qwen3.6-Plus API model. The 35-billion-parameter Mixture-of-Experts model activates only 3B parameters per token and was framed with the tagline Qwen3.6-35B-A3B uses a sparse MoE with 256 experts (8 routed plus 1 shared activating per token), a 40-layer stack of three Gated DeltaNet linear-attention layers followed by one Gated Attention layer, a hidden dimension of 2048, an expert intermediate dimension of 512, and Multi-Token Prediction training for faster sp