Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Pioneer logo

Model details

Qwen3.6 35B-A3B

Qwen3.6-35B-A3B is a sparse Mixture-of-Experts model in the Qwen family, carrying 35 billion total parameters while activating only 3 billion per token. It is the first open-weight variant of the Qwen3.6 series, released by the Qwen team to bring the line's stability and real-world coding focus to self-hosted and community workflows. The weights ship as a post-trained Hugging Face Transformers artifact compatible with major inference stacks, including vLLM, SGLang, and KTransformers, making deployment straightforward for teams that want a capable coding model without managing proprietary APIs.

The model is tuned for agentic coding, with the Qwen team highlighting stronger handling of frontend workflows and repository-level reasoning, as well as a new Thinking Preservation option that retains reasoning context from earlier messages to streamline iterative development. Despite its compact active footprint, it is reported to surpass its predecessor Qwen3.5-35B-A3B and rival considerably larger dense models such as Qwen3.5-27B and Gemma4-31B on agentic coding tasks. It continues to support multimodal thinking and non-thinking modes alongside text output, positioning it as a versatile open-source choice for developers building coding assistants, multimodal agents, and other productivity tools.

PioneerQwen/Qwen3.6-35B-A3Bqwen

Quick Info

Powered by
Provider
Pioneer
Model key
Qwen/Qwen3.6-35B-A3B
Release date
Apr 17, 2026
Last updated
Apr 17, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.14
Output token cost
$1.00

Limits

Output tokens
131,072 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.6 35B-A3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.6 35B-A3B

OpenRouter

CoverageBenchmark

The OpenRouter listing for qwen/qwen3.6-35b-a3b confirms the model as an open-weight multimodal release from Alibaba Cloud with 35 billion total parameters and 3 billion active per token, using a hybrid sparse MoE that combines Gated DeltaNet linear attention with standard gated attention layers for efficient inference The listing records the model's OpenRouter release date as April 27, 2026, with listed pricing of $0.05 per 1M input tokens and $0.70 per 1M output tokens. OpenRouter also displays a weighted-average effective price based on what customers actually pay across its routed providers, along with a price-history chart spann

OpenRouter

Coverage

On April 2, 2026, Alibaba's Qwen team open-sourced Qwen3.6-35B-A3B as the first open-weight variant of the Qwen3.6 generation, releasing it under Apache 2.0 on Hugging Face alongside the proprietary Qwen3.6-Plus API model. The page frames the release with the tagline "Agentic Coding Power, Now Open to All," noting the Architecturally, Qwen3.6-35B-A3B is a sparse MoE with 256 experts where 8 routed plus 1 shared expert activate per token, yielding 3B active parameters out of 35B total. The 40-layer stack uses a repeating block of three Gated DeltaNet (linear attention) layers followed by one Gated Attention layer, each paired with an

Videos about Qwen3.6 35B-A3B

More models around Qwen3.6 35B-A3B