Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
302.AI logo

Model details

Qwen3.6 35B-A3B

As the first open-weight variant of the Qwen3.6 family, this sparse mixture-of-experts language model pairs 35 billion total parameters with roughly 3 billion active parameters, producing strong performance at a modest inference footprint. It is built on a hybrid stack that layers Gated DeltaNet linear attention blocks with Gated Attention blocks wrapped around MoE layers, allowing it to handle long contexts while keeping per-token compute low. The release comes as post-trained weights in the Hugging Face Transformers format and is compatible with vLLM, SGLang, and KTransformers, making it straightforward to integrate into existing open-source serving pipelines. A vision encoder is attached, so the model can jointly process text, images, video frames, and audio alongside its causal language modeling path.

The model is shaped around agentic coding workflows, with notable gains in frontend handling and repository-level reasoning that let it perform in the same range as considerably larger dense competitors. To support iterative development, the Qwen team introduced a thinking-preservation option that retains reasoning context across turns, reducing redundant re-thinking during multi-step coding sessions. The model can operate in both multimodal thinking and non-thinking modes, giving developers a choice between deeper deliberation and faster responses depending on the task. Distributed through Hugging Face and Model Scope under the Qwen organization, it offers a practical path for teams that want a capable, multimodal, open-weights coding model without paying the cost or latency of a much larger dense system.

302.AIqwen3.6-35b-a3bqwen

Quick Info

Powered by
Provider
302.AI
Model key
qwen3.6-35b-a3b
Release date
Apr 17, 2026
Last updated
Apr 17, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.283
Output token cost
$1.705

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.6 35B-A3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.6 35B-A3B

302.AI

CoverageBenchmark

Millstone AI provides detailed FP8 inference benchmark data for Qwen3.6-35B-A3B, confirming the model is a 35B-parameter MoE (256 experts, 8 routed + 1 shared active per forward pass, 3B activated) with an integrated vision encoder accepting text, images, and video. The architecture combines Gated DeltaNet and Gated At Millstone's hardware testing shows meaningful throughput differences across configurations: 1x MI300X (192GB) achieves 256 tok/s peak with 1-3 concurrent users and 3 chatbot users at 32K context, while 1x RTX Pro 6000 Blackwell (96GB) reaches 449 tok/s peak with 1-5 concurrent users and 41 chatbot users at 32K context,

302.AI

Coverage

On April 2, 2026, Alibaba's Qwen team open-sourced Qwen3.6-35B-A3B as the first open-weight variant of the Qwen3.6 generation, releasing it under Apache 2.0 on Hugging Face alongside the proprietary Qwen3.6-Plus API model. The 35-billion-parameter Mixture-of-Experts model activates just 3B parameters per token (256 exp The model's 40-layer stack uses a repeating block of three Gated DeltaNet (linear attention) layers followed by one Gated Attention layer, each paired with an MoE feed-forward block, with a hidden dimension of 2048 and expert intermediate dimension of 512. It was trained with Multi-Token Prediction (popularized by Deep

Videos about Qwen3.6 35B-A3B

More models around Qwen3.6 35B-A3B