Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
QVAC logo

Model details

Qwen3.6 35B-A3B

Qwen3.6-35B-A3B is a fully open-source mixture-of-experts model with 35 billion total parameters and only 3 billion active parameters, positioning it as a sparse yet capable architecture optimized for efficient inference. Released by the Qwen team as part of the Qwen3.6 family, the model is available on Qwen Studio, through API access, and as open weights for community use on platforms like Hugging Face and ModelScope. Its design philosophy centers on delivering strong performance while maintaining a lightweight active parameter footprint during inference.

The model's standout strength lies in agentic coding capability, where it reportedly surpasses its predecessor Qwen3.5-35B-A3B by a wide margin and rivals much larger dense models including Qwen3.5-27B and Gemma4-31B. It supports both multimodal thinking and non-thinking modes, enabling flexible reasoning across text, image, video, and audio inputs while producing text outputs. This combination of efficient MoE architecture, multimodal versatility, and competitive coding performance makes it well-suited for developers building agentic systems that require sustained coding competence without the computational overhead of dense alternatives.

QVACqwen3.6-35b-a3bqwen

Quick Info

Powered by
Provider
QVAC
Model key
qwen3.6-35b-a3b
Release date
Apr 17, 2026
Last updated
Apr 17, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Latest news about Qwen3.6 35B-A3B

QVAC

Coverage

A practical Medium tutorial demonstrates running Qwen3.6-35B-A3B locally via llama.cpp on consumer hardware with just 6 GB of VRAM and 32 GB of RAM, achieving approximately 30 tokens per second. The article, published May 10, 2026, highlights that recent updates to llama.cpp combined with the model's sparse 35B/3B MoE The piece frames this as a significant improvement over prior setups where larger models required substantial VRAM and were impractical locally, while smaller models were limited to 32K–128K contexts. By leveraging Qwen3.6-35B-A3B's sparsity and recent llama.cpp optimizations, the author shows a path to running frontie

QVAC

CoverageBenchmark

BenchLM aggregates verified and provisional evidence for Qwen3.6-35B-A3B, recording category scores of 40.6 in agentic, 45.3 in coding, and 50.0 in multimodal. Reasoning, multilingual, and instruction following categories are listed as not measured. The dashboard ranks it 39th in multimodal and grounded workflows. The aggregator reports a 262K token context, median speed of 115 tokens per second, and a first-token latency of 48.94 seconds. Capability is scored at 43.6 out of 100 against a field median of 56.6, with no first-party API token rate available.

QVAC

CoverageBenchmark

Developer-reported benchmarks for Qwen3.6-35B-A3B list SWE-bench Verified at 73.4 percent resolved, SWE-bench Pro at 49.5 percent pass@1, Terminal-Bench 2.0 at 51.5 percent tasks resolved, LiveCodeBench at 80.4 percent pass@1, and MMLU-Pro at 85.2 percent accuracy. The model card attributes development to Alibaba's Qwen team. The model card reports a 262K token context window and an April 2026 release for the 35B/3B active MoE variant. A hardware fitting table shows FP16 weights near 70GB while IQ3 quantizations compress to about 14GB, enabling runs on consumer GPUs like the RTX 3060 12GB with partial CPU offload.

QVAC

Coverage

Alibaba's Tongyi Lab released Qwen3.6-35B-A3B as an open model on April 15, 2026, under the Apache 2.0 license. The launch followed the April 2, 2026 announcement of Qwen3.6-Plus and positions the variant for agentic coding workflows. Tongyi Lab markets it as stronger than Google's Gemma 4 open model suite. Qwen3.6-35B-A3B is a sparse mixture-of-experts model with 35 billion total parameters and only 3 billion active parameters, designed for efficient local inference. Simon Willison demonstrated the model drawing a pelican on a laptop that beat Claude Opus 4.7 in his qualitative test. Official details are documented on the qwen.ai blog linked from the coverage.

QVAC

Coverage

Alibaba's Qwen team open-sourced Qwen3.6-35B-A3B in April 2026 as the first open-weight variant of the Qwen3.6 generation, releasing weights under Apache 2.0 on Hugging Face alongside the proprietary Qwen3.6-Plus API model. The 35-billion-parameter sparse Mixture-of-Experts activates only 3B parameters per token and is Architecturally, Qwen3.6-35B-A3B uses 256 experts with 8 routed plus 1 shared expert activating per token in a 40-layer stack that repeats a 3:1 block of Gated DeltaNet linear-attention layers followed by a Gated Attention layer, each paired with an MoE feed-forward block, with a hidden dimension of 2048 and expert int

Videos about Qwen3.6 35B-A3B

More models around Qwen3.6 35B-A3B