Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
SiliconFlow (China) logo

Model details

Qwen/Qwen3.6-35B-A3B

As a smaller sibling within the Qwen3.6 family, this release is built as a sparse Mixture-of-Experts model with 35 billion total parameters but only around 3 billion active parameters per token, an efficiency-oriented design that lowers compute and memory demands without sacrificing capability. The open-weights distribution lets developers run the model locally or self-host it, with quantization-friendly GGUF builds that fit on consumer hardware such as 24GB-RAM Macs. It carries forward the lineage of Qwen3.5 while targeting coding and tool-driven workflows, positioning it as a practical option for repository-level reasoning and multi-step agentic tasks rather than as a general-purpose chatbot.

The model is tuned for agentic coding, combining a very large 262,144-token context window with native tool calling and reasoning parsing, which makes it well suited for long codebases, multi-file edits, and orchestrated developer assistants. Independent benchmark coverage highlights strong coding results, including a 73.4 score on SWE-bench Verified and a 51.5 score on Terminal-Bench 2.0, and notes that it can outperform dense models in its weight class while remaining competitive with much larger frontier systems. Recommended vLLM deployments use single-node tensor parallelism with features like auto tool-choice and reasoning-mode decoding, supporting flexible serving across a wide range of data-center GPUs as well as high-end workstation cards.

SiliconFlow (China)Qwen/Qwen3.6-35B-A3Bqwen

Quick Info

Powered by
Provider
SiliconFlow (China)
Model key
Qwen/Qwen3.6-35B-A3B
Release date
Apr 17, 2026
Last updated
Apr 17, 2026
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.23
Output token cost
$1.86

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen/Qwen3.6-35B-A3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen/Qwen3.6-35B-A3B

SiliconFlow (China)

CoverageBenchmark

An OpenVINO-toolkit Medium post details measured inference performance for Qwen3.6–35B-A3B when INT4-weight-compressed and run through OpenVINO GenAI on an Intel Core Ultra AI PC. The article confirms the model's architecture: a Mixture-of-Experts vision-language model with roughly 35B total parameters and about 3B act The post reports concrete deployment numbers including tokens per second, latency to first token, and how those metrics hold up as prompt length grows, providing developer-actionable data for running the model on Intel client hardware. It credits INT4 compression plus sparse activation as the combination that makes a 3

SiliconFlow (China)

Coverage

A developer-oriented analysis from TokenMix.ai (published May 25, 2026) covers the full Qwen 3.6 family tier lineup: Plus (2026-04-02), 35B-A3B (2026-04-16), Max-Preview (2026-04-20), 27B (2026-04-22), and Flash (April 2026). The 35B-A3B tier is highlighted as the open-weights Apache-2.0 variant with 262K context exten The article provides a routing and fallback playbook for the 35B-A3B variant, including a self-host versus API break-even analysis tied to its Apache-2.0 license, and routing guidance for chaining the cheaper 35B-A3B tier against the higher-cost Plus/Max-Preview tiers. It cautions that the "Preview" tag on Max-Preview

SiliconFlow (China)

Coverage

The news blog specialized in Japanese culture, odd news, gadgets and all other funny stuffs. Updated everyday.

SiliconFlow (China)

CoverageBenchmark

Qwen/Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba/Qwen with 35 billion total parameters and 3 billion active per token, using a hybrid sparse mixture-of-experts architecture that combines Gated DeltaNet linear attention with standard gated attention layers for efficient inference. It supports a 262K When routed through SiliconFlow on OpenRouter, Qwen3.6-35B-A3B is priced at $0.20 per 1M input tokens and $1.60 per 1M output tokens, with 1.27s P50 latency, 94 tokens/sec throughput, and 97.88% uptime. Across providers on OpenRouter, SiliconFlow sits in the mid-range for price and latency, while CoreWeave leads throug

Videos about Qwen/Qwen3.6-35B-A3B

More models around Qwen/Qwen3.6-35B-A3B