Currently listed through these providers:
Model details
Qwen3.6 35B-A3B
Qwen3.6-35B-A3B is a sparse Mixture-of-Experts language model from Alibaba's Qwen team and the first open-weight variant of the Qwen3.6 series, arriving after the February Qwen3.5 release and positioned for stability and real-world developer use. The architecture totals 35 billion parameters but activates only about 3 billion per token, a design highlighted in independent coverage as the reason the model can run on consumer laptops such as a MacBook. It is distributed on Hugging Face as a post-trained checkpoint compatible with Transformers, vLLM, SGLang, and KTransformers, and released under the Apache 2.0 license, which permits commercial use, modification, and redistribution without licensing fees.
The release is tuned around agentic coding workflows, with the Qwen team pointing to improved handling of frontend tasks and repository-level reasoning, plus a "Thinking Preservation" option that retains reasoning context from earlier messages to streamline iterative development. In reported benchmark numbers, it reaches 73.4% on SWE-Bench Verified, a result a third-party review framed as beating Claude Opus 4.7 in a local-inference setting and outpacing dense models like Gemma 4-31B. Together, those characteristics make the model a practical fit for developers who want strong coding and reasoning behavior from a self-hostable, open-weight checkpoint that fits on a single workstation.
Quick Info
Powered by- Provider
- DevPass (LLM Gateway)
- Model key
- qwen3.6-35b-a3b
- Release date
- Apr 17, 2026
- Last updated
- Apr 17, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.248
- Output token cost
- $1.485
Limits
- Output tokens
- 65,536 tokens
- Context window
- 262,144 tokens