Currently listed through these providers:
Model details
Qwen3.6 35B-A3B
Qwen3.6-35B-A3B is a sparse Mixture-of-Experts model in the Qwen3.6 family, carrying 35 billion total parameters with only 3 billion active per pass, which lets it deliver strong capability at modest inference cost. It is the first open-weight variant of Qwen3.6, distributed through community repositories in the Hugging Face Transformers format and compatible with vLLM, SGLang, and KTransformers for flexible serving. The architecture pairs a vision encoder with a causal language model whose hidden layout interleaves Gated DeltaNet linear attention layers and Gated Attention layers with MoE blocks, blending efficient long-context processing with selective expert routing.
Designed for hands-on developer workflows, the model targets agentic coding tasks such as frontend generation and repository-level reasoning, with claimed improvements over the Qwen3.5-35B-A3B predecessor and competitive performance against larger dense models like Qwen3.5-27B and Gemma4-31B. It supports both multimodal thinking and non-thinking modes, and a new Thinking Preservation option retains reasoning context across turns to reduce overhead in iterative sessions. For practitioners, this combination of open weights, coding-oriented tuning, and multimodal reasoning makes it well suited to assistants and automation pipelines that need fluent repository-aware behavior without the footprint of a fully dense large model.
Quick Info
Powered by- Provider
- Hugging Face
- Model key
- Qwen/Qwen3.6-35B-A3B
- Release date
- Apr 17, 2026
- Last updated
- Apr 17, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $0.95
Limits
- Output tokens
- 65,536 tokens
- Context window
- 262,144 tokens