Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
GreenPT logo

Model details

Qwen3.6 35B-A3B

Qwen3.6 35B-A3B continues the Qwen family with a focus on stability and real-world utility, framing itself as a more intuitive and responsive coding companion for developers. The model carries a 35B-scale footprint tagged as a mixture-of-experts variant on third-party listings, which is consistent with its "A3B" naming and its design goal of keeping active compute modest while preserving broader capability. It supports vision input alongside text, is trained for tool use, and offers explicit reasoning support, making it suitable for multimodal understanding, agent-style workflows, and code-related tasks where structured decisions and tool orchestration matter.

Community availability landed in mid-April 2026, when the model and an FP8 variant were discussed on the NVIDIA DGX Spark / GB10 developer forums as a fresh candidate to evaluate. On LM Studio, the model has already drawn significant traction with multi-million downloads and active forking, suggesting strong grassroots interest among local developers. A minimum system memory of around 20 GB is recommended for running it locally, placing it within reach of well-equipped workstations while still benefiting from the efficiency that an MoE-style 35B design brings, so it fits well for teams seeking a balanced, open-weights model for coding assistance, reasoning pipelines, and vision-aware tool calling without resorting to the largest frontier models.

GreenPTqwen3.6-35b-a3bqwen

Quick Info

Powered by
Provider
GreenPT
Model key
qwen3.6-35b-a3b
Release date
Apr 17, 2026
Last updated
Apr 17, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.342
Output token cost
$2.052

Limits

Output tokens
32,768 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.6 35B-A3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.6 35B-A3B

GreenPT

Coverage

A practitioner guide by Minyang Chen (May 10, 2026) demonstrates running Qwen3.6-35B-A3B via llama.cpp on low-end consumer hardware — specifically 6 GB VRAM and 32 GB RAM — achieving approximately 30 tokens per second with a 256K-token context window. The model, released by Alibaba in April 2026 as a sparse Mixture-of- The article confirms the model's efficiency advantages for local deployment, noting that recent llama.cpp updates combined with Qwen3.6-35B-A3B's sparse architecture make it possible to run larger models efficiently on older CPUs and limited VRAM, achieving speeds comparable to smaller 2B/4B/7B models while supporting

GreenPT

Coverage

Alibaba's Qwen team open-sourced Qwen3.6-35B-A3B on April 2, 2026, releasing it under Apache 2.0 on Hugging Face alongside the proprietary Qwen3.6-Plus API. The 35-billion-parameter Mixture-of-Experts model activates just 3B parameters per token, posting frontier-level scores on agentic coding and reasoning benchmarks, The architecture is a sparse MoE with 256 experts (8 routed plus 1 shared activated per token), 40 layers combining Gated DeltaNet linear-attention and Gated Attention blocks, hidden dimension 2048, and expert intermediate dimension 512. It uses Multi-Token Prediction and introduces thinking preservation to retain reas

Videos about Qwen3.6 35B-A3B

More models around Qwen3.6 35B-A3B