Currently listed through these providers:
Model details
Qwen3.6 35B A3B
Qwen3.6-35B-A3B is a sparse mixture-of-experts model built around a hybrid Gated DeltaNet architecture that combines linear attention heads with standard gated attention layers. This design distributes computation across 256 experts per layer while activating only a small subset per token, keeping active parameter count lean while maintaining the representational depth of a much larger dense model. The architecture also preserves reasoning traces across multi-turn conversations, making it well suited for iterative coding tasks where context from earlier steps informs later decisions. By supporting both multimodal perception and dual thinking modes, the model can shift between fast responses and deliberate step-by-step reasoning depending on the task demands.
The model emerges as an open-weight release following the Qwen3.5 series, with community feedback shaping priorities around stability and practical utility over raw benchmark chasing. Agentic coding receives particular emphasis, with the model handling frontend workflows and repository-level reasoning more fluently than its predecessor and competitive with significantly larger dense alternatives. Released under the Apache 2.0 license and compatible with popular inference stacks like vLLM, SGLang, and Hugging Face Transformers, it is designed for developers who want to deploy capable coding agents without enterprise lock-in. The thinking preservation feature and integrated function calling round out a toolset aimed at real development pipelines rather than demonstration benchmarks.
Quick Info
Powered by- Provider
- Scaleway
- Model key
- qwen3.6-35b-a3b
- Release date
- May 1, 2026
- Last updated
- May 22, 2026
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.25
- Output token cost
- $1.50
Limits
- Output tokens
- 16,384 tokens
- Context window
- 128,000 tokens