Currently listed through these providers:
Model details
Qwen3.5 35B-A3B
Qwen3.5 35B-A3B is built around a sparse Mixture-of-Experts architecture that activates just 3 billion of its 35 billion total parameters per token, making it remarkably efficient for deployment. The model combines Gated Delta Networks with this MoE design to deliver high-throughput inference with minimal latency and cost overhead. Early fusion training on multimodal tokens enables the model to achieve cross-generational parity with Qwen3 and outperform Qwen3-VL models across reasoning, coding, agents, and visual understanding benchmarks—a remarkable feat for a model engineered to use only a fraction of its total capacity at any given moment.
The training approach leverages reinforcement learning at scale, which has driven the model's ability to generalize across diverse tasks. This RL-based post-training strategy has produced a model that outperforms the previous generation's 235 billion parameter model on most benchmarks and even surpasses GPT-5 mini and Claude Sonnet 4.5 on knowledge and visual reasoning assessments. The practical implications are significant: the model runs comfortably on consumer-grade hardware with an 8GB GPU while maintaining benchmark performance that rivals models many times its size. Qwen3.5 35B-A3B is released under an open license, making it accessible for developers and enterprises seeking powerful multimodal reasoning and tool-use capabilities without the resource demands of dense models.
Quick Info
Powered by- Provider
- OpenRouter
- Model key
- qwen/qwen3.5-35b-a3b
- Release date
- Feb 23, 2026
- Last updated
- Feb 23, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.25
- Output token cost
- $1.25
Limits
- Output tokens
- 235,929 tokens
- Context window
- 262,144 tokens