Currently listed through these providers:
Model details
Qwen3.5-35B-A3B
Qwen3.5-35B-A3B is built on a Mixture-of-Experts architecture that routes through 256 specialized expert subnetworks while activating only 3 billion parameters at inference time, delivering strong capability with a fraction of the compute cost of dense models of comparable total size. The model uses early fusion training on multimodal tokens to unify vision and language understanding within a single foundation, and combines this with Gated Delta Networks and sparse MoE layers to keep throughput high and latency low. It supports tool use natively and pairs visual input with reasoning across coding, agents, and visual understanding tasks, consistently outperforming previous-generation models that are more than six times its activated parameter count.
The model benefits from reinforcement learning at scale applied during its post-training phase, which the Qwen team highlights as a key driver of its generalization capabilities across diverse tasks. Released under the Apache 2.0 license, it is available as open weights for self-hosted inference in frameworks like Hugging Face Transformers, vLLM, and SGLang, and it integrates smoothly with desktop runtimes such as LM Studio for local deployment. Its benchmark profile shows particular strength in tool calling, where it ranks among the top performers, alongside solid reasoning and coding results. This combination of open accessibility, architectural efficiency, and multimodal reasoning makes it well-suited for developers and teams looking to power agentic workflows, visual document understanding, or cost-sensitive production systems without relying on managed API infrastructure.
Quick Info
Powered by- Provider
- NovitaAI
- Model key
- qwen/qwen3.5-35b-a3b
- Release date
- Feb 26, 2026
- Last updated
- Feb 26, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.25
- Output token cost
- $2.00
Limits
- Output tokens
- 65,536 tokens
- Context window
- 262,144 tokens
Transparent token rates
Compare Qwen3.5-35B-A3B pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Qwen3.5-35B-A3B
No articles yet. Fetch the latest news to show it here.