Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
SiliconFlow (China) logo

Model details

Qwen/Qwen3.5-35B-A3B

Qwen3.5-35B-A3B is positioned as a reasoning vision-language model with native multimodal training, where vision and language tokens are learned jointly so the system can interpret images alongside text. Its design pairs linear attention with a sparse mixture-of-experts layout, yielding a hybrid architecture that the OpenRouter listing credits with higher inference efficiency. With 35B total parameters and only 3B activated per token, the model is engineered to deliver stronger reasoning and coding behavior than predecessors of substantially larger active size, and it is explicitly trained for tool use, making it well suited to agent-style workflows that combine perception, planning, and external calls.

For practitioners, this combination of an MoE routing strategy, a vision-aware early-fusion foundation, and a very long context window makes the model a strong fit for tasks such as document and chart understanding, visual question answering, multi-step reasoning, and tool-augmented assistants that need to read images and act on them. Because only a small fraction of the parameters fires on any given token, it can behave like a much larger model in capability while keeping per-request compute closer to a compact model, and the local runtime community has already packaged it with a 21GB system memory recommendation, reflecting a practical mid-range hardware target. The combination of long-context support, vision grounding, and tool use positions it as a versatile backbone for applications that blend seeing, reasoning, and doing.

SiliconFlow (China)Qwen/Qwen3.5-35B-A3Bqwen

Quick Info

Powered by
Provider
SiliconFlow (China)
Model key
Qwen/Qwen3.5-35B-A3B
Release date
Feb 25, 2026
Last updated
Feb 25, 2026
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.23
Output token cost
$1.86

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen/Qwen3.5-35B-A3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen/Qwen3.5-35B-A3B

SiliconFlow (China)

CoverageBenchmark

Per the OpenRouter listing, Qwen/Qwen3.5-35B-A3B is a native vision-language model with a hybrid architecture combining linear attention mechanisms and a sparse mixture-of-experts (MoE) design, offering 262K context and overall performance the aggregator describes as comparable to Qwen3.5-27B. The page records a releas The same page documents cross-provider routing on OpenRouter, including a SiliconFlow row at $0.24 input / $1.80 output per 1M tokens, 1.79s P50 latency, 54 tps throughput, and 98.89% uptime, alongside other hosts such as Darkbloom, DeepInfra, Parasail, Alibaba Cloud Int., AtlasCloud, CoreWeave, and Venice. CoreWeave s

Videos about Qwen/Qwen3.5-35B-A3B

More models around Qwen/Qwen3.5-35B-A3B