Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
DevPass (LLM Gateway) logo

Model details

Qwen3.5 35B-A3B

The model overview is being prepared.

DevPass (LLM Gateway)qwen3.5-35b-a3bqwen

Quick Info

Powered by
Provider
DevPass (LLM Gateway)
Model key
qwen3.5-35b-a3b
Release date
Feb 23, 2026
Last updated
Feb 23, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.25
Output token cost
$2.00

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.5 35B-A3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.5 35B-A3B

DevPass (LLM Gateway)

CoverageBenchmark

Alibaba's Qwen team released the Qwen3.5 Medium series around February 24, 2026, with Qwen3.5-35B-A3B as one of four new models, per VentureBeat. It ships under an Apache 2.0 license for commercial use and is available on Hugging Face and ModelScope. The model uses a hybrid architecture combining Gated Delta Networks with a sparse Mixture-of-Experts system, activating roughly 3B of its 35B parameters per pass. The Qwen3.5-35B-A3B delivers benchmark performance comparable to Claude Sonnet 4.5 and reportedly beats OpenAI's GPT-5-mini, according to the report. It supports agentic tool calling and maintains near-lossless accuracy under 4-bit weight and KV cache quantization. A standout technical milestone is exceeding 1 million tokens of context on consumer-grade GPUs with 32GB of VRAM, bringing frontier-level context windows to desktop hardware.

DevPass (LLM Gateway)

CoverageBenchmark

Qwen3.5-35B-A3B ranks 121st on the LLM Stats composite leaderboard with a score of 30.8 at a blended price of $0.18 per million tokens, positioning it in the mid-cost band. It places in the top 10% for legal (19/212) and finance (20/228) capability tiers, and shows strong multimodal results including 0.98/100 on CountBench (rank 4) and 0.97/1 on VLMsAreBlind (rank 2), with benchmark scores sourced from qwen.ai. Capability tier rankings place the model in the top half for chat (18/133) and long-context (29/119) but lower in coding (159/273), tool calling (117/200), and reasoning (111/369). Quality tracker data shows a +0.78σ stable shift in the 7-day window with 230 votes, and conversation-depth breakdown indicates the model performs best on chat (+2.07σ across 123 evaluations) and weakest on games (-1.28σ across 19 evaluations).

DevPass (LLM Gateway)

CoverageBenchmark

Qwen3.5-35B-A3B is a native vision-language model released on February 25, 2026, built on a hybrid architecture that combines linear attention mechanisms with a sparse mixture-of-experts design for higher inference efficiency. According to the catalog page, its overall performance is comparable to Qwen3.5-27B, and it supports a 262K-token context window with vision input and output modalities. The model is re-hosted by multiple infrastructure providers at listed prices ranging from $0.08 to $0.3125 per million input tokens and $0.75 to $1.80 per million output tokens, with weighted averages around $0.1594 input and $1.06 output per million tokens. Reported round-trip latency spans roughly 0.69s to 1.52s across hosts, with throughput up to 118 tokens per second; these figures reflect serving-provider behavior rather than intrinsic model capability.

DevPass (LLM Gateway)

Coverage

Apxml's reference entry details Qwen3.5-35B-A3B's architecture: 35B total parameters with 3B activated across 256 experts (9 active per token), 40 layers, Grouped-Query Attention with 16Q/2KV heads, head dimension 256, RoPE theta 10M, and SwiGLU activations. Released February 23, 2026 under Apache 2.0, the model has a 262K native context window extensible to 1M tokens with multimodal support via 450M auxiliary parameters. The model scores 85.3% on MMLU-Pro, 84.2% on GPQA Diamond, 69.2% on SWE-bench Verified, and 40.5% on Terminal-Bench 2.0, ranking 113 overall and 119 on coding. Apxml's VRAM calculator shows roughly 71GB required at FP16 for a 1K context, recommending 72-96GB GPUs like the RTX PRO 5000 Blackwell, A100, or M2 Max. Creator attribution is explicitly given to Alibaba Cloud.

Videos about Qwen3.5 35B-A3B

More models around Qwen3.5 35B-A3B