Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
302.AI logo

Model details

Qwen3.5 35B-A3B

Qwen3.5 35B-A3B is a 35-billion-parameter Mixture-of-Experts model that activates only 3 billion parameters per forward pass, drawing on 256 experts (8 routed plus 1 shared) distributed across 40 transformer layers. This sparse design is paired with a Gated Delta Networks hybrid architecture, which combines linear-attention-style gating with standard transformer layers to balance efficiency and expressivity. The result is a compact yet capable foundation model intended for reasoning, coding, agentic workflows, and visual understanding, released under the Apache 2.0 license for open-weight use and self-hosting.

Built as a unified vision-language model, Qwen3.5 35B-A3B processes text alongside images and video while supporting 201 languages, making it well suited for multilingual assistants and multimodal retrieval or analysis tasks. A native 256K-token context window, extendable to roughly 1M tokens via YaRN-based extrapolation, allows it to handle long documents, large codebases, and extended multi-turn conversations. The ecosystem also includes an FP8-quantized variant using fine-grained quantization with a block size of 128, which is reported to deliver performance close to the original while lowering memory and compute requirements for production deployments.

302.AIqwen3.5-35b-a3bqwen

Quick Info

Powered by
Provider
302.AI
Model key
qwen3.5-35b-a3b
Release date
Feb 23, 2026
Last updated
Feb 23, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.06
Output token cost
$0.46

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.5 35B-A3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.5 35B-A3B

302.AI

CoverageBenchmark

LLM Stats ranks Qwen3.5-35B-A3B 121st overall on its composite LLM Stats Score, with capability tier "Good (Top 10%)" only in Legal (19 of 212) and tier "Average" in Finance (20 of 228), Healthcare (22 of 246), Chat, Long Context, 3D, Math, Vision, and Reasoning — and tier "Bad (Below top half)" in Search, Games, Websi For long-context behavior, LLM Stats measures Qwen3.5-35B-A3B scores of 14.7 at turn 1, 14.6 across turns 2–10, 14.0 across turns 11–30 (rank 34 of 100), and 13.6 at turns 31+ (rank 33 of 80), a cumulative drop of −1.2 relative to turn 1 within confidence intervals between roughly 14.1 and 16.2. Benchmark-level scores

302.AI

CoverageBenchmark

Millstone's overview page for Qwen3.5-35B-A3B-FP8 specifies an MoE architecture with 35B total parameters, 256 experts (8 routed plus 1 shared active per forward pass), and 3B activated parameters across 40 transformer layers. It documents a hybrid Gated Delta Networks plus sparse MoE design, a native 256K-token contex The model is described as a unified vision-language foundation model for reasoning, coding, agents, and visual understanding, supporting 201 languages plus multimodal image and video inputs, and licensed Apache 2.0. The page also reports competitive throughput across hardware configurations, including 598 tok/s on 1x R

Videos about Qwen3.5 35B-A3B

More models around Qwen3.5 35B-A3B