Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Deep Infra logo

Model details

Qwen 3.5 35B A3B

Qwen 3.5 35B A3B is positioned as a multimodal mixture-of-experts model in the Qwen family, with 35 billion total parameters and roughly 3 billion activated parameters at inference. The architecture blends linear attention mechanisms with a sparse mixture-of-experts design, which the developers describe as a hybrid approach aimed at lifting inference efficiency while preserving reasoning capability. It is presented as a native vision-language model, meaning image and video understanding are integrated into the core architecture rather than bolted on as a separate adapter, and the model ships with a 262K token context window that supports long documents, extended conversations, and multi-step agent workflows.

Practically, the model is aimed at teams that want strong reasoning, coding, and agentic behavior alongside visual understanding, but without the compute footprint of a fully dense frontier model. Because only a small fraction of parameters activate per token, it is well suited to running on a single high-end consumer or workstation GPU, and its open weights make it attractive for self-hosted, fine-tuned, and privacy-sensitive deployments. The combination of long context, multimodal input, and mixture-of-experts efficiency makes it a natural fit for document analysis, code and tool-using agents, and multimodal assistants where balance between capability and cost matters more than chasing the absolute largest model.

Deep InfraQwen/Qwen3.5-35B-A3Bqwen

Quick Info

Powered by
Provider
Deep Infra
Model key
Qwen/Qwen3.5-35B-A3B
Release date
Feb 1, 2026
Last updated
Apr 20, 2026
Knowledge cutoff
2025-01
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.14
Output token cost
$1.00

Limits

Output tokens
81,920 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen 3.5 35B A3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen 3.5 35B A3B

Deep Infra

Coverage

The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. $0.1625 per million input tokens, $1.30 per million output tokens. 262,144 token context window, maxi

Deep Infra

CoverageBenchmark

Qwen3.5-35B-A3B is a multimodal Mixture-of-Experts model with 35 billion total parameters and 3 billion activated parameters. It combines strong reasoning, coding, agentic, and visual understanding performance with production-friendly efficiency and a native 262K context window.

Videos about Qwen 3.5 35B A3B

More models around Qwen 3.5 35B A3B