Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Jalapeno Cloud logo

Model details

Qwen3.5 35B-A3B

Qwen3.5 35B-A3B is designed as a native vision-language foundation model, using early fusion training on multimodal tokens to connect language, visual understanding, reasoning, coding, and agent-oriented work. The supplied evidence positions it as part of Qwen3.5’s effort to bring those capabilities together in one model, while its architecture combines Gated Delta Networks with a sparse Mixture-of-Experts design for efficient inference.

The model is a practical fit for applications that need multimodal understanding alongside efficient generation, particularly when deployed through compatible serving tools or managed inference services. The source material highlights scalable reinforcement-learning generalization and reports overall performance comparable to Qwen3.5-27B, while the architecture is intended to support high-throughput inference with limited latency and cost overhead. Its open-weight artifacts are available in Hugging Face Transformers format, with compatibility extending to vLLM, SGLang, and KTransformers.

Jalapeno CloudQwen3.5-35B-A3Bqwen

Quick Info

Powered by
Provider
Jalapeno Cloud
Model key
Qwen3.5-35B-A3B
Release date
Feb 23, 2026
Last updated
Feb 23, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.25
Output token cost
$2.00

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.5 35B-A3B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.5 35B-A3B

Jalapeno Cloud

CoverageBenchmark

The Millstone AI inference benchmark page documents Qwen3.5-35B-A3B-FP8 as a 35B-parameter Mixture-of-Experts model with 256 experts (8 routed plus 1 shared active per forward pass) and 3B activated parameters, organized across 40 transformer layers. It confirms the hybrid Gated Delta Networks architecture combined wit The page positions Qwen3.5-35B-A3B as competitive with much larger models such as Qwen3-235B-A22B across math, coding, and multilingual benchmarks, and provides measured throughput figures across hardware configurations including 1x and 2x RTX Pro 6000 Blackwell, 1x H100 SXM, and 1x H200 SXM, with peak throughput reach

Jalapeno Cloud

CoverageBenchmark

OpenRouter's listing for Qwen3.5-35B-A3B identifies it as a native vision-language model with a hybrid architecture combining linear attention mechanisms and a sparse mixture-of-experts design, with overall performance described as comparable to Qwen3.5-27B. Key specs surfaced on the page include a 262K context window, The weighted-average effective price across providers is reported at roughly $0.2143 input / $1.143 output per million tokens, with CoreWeave and Darkbloom dominating recent token share. A price-history panel tracks effective rates from June through early September, showing modest drift. Notably, Jalapeno Cloud is not

Videos about Qwen3.5 35B-A3B

More models around Qwen3.5 35B-A3B