Sulat.com
AI models
Get 10-25% off
Get 10-25% off from Qwen
Alibaba logo

Model details

Qwen3.5 122B-A10B

Qwen3.5-122B-A10B is built on a sparse Mixture-of-Experts architecture with 122 billion total parameters, routing through a pool of 256 experts to activate only 10 billion per token during inference. This design sits between the flagship 397B model and smaller variants, offering a practical balance of high capability and computational efficiency. The architecture combines Gated Delta Networks with sparse MoE layers arranged in a 3:1 ratio of linear attention to full attention, using Grouped-Query Attention with 2 key-value heads per 32 query heads across 48 layers, all conditioned by RoPE position embeddings and SwigLU activations. The model was trained with an early fusion multimodal approach that achieved cross-generational parity with the Qwen3 text-only baseline across reasoning, coding, agents, and visual understanding benchmarks.

The training pipeline emphasizes reinforcement learning scaled across million-agent environments with progressively complex task distributions, enabling robust real-world adaptability. Near-100% multimodal training efficiency compared to text-only training, alongside asynchronous RL frameworks supporting massive-scale deployment, underpins the model's practical strengths. With a native 262K context window that can extend to 1M+ via YaRN extrapolation, the model maintains high needle-in-haystack accuracy across full context lengths, making it particularly suited for complex long-horizon agentic workflows, large document analysis spanning hundreds of pages, and sustained reasoning over extended codebases. Support for 201 languages enables inclusive worldwide deployment, and the Apache 2.0 license facilitates commercial use across diverse applications.

Alibabaqwen3.5-122b-a10bqwen

Quick Info

Powered by
Provider
Alibaba
Model key
qwen3.5-122b-a10b
Release date
Feb 23, 2026
Last updated
Feb 23, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.40
Output token cost
$3.20

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.5 122B-A10B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.5 122B-A10B

Alibaba

CoverageBenchmark

⚡ Update: v2 (post #71) achieves 51 tok/s. v2.1 (post #104) adds a quick-start script. See those posts for the latest setup. Been chasing every last token/second out of Qwen3.5-122B-A10B on a single DGX Spark for the pa…

Alibaba

CoverageBenchmark

The news blog specialized in Japanese culture, odd news, gadgets and all other funny stuffs. Updated everyday.

Alibaba

CoverageBenchmark

The Roboflow Playground profile documents Qwen3.5-122B-A10B as a 122-billion-parameter Mixture-of-Experts model from Alibaba's Qwen team with roughly 10 billion parameters activated per token through sparse expert routing, released in February 2026 under the Apache 2.0 license. It is described as a multimodal model tha For vision capabilities, Roboflow's profile places Qwen3.5-122B-A10B among 118 evaluated models on its Visual Understanding suite, ranking 9th of 77 with a 76.12% pass rate across 67 tasks, which the page notes places it above roughly 86% of compared models. The listed supported vision tasks include Captioning, OCR, Do

Alibaba

CoverageBenchmark

Qwen3.5-122B-A10B is a multimodal Mixture-of-Experts model with 122 billion total parameters and 10 billion activated parameters. It combines strong reasoning, coding, long-context, and visual understanding performance with production-friendly efficiency and a native 262K context window.

Videos about Qwen3.5 122B-A10B

More models around Qwen3.5 122B-A10B