Alibaba
⚡ Update: v2 (post #71) achieves 51 tok/s. v2.1 (post #104) adds a quick-start script. See those posts for the latest setup. Been chasing every last token/second out of Qwen3.5-122B-A10B on a single DGX Spark for the pa…
Model details
Qwen3.5-122B-A10B is built on a sparse Mixture-of-Experts architecture with 122 billion total parameters, routing through a pool of 256 experts to activate only 10 billion per token during inference. This design sits between the flagship 397B model and smaller variants, offering a practical balance of high capability and computational efficiency. The architecture combines Gated Delta Networks with sparse MoE layers arranged in a 3:1 ratio of linear attention to full attention, using Grouped-Query Attention with 2 key-value heads per 32 query heads across 48 layers, all conditioned by RoPE position embeddings and SwigLU activations. The model was trained with an early fusion multimodal approach that achieved cross-generational parity with the Qwen3 text-only baseline across reasoning, coding, agents, and visual understanding benchmarks.
The training pipeline emphasizes reinforcement learning scaled across million-agent environments with progressively complex task distributions, enabling robust real-world adaptability. Near-100% multimodal training efficiency compared to text-only training, alongside asynchronous RL frameworks supporting massive-scale deployment, underpins the model's practical strengths. With a native 262K context window that can extend to 1M+ via YaRN extrapolation, the model maintains high needle-in-haystack accuracy across full context lengths, making it particularly suited for complex long-horizon agentic workflows, large document analysis spanning hundreds of pages, and sustained reasoning over extended codebases. Support for 201 languages enables inclusive worldwide deployment, and the Apache 2.0 license facilitates commercial use across diverse applications.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Alibaba
⚡ Update: v2 (post #71) achieves 51 tok/s. v2.1 (post #104) adds a quick-start script. See those posts for the latest setup. Been chasing every last token/second out of Qwen3.5-122B-A10B on a single DGX Spark for the pa…
Alibaba
The news blog specialized in Japanese culture, odd news, gadgets and all other funny stuffs. Updated everyday.
Alibaba
The Roboflow Playground profile documents Qwen3.5-122B-A10B as a 122-billion-parameter Mixture-of-Experts model from Alibaba's Qwen team with roughly 10 billion parameters activated per token through sparse expert routing, released in February 2026 under the Apache 2.0 license. It is described as a multimodal model tha For vision capabilities, Roboflow's profile places Qwen3.5-122B-A10B among 118 evaluated models on its Visual Understanding suite, ranking 9th of 77 with a 76.12% pass rate across 67 tasks, which the page notes places it above roughly 86% of compared models. The listed supported vision tasks include Captioning, OCR, Do
Alibaba
Qwen3.5-122B-A10B is a multimodal Mixture-of-Experts model with 122 billion total parameters and 10 billion activated parameters. It combines strong reasoning, coding, long-context, and visual understanding performance with production-friendly efficiency and a native 262K context window.