Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
DevPass (LLM Gateway) logo

Model details

Qwen3.5 122B-A10B

Qwen3.5 122B-A10B is a multimodal vision-language Mixture-of-Experts model released by Alibaba Cloud in February 2026, positioned as a mid-tier foundation model in the Qwen3.5 family. Its identifier and weights are publicly hosted on Hugging Face under the Qwen organization, which has enabled community re-derivations such as the wangzhang abliterated variant built on top of the upstream Qwen/Qwen3.5-122B-A10B repository. The MoE design with active experts denoted in the A10B naming is consistent with the family's broader move toward sparse architectures for efficient multimodal reasoning.

In practical terms, the model's open weights make it attractive for teams that want to self-host or fine-tune a large multimodal model rather than depend solely on a hosted API, while its Mixture-of-Experts structure aims to balance capability and throughput for text, image, and video workloads. The existence of an abliterated community variant reporting a 0.5% refusal rate with minimal capability drift suggests the base weights are robust and amenable to downstream alignment adjustments. Developers evaluating Qwen3.5 122B-A10B should weigh its mid-tier positioning against larger Qwen3.5 siblings when matching model scale to latency, cost, and multimodal requirements.

DevPass (LLM Gateway)qwen3.5-122b-a10bqwen

Quick Info

Powered by
Provider
DevPass (LLM Gateway)
Model key
qwen3.5-122b-a10b
Release date
Feb 23, 2026
Last updated
Feb 23, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.40
Output token cost
$3.20

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.5 122B-A10B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.5 122B-A10B

DevPass (LLM Gateway)

CoverageBenchmark

Community users on the NVIDIA DGX Spark forum report running Qwen3.5-122B-A10B on a single Spark node, achieving up to 51 tokens per second with vLLM patches and custom quick-start configurations. The deployment thread documents hands-on local inference performance for the MoE model on GB10 hardware. Thread participants compare vLLM 0.19.1-stable against 0.20.1 builds specifically for this model. Testing showed approximately a 9% performance regression when moving from vLLM 0.19.1 to 0.20.1 on Qwen3.5-122B-A10B, with Q&A dropping 8.3%, code 9.6%, and long-context 10.5%. Users noted vLLM 0.20.1 stabilized tool calls for stock templates but degraded MTP-enabled tool quality. Maintainers recommended staying on 0.19.1-stable pending further regression benchmarks.

DevPass (LLM Gateway)

CoverageBenchmark

An NVIDIA DGX Spark community thread by user Albond documents optimization work for Qwen3.5-122B-A10B, achieving up to 51 tokens per second on a single Spark node. The stack combines vLLM 0.19 with Intel AutoRound INT4, FlashInfer, hybrid INT4+FP8 quantization for shared expert dense layers using Qwen's official FP8 checkpoint, and MTP-1 speculative decoding reaching 95% acceptance rate. The thread notes that vLLM's default INC quantization config silently dispatched FP8 shared-expert layers through UnquantizedLinearMethod, effectively a bug fixed by a 95-line patch enabling CUTLASS FP8 execution. MTP head weights from Intel AutoRound's model_extra_tensors.safetensors (4.8 GB) can be re-enabled for vanilla checkpoints via --speculative-config, while hybrid checkpoints require the add-mtp-weights.py script to restore MTP mappings.

DevPass (LLM Gateway)

CoverageBenchmark

Alibaba's Qwen team released the Qwen3.5 Medium series on February 25, 2026, with Qwen3.5-122B-A10B as one of three open-weight models available under the Apache 2.0 license on Hugging Face and ModelScope. The model uses a hybrid architecture combining Gated Delta Networks with a sparse Mixture-of-Experts system, activating roughly 10 billion parameters per token from 122 billion total, and supports multimodal inputs alongside agentic tool calling. According to VentureBeat's launch coverage, Qwen3.5-122B-A10B delivers performance comparable to proprietary models from OpenAI and Anthropic, reportedly beating GPT-5-mini and Claude Sonnet 4.5 on third-party benchmarks even after aggressive quantization. The release emphasizes that frontier-level context windows and strong accuracy are preserved at 4-bit weight and KV cache quantization, enabling deployment on consumer-grade GPUs without server-grade infrastructure.

DevPass (LLM Gateway)

CoverageBenchmark

The Kilo Code benchmark page describes Qwen3.5-122B-A10B as a hybrid architecture combining a linear attention mechanism with a sparse mixture-of-experts model for higher inference efficiency. The native vision-language model supports a 262,144-token context window with up to 65,536 tokens of max output and multimodal input support. The page positions it for integration through the Kilo Code open-source coding agent. Enkrypt AI red-team evaluations report an overall risk score of 20.5/100 for Qwen3.5-122B-A10B, with a composite safety score of 28.2/100. Category breakdowns include bias at 59.7/100, CBRN at 29.8/100, and insecure code at 7.1/100, while harmful content measured 2.2/100 and toxicity 3.6/100. NIST and OWASP-weighted averages landed at 20.0 and 23.0 respectively, giving developers a third-party safety baseline.

DevPass (LLM Gateway)

CoverageBenchmark

According to the LLM Stats composite dashboard, Qwen3.5-122B-A10B ranks 97 overall and scores 35.0 against a blended price of $0.39 per million tokens. It places in the top 10% for healthcare, legal, and finance, while scoring below average on websites, games, tool calling, and 3D tasks. Conversation-depth analysis shows quality holds between 13.0 and 14.9 across turn 1 through turn 31+. On dataset-level benchmarks sourced from qwen.ai, Qwen3.5-122B-A10B achieves 0.97 on CountBench and 0.97 on VLMsAreBlind, with MMLU-Redux also listed. The quality tracker reports a stable +0.80σ trend over 177 seven-day votes, indicating consistent performance. Capability-tier rankings span Chat, Coding, Reasoning, and Vision categories against hundreds of competing models.

DevPass (LLM Gateway)

CoverageBenchmark

Qwen3.5-122B-A10B is a 122-billion-parameter multimodal Mixture-of-Experts model from Alibaba's Qwen team that activates roughly 10 billion parameters per token, balancing large-scale reasoning with efficient inference. The unified multimodal framework handles text and images for document understanding, chart interpretation, and visual question answering. It supports a native 256,000-token context, extendable via YaRN scaling. Released under the Apache 2.0 license, the open-weight model builds on earlier Qwen multimodal systems for developers tackling demanding multimodal reasoning. Roboflow Playground's vision benchmarks place it at 76.12% pass rate across 67 tasks, ranking 9 of 77 visual-understanding models. Per-task cost is reported at $0.0003, with token pricing at $0.26 per million input.

DevPass (LLM Gateway)

Coverage

Ollama's library page for qwen3.5:122b-a10b distributes the model under Apache License 2.0 as an 81GB Q4_K_M quantized package built on the qwen35moe architecture with approximately 125 billion parameters. The model card confirms multimodal vision support alongside tool calling and thinking capabilities, reflecting Qwen3.5's unified vision-language foundation designed for cross-modal reasoning, coding, and agent workflows. The Ollama listing documents Qwen3.5's core architectural enhancements, including Gated Delta Networks paired with sparse Mixture-of-Experts for high-throughput inference and early-fusion multimodal training. However, the benchmark table displayed on the page reports results for the larger Qwen3.5-397B-A17B sibling rather than the 122B-A10B variant, meaning direct performance numbers for this subject should not be cited from that table.

Videos about Qwen3.5 122B-A10B

More models around Qwen3.5 122B-A10B