Sulat.com
AI models
Hugging Face logo

Model details

Qwen3.5 122B-A10B

Qwen3.5 122B-A10B sits in Alibaba Cloud's mid-tier multimodal lineup as a vision-language Mixture-of-Experts foundation model that accepts text, image, and video inputs and produces text outputs. Its "A10B" designation indicates an active-parameter count of roughly 10 billion out of a 122 billion total, a sparsity pattern that keeps inference costs and latency tractable while preserving the capacity of a much larger dense model. This balance makes it attractive for teams that need strong multimodal reasoning without paying for full-parameter inference on every request.

Independent practitioners have already shown that the open-weights release runs comfortably on a single DGX Spark / GB10 workstation, with an NVIDIA community thread documenting throughput of up to 51 tokens per second alongside patches and a quick-start recipe for local deployment. A third-party inference benchmark on DeepInfra further characterizes the model's API latency, throughput, and cost profile. Together these reports suggest the model is well suited to on-prem multimodal assistants, document and chart understanding, and cost-sensitive production pipelines where a sparse MoE offers a favorable accuracy-per-dollar trade-off.

Hugging FaceQwen/Qwen3.5-122B-A10Bqwen

Quick Info

Powered by
Provider
Hugging Face
Model key
Qwen/Qwen3.5-122B-A10B
Release date
Feb 23, 2026
Last updated
Feb 23, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.40
Output token cost
$3.20

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Latest news about Qwen3.5 122B-A10B

Videos about Qwen3.5 122B-A10B

Recent tweets and retweets from Hugging Face

More models around Qwen3.5 122B-A10B