Currently listed through these providers:
Model details
Qwen3.5 122B-A10B
Qwen3.5 122B-A10B sits in Alibaba Cloud's mid-tier multimodal lineup as a vision-language Mixture-of-Experts foundation model that accepts text, image, and video inputs and produces text outputs. Its "A10B" designation indicates an active-parameter count of roughly 10 billion out of a 122 billion total, a sparsity pattern that keeps inference costs and latency tractable while preserving the capacity of a much larger dense model. This balance makes it attractive for teams that need strong multimodal reasoning without paying for full-parameter inference on every request.
Independent practitioners have already shown that the open-weights release runs comfortably on a single DGX Spark / GB10 workstation, with an NVIDIA community thread documenting throughput of up to 51 tokens per second alongside patches and a quick-start recipe for local deployment. A third-party inference benchmark on DeepInfra further characterizes the model's API latency, throughput, and cost profile. Together these reports suggest the model is well suited to on-prem multimodal assistants, document and chart understanding, and cost-sensitive production pipelines where a sparse MoE offers a favorable accuracy-per-dollar trade-off.
Quick Info
Powered by- Provider
- Hugging Face
- Model key
- Qwen/Qwen3.5-122B-A10B
- Release date
- Feb 23, 2026
- Last updated
- Feb 23, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.40
- Output token cost
- $3.20
Limits
- Output tokens
- 65,536 tokens
- Context window
- 262,144 tokens