Qwen3.5-122B-A10B is an open-weight large language model in Alibaba Cloud's Qwen family, built around a 122-billion-parameter architecture with an active-parameter design implied by the A10B suffix. Across four multimodal leaderboards, it consistently appears with a 262K-token context window and is attributed to Alibaba Cloud and the Qwen Team. The model accepts text, image, video, and audio inputs and produces text outputs, making it suitable for general-purpose multimodal assistants rather than a narrow single-task system.
The model distinguishes itself on demanding multimodal video and vision benchmarks, posting a 0.766 on MVBench (second place behind GLM-5.3-Flash at 0.778), a 0.829 on MMStar (second place behind Qwen3.6 Plus at 0.833), a 0.873 on MLVU for long-form video understanding (second behind Qwen3.7-Plus at 0.874), and a tied-leading 0.928 on MMBench-V1.1 alongside Qwen3.6-35B-A3B. These placements place it among the strongest open-weight entrants on tasks requiring temporal reasoning, spatial reasoning, and vision-grounded question solving, which makes it a practical fit for teams that need a self-hostable multimodal model with a long context window and competitive results on both short-clip and long-form video analysis.