Mixlayer
⚡ Update: v2 (post #71) achieves 51 tok/s. v2.1 (post #104) adds a quick-start script. See those posts for the latest setup. Been chasing every last token/second out of Qwen3.5-122B-A10B on a single DGX Spark for the pa…
Model details
Qwen3.5-122B A10B is a Mixture-of-Experts model designed to balance massive capability with deployment practicality. It carries 122 billion total parameters but activates only 10 billion per forward pass, using a hybrid Gated DeltaNet and MoE architecture where linear attention layers and full attention are interleaved in a 3-to-1 ratio for efficient long-context processing. Early fusion multimodal training gives it cross-generational parity with its text-only siblings across reasoning, coding, agents, and visual understanding benchmarks, and its vision performance even surpasses the larger Qwen3-VL-235B. As a mid-tier option in the Qwen3.5 family, it sits between the flagship 397B model and smaller variants, offering a practical sweet spot for teams that need frontier-level performance without flagship-level resource demands.
The model benefits from reinforcement learning scaled across million-agent environments with progressively complex task distributions, which the developers describe as supporting robust real-world adaptability. Alongside broad framework compatibility—Hugging Face Transformers, vLLM, SGLang, and KTransformers—it supports 201 languages and dialects for global deployment. Hardware benchmarks show meaningful throughput across different configurations, reaching over 2,000 tokens per second on 4× B200 GPUs with vLLM, and competitive performance on H100 and A100 setups as well. Available as open weights on Hugging Face, it serves developers and enterprises seeking a high-capability vision-language model that remains efficient enough for production use cases.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Mixlayer
⚡ Update: v2 (post #71) achieves 51 tok/s. v2.1 (post #104) adds a quick-start script. See those posts for the latest setup. Been chasing every last token/second out of Qwen3.5-122B-A10B on a single DGX Spark for the pa…
Mixlayer
Roboflow's Playground page describes Qwen3.5-122B-A10B as a high-capacity multimodal MoE model from Alibaba's Qwen team with 122B total parameters and roughly 10B activated per token through sparse expert routing. The model processes both text and visual inputs within a unified multimodal framework, enabling document u The page reports Roboflow's legacy vision evaluation showing a 76.12% pass rate across 67 visual understanding tasks (ranking 9th of 77 models, better than 86%). Playground usage statistics show 41 inferences in the past 30 days with average latency of 18.02s, and speed averaging 1.77s. Interactive demo capabilities in