Currently listed through these providers:
Model details
Qwen3.5 122B-A10B
Qwen3.5 122B-A10B is a sparse Mixture-of-Experts model from Alibaba Cloud's Qwen family that activates roughly 10B parameters out of a 122B total, a design that aims to keep the inference cost closer to a mid-size model while preserving the reasoning depth of a much larger one. It is positioned as a native multimodal agent, pairing image understanding with text and a very large working memory in the neighborhood of 262K tokens, which makes it well suited for long documents, multi-image prompts, and extended agent loops where the model has to keep a long conversation, retrieved evidence, or a growing scratchpad in view at once.
In practice, the model targets agentic and tool-using workloads: it offers a hybrid reasoning mode with extended thinking for harder problems, function calling for orchestrating external tools, and broad multilingual coverage across roughly 201 languages. The open Apache 2.0 license lowers the barrier for self-hosting and fine-tuning, and there is visible community interest in squeezing it onto single-node hardware such as NVIDIA's DGX Spark, where early benchmarks report throughput in the tens of tokens per second. That combination of MoE efficiency, long context, vision, and permissive licensing makes it a practical fit for teams building assistants, document-analysis pipelines, or multilingual agents who want a capable open-weight backbone without paying for full dense inference at 122B scale.
Quick Info
Powered by- Provider
- OpenRouter
- Model key
- qwen/qwen3.5-122b-a10b
- Release date
- Feb 23, 2026
- Last updated
- Feb 23, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.29
- Output token cost
- $2.40
Limits
- Output tokens
- 81,920 tokens
- Context window
- 262,144 tokens