SiliconFlow (China)
⚡ Update: v2 (post #71) achieves 51 tok/s. v2.1 (post #104) adds a quick-start script. See those posts for the latest setup. Been chasing every last token/second out of Qwen3.5-122B-A10B on a single DGX Spark for the pa…
Model details
Qwen3.5-122B-A10B is built around a sparse Mixture-of-Experts architecture that keeps inference efficient by activating only 10 billion parameters out of 122 billion total, drawn from 256 expert networks. Its 48-layer stack interleaves Gated DeltaNet blocks with Gated Attention in a 3-to-1 ratio, creating a hybrid design that balances depth and responsiveness. The model was trained with early fusion of visual and textual tokens, giving it native multimodality rather than bolting on vision as a separate stage—this lets it handle images, documents, and video alongside text seamlessly. Its architecture supports a native 262K context window that can be extended to 1 million tokens, making it practical for processing entire books, lengthy codebases, or massive log files without losing coherence.
The post-training approach leans heavily on reinforcement learning scaled across million-agent environments with progressively complex task distributions, which builds real-world adaptability rather than just static benchmark performance. Compared to its predecessor Qwen3, the 3.5 release adds adaptive thinking mode that switches between deep reasoning and quick responses depending on the query. On standard reasoning benchmarks, it posts leading numbers—86.7 on MMLU-Pro, 86.6 on GPQA Diamond, and strong scores on programming agents (BFCL-V4 at 72.2, TAU2-Bench at 79.5) and visual reasoning (MMMU-Pro at 76.9). With weights released openly and compatibility across Hugging Face Transformers, vLLM, SGLang, and KTransformers, teams can deploy it on their own infrastructure or through hosted APIs. The model covers 201 languages and dialects, positioning it for global applications where multilingual reasoning and agentic workflows matter.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
SiliconFlow (China)
⚡ Update: v2 (post #71) achieves 51 tok/s. v2.1 (post #104) adds a quick-start script. See those posts for the latest setup. Been chasing every last token/second out of Qwen3.5-122B-A10B on a single DGX Spark for the pa…
SiliconFlow (China)
The news blog specialized in Japanese culture, odd news, gadgets and all other funny stuffs. Updated everyday.
SiliconFlow (China)
AWS published a first-party "What's New" notice announcing that Qwen3.5-122B-A10B is now available on Amazon SageMaker JumpStart, alongside the LocateAnything-3B and Qwen-AgentWorld-35B-A3B models. The page title and URL explicitly name the exact Qwen3.5-122B-A10B variant, confirming the announcement covers this specif Distribution through SageMaker JumpStart gives AWS customers a managed path to deploy Qwen3.5-122B-A10B, expanding access to Alibaba's Qwen team's Mixture-of-Experts model for developers building on AWS infrastructure. While the visible scrape excerpt is dominated by the cookie banner, the first-party channel and expli
SiliconFlow (China)
Roboflow Playground's model page describes Qwen3.5-122B-A10B as a high-capacity multimodal Mixture-of-Experts model from Alibaba's Qwen team with 122B total parameters and roughly 10B activated per token via sparse expert routing, designed for unified text-and-vision reasoning across images, documents, charts, and natu For developers, the page positions Qwen3.5-122B-A10B as an open-weight multimodal model suitable for document understanding, diagram interpretation, and complex visual question answering, with foundation vision capabilities and the ability to combine LLMs with vision. While the benchmark coverage is thin and traffic on
SiliconFlow (China)
OpenRouter lists Qwen3.5-122B-A10B as a native multimodal vision-language MoE model from Alibaba's Qwen team with 122B total parameters (~10B activated), a 262K-token context window, and a release date of February 25, 2026. The listing positions the model as second only to Qwen3.5-397B-A17B in overall performance and n The aggregate effective pricing across providers on OpenRouter is $0.3064/M input and $2.404/M output, with AtlasCloud currently holding the largest token share (33.4%) and SiliconFlow second at 20.7%, followed by NovitaAI (19.1%), DeepInfra (16.4%), and Alibaba Cloud International (10.3%). Routing modes supported are
SiliconFlow (China)
Qwen3.5-122B-A10B is a multimodal Mixture-of-Experts model with 122 billion total parameters and 10 billion activated parameters. It combines strong reasoning, coding, long-context, and visual understanding performance with production-friendly efficiency and a native 262K context window.