Sulat.com
AI models
SiliconFlow logo

Model details

Qwen3.5 122B-A10B

Qwen3.5-122B-A10B is a post-trained model in the Qwen3.5 family, distributed through the Hugging Face repository Qwen/Qwen3.5-122B-A10B with weights and configuration files in the standard Transformers format. The repository documentation confirms compatibility with widely used inference stacks, including Hugging Face Transformers, vLLM, SGLang, and KTransformers, which makes it straightforward to drop into existing serving pipelines. The naming convention "122B-A10B" indicates a sparse Mixture-of-Experts design with a large total parameter pool and a much smaller set of active parameters per token, reflecting the family's emphasis on throughput-efficient inference.

According to the published model card, Qwen3.5 introduces an efficient hybrid architecture that pairs Gated Delta Networks with sparse Mixture-of-Experts routing, aiming to deliver high-throughput generation with minimal latency overhead. The card also highlights a unified vision-language foundation built on early-fusion multimodal training, positioned as reaching parity with prior-generation Qwen3 reasoning and coding models while extending coverage to image understanding. In independent community testing on NVIDIA DGX Spark (GB10) hardware, users have reported running the model on a single Spark and observing throughput in the range described as up to roughly 51 tokens per second, suggesting the MoE design translates into practical single-node performance for long-context workloads.

SiliconFlowQwen/Qwen3.5-122B-A10Bqwen

Quick Info

Powered by
Provider
SiliconFlow
Model key
Qwen/Qwen3.5-122B-A10B
Release date
Feb 23, 2026
Last updated
Feb 23, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.26
Output token cost
$2.08

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Latest news about Qwen3.5 122B-A10B

Videos about Qwen3.5 122B-A10B

Recent tweets and retweets from SiliconFlow

More models around Qwen3.5 122B-A10B