Currently listed through these providers:
Model details
Qwen3.5 122B-A10B
Qwen3.5-122B-A10B is a post-trained model in the Qwen3.5 family, distributed through the Hugging Face repository Qwen/Qwen3.5-122B-A10B with weights and configuration files in the standard Transformers format. The repository documentation confirms compatibility with widely used inference stacks, including Hugging Face Transformers, vLLM, SGLang, and KTransformers, which makes it straightforward to drop into existing serving pipelines. The naming convention "122B-A10B" indicates a sparse Mixture-of-Experts design with a large total parameter pool and a much smaller set of active parameters per token, reflecting the family's emphasis on throughput-efficient inference.
According to the published model card, Qwen3.5 introduces an efficient hybrid architecture that pairs Gated Delta Networks with sparse Mixture-of-Experts routing, aiming to deliver high-throughput generation with minimal latency overhead. The card also highlights a unified vision-language foundation built on early-fusion multimodal training, positioned as reaching parity with prior-generation Qwen3 reasoning and coding models while extending coverage to image understanding. In independent community testing on NVIDIA DGX Spark (GB10) hardware, users have reported running the model on a single Spark and observing throughput in the range described as up to roughly 51 tokens per second, suggesting the MoE design translates into practical single-node performance for long-context workloads.
Quick Info
Powered by- Provider
- SiliconFlow
- Model key
- Qwen/Qwen3.5-122B-A10B
- Release date
- Feb 23, 2026
- Last updated
- Feb 23, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.26
- Output token cost
- $2.08
Limits
- Output tokens
- 262,144 tokens
- Context window
- 262,144 tokens