Sulat.com
AI models
OrcaRouter logo

Model details

Qwen3.5 122B-A10B

Qwen3.5-122B-A10B sits in the middle of the Qwen3.5 family, which was released alongside the flagship Qwen3.5-397B-A17B and also includes a smaller 35B-A3B MoE variant and a dense 27B option. The model carries 122 billion total parameters but activates only about 10 billion per token, giving it a balance between the flagship's capacity and the lighter variants' efficiency. It inherits the family's hybrid Gated DeltaNet plus MoE architecture, where linear attention layers are interleaved with full attention in roughly a three-to-one ratio to keep long-context processing affordable. The base variant is also available for teams that want to fine-tune the weights themselves.

Independent benchmarking on Lambda's inference stack shows how the model behaves across hardware: on 4× B200 it reaches about 2,197 tokens per second of generation throughput under SGLang, while 8× H100 delivers 1,585 tokens per second and 8× A100 produces 930 tokens per second, with sub-30ms inter-token latency on all three setups. That kind of throughput profile makes the 122B-A10B a pragmatic choice for workloads that need stronger reasoning than the 35B tier but do not require the full 397B flagship, especially when paired with modern serving engines. Its open-weight availability also makes it well suited for teams that want to self-host, customize, or experiment with long-context multimodal inputs.

OrcaRouterqwen/qwen3.5-122b-a10bqwen

Quick Info

Powered by
Provider
OrcaRouter
Model key
qwen/qwen3.5-122b-a10b
Release date
Feb 23, 2026
Last updated
Feb 23, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.115
Output token cost
$0.917

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Latest news about Qwen3.5 122B-A10B

Videos about Qwen3.5 122B-A10B

More models around Qwen3.5 122B-A10B