Currently listed through these providers:
Model details
Qwen3.5 122B-A10B
Qwen3.5-122B-A10B sits in the middle of the Qwen3.5 family, which was released alongside the flagship Qwen3.5-397B-A17B and also includes a smaller 35B-A3B MoE variant and a dense 27B option. The model carries 122 billion total parameters but activates only about 10 billion per token, giving it a balance between the flagship's capacity and the lighter variants' efficiency. It inherits the family's hybrid Gated DeltaNet plus MoE architecture, where linear attention layers are interleaved with full attention in roughly a three-to-one ratio to keep long-context processing affordable. The base variant is also available for teams that want to fine-tune the weights themselves.
Independent benchmarking on Lambda's inference stack shows how the model behaves across hardware: on 4× B200 it reaches about 2,197 tokens per second of generation throughput under SGLang, while 8× H100 delivers 1,585 tokens per second and 8× A100 produces 930 tokens per second, with sub-30ms inter-token latency on all three setups. That kind of throughput profile makes the 122B-A10B a pragmatic choice for workloads that need stronger reasoning than the 35B tier but do not require the full 397B flagship, especially when paired with modern serving engines. Its open-weight availability also makes it well suited for teams that want to self-host, customize, or experiment with long-context multimodal inputs.
Quick Info
Powered by- Provider
- OrcaRouter
- Model key
- qwen/qwen3.5-122b-a10b
- Release date
- Feb 23, 2026
- Last updated
- Feb 23, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.115
- Output token cost
- $0.917
Limits
- Output tokens
- 65,536 tokens
- Context window
- 262,144 tokens