Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Ofox logo

Model details

Qwen3.5 122B-A10B

Qwen3.5-122B-A10B is a multimodal model from Alibaba's Qwen team, built as part of the broader Qwen3.5 family with a Mixture-of-Experts architecture that scales to 122 billion total parameters while activating roughly 10 billion per token. This sparse routing design is meant to preserve large-scale reasoning ability while keeping inference more efficient than dense models of comparable size. The model processes text and visual inputs within a unified framework, making it well suited to tasks that blend language with images, documents, and charts, including document understanding, diagram interpretation, and complex visual question answering.

Beyond its multimodal reach, Qwen3.5-122B-A10B offers a native context window of about 256,000 tokens that can be pushed further with YaRN-style extensions for very long-context workloads. It is distributed as an open-weight release under the Apache 2.0 license, with the upstream Qwen/Qwen3.5-122B-A10B repository publicly hosted and already serving as a base for community forks such as abliterated derivatives. Practically, this combination of an open license, efficient MoE inference, and broad multimodal coverage positions the model as a flexible option for developers who need strong visual reasoning and long-context handling without committing to a closed proprietary system.

Ofoxqwen/qwen3.5-122b-a10bqwen

Quick Info

Powered by
Provider
Ofox
Model key
qwen/qwen3.5-122b-a10b
Release date
Feb 23, 2026
Last updated
Feb 23, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.29
Output token cost
$2.29

Limits

Output tokens
64,000 tokens
Context window
256,000 tokens

Transparent token rates

Compare Qwen3.5 122B-A10B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.5 122B-A10B

SiliconFlow

CoverageRelease Notes

An AWS "What's New" announcement adds Qwen3.5-122B-A10B to Amazon SageMaker JumpStart alongside LocateAnything-3B and Qwen-AgentWorld-35B-A3B, giving AWS customers a managed path to deploy and fine-tune the exact 122B-A10B variant of Qwen3.5 without standing up their own inference stack. The page's URL slug places the The provided scrape is dominated by cookie-preference boilerplate, so region, instance, and pricing tier details are not visible in the excerpt, and the model itself remains attributed to Qwen/Alibaba rather than AWS. For SiliconFlow-served users the practical signal is that Qwen3.5-122B-A10B is now broadly available a

OpenRouter

CoverageBenchmark

Community oMLX benchmark results dated May 10, 2026 show Qwen3.5-122B-A10B running at 4-bit quantization on an M4 Max with 128 GB unified memory and 40 GPU cores under oMLX v0.3.8 on macOS 26.4. At a 1K context, it sustains 526.8 prompt tokens/s and 57.5 generation tokens/s with peak memory of 65.6 GB. Throughput declines with context length, dropping to 263.9 PP tok/s and 25.4 TG tok/s at 128K tokens (87.2 GB peak memory). Batching tests on the same M4 Max setup show generation throughput scaling from 57.5 tok/s at batch 1 to 84.8 tok/s at batch 2 (1.47x) and 112.6 tok/s at batch 4 (1.96x speedup). Peak memory climbs modestly with context, from 65.6 GB at 1K to 76.8 GB at 64K, indicating the 4-bit quantized model fits comfortably within 128 GB even at long contexts. These figures are from a single community submission and may not transfer to full-precision deployments.

OpenRouter

CoverageBenchmark

BenchLM's September 25, 2026 tracker places Qwen3.5-122B-A10B at an overall capability score of 40.1/100 (ranked 113 of 194 models), with the strongest category being instruction following at rank 11 of 124 (score 91.6, 92nd percentile). The model reports a 262K token context and a measured throughput of 129 tok/s against a field median of 91 tok/s, placing it among faster open-weights options. Multilingual performance ranks it 10th of 12 tracked models in that category. Per BenchLM, the model's published evidence includes 15 displayable benchmark rows, with verified scores in agentic (23.1, rank 84/105), coding (35.7, rank 73/135), reasoning (49.8), multimodal (57.0, rank 34/50), and knowledge (41.5, rank 86/158). No comparable first-party hosted token rate is published, and math benchmarks are not measured for this variant. The tracker flags the context-length figure as lacking a stored direct source link.

OpenRouter

Coverage

Qwen3.5-122B-A10B is Alibaba Cloud's mid-tier multimodal MoE foundation model, released February 23, 2026, with 122B total parameters and 10B activated per token across 256 experts (apxml.com specs page). It pairs a 262K native context window, extendable to 1M tokens, with grouped-query attention (32Q/2KV heads, head dim 256) and SwiGLU feed-forward layers. The page lists benchmark scores including MMLU-Pro 0.867 and GPQA 0.866, and recommends 3x RTX 6000 Blackwell or a single MI325X for full FP16 hosting. According to apxml.com, the model is distributed under Apache 2.0 and reports active parameter count of 122B with 450M auxiliary parameters, supporting multimodal inputs across vision and language. A 128K-context FP16 inference workload requires roughly 245.72 GB VRAM, dominated by expert weights. Its strongest published rankings are in knowledge (rank 86 of 158) and coding (rank 96), with an overall catalog rank of 100 among tracked models.

Kilo Gateway

CoverageBenchmark

Roboflow's Playground page documents Qwen3.5-122B-A10B as a high-capacity multimodal Mixture-of-Experts model from Alibaba's Qwen team, with 122B total parameters activating approximately 10B per token via sparse expert routing. The page specifies the model supports both text and visual inputs in a unified framework fo The entry includes independent usage telemetry showing 41 inferences in the trailing 30 days with an average latency of 18.02 seconds, and reports a legacy Vision Evals pass rate of 76.12% across 67 tasks, placing the model 9th of 77 in that evaluation and above 86% of peers on that benchmark. Methodology details for t

Videos about Qwen3.5 122B-A10B

More models around Qwen3.5 122B-A10B