Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Jalapeno Cloud logo

Model details

Qwen3.5 122B-A10B

Qwen3.5-122B-A10B is a large open-weight model in the Qwen family whose name encodes its MoE design, pairing 122 billion total parameters with about 10 billion active per token, a configuration that delivers frontier-class capacity while keeping per-query compute closer to a much smaller dense model. The availability of community-quantized builds, including the RedHatAI NVFP4 variant referenced in deployment threads, shows the checkpoints can be redistributed and run efficiently on single high-memory accelerators, reinforcing the open-weight lineage inherited from earlier Qwen3 generations. That combination of substantial total capacity, sparse expert routing, and freely available weights makes the model attractive for teams that want to self-host serious reasoning workloads without licensing friction.

Practically, the model is aimed at long-horizon reasoning, document and multimedia analysis, and tool-augmented assistants, where its wide context budget and hybrid expert routing help it handle extended prompts and complex multi-step problems. Forum benchmarks for a quantized build on a single DGX Spark reaching roughly 51 tokens per second suggest it can sustain responsive interactive use even when squeezed onto one device, hinting at the efficiency headroom available to hosted deployments with more memory bandwidth. The fit is strongest for engineering teams, researchers, and product builders who need strong general reasoning plus multimodal understanding in a model they can inspect, fine-tune, or deploy under flexible terms.

Jalapeno CloudQwen3.5-122B-A10Bqwen

Quick Info

Powered by
Provider
Jalapeno Cloud
Model key
Qwen3.5-122B-A10B
Release date
Feb 23, 2026
Last updated
Feb 23, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.40
Output token cost
$3.20

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.5 122B-A10B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.5 122B-A10B

Jalapeno Cloud

CoverageBenchmark

Roboflow's Playground page describes Qwen3.5-122B-A10B as a multimodal Mixture-of-Experts model from Alibaba's Qwen team with 122B total parameters activating roughly 10B per token, designed for unified text-and-vision reasoning across images, documents, charts, and natural language. It states a native context window o On Roboflow's legacy Vision Evals (67 visual-understanding tasks across 77 models), Qwen3.5-122B-A10B achieves a 76.12% pass rate, ranking 9 of 77 and beating 86% of compared models, with an average response time of 1.77 seconds per task and a reported cost of $0.0003 per task at $0.260 in / $2.08 out per 1M tokens. Th

Videos about Qwen3.5 122B-A10B

More models around Qwen3.5 122B-A10B