Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
NovitaAI logo

Model details

Qwen3.5-122B-A10B

Qwen3.5-122B-A10B is a native vision-language model from the Qwen family that pairs text, image, and video inputs with text outputs, making it suitable for assistants that need to read documents, describe photos, or parse video frames in a single conversation. Its design fuses a linear attention mechanism with a sparse mixture-of-experts structure, which is intended to keep inference efficient while letting the model draw on a very large pool of parameters for nuanced understanding. The full parameter count sits in the 122-billion range with roughly 10 billion activated per token, so the system behaves like a heavyweight model in capability but routes computation more selectively than a dense transformer of comparable size.

Practically, the model fits well into agent systems, chatbots, retrieval-augmented pipelines, and other AI-powered applications where long context and multimodal grounding matter, and a 262,144-token context window supports extended document or transcript analysis. Independent listings on Roboflow's Playground surface it as a vision model for tasks such as image captioning and OCR, while an NVIDIA-published NVFP4 quantized variant on Hugging Face, distributed under Apache License 2.0, shows the ecosystem is investing in deployment-friendly formats for production use. The combination of a hybrid attention-plus-MoE backbone with open-weight availability makes the model a flexible choice for teams building multimodal assistants that require both broad reasoning capacity and cost-aware serving options.

NovitaAIqwen/qwen3.5-122b-a10bqwen

Quick Info

Powered by
Provider
NovitaAI
Model key
qwen/qwen3.5-122b-a10b
Release date
Feb 26, 2026
Last updated
Feb 26, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.40
Output token cost
$3.20

Limits

Output tokens
65,536 tokens
Context window
262,144 tokens

Transparent token rates

Compare Qwen3.5-122B-A10B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3.5-122B-A10B

NovitaAI

CoverageRelease Notes

Amazon Web Services announced that Qwen3.5-122B-A10B is among the models now available on Amazon SageMaker JumpStart, alongside LocateAnything-3B and Qwen-AgentWorld-35B-A3B. The announcement page explicitly names Qwen3.5-122B-A10B as part of the JumpStart catalog update, giving developers an AWS-managed path to deploy As an Apache 2.0 licensed 122B-parameter MoE with roughly 10B active parameters and a ~256K token context, Qwen3.5-122B-A10B is designed for text and visual reasoning across documents, charts, and natural language. Its inclusion in SageMaker JumpStart lowers the operational barrier for enterprises wanting to run the mo

NovitaAI

CoverageBenchmark

Page 20 of the same NVIDIA DGX Spark forum thread captures a May 2026 discussion benchmarking vLLM 0.19.1 against 0.20.1 on Qwen3.5-122B-A10B (hybrid INT4+FP8 + INT8 LM Head + MTP-2) on a single Spark. The author reports a ~7.7–10.5% regression across Q&A, Code, JSON, Math, and LongCode workloads, attributing the drop Community contributors note that the Qwen3.5-122B-A10B gdn attention core (DeltaNet linear attention, 36 of 48 layers) is now split out rather than stitched into one graph, and on a 273 GB/s UMA Spark with low SM count each extra boundary is real overhead; on H100/H200 (where vLLM CI runs) the effect is reportedly invi

NovitaAI

CoverageBenchmark

Roboflow's Playground page for Qwen3.5-122B-A10B confirms the model's core specifications: a high-capacity multimodal Mixture-of-Experts architecture with 122 billion total parameters and roughly 10 billion activated per token, built by Alibaba's Qwen team as part of the Qwen3.5 family. It supports unified text-and-vis On Roboflow's legacy Vision Evals benchmark, Qwen3.5-122B-A10B achieves a 76.12% pass rate, ranking 9th out of 77 evaluated models and outperforming 86% of the field across 67 visual-understanding tasks spanning captioning, classification, document QA, OCR, and object detection. While the Roboflow page is an aggregator

NovitaAI

CoverageBenchmark

Benchmark Qwen3.5 122B A10B API latency, throughput, and cost efficiency. Compare response speed, token output, and pricing for production AI workloads.

NovitaAI

Coverage

The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. $0.26 per million input tokens, $2.08 per million output tokens. 262,144 token context window, maximum output o

Videos about Qwen3.5-122B-A10B

More models around Qwen3.5-122B-A10B