Sulat.com
AI models
Get 10-25% off
Get 10-25% off from Qwen
Alibaba (China) logo

Model details

Qwen3-VL 235B-A22B

Qwen3-VL 235B-A22B is a vision-language model in the Qwen3 family that uses a sparse Mixture-of-Experts design, with 235 billion total parameters and 22 billion activated per inference. That architecture makes it well suited to multimodal chat, single- and multi-image reasoning, OCR-style visual question answering, and long-context generation, with hosted documentation pointing to context windows reaching into the hundreds of thousands of tokens for extended document and conversation workloads. The open-weight release and MoE structure position it as a flexible foundation for teams that want frontier-scale visual understanding without paying the full compute cost of a dense model at every step.

In practical deployment, the model benefits from a maturing serving ecosystem. vLLM Ascend provides a dedicated tutorial covering single-node and multi-node deployment, Prefill-Decode disaggregation, and accuracy and performance tuning, while managed platforms expose the weights for on-demand inference with image input, function calling, and calibration support. The combination of image understanding, long context, and function calling makes Qwen3-VL 235B-A22B a practical choice for vision-grounded assistants, document and chart analysis pipelines, and agent workflows that need to reason over rich visual inputs alongside text.

Alibaba (China)qwen3-vl-235b-a22bqwen

Quick Info

Powered by
Provider
Alibaba (China)
Model key
qwen3-vl-235b-a22b
Release date
Apr 1, 2025
Last updated
Apr 1, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.286705
Output token cost
$1.14682

Limits

Output tokens
32,768 tokens
Context window
131,072 tokens

Transparent token rates

Compare Qwen3-VL 235B-A22B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3-VL 235B-A22B

Alibaba (China)

CoverageBenchmark

The Roboflow Playground benchmark page for Qwen3 VL 235B A22B Instruct presents the model as a flagship multimodal vision-language model developed by Qwen (Alibaba Cloud) under an Apache 2.0 license, with 235B parameters and a 256K token context window. It confirms the MoE architecture with about 22B active parameters, Roboflow's Vision Evals benchmark, updated September 5, 2026, places Qwen3 VL 235B A22B Instruct at an overall score of 65.9 percent, ranking it 35 of 53 models evaluated across six real-world vision tasks. The page records 75 inferences in the past 30 days and lists the model's release date as September 2025, an earli

Videos about Qwen3-VL 235B-A22B

More models around Qwen3-VL 235B-A22B