Sulat.com
AI models
Get 10-25% off
Get 10-25% off from Qwen
Alibaba logo

Model details

Qwen3-VL 235B-A22B

Qwen3-VL 235B-A22B represents a significant advancement in the Qwen series, functioning as a versatile vision-language model that unifies high-level text generation with deep visual perception. Designed to handle both dense and mixture-of-experts architectures, the model is built to scale effectively across diverse environments, from edge devices to cloud infrastructure. Its core strength lies in its ability to perform complex multimodal reasoning, including 2D and 3D spatial grounding, which allows it to interpret object positions, viewpoints, and occlusions with high precision. By integrating seamless text-vision fusion, the model achieves text understanding performance comparable to pure language models while maintaining specialized capabilities in document parsing, chart extraction, and multilingual optical character recognition.

The model benefits from a comprehensive training approach that emphasizes broad, high-quality pretraining, enabling it to recognize a vast array of real-world categories ranging from landmarks and products to rare characters. Through its instruct-tuned lineage, the model excels as a visual agent capable of operating PC and mobile interfaces by identifying GUI elements and invoking tools to complete tasks. Its architecture supports long-form visual comprehension, allowing it to process hours-long video content with second-level indexing and native long-context recall. These features, combined with its ability to generate code from visual mockups and perform logical, evidence-based reasoning in STEM fields, position the model as a robust solution for production-grade document AI, embodied robotics, and sophisticated software assistance.

Alibabaqwen3-vl-235b-a22bqwen

Quick Info

Powered by
Provider
Alibaba
Model key
qwen3-vl-235b-a22b
Release date
Apr 1, 2025
Last updated
Apr 1, 2025
Knowledge cutoff
2025-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.70
Output token cost
$2.80

Limits

Output tokens
32,768 tokens
Context window
131,072 tokens

Transparent token rates

Compare Qwen3-VL 235B-A22B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen3-VL 235B-A22B

Alibaba

CoverageBenchmark

The Roboflow Playground page documents Qwen3-VL 235B A22B Instruct as a flagship multimodal vision-language model from Qwen (Alibaba Cloud), featuring a mixture-of-experts architecture with approximately 22B active parameters out of 235B total, a 256K-token context window, and support for interleaved text and image inp Roboflow's Vision Evals, updated September 5, 2026, report an overall score of 65.9% (rank 35 of 53) for Qwen3-VL 235B A22B Instruct across six vision tasks, with an average latency of 10.54 seconds and approximately 75 inferences over the prior 30 days. The page also lists a maximum output token setting of 32,768, and

Videos about Qwen3-VL 235B-A22B

More models around Qwen3-VL 235B-A22B