Sulat.com
AI models
Vercel AI Gateway logo

Model details

Qwen3 VL 235B A22B Instruct

Qwen3 VL 235B A22B Instruct is the most powerful vision-language model in the Qwen series to date, built on a Mixture-of-Experts architecture that activates 22 billion parameters while maintaining 235 billion total parameters. The model was designed to excel at visual perception and reasoning tasks, with comprehensive upgrades across visual coding, spatial understanding, and multimodal reasoning. Its Visual Agent capabilities allow it to operate across PC and mobile GUIs—recognizing interface elements, understanding their functions, and invoking tools to complete tasks autonomously. Advanced spatial perception gives it the ability to judge object positions, viewpoints, and occlusions, enabling both stronger 2D grounding and 3D grounding for spatial reasoning and embodied AI applications. The model also features a Visual Coding Boost, generating structured outputs like Draw.io diagrams, HTML, CSS, and JavaScript directly from images or videos.

The Qwen3 VL series was trained with extended context in mind, supporting native 256K token contexts with expandability to 1M for handling books and hours-long video content with full recall and second-level indexing. Its OCR capabilities were substantially upgraded to recognize text across 32 languages, handling challenging conditions like low light, blur, and tilt with improved accuracy for rare characters and ancient scripts. Benchmarks show the model performing exceptionally well on document understanding tasks (DocVQA rank 1), multimodal multi-turn instruction following (MM-MT-Bench rank 2), and GUI grounding (ScreenSpot rank 3). The model excels in STEM and math reasoning, delivering causal analysis and logical, evidence-based answers. Organizations can customize the model through fine-tuning using LoRA, enabling efficient adaptation to specific domains while maintaining the base model's strong visual and reasoning capabilities.

Vercel AI Gatewayalibaba/qwen3-vl-235b-a22b-instructqwen

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
alibaba/qwen3-vl-235b-a22b-instruct
Release date
Sep 23, 2025
Last updated
Sep 23, 2025
Knowledge cutoff
2025-03-31
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.40
Output token cost
$1.60

Limits

Output tokens
129,024 tokens
Context window
131,072 tokens

Latest news about Qwen3 VL 235B A22B Instruct

Videos about Qwen3 VL 235B A22B Instruct

Recent tweets and retweets from Vercel AI Gateway

More models around Qwen3 VL 235B A22B Instruct