Currently listed through these providers:
Model details
Qwen3 VL 235B A22B Instruct
Qwen3 VL 235B A22B Instruct is the most powerful vision-language model in the Qwen series to date, built on a Mixture-of-Experts architecture that activates 22 billion parameters while maintaining 235 billion total parameters. The model was designed to excel at visual perception and reasoning tasks, with comprehensive upgrades across visual coding, spatial understanding, and multimodal reasoning. Its Visual Agent capabilities allow it to operate across PC and mobile GUIs—recognizing interface elements, understanding their functions, and invoking tools to complete tasks autonomously. Advanced spatial perception gives it the ability to judge object positions, viewpoints, and occlusions, enabling both stronger 2D grounding and 3D grounding for spatial reasoning and embodied AI applications. The model also features a Visual Coding Boost, generating structured outputs like Draw.io diagrams, HTML, CSS, and JavaScript directly from images or videos.
The Qwen3 VL series was trained with extended context in mind, supporting native 256K token contexts with expandability to 1M for handling books and hours-long video content with full recall and second-level indexing. Its OCR capabilities were substantially upgraded to recognize text across 32 languages, handling challenging conditions like low light, blur, and tilt with improved accuracy for rare characters and ancient scripts. Benchmarks show the model performing exceptionally well on document understanding tasks (DocVQA rank 1), multimodal multi-turn instruction following (MM-MT-Bench rank 2), and GUI grounding (ScreenSpot rank 3). The model excels in STEM and math reasoning, delivering causal analysis and logical, evidence-based answers. Organizations can customize the model through fine-tuning using LoRA, enabling efficient adaptation to specific domains while maintaining the base model's strong visual and reasoning capabilities.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- alibaba/qwen3-vl-235b-a22b-instruct
- Release date
- Sep 23, 2025
- Last updated
- Sep 23, 2025
- Knowledge cutoff
- 2025-03-31
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.40
- Output token cost
- $1.60
Limits
- Output tokens
- 129,024 tokens
- Context window
- 131,072 tokens