Currently listed through these providers:
Model details
Qwen3 VL 235B A22B Thinking
Qwen3-VL-235B-A22B-Thinking is the flagship reasoning-enhanced vision-language model in the Qwen3 generation, designed to push the boundaries of multimodal understanding and complex problem-solving. The model uses a Mixture-of-Experts architecture with 235 billion total parameters and 22 billion active parameters, allowing it to route tasks to specialized expert subnetworks efficiently. This Thinking edition extends chain-of-thought reasoning specifically for intricate visual and multimodal challenges, making it particularly strong at causal analysis, logical deduction, and providing evidence-based answers in STEM and mathematics domains. Its capabilities span visual agent tasks like operating desktop and mobile interfaces, generating code from diagrams and images, perceiving spatial relationships in both 2D and 3D environments, and handling long-context inputs up to 256K tokens natively with expandability to 1 million tokens for processing books or hours of video content with precise second-level indexing.
The model was developed with comprehensive training across visual recognition, OCR, and text understanding, achieving performance on par with pure language models while seamlessly fusing vision and text modalities. Visual recognition was expanded significantly to identify celebrities, products, landmarks, flora and fauna, while OCR now supports 32 languages with robust handling of low-light, blur, and tilted inputs alongside rare characters and domain-specific jargon. Available under the Apache 2.0 license, the model runs well on local deployments through Ollama, with Azure Foundry providing FP8-optimized cloud serving and vLLM powering the inference stack. Benchmark scores show strong performance on ZebraLogic logical reasoning puzzles, DocVQA document understanding, MM-MT-Bench multi-turn instruction following, and ScreenSpot GUI grounding tasks, positioning this as a versatile foundation for applications ranging from document analysis and visual automation to multimodal reasoning systems and long-form video understanding.
Quick Info
Powered by- Provider
- DevPass (LLM Gateway)
- Model key
- qwen3-vl-235b-a22b-thinking
- Release date
- Sep 15, 2025
- Last updated
- Sep 15, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.98
- Output token cost
- $3.95
Limits
- Output tokens
- 8,192 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare Qwen3 VL 235B A22B Thinking pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Qwen3 VL 235B A22B Thinking
No articles yet. Fetch the latest news to show it here.