Currently listed through these providers:
Model details
Qwen3 VL 235B A22B Instruct
Qwen3 VL 235B A22B is a massive open-weight vision-language model built on a Mixture-of-Experts architecture that activates 22 billion parameters while utilizing 235 billion total, enabling it to deliver frontier-level multimodal intelligence without proportional compute costs. The model was designed to unify strong text generation with deep visual understanding, supporting tasks ranging from general visual question-answering to sophisticated 2D and 3D spatial grounding. Its capabilities extend to visual agent workflows—operating PC and mobile GUIs, recognizing interface elements, and invoking tools to complete multi-step tasks—as well as visual coding that transforms mockups and whiteboard sketches directly into functional code. With native support for 256K token context (expandable to 1M) and 32-language OCR, it handles long documents, hours-long video, and multilingual content with full recall and second-level indexing.
As an open-weights model released under Apache 2.0, Qwen3 VL provides accessibility for researchers and developers who need to run, fine-tune, or self-host a capable vision-language system. The Instruct variant reflects a post-training lineage aimed at following complex instructions and maintaining coherent multi-turn dialogues, while the broader Qwen3-VL family includes reasoning-enhanced Thinking editions for deeper chain-of-thought exploration. Benchmarks show Qwen3 VL significantly outperforming comparable proprietary models on visual reasoning tasks like OSWorld, while costing a fraction of the price—making it particularly attractive for high-volume applications where quality and cost efficiency both matter. Its combination of text-on-par performance, enhanced visual recognition across real-world and synthetic categories, and strong agent interaction capabilities positions it well for demanding workflows in document intelligence, GUI automation, and multimodal research environments.
Quick Info
Powered by- Provider
- DevPass (LLM Gateway)
- Model key
- qwen3-vl-235b-a22b-instruct
- Release date
- Sep 15, 2025
- Last updated
- Sep 15, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.20
- Output token cost
- $0.88
Limits
- Output tokens
- 8,192 tokens
- Context window
- 262,144 tokens
Transparent token rates
Compare Qwen3 VL 235B A22B Instruct pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Qwen3 VL 235B A22B Instruct
No articles yet. Fetch the latest news to show it here.