Currently listed through these providers:
Model details
Qwen-VL Plus
Qwen-VL Plus is an enhanced large visual language model built to bridge the gap between textual understanding and visual perception. Designed as a significant upgrade within the Qwen family, it focuses on improving generalization across diverse multimodal tasks. The architecture is specifically optimized for high-definition visual processing, supporting images with resolutions exceeding one million pixels and accommodating a wide variety of aspect ratios. This design intent makes the model particularly effective at tasks requiring granular attention, such as extracting and analyzing complex details from both images and the text embedded within them.
The development of this model leverages unified multimodal pretraining, building upon the foundational capabilities of the broader Qwen language model series. By refining its visual reasoning and instruction-following performance, the model provides a robust tool for users who need to interpret intricate visual data. Its practical strengths lie in its ability to handle high-resolution inputs with precision, making it well-suited for applications that demand deep cognitive understanding of visual content. As an advancement in the Qwen-VL lineage, it serves as a powerful solution for complex visual tasks that require both broad contextual awareness and specific, detail-oriented analysis.
Quick Info
Powered by- Provider
- Alibaba (China)
- Model key
- qwen-vl-plus
- Release date
- Jan 25, 2024
- Last updated
- Aug 15, 2025
- Knowledge cutoff
- 2024-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.115
- Output token cost
- $0.287
Limits
- Output tokens
- 8,192 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare Qwen-VL Plus pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Qwen-VL Plus
No articles yet. Fetch the latest news to show it here.