Model details
QVQ Max
QVQ-Max is a visual reasoning model from Qwen (Alibaba) built to move beyond basic image recognition. Rather than simply identifying objects or scenes, it combines visual perception with logical reasoning, enabling it to tackle complex tasks such as mathematical reasoning over images, multi-image analysis, and video understanding. This approach positions the model as a tool for applications where interpreting visual data requires genuine comprehension rather than pattern matching alone.
The model targets developers and teams building applications that need sophisticated visual understanding, from document and diagram processing to scientific image analysis. With an extensive context window and support for multi-step reasoning, it handles tasks that demand both visual input and layered thinking. Its design fits use cases where users need answers derived from analyzing image sequences or extracting structured information from complex visual content.
Quick Info
Powered by- Provider
- Alibaba
- Model key
- qvq-max
- Release date
- Mar 25, 2025
- Last updated
- Mar 25, 2025
- Knowledge cutoff
- 2024-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.20
- Output token cost
- $4.80
Limits
- Output tokens
- 8,192 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare QVQ Max pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about QVQ Max
No articles yet. Fetch the latest news to show it here.