Model details
Qwen2.5-VL-72B-Instruct
Qwen2.5-VL-72B-Instruct serves as a flagship vision-language model designed to bridge the gap between visual perception and complex reasoning. Beyond standard object recognition, the model is engineered to analyze intricate visual data including charts, icons, graphics, and document layouts. Its architecture is built to function as a visual agent, enabling it to direct tools for computer and phone use. A significant design advancement is the implementation of dynamic resolution and frame rate training, which extends dynamic resolution to the temporal dimension through dynamic FPS sampling. This allows the model to process videos exceeding one hour in length while pinpointing specific events with high precision.
The model demonstrates robust capabilities in visual localization, generating accurate bounding boxes and coordinates for specific image elements. It is optimized for structured data extraction, making it a practical choice for digitizing invoices, forms, and tables into stable JSON formats. By leveraging updated mRoPE in the time dimension, the model maintains high performance across varied temporal sampling rates. Its design supports flexible integration, allowing for fine-tuning via methods like LoRA to adapt the model to specialized datasets. This combination of agentic reasoning, long-form video analysis, and structured output generation positions it as a versatile tool for professional applications in finance, commerce, and automated task execution.
Quick Info
Powered by- Provider
- OVHcloud AI Endpoints
- Model key
- qwen2.5-vl-72b-instruct
- Release date
- Mar 31, 2025
- Last updated
- Mar 31, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.01
- Output token cost
- $1.01
Limits
- Output tokens
- 32,768 tokens
- Context window
- 32,768 tokens
Latest news about Qwen2.5-VL-72B-Instruct
No articles yet. Fetch the latest news to show it here.