Currently listed through these providers:
Model details
Qwen: Qwen2.5 VL 72B Instruct
Qwen2.5-VL-72B-Instruct is a 72-billion-parameter multimodal model built to bridge the gap between visual perception and complex reasoning. It is designed to move beyond simple object recognition, demonstrating high proficiency in interpreting intricate visual data such as charts, icons, graphics, and document layouts. By functioning as a visual agent, the model is capable of dynamic tool use, enabling it to perform computer and phone-based tasks. Its architecture is specifically optimized for structured data extraction, making it a reliable tool for processing invoices, forms, and tables in professional environments.
The model incorporates significant architectural advancements, including dynamic resolution and frame rate training that extends to the temporal dimension. By adopting dynamic FPS sampling and updating the mRoPE mechanism in the time dimension, the model achieves a robust ability to comprehend videos exceeding one hour in length, including the precise pinpointing of relevant events. These design choices allow the model to provide stable, structured outputs for coordinates and attributes, positioning it as a versatile solution for applications requiring both deep visual localization and long-form video analysis.
Quick Info
Powered by- Provider
- Kilo Gateway
- Model key
- qwen/qwen2.5-vl-72b-instruct
- Release date
- Feb 1, 2025
- Last updated
- Feb 1, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.80
- Output token cost
- $1.00
Limits
- Output tokens
- 115,200 tokens
- Context window
- 128,000 tokens
Transparent token rates
Compare Qwen: Qwen2.5 VL 72B Instruct pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.