Sulat.com
AI models
Get 10-25% off
Get 10-25% off from Qwen
Alibaba logo

Model details

Qwen2.5-VL 72B Instruct

Qwen2.5-VL-72B-Instruct represents a significant evolution in vision-language architecture, designed to function as a versatile visual agent. Beyond basic object recognition, the model excels at interpreting complex visual data, including charts, icons, technical layouts, and dense text. Its design intent centers on high-level reasoning and tool orchestration, enabling it to perform tasks like computer and phone interaction. By integrating advanced visual localization, it can pinpoint specific elements within an image or video and generate precise, structured outputs, making it particularly effective for data extraction from invoices, forms, and complex documents.

The model benefits from architectural refinements that extend its capabilities into the temporal domain, specifically through dynamic resolution and frame rate training. By adopting dynamic FPS sampling and updating the mRoPE mechanism for time, the model can process videos exceeding one hour in length, identifying and capturing relevant events with high accuracy. These advancements in training lineage allow the model to maintain stability when generating coordinates and attributes in JSON format. As a result, it is well-suited for professional applications in finance and commerce that require reliable, structured analysis of both static imagery and long-duration video content.

Alibabaqwen2-5-vl-72b-instructqwen

Quick Info

Powered by
Provider
Alibaba
Model key
qwen2-5-vl-72b-instruct
Release date
Sep 1, 2024
Last updated
Sep 1, 2024
Knowledge cutoff
2024-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$2.80
Output token cost
$8.40

Limits

Output tokens
8,192 tokens
Context window
131,072 tokens

Transparent token rates

Compare Qwen2.5-VL 72B Instruct pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen2.5-VL 72B Instruct

Alibaba

Coverage

The Qwen team's Hugging Face model card for Qwen2.5-VL-72B-Instruct describes it as the 72-billion-parameter, instruction-tuned entry in the Qwen2.5-VL vision-language family, released alongside 3B and 7B variants. Key claimed capabilities include visual recognition of common objects plus analysis of text, charts, icon Architecture updates detailed on the card include extending dynamic-resolution processing to the temporal dimension through dynamic FPS sampling, updating mRoPE with absolute-time alignment so the model can learn temporal sequence and speed, and a streamlined vision encoder that adds window attention to the ViT alongsi

Videos about Qwen2.5-VL 72B Instruct

More models around Qwen2.5-VL 72B Instruct