Sulat.com
AI models
Get 10-25% off
Get 10-25% off from Qwen
Alibaba (China) logo

Model details

Qwen2.5-VL 72B Instruct

Qwen2.5-VL represents a substantial evolution of the Qwen vision-language lineage, designed as a flagship model that bridges visual perception with language reasoning. The architecture extends dynamic resolution processing into the temporal dimension through dynamic FPS sampling, enabling the model to process video content with varying frame rates while maintaining spatial detail. As a multimodal model, it handles image inputs alongside text and can generate structured JSON outputs for coordinates, attributes, and document contents. Its visual agent capabilities allow it to reason about what it perceives and dynamically direct tools, supporting use cases such as computer operation and phone interaction. The 72B parameter scale provides the capacity to handle complex visual understanding tasks, from analyzing charts and layouts to identifying objects within images and providing precise bounding box localizations.

The development of Qwen2.5-VL incorporated feedback gathered over five months following the Qwen2-VL release, during which numerous developers built upon the earlier vision-language models. The model family spans three sizes—3B, 7B, and 72B parameters—with both base and instruct variants released openly on Hugging Face and ModelScope. Key advancements include the ability to comprehend videos exceeding one hour in length and to capture events by pinpointing relevant segments within that content. For document-heavy workflows, the model excels at extracting structured information from invoices, forms, and tables, making it particularly useful in finance and commerce applications. The combination of open weights, vision-language capabilities, and structured output generation positions this model for developers seeking to build multimodal pipelines that require both visual understanding and reliable machine-readable results.

Alibaba (China)qwen2-5-vl-72b-instructqwen

Quick Info

Powered by
Provider
Alibaba (China)
Model key
qwen2-5-vl-72b-instruct
Release date
Sep 1, 2024
Last updated
Sep 1, 2024
Knowledge cutoff
2024-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$2.294
Output token cost
$6.881

Limits

Output tokens
8,192 tokens
Context window
131,072 tokens

Transparent token rates

Compare Qwen2.5-VL 72B Instruct pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen2.5-VL 72B Instruct

Alibaba (China)

CoverageBenchmark

The llm-stats leaderboard page for Qwen2.5-VL 72B Instruct places the model at overall rank 250 with a "Bad" capability tier, and shows mixed standing across categories: Long Context ranks 78 of 119, Vision 130 of 208, Healthcare 147 of 244, Reasoning 241 of 363, and Math 263 of 327. The page aggregates scores across 2 Notable benchmark results include DocVQA at rank 1 with a score of 0.96, Android Control Low EM at rank 1 with 0.94, ChartQA at rank 3 with 0.90, and OCRBench at rank 10 with 0.89, all sourced via huggingface.co. These figures highlight strength on document visual question answering, mobile agent control, chart reasoni

Alibaba (China)

Coverage

The official Qwen Hugging Face model card introduces Qwen2.5-VL-72B-Instruct as the instruction-tuned 72B entry in the Qwen2.5-VL vision-language family, released as the successor to Qwen2-VL. Key enhancements include strong visual recognition of objects, text, charts, icons, and layouts; agentic capabilities enabling Architecturally, Qwen2.5-VL extends dynamic resolution to the temporal dimension through dynamic FPS sampling, updating mRoPE with time-dimension IDs and absolute time alignment so the model can learn temporal sequence, speed, and pinpoint specific moments. The vision encoder was streamlined with window attention in th

Videos about Qwen2.5-VL 72B Instruct

More models around Qwen2.5-VL 72B Instruct