Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
OpenRouter logo

Model details

Qwen2.5 VL 72B Instruct

Qwen2.5 VL 72B Instruct is a flagship multimodal model designed to bridge the gap between visual perception and complex reasoning. Beyond standard object recognition, it excels at interpreting intricate visual data such as charts, technical graphics, and document layouts. Its architecture is built to function as a visual agent, enabling it to dynamically direct tools for computer and phone use. A key design innovation is the integration of dynamic resolution and frame rate training, which allows the model to process video content spanning over an hour while pinpointing specific events with high precision.

The model benefits from a refined training lineage that emphasizes temporal understanding through dynamic FPS sampling and updated temporal mRoPE mechanisms. This foundation supports its ability to perform precise visual localization, generating accurate bounding boxes and structured JSON outputs for coordinates and attributes. These capabilities make it a practical choice for data-heavy workflows in finance and commerce, such as processing invoices or complex forms. As an agentic tool, it is well-positioned for future-facing applications that require autonomous interaction with digital environments and deep analysis of long-form visual media.

OpenRouterqwen/qwen2.5-vl-72b-instructqwen

Quick Info

Powered by
Provider
OpenRouter
Model key
qwen/qwen2.5-vl-72b-instruct
Release date
Feb 1, 2025
Last updated
Feb 1, 2025
Knowledge cutoff
2024-06-30
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.80
Output token cost
$1.00

Limits

Output tokens
115,200 tokens
Context window
128,000 tokens

Transparent token rates

Compare Qwen2.5 VL 72B Instruct pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen2.5 VL 72B Instruct

OpenRouter

CoverageBenchmark

The LLM-Stats aggregator page for Qwen2.5 VL 72B Instruct places the model at composite rank 249 overall, with capability-tier standings of 78/119 in Long Context, 130/208 in Vision, 147/244 in Healthcare, 240/362 in Reasoning, and 263/327 in Math. The page also tracks how the model holds up as conversation length grow Per-benchmark scores sourced from the model's own scorecard, paper, or official blog posts include DocVQA rank 1 at 0.96, Android Control Low EM rank 1 at 0.94, ChartQA rank 3 at 0.90, and OCRBench rank 10 at 0.89, with 26 additional benchmarks listed. LLM-Stats notes that these scores are self-reported by the model pr

OpenRouter

Coverage

The Qwen team has published the official Hugging Face model card for Qwen2.5-VL-72B-Instruct, describing the 72-billion-parameter vision-language model in the three-size Qwen2.5-VL lineup (3B/7B/72B). The card documents key capability enhancements over Qwen2-VL: richer visual understanding of objects, text, charts, ico Architecture updates include extending dynamic resolution to the temporal dimension via dynamic FPS sampling, updating mRoPE's time dimension with IDs and absolute-time alignment so the model can learn temporal sequence and speed, and streamlining the ViT with window attention plus SwiGLU and RMSNorm to align with Qwen

Videos about Qwen2.5 VL 72B Instruct

More models around Qwen2.5 VL 72B Instruct