Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
NovitaAI logo

Model details

Qwen2.5 VL 72B Instruct

Qwen2.5 VL 72B Instruct is a flagship multimodal model engineered to bridge the gap between visual perception and complex reasoning. Beyond standard object recognition, the model is built to interpret intricate visual data, including charts, icons, graphics, and document layouts. Its design intent centers on agentic functionality, allowing it to act as a visual agent capable of computer and phone use. By integrating dynamic resolution and frame rate training, the model achieves a sophisticated understanding of temporal data, enabling it to process videos longer than an hour and pinpoint specific events with high precision.

The model benefits from architectural refinements such as the adoption of dynamic FPS sampling and updates to mRoPE in the temporal dimension, which enhance its ability to handle diverse video inputs. It is designed to support practical, structured workflows, offering the ability to generate stable JSON outputs for coordinates and attributes, which is particularly useful for automating tasks like invoice and form processing. With its capacity for visual localization through bounding boxes or points, the model serves as a robust tool for developers looking to build systems that require deep integration between visual analysis and structured data generation.

NovitaAIqwen/qwen2.5-vl-72b-instructqwen

Quick Info

Powered by
Provider
NovitaAI
Model key
qwen/qwen2.5-vl-72b-instruct
Release date
Mar 25, 2025
Last updated
Mar 25, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.80
Output token cost
$0.80

Limits

Output tokens
32,768 tokens
Context window
32,768 tokens

Transparent token rates

Compare Qwen2.5 VL 72B Instruct pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen2.5 VL 72B Instruct

OpenRouter

CoverageBenchmark

The LLM-Stats aggregator page for Qwen2.5 VL 72B Instruct places the model at composite rank 249 overall, with capability-tier standings of 78/119 in Long Context, 130/208 in Vision, 147/244 in Healthcare, 240/362 in Reasoning, and 263/327 in Math. The page also tracks how the model holds up as conversation length grow Per-benchmark scores sourced from the model's own scorecard, paper, or official blog posts include DocVQA rank 1 at 0.96, Android Control Low EM rank 1 at 0.94, ChartQA rank 3 at 0.90, and OCRBench rank 10 at 0.89, with 26 additional benchmarks listed. LLM-Stats notes that these scores are self-reported by the model pr

OpenRouter

Coverage

The Qwen team has published the official Hugging Face model card for Qwen2.5-VL-72B-Instruct, describing the 72-billion-parameter vision-language model in the three-size Qwen2.5-VL lineup (3B/7B/72B). The card documents key capability enhancements over Qwen2-VL: richer visual understanding of objects, text, charts, ico Architecture updates include extending dynamic resolution to the temporal dimension via dynamic FPS sampling, updating mRoPE's time dimension with IDs and absolute-time alignment so the model can learn temporal sequence and speed, and streamlining the ViT with window attention plus SwiGLU and RMSNorm to align with Qwen

Videos about Qwen2.5 VL 72B Instruct

More models around Qwen2.5 VL 72B Instruct