Sulat.com
AI models
Get 10-25% off
Get 10-25% off from Qwen
Alibaba logo

Model details

Qwen-VL OCR

Qwen-VL OCR is a specialized vision-language model built for the precise task of extracting readable text from images. It handles a wide variety of real-world image types—from photographs and screenshots to scanned receipts, street signs, and code snippets—converting visual text into clean, usable output. The model accepts image inputs alongside optional text instructions and returns accurate text transcriptions, making it a practical tool for workflows that need to digitize document content or capture text from visual sources without manual retyping.

The Qwen-VL OCR model fits into broader document processing and automation pipelines, particularly where image-to-text conversion is needed at scale. Community tools like the OCR Skill available through LobeHub demonstrate how developers integrate the model into lightweight CLI workflows, processing image files with a single command and optionally routing the extracted text to plain text files. This accessibility makes it well-suited for tasks ranging from quick screenshot capture to batch document digitization, supporting formats including JPEG, PNG, GIF, and WebP.

Alibabaqwen-vl-ocrqwen

Quick Info

Powered by
Provider
Alibaba
Model key
qwen-vl-ocr
Release date
Oct 28, 2024
Last updated
Apr 13, 2025
Knowledge cutoff
2024-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.72
Output token cost
$0.72

Limits

Output tokens
4,096 tokens
Context window
34,096 tokens

Transparent token rates

Compare Qwen-VL OCR pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen-VL OCR

Alibaba (China)

Official sourceRelease Notes

The Model Studio changelog ("Model Platform Feature Update Announcements") explicitly lists two entries named "Qwen-VL-OCR Model Update" — one under February 2026 and one under November 2025 — providing name-exact, time-stamped evidence that Alibaba Cloud has shipped at least two discrete updates to the Qwen-VL-OCR mod The scraped excerpt for this changelog index contains only the entry titles and dates; the bodies of the February 2026 and November 2025 Qwen-VL-OCR update entries (which would describe capability, API, or pricing changes) are not included in the supplied excerpt, so any further technical detail cannot be grounded from

Videos about Qwen-VL OCR

More models around Qwen-VL OCR