Sulat.com
AI models
Get 10-25% off
Get 10-25% off from Qwen
Alibaba (China) logo

Model details

Qwen-VL OCR

Qwen-VL OCR is a specialized visual language model developed as part of the broader Qwen series. It is engineered specifically for optical character recognition and document text extraction, serving as a dedicated tool for converting visual information into structured text. By leveraging the multimodal capabilities of the Qwen architecture, the model is designed to interpret and digitize content from a variety of sources, including photos, screenshots, receipts, signs, and code snippets.

The model benefits from a lineage of large-scale multilingual and multimodal pre-training, followed by post-training on quality data to align its outputs with human preferences. Its practical strengths lie in its ability to handle diverse image formats such as JPEG, PNG, GIF, and WebP, making it a reliable choice for developers building automated digitization workflows. With its capacity for high-volume text extraction, the model is well-suited for integration into command-line tools and automated pipelines that require precise and efficient text capture from visual media.

Alibaba (China)qwen-vl-ocrqwen

Quick Info

Powered by
Provider
Alibaba (China)
Model key
qwen-vl-ocr
Release date
Oct 28, 2024
Last updated
Sep 11, 2026
Knowledge cutoff
2024-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.043
Output token cost
$0.072

Limits

Output tokens
4,096 tokens
Context window
34,096 tokens

Transparent token rates

Compare Qwen-VL OCR pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen-VL OCR

Alibaba (China)

Official sourceRelease Notes

The Model Studio changelog ("Model Platform Feature Update Announcements") explicitly lists two entries named "Qwen-VL-OCR Model Update" — one under February 2026 and one under November 2025 — providing name-exact, time-stamped evidence that Alibaba Cloud has shipped at least two discrete updates to the Qwen-VL-OCR mod The scraped excerpt for this changelog index contains only the entry titles and dates; the bodies of the February 2026 and November 2025 Qwen-VL-OCR update entries (which would describe capability, API, or pricing changes) are not included in the supplied excerpt, so any further technical detail cannot be grounded from

Videos about Qwen-VL OCR

More models around Qwen-VL OCR