SiliconFlow (China)
PaddleOCR-VL 1.5: A Deep Dive into the 0.9B Model That Outperforms GPT-4o on Document Parsing Updated March 2026 — Technical breakdown from an AI Engineer’s perspective When Baidu’s …
Model details
PaddleOCR-VL-1.5 is a compact 0.9B-parameter vision-language model from Baidu's PaddlePaddle team, released under the Apache 2.0 license and framed as a multi-task system for in-the-wild document parsing. It is hosted openly on Hugging Face, with the upstream development tracked in the PaddlePaddle/PaddleOCR GitHub repository, and registers on the Hub with an image-text-to-text pipeline tag. The card positions it as the next generation of the PaddleOCR-VL line, taking a small specialist VLM approach rather than a general-purpose large model, in order to target real-world document understanding tasks where layout, text, and visual structure all need to be reasoned over together.
Third-party reporting highlights PaddleOCR-VL-1.5 as a strong small-model performer for document AI, with code aggregation sites citing a 94.5% score on OmniDocBench v1.5 and external commentary claiming it outperforms GPT-4o on document parsing despite its much smaller footprint. That combination of compact size, open weights, and benchmark-leading behaviour makes it a practical fit for teams that need to run document understanding on their own infrastructure, from enterprise PDF pipelines to embedded or edge-style deployments where a 0.9B model is easier to host than a frontier-scale VLM. Practitioners exploring the line should note that Hugging Face now advertises a newer PaddleOCR-VL-1.6 revision, so this 1.5 entry is a snapshot of an actively evolving series.
A provider subscription or plan supersedes token-based pricing for this model.
SiliconFlow (China)
PaddleOCR-VL 1.5: A Deep Dive into the 0.9B Model That Outperforms GPT-4o on Document Parsing Updated March 2026 — Technical breakdown from an AI Engineer’s perspective When Baidu’s …
SiliconFlow (China)
The Hugging Face paper page for PaddleOCR-VL-1.5 (arXiv:2601.21957, published Jan 29, 2026) presents the model as a 0.9B-parameter multi-task Vision-Language Model from Baidu's PaddlePaddle team, led by author Cheng Cui and collaborators. The abstract reports a new state-of-the-art accuracy of 94.5% on OmniDocBench v1. This candidate establishes the model's core technical identity and benchmark claims directly from the paper record rather than third-party summaries. By naming the exact PaddleOCR-VL-1.5 variant (not the 1.0 sibling or the broader PaddleOCR-VL family) and explicitly attributing the work to Baidu's PaddlePaddle team, th
SiliconFlow (China)
The vLLM Recipes page for PaddlePaddle/PaddleOCR-VL-1.5 confirms the model is a 0.9B-parameter Vision-Language Model built on a NaViT-style dynamic-resolution vision encoder paired with an ERNIE-4.5-0.3B language model. It documents installation requirements (vLLM 0.11.1+, paddlepaddle-gpu 3.2.1, paddleocr[doc-parser], The page reports PaddleOCR-VL-1.5 reaches 94.5% on OmniDocBench v1.5, establishing a new state-of-the-art, and adds new capabilities relative to its predecessor: line-level text spotting, seal recognition with SOTA results, polygonal detection of irregular shapes under skew or warping, additional multilingual support f
SiliconFlow (China)
The MDPBench leaderboard (updated May 6, 2026) ranks PaddlePaddle/PaddleOCR-VL-1.5 at #2 with a score of 78.3, behind rednote-hilab/dots.mocr (80.5) and ahead of the original PaddleOCR-VL (69.6). The benchmark evaluates document parsing across 3,400 images in 17 languages, including photographed documents with real-wor The benchmark analysis reveals that specialized VLMs including PaddleOCR-VL-1.5 generally outperform general VLMs but lose accuracy on photographed versus digital documents, and that non-Latin scripts (Arabic, Hindi, Thai) see average performance drops of 14.0% across all model types. A noted caveat is that models spec