Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
NovitaAI logo

Model details

PaddleOCR-VL

PaddleOCR-VL is purpose-built as an ultra-compact vision-language model aimed squarely at the messy reality of document understanding, where pages mix printed prose, tables, equations, charts, and many languages. The design centers on a NaViT-style dynamic-resolution visual encoder paired with a lightweight language stack derived from ERNIE-4.5-0.3B, giving the roughly 0.9B parameter system an encoder-decoder shape that can ingest an image and emit structured text without ballooning into a general-purpose giant. That compact footprint is the defining design choice: the team prioritized resource efficiency so the model can run where heavier parsers cannot, while still handling the wide range of layouts a real-world document pipeline encounters, including low-resolution scans and densely packed multilingual pages spanning 109 languages. The intent is a small specialist that feels closer to a perception engine than a chat model, focused on turning pixels into clean, machine-readable structure.

The PaddlePaddle team at Baidu trained this model with progressive post-training and region-refinement techniques aimed specifically at under-optimized document regions, a recipe reflected in later checkpoints like PaddleOCR-VL-1.5 and PaddleOCR-VL-1.6 that continue to push the frontier of compact document parsing. On the OmniDocBench V1.5 benchmark the 0.9B model reaches 92.56, outperforming substantially larger systems including GPT-4o and DeepSeek-OCR on document parsing, which is a striking result given the parameter budget. In practice that combination of small size, open weights, and strong benchmark performance makes the model attractive for retrieval-augmented pipelines that need reliable extraction of tables, formulas, and multilingual text on commodity hardware. The lineage of incremental post-training releases points to a model that is expected to keep improving through targeted optimization rather than sheer scale, fitting neatly into RAG workflows and production document automation where cost, latency, and accuracy all matter.

NovitaAIpaddlepaddle/paddleocr-vl

Quick Info

Powered by
Provider
NovitaAI
Model key
paddlepaddle/paddleocr-vl
Release date
Oct 22, 2025
Last updated
Oct 22, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.02
Output token cost
$0.02

Limits

Output tokens
16,384 tokens
Context window
16,384 tokens

Latest news about PaddleOCR-VL

NovitaAI

CoverageAnalysis

PaddleOCR-VL 1.5: A Deep Dive into the 0.9B Model That Outperforms GPT-4o on Document Parsing Updated March 2026 — Technical breakdown from an AI Engineer’s perspective When Baidu’s …

Videos about PaddleOCR-VL