Clarifai
Deepseek has unveiled a vision encoder that processes image information based on meaning rather than position. The approach uses far fewer tokens and improves document recognition.
Model details
DeepSeek-OCR is introduced in the paper "Contexts Optical Compression" as an initial investigation into compressing long contexts by rendering them as 2D images, an approach that shifts heavy text representation into a compact visual format. The architecture pairs a purpose-built DeepEncoder with a DeepSeek3B-MoE-A570M decoder, where the encoder is designed to keep activations low under high-resolution input while reaching high compression ratios so that downstream vision tokens remain manageable. This split lets the model behave like an OCR system whose core innovation is how it represents documents, not how it transcribes them.
The compression-driven design shows clear practical strengths: when text tokens are within ten times the number of vision tokens, decoding precision reaches about 97%, and even at a 20x compression ratio the OCR accuracy stays near 60%, suggesting graceful degradation as documents grow denser. Because the visual tokens carry the bulk of the context, the approach is well suited to very long documents, historical archives, and structured forms where layout matters as much as words. A successor line, DeepSeek-OCR 2, extends this direction with a causal visual flow encoder aimed at layout-aware understanding, signaling continued progress toward more human-like machine vision for document AI.
Clarifai
Deepseek has unveiled a vision encoder that processes image information based on meaning rather than position. The approach uses far fewer tokens and improves document recognition.
Clarifai
DeepSeek AI Releases DeepSeek-OCR 2 with Causal Visual Flow Encoder for Layout Aware Document Understanding
Clarifai
Chinese AI startup DeepSeek on Tuesday released a research paper and open-sourced its latest optical character recognition (OCR) model, DeepSeek-OCR 2,