Model details
DeepSeek-OCR
DeepSeek-OCR is a specialized vision-language model designed to redefine document intelligence through a unique architectural approach known as context optical compression. By mapping large 2D visual inputs into highly compressed vision tokens, the model bypasses the limitations of traditional tokenization, allowing it to handle extensive documents and complex visual data with significantly improved efficiency. This design intent focuses on mitigating the computational overhead typically associated with long-context processing, positioning the model as a powerful tool for tasks ranging from structured data extraction to the interpretation of handwritten notes and intricate charts.
The model is built upon a 3B parameter architecture, reflecting a focused effort to balance performance with resource optimization. While the supplied evidence does not detail the specific post-training recipes such as SFT or GRPO, it highlights that the model represents a paradigm shift in how information is stored and retrieved. By prioritizing the compression of visual information, the development lineage emphasizes a departure from standard language model conventions, aiming to solve the problem of context rot where models struggle to maintain coherence over lengthy interactions.
Quick Info
Powered by- Provider
- NovitaAI
- Model key
- deepseek/deepseek-ocr
- Release date
- Oct 24, 2025
- Last updated
- Oct 24, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.03
- Output token cost
- $0.03
Limits
- Output tokens
- 8,192 tokens
- Context window
- 8,192 tokens