Currently listed through these providers:
Model details
DeepSeek V4 Flash Vision Exp
DeepSeek V4 Flash Vision Exp extends the V4 Flash lineage into multimodal territory by adding visual modules and continued training on top of the existing text-and-agent foundation. The model is positioned by DeepSeek as an experimental entry in the V4 family, intended to handle image-augmented tasks such as describing pictures, reading text from screenshots, and analyzing charts while preserving the agent and reasoning strengths of its text-only predecessor. Because it shares architectural roots with V4 Flash, it is best understood as a vision-capable variant of a model that was already tuned for agentic workflows rather than a standalone vision specialist.
In practice, the model accepts image inputs alongside text using DeepSeek's standard OpenAI-compatible Chat Completions interface, where content is supplied as an array of blocks; the same three image-delivery methods are also available through the Responses API via input image content parts. DeepSeek's official documentation lists JPEG, PNG, GIF, and WebP as supported formats, with format detection based on the actual file bytes rather than the filename or declared MIME type, which makes integration straightforward for pipelines that already handle multimodal content. This combination of a familiar API surface, broad format support, and a lineage tuned for agent and tool-calling behavior makes the experimental vision model a natural fit for developers who want to add screenshot or document understanding to existing text-driven DeepSeek workflows.
Quick Info
Powered by- Provider
- Deep Infra
- Model key
- deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
- Release date
- Aug 21, 2026
- Last updated
- Aug 21, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.44
- Output token cost
- $1.32
Limits
- Output tokens
- 384,000 tokens
- Context window
- 1,048,576 tokens
Transparent token rates
Compare DeepSeek V4 Flash Vision Exp pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about DeepSeek V4 Flash Vision Exp
No articles yet. Fetch the latest news to show it here.
Videos about DeepSeek V4 Flash Vision Exp
More models around DeepSeek V4 Flash Vision Exp
This exact model name is also listed by 18 other providers.