Sulat.com
AI models
CrofAI logo

Model details

DeepSeek V4 Flash Vision Exp

DeepSeek V4 Flash Vision Exp was released on August 21, 2026 as an experimental multimodal API model from the same family as V4 Flash, extending that line with image input while preserving its text-focused behavior. The release was framed as a measured addition rather than a full architecture overhaul, with vision support layered onto the existing text and agent foundation. Input accepts images alongside text, opening use cases such as screenshots, charts, documents, and other visual agent workflows, while outputs continue to flow through the same reasoning and tool-calling paths that defined the V4 Flash series.

For builders, the practical story is that V4 Flash Vision Exp offers a single endpoint where text reasoning and visual understanding meet, so workflows that previously needed a separate vision model can stay within one stack. Community and third-party reports describe a long context window and reported multimodal benchmark results near Claude Opus 4.8 on agent-style tasks, though those benchmark figures and the underlying methodology have not been independently confirmed. Given its experimental status and thin surrounding evidence, the model is best suited for pilots and workload-specific evaluation against real documents and screenshots before committing to production pipelines.

CrofAIdeepseek-v4-flash-vision-expdeepseek-flash

Quick Info

Powered by
Provider
CrofAI
Model key
deepseek-v4-flash-vision-exp
Release date
Aug 21, 2026
Last updated
Aug 21, 2026
AI SDK package
@ai-sdk/openai-compatible
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.08
Output token cost
$0.20

Limits

Output tokens
131,072 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare DeepSeek V4 Flash Vision Exp pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4 Flash Vision Exp

CrofAI

CoverageRelease Notes

AI Weekly reports that DeepSeek published a 305B-parameter experimental multimodal model, DeepSeek-V4-Flash-Vision-Exp, to Hugging Face under an MIT license, marking the first vision model in the V4-Flash family. The article describes the model as adding a vision encoder and aligner on top of the text model's DFlash at The piece highlights benchmark results where Vision-Exp scores 36.5 on ApexBench Pass@1 (up from 26.2 for text-only V4-Flash-0731), 27.3 on Agents' Last Exam, 35.0 on ZeroBench Pass@5, and 64.3 on Chartography, positioning it close to but not matching Opus-4.8's scores. Text-only agent task scores remain narrower in ga

CrofAI

Coverage

DeepSeek's official API documentation change log confirms that on August 21, 2026, the company released DeepSeek-V4-Flash-Vision-Exp as a new experimental multimodal vision understanding model accessible via the DeepSeek API by setting model='deepseek-v4-flash-vision-exp'. The entry provides the complete benchmark tabl According to the same first-party entry, DeepSeek-V4-Flash-Vision-Exp maintains on-par performance with DeepSeek-V4-Flash on pure-text capabilities including agent, reasoning, and world knowledge tasks, while delivering a significant leap on agent benchmarks that require visual understanding, bringing its multimodal ag

Videos about DeepSeek V4 Flash Vision Exp

More models around DeepSeek V4 Flash Vision Exp