Sulat.com
AI models
Hugging Face logo

Model details

DeepSeek V4 Flash Vision Exp

DeepSeek V4 Flash Vision Exp extends the Flash family with experimental multimodal vision understanding, accepting images alongside text so users can describe pictures, read text from screenshots, and analyze charts. It supports JPEG, PNG, GIF, and WebP inputs, with format detection taken from the actual file content rather than filenames or declared MIME types, and it slots into the existing OpenAI-compatible Chat Completions and Responses APIs against the standard DeepSeek endpoint. The vision guide documents three ways to feed images, including base64-encoded inline data within content-block arrays, so existing text-based integrations can be extended to visual inputs without changing their overall request shape.

According to the release notes, this experimental variant is positioned for agent-style and visual reasoning workloads, with reported scores such as Terminal Bench 2.1 at 83.9, DeepSWE at 59.3, DSBench-Hard at 63.6, ApexBench Pass@1 at 36.5, Chartography at 64.3, ZeroBench Pass@5 at 35.0, and Agents' Last Exam at 27.3, alongside NL2Repo at 57.7 and AutomationBench Public at 25.7. The release framing notes that pure-text capabilities remain on par with the base Flash model, while benchmarks that require visual understanding show a substantial lift, closing the gap to higher-tier reasoning systems on multimodal agent tasks. Practically, it is best suited to developers who want to add image-aware tool calling and structured outputs to agent pipelines without leaving the DeepSeek API surface.

Hugging Facedeepseek-ai/DeepSeek-V4-Flash-Vision-Expdeepseek-flash

Quick Info

Powered by
Provider
Hugging Face
Model key
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
Release date
Aug 21, 2026
Last updated
Aug 21, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.44
Output token cost
$1.32

Limits

Output tokens
384,000 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare DeepSeek V4 Flash Vision Exp pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4 Flash Vision Exp

Hugging Face

Official sourceDiscourse

A community pull request (#12) titled "Update Model Card for SGLang support" was opened against the DeepSeek-V4-Flash-Vision-Exp repository by contributor zjxia and merged roughly an hour later by GeeeekExplorer from the DeepSeek org. The change added 13 lines and removed 0 lines in the model card, signaling that SGLan This SGLang integration matters for developers serving V4-Flash-Vision-Exp at scale, since it expands the runnable serving options beyond the reference PyTorch inference implementation shipped in the repo. The fast turnaround from PR opening to merge by a DeepSeek org member indicates active upstream engagement with th

Hugging Face

Official sourceOfficial

DeepSeek launched DeepSeek-V4-Flash-Vision-Exp as the first experimental multimodal model in the DeepSeek-V4 family, built on the V4-Flash text architecture with added visual modules and continued training. According to the official model card on Hugging Face, the model delivers substantial improvements on multimodal a The model card also highlights multimodal agent gains, with ApexBench (Pass@1) reaching 36.5 versus 26.2 for V4-Flash-0731, Agents' Last Exam at 27.3, Chartography at 64.3, and ZeroBench (Pass@5) at 35.0. The repository ships a tokenizer, prompt-encoding reference, and minimal PyTorch inference implementation that cove

Hugging Face

CoverageAnalysis

According to a Kie.ai write-up, DeepSeek released V4 Flash Vision Exp on August 21, 2026, as an API-first experimental multimodal model that adds native image understanding to the V4 Flash line while matching V4 Flash on agents, reasoning, and world knowledge. The official vision guide confirms support for JPEG, PNG, G The piece frames the release as capability-continuity plus a new modality rather than a new architecture, and it cautions that the "Exp" suffix means stable behavior, permanent model identifiers, fixed limits, and final pricing are not guaranteed. Multimodal-agent performance is described as close to Claude Opus 4.8 ra

Videos about DeepSeek V4 Flash Vision Exp

More models around DeepSeek V4 Flash Vision Exp