Sulat.com
AI models
AMD logo

Model details

DeepSeek V4 Flash Vision Exp

DeepSeek V4 Flash Vision Exp is an experimental multimodal variant of the V4 Flash agent model, designed to add image understanding on top of the existing text capabilities. According to the official changelog, pure-text strengths in agents, reasoning, and world knowledge remain on par with the standard DeepSeek V4 Flash, while the new release extends use cases to workflows that require visual comprehension. The model is positioned for developers building agents that must interpret screenshots, diagrams, or other visual inputs without giving up the lightweight footprint of the Flash line.

Benchmark figures cited from the official announcement show the model reaching 83.9 on Terminal Bench 2.1, 59.3 on DeepSWE, and 63.6 on DSBench-Hard, with multimodal agent performance described as close to Claude Opus 4.8. The release supports JPEG, PNG, GIF, and WebP images delivered through Base64, public URLs, or the Files API, fitting cleanly into OpenAI-style chat completion patterns. This combination of competitive agent scores and broad image support makes V4 Flash Vision Exp a practical choice for builders who want vision-aware reasoning at the lower cost tier of the DeepSeek family, especially for prototyping multimodal assistants before committing to larger frontier models.

AMDDeepSeek-V4-Flash-Vision-Expdeepseek-flash

Quick Info

Powered by
Provider
AMD
Model key
DeepSeek-V4-Flash-Vision-Exp
Release date
Aug 21, 2026
Last updated
Aug 21, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.14
Output token cost
$0.28

Limits

Output tokens
384,000 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare DeepSeek V4 Flash Vision Exp pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4 Flash Vision Exp

AMD

CoverageAnalysis

A third-party analysis published on Aug 29, 2026 on Kie.ai summarizes the August 21 announcement and adds developer-facing implementation details for DeepSeek-V4-Flash-Vision-Exp. According to this coverage, the model supports image input in JPEG, PNG, GIF, and WebP formats, with images entering the system via Base64 e The same analysis reports community-sourced specifications — a 1M-token context window, a 384K-token maximum output, and a 384-token image billing cap — while noting these are community reports rather than official figures. It describes the release as a capability continuation plus a new modality rather than a new arch

AMD

Coverage

The DeepSeek API change log dated 2026-08-21 officially announces DeepSeek-V4-Flash-Vision-Exp as a new multimodal vision understanding model available on the DeepSeek API platform under the exact model identifier "deepseek-v4-flash-vision-exp". The entry states that pure-text capabilities — covering agents, reasoning, The same official change log lists nine public benchmark scores for the model's Code Agent text tasks: Terminal Bench 2.1 at 83.9, NL2Repo at 57.7, DeepSWE at 59.3, DSBench-Hard at 63.6, AutomationBench (Public) at 25.7, ApexBench (Pass@1) at 36.5, Agents' Last Exam at 27.3, Chartography at 64.3, and ZeroBench (Pass@5)

Videos about DeepSeek V4 Flash Vision Exp

More models around DeepSeek V4 Flash Vision Exp