Currently listed through these providers:
Model details
DeepSeek V4 Flash
DeepSeek V4 Flash sits in the Flash branch of the DeepSeek V4 family and is distributed as an open-weights text model from the deepseek-ai organization on Hugging Face under an MIT license. The Flash-0731 release is described as the official successor to an earlier preview, with substantially enhanced agentic capabilities and a model structure shared with the sibling DSpark variant, which attaches a speculative decoding module for faster inference. A linked technical report for the V4 family is available on arXiv, giving practitioners a deeper reference for the architecture and training approach behind the Flash line. This combination of speculative decoding and an open release suggests a design aimed at efficient, reproducible deployment rather than maximum raw scale.
In published benchmark comparisons, DeepSeek V4 Flash-0731 is reported to outperform the larger DeepSeek V4 Pro Preview on agent-focused evaluations such as Terminal Bench 2.1, NL2Repo, Cybergym, DeepSWE, and Toolathlon-Verified, while its activated parameter footprint stays comparatively small. The card positions it as broadly competitive with strong proprietary models on these tasks, which indicates it is tuned for tool use and multi-step agentic workflows rather than pure chat. Practically, this makes the Flash variant a sensible choice for coding agents, repository-level automation, and other agentic pipelines where latency and cost efficiency matter, and where an open release with a speculative-decoding design can be deployed and inspected on local or self-hosted infrastructure.
Quick Info
Powered by- Provider
- ClinePass
- Model key
- cline-pass/deepseek-v4-flash
- Release date
- Apr 24, 2026
- Last updated
- Apr 24, 2026
- Knowledge cutoff
- 2025-05
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.14
- Output token cost
- $0.28
Limits
- Output tokens
- 384,000 tokens
- Context window
- 1,000,000 tokens
Transparent token rates
Compare DeepSeek V4 Flash pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about DeepSeek V4 Flash
Videos about DeepSeek V4 Flash
More models around DeepSeek V4 Flash
This exact model name is also listed by 17 other providers.