Currently listed through these providers:
Model details
Deepseek V4 Flash
DeepSeek V4 Flash is positioned as a fast, agentic text model built for tool-driven workflows. The official Hugging Face model card identifies the 0731 release as the stable successor to an earlier preview, with substantially enhanced agentic capabilities, and notes that it shares the same architecture as DeepSeek-V4-Flash-DSpark, including an attached speculative decoding module designed to accelerate inference. An accompanying technical report on arXiv (2606.19348) documents the broader DeepSeek-V4 family lineage that this model extends, situating the Flash variant as a compact but capability-focused member of that line.
In benchmark comparisons shared on the model card, DeepSeek V4 Flash posts strong agentic and coding results despite a far smaller activated parameter count, surpassing DeepSeek-V4-Pro Preview on Terminal Bench 2.1, NL2Repo, Cybergym, DeepSWE, and Toolathlon-Verified, while remaining broadly competitive with leading proprietary models. The combination of speculative decoding, an emphasis on tool use, and a coding-heavy benchmark profile makes it a practical fit for developers building autonomous agents, repository-level code tasks, and other multi-step workflows where both responsiveness and reliable tool orchestration matter.
Quick Info
Powered by- Provider
- DigitalOcean
- Model key
- deepseek-4-flash
- Release date
- May 27, 2026
- Last updated
- May 29, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.0679
- Output token cost
- $0.168
Limits
- Output tokens
- 384,000 tokens
- Context window
- 1,048,576 tokens