Currently listed through these providers:
Model details
DeepSeek V4 Flash 0731
DeepSeek V4 Flash 0731 represents the official release of the Flash line, succeeding an earlier preview with substantially enhanced agentic capabilities. It is built as a 284B-parameter Mixture-of-Experts transformer that activates only 13B parameters per pass, an efficiency-oriented design aimed at keeping inference costs low while preserving the reasoning depth expected of larger DeepSeek models. The model is positioned around coding, terminal-style tool use, and multi-step automation, reflecting DeepSeek's broader emphasis on practical, agent-driven tasks rather than general chat alone.
The model's open-weight release under an MIT-style license makes it attractive for teams that want to self-host or fine-tune locally, with LM Studio noting a roughly 156 GB memory footprint for the smallest deployment. A 1,048,576-token context window supports long-horizon coding sessions and extended agent traces, while confirmed support for reasoning and function calling lets developers build structured tool pipelines. DeepSeek reports that this official Flash release outperforms both the Flash preview and the V4 Pro preview on its coding and agentic benchmark suite, signaling that the smaller active parameter count did not come at the cost of capability in its target domains.
Quick Info
Powered by- Provider
- Volcengine Ark
- Model key
- deepseek-v4-flash-ga-260731
- Release date
- Jul 31, 2026
- Last updated
- Jul 31, 2026
- Knowledge cutoff
- 2025-05
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.4453
- Output token cost
- $1.3359
Limits
- Output tokens
- 384,000 tokens
- Context window
- 1,000,000 tokens
Latest news about DeepSeek V4 Flash 0731
Videos about DeepSeek V4 Flash 0731
More models around DeepSeek V4 Flash 0731
This exact model name is also listed by 38 other providers.