Currently listed through these providers:
Model details
DeepSeek V4 Flash 0731
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model that activates roughly 13B parameters out of a much larger 284B total, a design that keeps inference costs low while preserving the capacity of a very wide model. Described as a re-post-trained revision and the GA release of the DeepSeek V4 Flash family, it is explicitly positioned for coding, reasoning, and agent workflows. Open weights are mirrored on Hugging Face under the deepseek-ai namespace, which makes it accessible to teams that want to self-host or fine-tune rather than rely solely on a managed endpoint.
Practically, the model is a good fit for developers building agentic systems and code assistants who want strong reasoning without paying for a dense flagship, and its listing on OpenRouter signals multi-provider availability with routing options that prioritize price, speed, or tool-calling accuracy. The very wide context window supports long-document reasoning, multi-file code analysis, and extended tool transcripts, which aligns with its agent-oriented design. Independent developer interest in deploying it via NVIDIA NIM further suggests it is being adopted across heterogeneous inference stacks rather than locked to a single host.
Quick Info
Powered by- Provider
- Alibaba
- Model key
- deepseek-v4-flash-0731
- Release date
- Jul 31, 2026
- Last updated
- Jul 31, 2026
- Knowledge cutoff
- 2025-05
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.20
- Output token cost
- $0.40
Limits
- Output tokens
- 384,000 tokens
- Context window
- 1,000,000 tokens
Latest news about DeepSeek V4 Flash 0731
Videos about DeepSeek V4 Flash 0731
Recent tweets and retweets from Alibaba
More models around DeepSeek V4 Flash 0731
This exact model name is also listed by 38 other providers.