DeepSeek V4 Flash 0731 is an open-weight mixture-of-experts model built around a 284B-parameter architecture that activates only 13B parameters per token, paired with hybrid attention designed to keep inference economical even at very long context lengths. A speculative decoding module is attached at release, the same structural configuration used in the companion DeepSeek-V4-Flash-DSpark variant, which speeds generation by predicting likely continuations before the main model commits to them. The weights are released under the MIT license, placing the model firmly in the open-weight category and giving developers freedom to self-host, fine-tune, or distill from it without restrictive licensing constraints.
Positioned as the official release that supersedes an earlier April preview, DeepSeek V4 Flash 0731 places clear emphasis on agentic capability rather than raw general chat quality. Reported benchmark numbers reflect that focus: Terminal Bench 2.1 reaches 82.7, NL2Repo climbs to 54.2, Cybergym hits 76.7, DeepSWE lands at 54.4, and Toolathlon-Verified posts 70.3, with the model outperforming the larger DeepSeek-V4-Pro Preview on these same benchmarks despite a much smaller active parameter count. In practical terms, the model fits coding assistants, repository-level generation, security research agents, and multi-step tool workflows where long context and rapid function calls matter more than open-ended conversation.