Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Eden AI logo

Model details

DeepSeek V4 Flash 0731 (Alibaba)

DeepSeek V4 Flash 0731 is an open-weight mixture-of-experts model built around a 284B-parameter architecture that activates only 13B parameters per token, paired with hybrid attention designed to keep inference economical even at very long context lengths. A speculative decoding module is attached at release, the same structural configuration used in the companion DeepSeek-V4-Flash-DSpark variant, which speeds generation by predicting likely continuations before the main model commits to them. The weights are released under the MIT license, placing the model firmly in the open-weight category and giving developers freedom to self-host, fine-tune, or distill from it without restrictive licensing constraints.

Positioned as the official release that supersedes an earlier April preview, DeepSeek V4 Flash 0731 places clear emphasis on agentic capability rather than raw general chat quality. Reported benchmark numbers reflect that focus: Terminal Bench 2.1 reaches 82.7, NL2Repo climbs to 54.2, Cybergym hits 76.7, DeepSWE lands at 54.4, and Toolathlon-Verified posts 70.3, with the model outperforming the larger DeepSeek-V4-Pro Preview on these same benchmarks despite a much smaller active parameter count. In practical terms, the model fits coding assistants, repository-level generation, security research agents, and multi-step tool workflows where long context and rapid function calls matter more than open-ended conversation.

Eden AIqwen/deepseek-v4-flash-0731deepseek-flash

Quick Info

Powered by
Provider
Eden AI
Model key
qwen/deepseek-v4-flash-0731
Release date
Jul 31, 2026
Last updated
Jul 31, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.22
Output token cost
$0.66

Limits

Output tokens
384,000 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare DeepSeek V4 Flash 0731 (Alibaba) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4 Flash 0731 (Alibaba)

Eden AI

Coverage

This NxCode analysis confirms that DeepSeek's July 31 release of V4-Flash-0731 kept the same architecture, size, MoE design, and context window as the April preview—the model was "only re-post-trained." Despite no architectural changes, benchmark scores rose significantly: Terminal Bench 2.1 to 82.7, DeepSWE to 54.4, a The release is notable for improving agent performance within long tool loops—where one request generates dozens of commands, edits, tests, and retries—without altering the base model. V4-Flash now supports the Responses API protocol with a published direct Codex configuration, enabling engineering teams to swap provid

Eden AI

Coverage

This third-party Medium article by Mehul Gupta reports that DeepSeek officially announced DeepSeek-V4-Flash-0731 on July 31, 2026 as the production/GA release of its Flash model. The piece details that the model retains the same architecture as the April preview—284 billion total parameters with 13 billion activated pe The article frames DeepSeek-V4-Flash-0731 as the official General Availability release of DeepSeek's fast flagship language model, noting the deliberate decision to invest in post-training rather than architectural changes. It highlights that improvements in intelligence are increasingly being driven by better training

Eden AI

CoverageBenchmark

This BenchLM benchmark aggregator entry tracks DeepSeek V4 Flash 0731 as released on July 31, 2026 with 40 sourced benchmark rows across agentic (11/11 verified), coding (15/15 verified), reasoning (2/2 verified), knowledge (8/8 verified), and math (4/4 verified) categories, though the model is not yet publicly ranked. The aggregator records a capability field median of 58.1, price blended at $0.21, and speed of 140 tok/s with a first-token latency of 15.48 seconds, compared to field medians of 92 tok/s and 256,000 context tokens respectively. Maximum output length and knowledge cutoff remain marked as not yet sourced, reflecting the

Eden AI

Coverage

This official DeepSeek API Documentation Change Log entry provides first-party context for the DeepSeek V4 model family timeline. The 2026-08-13 entry announces the GA release of DeepSeek-V4-Pro with significantly enhanced agent capabilities across benchmarks including Terminal Bench 2.1 (87.9), NL2Repo (61.5), and Too The changelog confirms that V4-Pro and V4-Flash now support three thinking effort levels—low, high, and max—giving developers more flexible control over reasoning depth. The V4-Flash-Vision-Exp model can be accessed via model='deepseek-v4-flash-vision-exp' on the API platform, and DeepSeek-V4-Flash serves as its text-c

Videos about DeepSeek V4 Flash 0731 (Alibaba)

More models around DeepSeek V4 Flash 0731 (Alibaba)