Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vultr logo

Model details

DeepSeek V4 Flash 0731

DeepSeek V4 Flash 0731 arrives as a July 2026 entry in the DeepSeek Flash line, extending the family with a refresh-style version tag that suggests iterative improvements over earlier Flash checkpoints rather than a completely new architecture. The naming convention aligns it with prior DeepSeek-V3.x-Flash revisions, implying continuity in design philosophy around efficient inference for production text workloads. Because the supplementary sources available do not provide a model card or technical report for this exact build, the practical interpretation of the 0731 tag remains a release-dated snapshot within the Flash family's pattern of stamped refreshes.

For practitioners, the model is positioned for general text generation tasks where a balance of latency and quality matters, fitting use cases such as conversational agents, summarization, and tool-augmented assistants. Without published benchmarks or architecture disclosures for this specific revision, the most reliable framing is that V4 Flash 0731 is a same-family successor intended to maintain the Flash line's emphasis on cost-aware, high-throughput text serving. Teams evaluating it should weigh it against other Flash revisions they have already benchmarked, since independent performance data for the 0731 build has not yet been documented in the reviewed sources.

Vultrdeepseek-v4-flash-0731deepseek-flash

Quick Info

Powered by
Provider
Vultr
Model key
deepseek-v4-flash-0731
Release date
Jul 31, 2026
Last updated
Jul 31, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.25

Limits

Output tokens
384,000 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare DeepSeek V4 Flash 0731 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4 Flash 0731

Melious

CoverageBenchmark

On July 31, 2026, DeepSeek shipped DeepSeek-V4-Flash-0731, the official release of V4-Flash that supersedes the April preview at the same launch price (later shifted to peak/off-peak rates on August 16, 2026). The release posts a 82.7 score on Terminal Bench 2.1, up from 61.8 for the preview and above the larger V4-Pro DeepSeek-V4-Flash carries 284 billion total parameters with 13 billion active per token, a native one-million-token context window, and is distributed under the MIT license on Hugging Face, through the DeepSeek API, and via chat.deepseek.com. At one-million-token inputs, V4-Flash uses roughly 10% of the single-token in

IteraCompute

CoverageBenchmark

DeepSeek released DeepSeek-V4-Flash-0731 on July 31, 2026 as an open model under a commercially usable license, available on Hugging Face and ModelScope under deepseek-ai/DeepSeek-V4-Flash-0731. Per DeepSeek's own X post (quoted in the article), the 0731 build keeps the exact same MoE architecture and parameter counts On Artificial Analysis intelligence evaluations, DeepSeek-V4-Flash-0731 is rated on par with Google's Gemini 3.6 Flash and ranked third among open models behind Kimi K3 and GLM-5.2. Coding benchmarks place it above GLM-5.2, and agent benchmarks show it significantly outperforming Gemini 3.6 Flash while edging out OpenA

Charm Hyper

CoverageBenchmark

MorphLLM's reference (verified 7 September 2026) lists the current V4-Flash checkpoint as V4-Flash-0731, GA on 31 July 2026, with 284B total / 13B active parameters, a 1,000,000-token context and 384K max output under MIT. Off-peak API pricing is given as $0.22 input / $0.66 output per 1M with a $0.007 cache-hit rate, It contrasts Flash against V4-Pro-0813 (1.6T/49B, off-peak $0.66/$1.98, cache $0.022) and lists first-party and partner runtimes including api.deepseek.com, OpenRouter, Cloudflare Workers AI, DeepInfra, Together, Lightning, NVIDIA build and NIM, LM Studio GGUF builds, and Morph's own morph-dsv4flash endpoint at 16-bit

TensorX

CoverageBenchmark

DeepSeek released V4-Flash-0731 on July 31, 2026 as a retrained checkpoint of the April-previewed V4-Flash, keeping the 284B-parameter MoE backbone with 13B active parameters per token. The delta is entirely in post-training: an expanded reinforcement-learning stage targeting instruction-following fidelity, tool-call a The release is reported to beat DeepSeek's own V4-Pro-Preview on nine published benchmarks, combining open weights with what the article describes as frontier-level performance, though the piece itself flags reproducibility asterisks on the headline numbers. The economic framing emphasizes what happens when a $0.28/M-o

GreenPT

CoverageRelease Notes

DeepSeek released DeepSeek-V4-Flash-0731 on July 31, 2026, publishing the weights on Hugging Face under an MIT licence and moving the official V4-Flash API into public beta. The model is a 284 billion parameter Mixture-of-Experts that activates 13 billion parameters per token, supports a 1 million token context window, DeepSeek-V4-Flash-0731 is a decoder-only MoE model with 1 shared expert and 256 routed experts per MoE layer (intermediate dimension 2048), and 6 routed experts fire per token. The deepseek-v4-flash endpoint natively supports the Responses API format and has been adapted for Codex, while the V4-Pro API and DeepSeek app

Alibaba Token Plan

Coverage

A NVIDIA NIM user request thread dated 2 August 2026 asks NVIDIA to update the hosted DeepSeek V4 Flash to deepseek-ai/DeepSeek-V4-Flash-0731, characterizing it as a significant post-training upgrade with stronger coding/software engineering, better agentic reasoning and tool use, and improved coding/agent benchmark pe A subsequent reply dated 8 August 2026 states that fetching the NIM v1/models endpoint returns a list including deepseek flash 0731, indicating downstream platform adoption of the 0731 build. The thread concerns NVIDIA NIM catalog rollout, not Alibaba Token Plan, and identifies DeepSeek as the model creator.

TensorX

CoverageBenchmark

On July 31, 2026, the DeepSeek research team published the open-weight repository for DeepSeek-V4-Flash-0731 (deepseek-ai/DeepSeek-V4-Flash-0731) on Hugging Face, transitioning the high-speed model variant from preview to a production-ready release. The base architecture retains the 284B total / 13B active MoE configur The post-training overhaul reportedly boosted scores in coding, mathematics, and agentic tool-use benchmarks beyond even larger variants like the earlier V4-Pro Preview. By activating only 13B parameters per token through MoE and using Multi-Head Latent Attention, output speeds exceed traditional dense models while low

TensorX

CoverageBenchmark

On July 31, 2026, DeepSeek's API Changelog introduced DeepSeek-V4-Flash-0731 as a Public Beta while keeping the API alias deepseek-v4-flash, updating only the V4-Flash API and leaving the web interface and V4-Pro API unchanged. Architecturally it matches the April Preview with 284B total / 13B active MoE parameters and In Artificial Analysis's Intelligence Index v4.1, DeepSeek-V4-Flash-0731 scored 50 points at Max effort, up 10 from the Preview's 40, ranking second among open-weight models behind Kimi K3 max at 57. This places it in the commercial mid-tier alongside Gemini 3.6 Flash (50), GPT-5.6 Luna (51), and GLM-5.2 (51), undercut

EmpirioLabs AI

Coverage

DeepSeek-V4-Flash-0731 was officially announced on July 31, 2026 as the General Availability (GA) release of DeepSeek's Flash model, according to a third-party Medium explainer by Mehul Gupta. The model retains the same architecture as the prior April Preview: 284 billion total parameters with 13 billion active per tok The core change in DeepSeek-V4-Flash-0731 is a significantly improved post-training pipeline targeting coding, AI agents, reasoning, and tool use. Per the announcement as summarized in the Medium article, this re-training delivers substantially better performance while keeping the same architecture, API endpoint, laten

Alibaba Token Plan

Coverage

A community thread posted 31 July 2026 announces DeepSeek-V4-Flash-0731 as a new open-weight model from DeepSeek, with the Flash variant sized at 284B parameters (13B active) and V4-Pro at 1.6T parameters (49B active). The thread notes the 0731 build delivers the best performance for its size and outperforms V4-Pro (Pr For local serving, the thread documents that full-precision lossless Q8 (UD-Q8_K_XL) quantization of DeepSeek-V4-Flash-0731 runs at 162GB, only 7GB larger than the Q4 (UD-Q4_K_XL) build. Discussion focuses on DGX Spark / GB10 hardware deployment rather than Alibaba Token Plan hosting, and the model creator is identifie

IteraCompute

CoverageRelease Notes

DeepSeek pushed the DeepSeek-V4-Flash-0731 API into public beta on July 31, 2026, describing it as a substantial upgrade to agentic capabilities over the April preview, with the V4-Flash API now natively supporting the Responses format and adapted for Codex. DeepSeek has been explicit that V4-Flash-0731 retains the sam The benchmark jump from preview to release is steep: Terminal-Bench 2.1 climbs to 82.7 (from 61.8 preview, above V4-Pro-Preview's 72.1, ahead of GLM-5.2's 81.0, within range of Claude Opus 4.8's 85.0); NL2Repo reaches 54.2 from 39.4; Cybergym jumps to 76.7 from 38.7; DeepSWE goes from 7.3 to 54.4; Toolathlon-Verified r

GreenPT

Coverage

On July 31, 2026, DeepSeek released DeepSeek-V4-Flash-0731 as the official release of DeepSeek V4 Flash, in public beta on the API, superseding the April 24 Preview with the same architecture and size and changing post-training only. The stable API model ID deepseek-v4-flash did not change, so existing API users starte The release status uses two labels at once: this is the official release of V4 Flash, and the API serving it is in public beta, and neither first-party source calls 0731 GA, generally available, or production-stable. DeepSeek's own logged-out homepage banner uses the same public-beta wording, observed directly on deeps

Videos about DeepSeek V4 Flash 0731

More models around DeepSeek V4 Flash 0731