Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
SiliconFlow logo

Model details

DeepSeek V4 Flash 0731

DeepSeek V4 Flash 0731 is an autoregressive Mixture-of-Experts Transformer built for advanced reasoning and agentic work. Its architecture combines Compressed Sparse Attention and Heavily Compressed Attention with Manifold-Constrained Hyper-Connections, while placing DeepSeek’s DSpark speculative decoding module in the same checkpoint. The MIT-licensed release checkpoint is designed to support demanding text tasks, including tool-using assistants, software engineering, and complex problem solving.

The model’s practical profile centers on reasoning-intensive agents and tool use rather than multimodal work. Published text-agent evaluations cover terminal tasks, repository generation, software engineering, tool interaction, and automation, giving useful evidence for workflows that need multi-step execution. Its sparse hybrid attention and bundled speculative decoding make it a strong fit for applications that value capable reasoning alongside efficient text generation.

SiliconFlowdeepseek-ai/DeepSeek-V4-Flash-0731deepseek-flash

Quick Info

Powered by
Provider
SiliconFlow
Model key
deepseek-ai/DeepSeek-V4-Flash-0731
Release date
Jul 31, 2026
Last updated
Jul 31, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.22
Output token cost
$0.66

Limits

Output tokens
384,000 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare DeepSeek V4 Flash 0731 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4 Flash 0731

Nebius Token Factory

CoverageBenchmark

GIGAZINE reported on August 3, 2026 that DeepSeek released DeepSeek-V4-Flash-0731 as an open model under a license permitting commercial use, with weights hosted on Hugging Face and ModelScope. The piece describes it as an updated version of DeepSeek-V4-Flash from April 2026, retaining the same MoE architecture with 28 Citing Artificial Analysis intelligence tests, the article places DeepSeek-V4-Flash-0731 at performance equivalent to Google's Gemini 3.6 Flash, ranking third among open models behind Kimi K3 and GLM-5.2. Coding performance is reported to surpass GLM-5.2, and agent performance is shown to significantly outperform Gemin

Nebius Token Factory

Coverage

A Medium summary dated August 1, 2026 reports that DeepSeek officially announced DeepSeek-V4-Flash-0731 on July 31, 2026 as the production-ready version of its Flash model. The piece emphasizes that the release keeps the same 284 billion total parameters, 13 billion active parameters per token via Mixture-of-Experts, 1 According to the article, the gains come from post-training improvements rather than architectural changes, including improved reinforcement learning, better supervised fine-tuning, multi-teacher distillation, and stronger agent-focused training datasets. DeepSeek claims significantly better reasoning, coding, and AI a

Hugging Face

Coverage

On July 31, 2026, DeepSeek officially announced DeepSeek-V4-Flash-0731 as the production (General Availability) release of its Flash model, according to a third-party write-up by Mehul Gupta. The article states the model retains the same architecture as the earlier April Preview: 284 billion total parameters with 13 bi The piece frames DeepSeek-V4-Flash-0731 as evidence of an industry shift toward better post-training pipelines as a lever for capability gains, rather than simply enlarging model size. It is authored by an independent Medium commentator who paraphrases DeepSeek's announcement rather than linking to a primary release pa

Nebius Token Factory

CoverageRelease Notes

OfficeChai reported on July 31, 2026 that DeepSeek pushed the official API for DeepSeek-V4-Flash-0731 into public beta, with the company making a substantial upgrade to the model's agentic capabilities and benchmark scores now ahead of its own V4-Pro-Preview. The API now natively supports the Responses format and has b The article highlights that V4-Flash-0731 keeps the same architecture and parameter count as the preview, with the update applying only to the V4-Flash API while V4-Pro and the consumer app/web versions remain on their existing builds. Benchmark jumps include Terminal Bench 2.1 to 82.7 (from 61.8 for the preview), NL2R

Nebius Token Factory

CoverageAnalysis

Artificial Analysis published a benchmark analysis on July 31, 2026 showing DeepSeek V4 Flash 0731 scoring 50 on the Artificial Analysis Intelligence Index, a 10-point jump over the earlier DeepSeek V4 Flash (40) and 6 points above DeepSeek V4 Pro. The model keeps identical architecture and pricing as the prior Flash, The analysis confirms DeepSeek V4 Flash 0731 retains a 1M token context window with 284B total parameters and 13B active at inference. Key gains are in agentic work: GDPval-AA v2 Elo climbs to 1559 (from 1189), Terminal-Bench 2.1 rises 17 points to 79%, and τ³-Bench Banking rises 8 points to 31%. AA-Omniscience Index i

Nvidia

CoverageBenchmark

DeepSeek shipped DeepSeek-V4-Flash-0731 on 31 July 2026 as the official GA release of the V4-Flash variant, superseding the April preview at the same initial price (peak/off-peak pricing was later introduced on 16 August 2026). The model is an open-source Mixture-of-Experts release with 284 billion total parameters and DeepSeek's own model card places V4-Flash-0731 at 82.7 on Terminal Bench 2.1, up from 61.8 for the preview and above the larger V4-Pro preview's 72.1, while Artificial Analysis moved the Flash from 40 to 50 on its Intelligence Index, placing it third in the open-weight field. Note: this item describes the 0731 checkpoi

Nebius Token Factory

CoverageBenchmark

A Hacker News discussion thread on the Artificial Analysis analysis of DeepSeek V4 Flash 0731 attracted 594 points and 312 comments, providing community signal around the July 31, 2026 release. The thread links to the open-weights repository at huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731 and discusses the model's Commenters noted that DeepSeek's evaluation used the minimal mode of DeepSeek Harness (to be released) as the agent framework with max reasoning effort, fueling speculation that DeepSeek may release an optimized coding agent harness alongside the open weights. The discussion positions V4 Flash 0731 as a strong daily-dr

RunInfra

CoverageBenchmark

OpenRouter's page for deepseek/deepseek-v4-flash-0731 confirms the July 31, 2026 release of DeepSeek V4 Flash 0731 as a sparse mixture-of-experts model with 284B total parameters and 13B active, a 1M token context window, and a stated base price of $0.05 input / $0.16 output per 1M tokens with text in/out modalities. D As a router, OpenRouter enumerates many downstream inference hosts (OpenInference, Relace, DigitalOcean, DeepInfra, Together, Baseten, CoreWeave, Fireworks, Cloudflare, DeepSeek first-party, and others) with per-host latency, throughput, uptime, and pricing tiers for serving DeepSeek V4 Flash 0731. These figures reflec

Videos about DeepSeek V4 Flash 0731

More models around DeepSeek V4 Flash 0731