Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Modelis logo

Model details

DeepSeek V4 Flash

DeepSeek V4 Flash sits within DeepSeek's flash family as a compact open-weights text model oriented toward fast reasoning and agent-style workloads. Its officially distributed checkpoint, DeepSeek-V4-Flash-0731, shares its underlying model structure with DeepSeek-V4-Flash-DSpark, which adds a speculative decoding module to accelerate inference while keeping activated parameters modest. The release is published under the MIT license, with an accompanying technical report linked on arXiv, so developers can both self-host the weights and consult the methodology behind them.

Qualitative benchmark signals for the V4-Flash-0731 build point to unusually strong agentic behavior for its size, with reported scores on Terminal Bench 2.1, NL2Repo, Cybergym, DeepSWE, and Toolathlon-Verified that exceed DeepSeek-V4-Pro Preview and stay broadly competitive with leading proprietary systems. That profile makes V4 Flash a practical fit for tool-using assistants, code automation, and other latency-sensitive agent workflows where open deployment matters, while leaving room for forward-looking improvements as the broader V4 architecture, including later asymmetric MoE variants with reduced KV cache footprints, continues to mature.

Modelisdeepseek-v4-flashdeepseek-flash

Quick Info

Powered by
Provider
Modelis
Model key
deepseek-v4-flash
Release date
Apr 24, 2026
Last updated
Apr 24, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.0983
Output token cost
$0.1966

Limits

Output tokens
384,000 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare DeepSeek V4 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4 Flash

Eden AI

Official sourceAnnouncement

DeepSeek's official API documentation announces that V4-Flash and V4-Flash-Vision-Exp are retired, and that the legacy "deepseek-v4-flash" and "deepseek-v4-flash-vision-exp" model names temporarily route to the newly released V4.1-Flash for backward compatibility. This is the only candidate that explicitly addresses th The same official post details that V4.1-Flash (the live target behind the retired deepseek-v4-flash alias) is a 552B-parameter MoE model with a new Causal Encoder–Decoder architecture, 8B active parameters for input and 16B for output, and a significantly smaller KV cache using 1/4 the HBM and 1/8 the SSD storage comp

Vercel AI Gateway

Official sourceAnnouncement

DeepSeek announced DeepSeek-V4.1-Flash on September 9, 2026 as the new smallest model in its architecture family, but the post explicitly retires V4-Flash and V4-Flash-Vision-Exp, routing deepseek-v4-flash and deepseek-v4-flash-vision-exp to V4.1-Flash for compatibility. As of the announcement, V4-Flash is no longer se DeepSeek-V4.1-Flash is a 552-billion-parameter MoE with a new Causal Encoder–Decoder architecture using just 8 billion active parameters for input and 16 billion for output, native multimodal support, and a KV cache that needs roughly one-quarter the HBM and one-eighth the SSD storage of the previous generation. V4.1-F

SenseNova (China)

CoverageBenchmark

The InferenceX model page for DeepSeek V4 provides technically substantive coverage of the V4 series, explicitly naming DeepSeek-V4-Flash as a 284B-total / 13B-active parameter MoE model co-released with the flagship DeepSeek-V4-Pro (1.6T total / 49B active) on April 23, 2026. The page cites DeepSeek's own API news rel While the page's architectural deep-dive focuses on V4-Pro (hybrid CSA + HCA attention, MLA-style shared-KV MQA, hash-routed MoE, FP4 experts with FP8 mixed precision, Muon optimizer, YaRN RoPE for 1M context), the explicit naming of V4-Flash with its parameter counts and shared release context provides direct model-le

Neuralwatt

Coverage

According to a third-party API guide dated August 26, 2026, DeepSeek's hosted API exposes `deepseek-v4-flash` as a current production route, mapping via the official quick start to the snapshot build DeepSeek-V4-Flash-0731. The guide recommends V4 Flash as the default for text input and routine agent workloads, with V4 The guide explicitly distinguishes the DeepSeek-hosted `deepseek-v4-flash` route from the DeepSeek app, Hugging Face repositories, and similarly named gateway resales, framing API access as a separate contract that does not transfer to self-hosted quantizations. It also notes that the sibling `deepseek-v4-flash-vision-

Vivgrid

CoverageBenchmark

Better Stack's technical guide confirms that DeepSeek-V4-Flash-0731 was published on July 31, 2026, moving the official V4-Flash API into public beta the same day. The model uses a sparse 284B-parameter MoE architecture with 13B parameters active per forward pass, a 1-million-token context window, a maximum output of 3 The guide documents launch pricing of $0.14 per million input tokens on a cache miss, $0.0028 per million on a cache hit, and $0.28 per million output tokens, and positions V4-Flash as occupying a useful point on the cost-intelligence curve rather than as a frontier-capability model. It includes a practical developer w

Neuralwatt

CoverageBenchmark

An independent model release tracker lists 22 DeepSeek releases between November 2023 and August 13, 2026, including the exact V4-Flash lineage relevant to the `deepseek-v4-flash` subject: DeepSeek-V4-Flash shipped on April 24, 2026, and was updated to DeepSeek-V4-Flash-0731 on July 31, 2026. Both entries are recorded The tracker surfaces benchmark context that contextualizes V4 Flash within the V4 family, even though Flash-specific scores are not broken out: V4-Pro-0813 records 80.6% on SWE-Bench Verified, 31.8% on AutomationBench, 87.9% on Terminal-Bench 2.1, and 60% on Humanity's Last Exam, with the earlier V4-Pro posting 93.5% o

Vivgrid

CoverageBenchmark

DeepSeek published V4-Flash-0731 on July 31, 2026, a retrained checkpoint of the 284-billion-parameter Mixture-of-Experts model first previewed in April, with the architecture unchanged at 13 billion active parameters per token. According to the article, the improvements come entirely from an expanded reinforcement-lea The article reports that V4-Flash-0731 beats DeepSeek's own V4-Pro-Preview on nine published agent and coding benchmarks, with pricing positioned at $0.14 per million input tokens and $0.28 per million output tokens. The author flags reproducibility asterisks on the headline benchmark numbers and frames the broader sig

above.dev

CoverageRelease Notes

DeepSeek released DeepSeek-V4-Flash-0731 on July 31, 2026, as the official V4-Flash build superseding the April 2026 preview. It is a 284-billion-parameter Mixture-of-Experts model with 13 billion activated parameters per token and a 1 million-token context window. The checkpoint on Hugging Face reports 304 billion par DeepSeek reports the 0731 build beats the larger V4-Pro preview on every published agentic benchmark, with a Terminal-Bench 2.1 score of 82.7 (up from 61.8 on the April preview). The deepseek-v4-flash API endpoint now natively supports the Responses API format and has been adapted for Codex, while the V4-Pro API and th

above.dev

Coverage

DeepSeek V4-Flash entered official public beta on July 31, 2026 as build V4-Flash-0731, retaining the same 284B-parameter architecture, context window, and price as the April preview. Across all nine agent and coding benchmarks DeepSeek published, the build outscores its own V4-Pro preview, including a jump from 7.3 to The rollout covers only the hosted API and not the public Hugging Face repository, which still hosts April's weaker preview build, leaving self-hosters on the older checkpoint. V4-Pro's official release is expected in August 2026, and DeepSeek's API changelog notes the calling method is unchanged, with deepseek-v4-flas

OpenCode Zen

CoverageBenchmark

DeepSeek released the official DeepSeek-V4-Flash API into public beta on July 31, 2026, under the model identifier "deepseek-v4-flash," with the release labeled V4-Flash-0731. According to the article, V4-Flash-0731 scored 82.7 on Terminal Bench 2.1, reportedly outperforming DeepSeek's own V4-Pro-Preview (72.1) by 14.7 Pricing is reported at $0.14 per million input tokens, which the article frames as the most cost-effective frontier agent model on the market, with full backward-compatible API access via the "deepseek-v4-flash" model name. The post discusses the implications for AI agent development and notes support for the Responses

above.dev

CoverageBenchmark

DeepSeek-V4-Flash-0731 was released on July 31, 2026, 98 days after the original V4-Flash, and is tracked against benchmarks including SWE-Bench Verified, Terminal-Bench 2.1, BrowseComp, and Toolathlon-Verified. The release tracker documents verified API pricing pulled from api-docs.deepseek.com, distinguishing off-pea The page records peak-hour pricing (01:00–04:00 and 06:00–10:00 UTC) at $0.44 input / $0.014 cached / $1.32 output per million tokens, with off-peak rates at $0.22 / $0.007 / $0.66. Thinking and non-thinking modes use the same published price, and rates were verified against DeepSeek's official API documentation on Aug

Alibaba Token Plan

CoverageAnalysis

DeepSeek V4 Flash 0731 scored 50 on the Artificial Analysis Intelligence Index, a 10-point jump over the original DeepSeek V4 Flash released in April 2026, and 6 points above DeepSeek V4 Pro. It lands within 1 point of GLM-5.2 (max) and Gemini 3.6 Flash (50), placing it near the frontier of cost-efficient open-weight m Agentic gains are the headline: GDPval-AA v2 Elo rose to 1559 from 1189 (second-highest open-weights score, behind Kimi K3 and ahead of GLM-5.2), Terminal-Bench 2.1 climbed 17 points to 79%, and τ³-Bench Banking improved 8 points to 31%. Omniscience gains came from reduced hallucination: the AA-Omniscience Index improv

Videos about DeepSeek V4 Flash

More models around DeepSeek V4 Flash