Sulat.com
AI models
DeepSeek logo

Model details

DeepSeek V4 Flash

DeepSeek V4 Flash is an autoregressive Mixture-of-Experts language model from DeepSeek AI, designed for advanced reasoning, agentic applications, tool use, and complex problem-solving in mathematics, software engineering, and enterprise assistants. Its architecture uses an optimized Transformer design; later-family evidence specifically describes hybrid attention mechanisms, Manifold-Constrained Hyper-Connections, and a DSpark speculative-decoding module, while treating these as a clearly named revision rather than silently generalizing every detail to the base checkpoint.

The April 2026 release earned a score of 40 on the Artificial Analysis Intelligence Index, while the later 0731 revision reached 50, reflecting a substantial improvement in agentic evaluations and reduced hallucination rates. The base model is therefore a practical fit for text-oriented reasoning and agent workflows, while exact performance, behavior, and architecture claims should be kept tied to the model version being evaluated rather than the broader Flash family.

DeepSeekdeepseek-v4-flashdeepseek-flash

Quick Info

Powered by
Provider
DeepSeek
Model key
deepseek-v4-flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.60

Limits

Output tokens
384,000 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare DeepSeek V4 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4 Flash

DeepSeek

Official sourceAnnouncement

On August 21, 2026, DeepSeek launched DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal extension of the V4 Flash model that matches the text-only Flash on agents, reasoning, and world knowledge while adding vision capabilities. According to DeepSeek, multimodal agent benchmark performance makes a major leap ove The multimodal API supports Chat Completions, Messages, and Responses endpoints, accepting mixed text-and-image input via base64, external URLs, or the newly launched Files API, which lets developers upload an image once for free and reference it by file id across requests. Images are tokenized for billing at up to 384

DeepSeek

Coverage

TextQL announced on August 4, 2026 that DeepSeek V4 Flash 0731 is now available in its "Ana" analytics platform, joining the Claude, GPT, and Kimi model families in Ana's model picker. The post positions V4 Flash as a low-cost option for routine Ana workloads such as scheduled reports, straightforward tool calls, and r All Ana users can enable DeepSeek V4 Flash by toggling it under Settings → Models, with full model configuration options documented in TextQL's docs. While the announcement is light on benchmarks or pricing, it confirms downstream enterprise analytics tooling has integrated the 0731 checkpoint, broadening developer acc

DeepSeek

Coverage

An NVIDIA DGX Spark forum post by user "entrpi" on August 1, 2026 details a forked CUDA inference engine ("ds4") that delivers roughly 1,000 tok/s prefill and ~59 tok/s multi-agent serving for DeepSeek V4 Flash 0731 on a single GB10. The fork, three months and 486 commits in, adds continuous batching, prefix caching wi Measured throughput on one GB10 includes 517,963 tokens ingested in 667 seconds at 776 tok/s sustained, ~960 tok/s prefill at 2k context, ~28 tok/s chat decode at 12k context with speculation, and ~22 tok/s at 240k context, with structured-output decode peaking higher. The build runs fully on-device and fully in memory

DeepSeek

CoverageBenchmark

A July 31, 2026 third-party post (Flowtivity) reports that DeepSeek pushed a re-trained V4-Flash build labeled V4-Flash-0731 into public beta on the same day, retaining the `deepseek-v4-flash` model identifier and backward-compatible API while pricing input tokens at $0.14 per million. The article attributes the jump i The post provides a comparative agent benchmark table placing V4-Flash-0731 ahead of V4-Pro-Preview and GLM-5.2 across Terminal Bench 2.1, NL2Repo, Cybergym, DeepSWE, Toolathlon Verified, Agents' Last Exam, AutomationBench and DSBench-FullStack/Hard, while trailing Anthropic's Opus-4.8 on most metrics. Because these nu

Eden AI

CoverageAnalysis

DeepSeek V4 Flash received a mid-cycle refresh designated "V4 Flash 0731," which Artificial Analysis reports scoring 50 on its Intelligence Index, a 10-point jump over the April 2026 V4 Flash release and one point behind GPT-5.6 Luna (max, 51). The updated model retains the same architecture and pricing as the earlier Agentic capabilities saw the largest gains: GDPval-AA v2 Elo rose to 1559 from 1189, Terminal-Bench 2.1 climbed 17 points to 79%, and τ³-Bench Banking improved 8 points to 31%. AA-Omniscience Index improved by 7 points to -16, driven entirely by a reduced hallucination rate of 84% (down 12 points) rather than accuracy

DeepSeek

CoverageBenchmark

Compare DeepSeek V4 Flash, Qwen3.6 35B A3B, and GLM-4.6 benchmarks, inference speed, and API costs on DeepInfra.

DeepSeek

CoverageRelease Notes

Releasebot's third-party tracker mirrors DeepSeek's official changelog entry for April 24, 2026, confirming the same V4-Pro and V4-Flash API availability through OpenAI ChatCompletions and Anthropic interfaces, the same "deepseek-v4-pro" / "deepseek-v4-flash" model parameter names, and the same July 24, 2026 retirement Beyond the V4 launch, the same aggregator surfaces earlier DeepSeek release notes (V3.2 on December 1, 2025; V3.2-Speciale on a temporary endpoint expiring December 15, 2025; V3.2-Exp on September 29, 2025), providing useful context for the model lineage leading up to V4-Flash. Because it duplicates the official change

DeepSeek

Official sourcePreview

DeepSeek's official API documentation announced the DeepSeek V4 Preview release, introducing the V4 family with a default 1M context length. DeepSeek-V4-Flash is specified as 284B total parameters with 13B active, positioned as the fast, efficient, and economical option alongside V4-Pro (1.6T total / 49B active). The r API access for V4-Flash was made available the same day under the model name `deepseek-v4-flash`, with backward compatibility for both OpenAI ChatCompletions and Anthropic APIs. Both V4-Flash and V4-Pro support 1M context and dual Thinking / Non-Thinking modes. The notice also stated that the legacy `deepseek-chat` and

Videos about DeepSeek V4 Flash

More models around DeepSeek V4 Flash