Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Deep Infra logo

Model details

DeepSeek V4 Flash

DeepSeek V4 Flash is an efficiency-oriented variant in the broader DeepSeek lineup, positioned as a leaner alternative to the heavier V4-Pro tier. According to its public catalog description, it uses a Mixture-of-Experts architecture with 284 billion total parameters but only 13 billion activated per token, which is what enables its focus on fast inference and high throughput. The model also incorporates hybrid attention to keep long-context processing efficient, and its open weights are published under the deepseek-ai/DeepSeek-V4-Flash repository on Hugging Face, making it accessible for self-hosting, fine-tuning, and research use beyond hosted endpoints.

In practical terms, DeepSeek V4 Flash is aimed at developers who want responsive language-model behavior without paying flagship-tier prices. The catalog frames it as well suited for coding assistants, chat systems, and agent workflows where latency and cost efficiency matter, while still preserving strong reasoning and coding quality. It supports configurable reasoning effort with “high” and “xhigh” levels, where the latter maps to maximum reasoning depth, giving applications a dial between speed and deliberation. Combined with a very large context window suitable for long documents and multi-turn agent sessions, this makes V4 Flash a practical middle ground for production deployments that need both open-weight flexibility and steady reasoning performance.

Deep Infradeepseek-ai/DeepSeek-V4-Flashdeepseek-flash

Quick Info

Powered by
Provider
Deep Infra
Model key
deepseek-ai/DeepSeek-V4-Flash
Release date
Apr 24, 2026
Last updated
Apr 24, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.09
Output token cost
$0.18

Limits

Output tokens
16,384 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare DeepSeek V4 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4 Flash

Nvidia

Coverage

DeepSeek released an official V4 Flash API retrain tagged DeepSeek-V4-Flash-0731, which keeps the same architecture and size as the preview build but has been re-post-trained to dramatically improve agentic capabilities. The retrain adds native Responses API support and is adapted for Codex, and developers who already DeepSeek published self-reported agentic benchmark scores for the 0731 build showing strong gains over the larger V4-Pro-Preview, including Terminal Bench 2.1 at 82.7, Cybergym at 76.7, Toolathlon (verified) at 70.3, DSBench-FullStack at 68.7, and DSBench-Hard at 59.6, alongside NL2Repo at 54.2, DeepSWE at 54.4, Agent

SiliconFlow

Coverage

As of August 26, 2026, the current DeepSeek API lineup exposes three routes, with deepseek-v4-flash mapped to the DeepSeek-V4-Flash-0731 checkpoint. The source describes V4 Flash as the practical default for text and routine agent workloads, recommending V4 Pro when hard reasoning, coding, or multi-step tool orchestrat The current models table gives deepseek-v4-flash a 1-million-token context window and a 384,000-token maximum output, with JSON output, tool calls, Responses API support, and an Anthropic-compatible interface listed for each route. The same page documents the August 13, 2026 general availability of V4 Pro across app, w

SiliconFlow

CoveragePreview

DeepSeek's official API documentation announced the DeepSeek V4 Preview release on April 24, 2026, with deepseek-v4-flash explicitly named as a 284B-total / 13B-active-parameter mixture-of-experts model. The release note establishes the model as the economical, fast text route within the V4 family, positioning it below The same page documents that legacy deepseek-chat and deepseek-reasoner routes were retired after July 24, 2026, with traffic routed to deepseek-v4-flash (non-thinking and thinking modes). DeepSeek notes V4's integration with agent tools including Claude Code, OpenClaw, and OpenCode, and points to the open-weight relea

Deep Infra

Coverage

The official DeepSeek API changelog dated April 24, 2026 confirms that V4-Pro and V4-Flash are served via both the OpenAI ChatCompletions interface and an Anthropic-compatible interface, accessed by setting the model parameter to `deepseek-v4-pro` or `deepseek-v4-flash`. This is the authoritative primary source for API The same changelog entry flags a time-sensitive migration: legacy aliases `deepseek-chat` and `deepseek-reasoner` will be discontinued on July 24, 2026 — five days after the current date — and during the transition they alias to V4-Flash's non-thinking and thinking modes, respectively. Developers targeting DeepInfra's

Videos about DeepSeek V4 Flash

More models around DeepSeek V4 Flash