Sulat.com
AI models
ai& logo

Model details

DeepSeek V4 Flash

DeepSeek V4 Flash sits in the DeepSeek Flash family as a text-only language model designed for interactive assistants and agentic pipelines. Its broad context window, reaching around one million tokens on the Cloudflare Workers AI deployment of the related 0731 revision, makes it well suited to long-form code analysis, multi-document summarization, and tool-augmented reasoning workflows that need to keep large amounts of intermediate state in memory. The model is documented as supporting function calling and step-by-step reasoning, which lines up with typical Flash-family design goals of balancing inference speed against the structured thinking required for planning, API orchestration, and retrieval-heavy applications.

Being distributed as open weights allows DeepSeek V4 Flash to be self-hosted on infrastructure such as the NVIDIA NGC catalog entry for the model, giving teams flexibility around latency, throughput, and data handling compared with relying solely on a hosted endpoint. Pricing on Cloudflare's deployment of the 0731 revision is positioned for cost-sensitive production use, with reduced rates for cached input tokens that reward prompt reuse and caching strategies common in agent loops. Practically, the model fits teams that want a reasoning-capable, tool-friendly text model they can run locally for fast prototyping while still being able to fall back to a managed endpoint for scale.

ai&deepseek-ai/deepseek-v4-flashdeepseek-flash

Quick Info

Powered by
Provider
ai&
Model key
deepseek-ai/deepseek-v4-flash
Release date
Apr 24, 2026
Last updated
Apr 24, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.25

Limits

Output tokens
384,000 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare DeepSeek V4 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4 Flash

Nvidia

Coverage

DeepSeek released an official V4 Flash API retrain tagged DeepSeek-V4-Flash-0731, which keeps the same architecture and size as the preview build but has been re-post-trained to dramatically improve agentic capabilities. The retrain adds native Responses API support and is adapted for Codex, and developers who already DeepSeek published self-reported agentic benchmark scores for the 0731 build showing strong gains over the larger V4-Pro-Preview, including Terminal Bench 2.1 at 82.7, Cybergym at 76.7, Toolathlon (verified) at 70.3, DSBench-FullStack at 68.7, and DSBench-Hard at 59.6, alongside NL2Repo at 54.2, DeepSWE at 54.4, Agent

ai&

CoverageBenchmark

The OpenRouter page for DeepSeek V4 Flash Latest (alias deepseek/deepseek-v4-flash-latest) confirms current pricing at $0.04998 per million input tokens and $0.09996 per million output tokens, with cache reads at $0.009996 per million tokens. The model is listed with a 1,310,720 token context window, up to 131,072 comp The page lists related DeepSeek text models, including DeepSeek V4 Flash Vision Exp, DeepSeek V4 Pro 0813, and DeepSeek V4 Pro 0423, and indicates that DeepSeek V4 Flash Latest was released on August 1, 2026. It positions OpenRouter's API as OpenAI-compatible, noting that most SDKs work by swapping the base URL and onl

ai&

Coverage

The Lambda inference page documents DeepSeek-V4-Flash as a 284 billion parameter sparse Mixture-of-Experts language model with 13 billion parameters active per forward pass, the smaller sibling of DeepSeek-V4-Pro (1.6T / 49B active). Both ship with the same native 1 million token context window and the same three reaso Lambda's first-party deployment benchmarks report throughput on NVIDIA HGX B200 (native FP4+FP8): 1,222 tok/s generation, 38 tok/s per-user, 11,000 tok/s total, 1,701 ms TTFT, and 66 ms ITL; on NVIDIA HGX H100 (FP8-quantized): 1,262 tok/s generation, 39 tok/s per-user, 11,361 tok/s total, 2,463 ms TTFT, and 60 ms ITL,

Videos about DeepSeek V4 Flash

More models around DeepSeek V4 Flash