Sulat.com
AI models
OpenCode Zen logo

Model details

DeepSeek V4 Flash

DeepSeek V4 Flash sits inside the DeepSeek V4 Preview family as the lighter, faster counterpart to V4 Pro, built on a 284B-total / 13B-active mixture-of-experts design that keeps inference cost and latency low while preserving much of the larger model's reasoning quality. The V4 Preview release explicitly positions the Flash variant as "Your fast, efficient, and economical choice," and groups it under the same 1M-context-length umbrella as V4 Pro, signaling a shared architecture lineage aimed at long-context workloads rather than a stripped-down toy model. Weights are published openly on Hugging Face under the deepseek-v4 collection, which means the same model that is served behind an API can also be self-hosted, fine-tuned, or distilled by downstream teams.

In practice, V4 Flash is meant for high-throughput chat, coding assistance, tool-augmented agents, and other production pipelines where responsiveness and price per token matter more than squeezing the last few points of benchmark accuracy. According to the Preview announcement, its reasoning capability approaches that of V4 Pro and it lands on par with V4 Pro on simpler agent tasks, which makes it well suited to retrieval-heavy assistants, multi-step tool calling, and structured-output workflows over very long documents. NVIDIA also lists V4 Flash in its NGC catalog under the deepseek-ai team, so the same model is available both as open weights for self-deployment and as an optimized NIM option for teams that prefer a managed runtime.

OpenCode Zendeepseek-v4-flashdeepseek-flash

Quick Info

Powered by
Provider
OpenCode Zen
Model key
deepseek-v4-flash
Release date
Jul 31, 2026
Last updated
Jul 31, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.14
Output token cost
$0.28

Limits

Output tokens
384,000 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare DeepSeek V4 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4 Flash

Videos about DeepSeek V4 Flash

More models around DeepSeek V4 Flash