Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Nvidia logo

Model details

DeepSeek V4 Flash

DeepSeek V4 Flash is built on a Mixture-of-Experts architecture that activates only 13 billion parameters per forward pass while maintaining access to 1.6 trillion total parameters across the model. This design choice enables the model to keep a vast knowledge reservoir without the computational burden of activating everything at once. The architecture incorporates hybrid attention mechanisms combining Compressed Sparse Attention and Heavily Compressed Attention, alongside Manifold-Constrained Hyper-Connections that support efficient long-context reasoning. Dual-mode operation lets users choose between explicit reasoning traces and direct response generation, making the model versatile across different task types—particularly for advanced reasoning, software engineering, and complex problem-solving in agentic AI applications.

Released under the MIT license in April 2026, DeepSeek V4 Flash continues DeepSeek's open-weight development philosophy. Benchmark results show it outperforming Claude Opus 4.6 across multiple evaluation sets, with especially strong performance on mathematics and software engineering benchmarks. The architecture supports quantization pipelines—NVIDIA's Model Optimizer has produced an NVFP4 quantized variant—making deployment feasible across varied hardware. DeepSeek V4 Flash integrates with popular agent frameworks and coding assistants, enabling direct backend use without custom integration work. Its tool calling support, structured output capabilities, and agent-friendly design position it well for building AI assistants that interact with external systems and execute multi-step workflows in enterprise and developer settings.

Nvidiadeepseek-ai/deepseek-v4-flashdeepseek-flashdeprecated

Quick Info

Powered by
Provider
Nvidia
Model key
deepseek-ai/deepseek-v4-flash
Release date
Apr 24, 2026
Last updated
Apr 24, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.14
Output token cost
$0.28

Limits

Output tokens
393,216 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare DeepSeek V4 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4 Flash

Nvidia

Coverage

DeepSeek released an official V4 Flash API retrain tagged DeepSeek-V4-Flash-0731, which keeps the same architecture and size as the preview build but has been re-post-trained to dramatically improve agentic capabilities. The retrain adds native Responses API support and is adapted for Codex, and developers who already DeepSeek published self-reported agentic benchmark scores for the 0731 build showing strong gains over the larger V4-Pro-Preview, including Terminal Bench 2.1 at 82.7, Cybergym at 76.7, Toolathlon (verified) at 70.3, DSBench-FullStack at 68.7, and DSBench-Hard at 59.6, alongside NL2Repo at 54.2, DeepSWE at 54.4, Agent

SiliconFlow

Coverage

As of August 26, 2026, the current DeepSeek API lineup exposes three routes, with deepseek-v4-flash mapped to the DeepSeek-V4-Flash-0731 checkpoint. The source describes V4 Flash as the practical default for text and routine agent workloads, recommending V4 Pro when hard reasoning, coding, or multi-step tool orchestrat The current models table gives deepseek-v4-flash a 1-million-token context window and a 384,000-token maximum output, with JSON output, tool calls, Responses API support, and an Anthropic-compatible interface listed for each route. The same page documents the August 13, 2026 general availability of V4 Pro across app, w

Nvidia

Coverage

DeepSeek v4 ships two MIT-licensed models with 10x cheaper inference. The cost case is real. The capability story is still developing.

Nvidia

Coverage

According to @deepseek_ai, the DeepSeek API now supports the new deepseek-v4-pro and deepseek-v4-flash models with 1M context windows and dual Thinking and...

SiliconFlow

CoveragePreview

DeepSeek's official API documentation announced the DeepSeek V4 Preview release on April 24, 2026, with deepseek-v4-flash explicitly named as a 284B-total / 13B-active-parameter mixture-of-experts model. The release note establishes the model as the economical, fast text route within the V4 family, positioning it below The same page documents that legacy deepseek-chat and deepseek-reasoner routes were retired after July 24, 2026, with traffic routed to deepseek-v4-flash (non-thinking and thinking modes). DeepSeek notes V4's integration with agent tools including Claude Code, OpenClaw, and OpenCode, and points to the open-weight relea

Nvidia

Official sourceDocumentation

DeepSeek-V4-Flash Overview Description: DeepSeek-V4-Flash is a Mixture-of-Experts (MoE) language model with 284 billion total parameters and 13 billion activated parameters. DeepSeek-V4-Flash was developed by DeepSeek as a part of DeepSeek-V4 collection. This model is ready for commercial/non-commer...

Nvidia

Official sourceOfficial

DeepSeek V4 Flash is a 284B MoE model with 1M-token context optimized for fast coding and agents.

Videos about DeepSeek V4 Flash

More models around DeepSeek V4 Flash