Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Merge Gateway logo

Model details

DeepSeek V4 Flash 0423

DeepSeek V4 Flash 0423 is an efficiency-focused Mixture-of-Experts large language model from DeepSeek, designed to deliver fast inference without sacrificing long-context capability. Third-party cataloging describes the architecture as having 284 billion total parameters with 13 billion activated per token, a sparse activation pattern that aims to keep compute and latency low while preserving the reasoning depth of a much larger dense model. The variant targets practical deployment scenarios where high throughput and responsive generation matter more than maximal parameter count, making it well-suited to production chat systems, document analysis, and tool-augmented agents running at scale.

Within the DeepSeek V4 family, the Flash 0423 variant occupies the efficiency-oriented tier, prioritizing speed and economy over the heavier capabilities of sibling releases. Its sparse expert design suggests a forward-looking approach to scaling language models where most parameters remain inactive for any given input, allowing DeepSeek to expand model capacity while keeping per-request cost competitive. For practitioners, the model fits use cases such as long document summarization, multi-turn assistants, and pipeline stages where a fast MoE backbone is preferable to a slower, denser alternative, and it reflects DeepSeek's broader trend of releasing specialized checkpoints optimized along different cost-quality tradeoffs.

Merge Gatewaydeepseek/deepseek-v4-flash-0423deepseek-flash

Quick Info

Powered by
Provider
Merge Gateway
Model key
deepseek/deepseek-v4-flash-0423
Release date
Apr 23, 2026
Last updated
Apr 23, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.139
Output token cost
$0.278

Limits

Output tokens
32,768 tokens
Context window
131,072 tokens

Transparent token rates

Compare DeepSeek V4 Flash 0423 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4 Flash 0423

No articles yet. Fetch the latest news to show it here.

Videos about DeepSeek V4 Flash 0423

More models around DeepSeek V4 Flash 0423