Currently listed through these providers:
Model details
DeepSeek V4 Flash (Tencent Cloud)
DeepSeek V4 Flash is the efficiency-oriented member of the DeepSeek V4 preview family, built as a Mixture-of-Experts language model with 284B total parameters and 13B activated per token. It is released as an open-source model alongside its larger sibling, and it inherits the same the cataloged API limit that anchors the V4 series. The smaller active footprint is positioned as a deliberate trade-off, delivering quicker response times and a lower API price point while keeping the core reasoning ability close to the flagship Pro variant.
Positioned as a practical, day-to-day workhorse, V4 Flash is described by its creators as approaching the reasoning quality of the V4 Pro model and performing on par with it on simpler agent tasks. The V4 series as a whole introduces a hybrid attention design that combines Compressed Sparse Attention and Heavily Compressed Attention to dramatically reduce the compute and memory cost of long contexts, and Flash benefits directly from that architectural work. This makes it well suited for applications that need long-context understanding, coding assistance, and lightweight agent workflows where speed and cost matter more than maximum reasoning depth.
Quick Info
Powered by- Provider
- LLM Gateway
- Model key
- tencent/deepseek-v4-flash
- Release date
- Apr 24, 2026
- Last updated
- Apr 24, 2026
- Knowledge cutoff
- 2025-05
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.14
- Output token cost
- $0.28
Limits
- Output tokens
- 393,216 tokens
- Context window
- 1,000,000 tokens
Transparent token rates
Compare DeepSeek V4 Flash (Tencent Cloud) pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about DeepSeek V4 Flash (Tencent Cloud)
No articles yet. Fetch the latest news to show it here.