Currently listed through these providers:
Model details
DeepSeek V4 Flash
DeepSeek V4 Flash is a 284-billion-parameter mixture-of-experts model with 13 billion active parameters, designed as the streamlined sibling to the V4 Pro variant within the broader V4 family. Its MoE architecture routes queries through only a fraction of the total weights per token, which is how the model delivers faster response times while keeping compute costs low. The variant was introduced alongside V4 Pro in the April 2026 preview release and ships as an open-weight model, making it attractive to teams that want to self-host or fine-tune rather than rely exclusively on a closed API. By leaning on a sparse-experts design rather than a dense transformer, the model can carry substantial world knowledge and reasoning capacity without the inference latency of a similarly sized dense network.
The model's practical positioning centers on bringing near-Pro reasoning quality to workloads where latency and cost matter more than squeezing out the last few accuracy points. DeepSeek's release notes describe its reasoning capabilities as closely approaching V4 Pro, while matching V4 Pro on simpler agentic tasks, and it benefits from a 1M-token context window that supports long-document reasoning and multi-turn agent workflows. Its combination of open weights, efficient sparse activation, and competitive agentic coding performance makes it well suited for production assistants, tool-using agents, and cost-sensitive deployments where developers want a capable but economical engine.
Quick Info
Powered by- Provider
- Weights & Biases
- Model key
- deepseek-ai/DeepSeek-V4-Flash
- Release date
- Apr 24, 2026
- Last updated
- Apr 24, 2026
- Knowledge cutoff
- 2025-05
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.14
- Output token cost
- $0.28
Limits
- Output tokens
- 1,048,576 tokens
- Context window
- 1,048,576 tokens