DeepSeek-V4-Flash is a Mixture-of-Experts language model developed by DeepSeek as part of the broader DeepSeek-V4 collection. It carries 284B total parameters with only 13B activated per token, a configuration that aims to deliver strong capability while keeping inference costs and latency manageable. The model is positioned as an efficiency-optimized variant of the V4 line, designed for fast inference and high-throughput workloads without sacrificing core reasoning quality. Its open weights are published under the deepseek-ai organization on Hugging Face, making it accessible for self-hosting, fine-tuning, and research use.
DeepSeek-V4-Flash supports a one-the cataloged API limit and incorporates hybrid attention to handle long inputs more efficiently, which makes it well suited for coding assistants, conversational chat systems, and agent workflows where long documents and sustained responsiveness matter. Configurable reasoning effort is available, with both high and xhigh settings supported and xhigh mapping to maximum reasoning depth. On standardized evaluations listed through OpenRouter, the model posts an Intelligence Index of 24.2, a Coding Index of 56.2, and an Agentic Index of 22.2 under max-effort reasoning, indicating a balanced profile rather than a single-axis specialist.