Sulat.com
AI models
Alibaba Token Plan (China) logo

Model details

DeepSeek V4 Flash

DeepSeek V4 Flash is the efficiency-focused sibling in the V4 Preview lineup, designed for fast and economical inference at long context. Its Mixture-of-Experts design pairs a large overall parameter pool with a much smaller active subset, letting it keep response latency low while still drawing on broad learned capacity. The model was released alongside a larger Pro variant as part of a coordinated V4 Preview launch, sharing the same generation's context and architectural innovations rather than being a stripped-down older release.

For practical fit, V4 Flash targets workloads where cost per token and throughput matter more than squeezing out the last fraction of benchmark accuracy, such as high-volume chat, tool-mediated agentic tasks, and retrieval-heavy pipelines over very long inputs. DeepSeek positions its reasoning as closely approaching the Pro variant while matching it on simpler agent tasks, making Flash a strong default when developers want V4 quality at lower latency and price. Open weights on Hugging Face also let teams self-host for data-sensitive or latency-critical deployments.

Alibaba Token Plan (China)deepseek-v4-flashdeepseek-flash

Quick Info

Powered by
Provider
Alibaba Token Plan (China)
Model key
deepseek-v4-flash
Release date
Apr 24, 2026
Last updated
Apr 24, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
384,000 tokens
Context window
1,000,000 tokens

Latest news about DeepSeek V4 Flash

Alibaba Token Plan (China)

CoverageAnalysis

DeepSeek V4 Flash 0731, evaluated by Artificial Analysis on July 31, 2026, scores 50 on the Artificial Analysis Intelligence Index—a 10-point jump over the April 2026 release of DeepSeek V4 Flash and 6 points ahead of DeepSeek V4 Pro. The model retains the same 1M-token context window and identical architecture/pricing On agentic workloads, V4 Flash 0731 reaches an Elo of 1559 on GDPval-AA v2 (up from 1189), Terminal-Bench 2.1 rises 17 points to 79%, and τ³-Bench Banking improves 8 points to 31%. AA-Omniscience gains stem from a 12-point drop in hallucination rate (to 84%) with accuracy unchanged, and DeepSeek's 98% cache-hit discoun

Videos about DeepSeek V4 Flash

More models around DeepSeek V4 Flash