Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Ollama Cloud logo

Model details

deepseek-v4-flash

DeepSeek-V4-Flash is a Mixture-of-Experts language model developed by DeepSeek as part of the broader DeepSeek-V4 collection. It carries 284B total parameters with only 13B activated per token, a configuration that aims to deliver strong capability while keeping inference costs and latency manageable. The model is positioned as an efficiency-optimized variant of the V4 line, designed for fast inference and high-throughput workloads without sacrificing core reasoning quality. Its open weights are published under the deepseek-ai organization on Hugging Face, making it accessible for self-hosting, fine-tuning, and research use.

DeepSeek-V4-Flash supports a one-the cataloged API limit and incorporates hybrid attention to handle long inputs more efficiently, which makes it well suited for coding assistants, conversational chat systems, and agent workflows where long documents and sustained responsiveness matter. Configurable reasoning effort is available, with both high and xhigh settings supported and xhigh mapping to maximum reasoning depth. On standardized evaluations listed through OpenRouter, the model posts an Intelligence Index of 24.2, a Coding Index of 56.2, and an Agentic Index of 22.2 under max-effort reasoning, indicating a balanced profile rather than a single-axis specialist.

Ollama Clouddeepseek-v4-flashdeepseek-flash

Quick Info

Powered by
Provider
Ollama Cloud
Model key
deepseek-v4-flash
Release date
Apr 24, 2026
Last updated
Apr 24, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.22
Output token cost
$0.66

Limits

Output tokens
1,048,576 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare deepseek-v4-flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about deepseek-v4-flash

InferX

Coverage

The AI Free API guide, dated August 26, 2026, maps DeepSeek's current official API surface and assigns the `deepseek-v4-flash` text route to the DeepSeek-V4-Flash-0731 build, with `deepseek-v4-pro` mapped to DeepSeek-V4-Pro-0813 and `deepseek-v4-flash-vision-exp` released on the API platform on August 21 as an experime The specification section states that all three current routes share a 1-million-token context window and a 384,000-token maximum output, with JSON output, tool calls, Responses API support, and an Anthropic-compatible interface listed for each, while FIM completion is limited (excerpt cuts off). The page also notes th

Ollama Cloud

CoverageRelease Notes

Swipeer's v5.8.0 changelog (Jul 31, 2026) states that it "Upgraded DeepSeek V4 Flash to the latest 0731 revision for improved coding, reasoning, and agentic workflows." This is one of several downstream confirmations that the 0731 revision of deepseek-v4-flash is a real release being picked up by third-party OpenRouter The entry is qualitative — no benchmarks, token counts, or API changes are provided — and Swipeer is not an Ollama Cloud source, so its relevance is limited to confirming that the 0731 revision exists and has propagated. It usefully indicates improved agentic behavior as a stated reason for the upgrade, which is a deve

Ollama Cloud

Coverage

TextQL announced on August 4, 2026 that the "DeepSeek V4 Flash 0731" checkpoint—DeepSeek's efficiency-focused model—is now available as a selectable option in the Ana platform's model picker, joining Claude, GPT, and Kimi families. The post positions V4 Flash as well-suited for routine workloads such as scheduled repor Enabling instructions state that V4 Flash is available to all Ana users and can be activated organization-wide by navigating to Settings → Models and toggling "Enabled" next to Deepseek v4 Flash, with full configuration options documented separately. The announcement is part of TextQL's broader rapid-release cadence an

Ollama Cloud

CoverageAnalysis

Powerdrill's August 3, 2026 analysis provides the most technically detailed treatment of DeepSeek V4 Flash currently serving traffic, confirming that the deepseek-v4-flash endpoint was updated to the "DeepSeek-V4-Flash-0731" checkpoint on July 31, 2026—the version accessible via Ollama Cloud's API. It documents that V4 The piece also flags an upcoming peak/off-peak pricing policy in which prices would double during 9:00–12:00 and 14:00–18:00 Beijing time daily across all billing items, with the effective date still pending an official DeepSeek announcement—an important operational consideration for developers scheduling batch jobs ag

Ollama Cloud

Coverage

An NVIDIA DGX Spark forum thread (Aug 1, 2026) details a forked CUDA/MLX inference engine (ds4, originally antirez's) tuned for running DeepSeek-V4-Flash-0731 on a single GB10. The maintainer reports v0.5 results on the new 0731 weights including ~1,000 tok/s prefill at best frontiers (~960 tok/s at 2k context, ~933 to The fork adds continuous batching, prefix caching with warm starts and fork-by-copy for parallel agent branches, disk KV persistence, lossless speculative decode, and an OpenAI-compatible server — all running on a single DGX Spark fully on-device. While these numbers are self-reported on a single consumer Blackwell SKU

InferX

CoverageBenchmark

The BenchLM profile tracks the most likely exact variant behind InferX's `deepseek-v4-flash` key: DeepSeek V4 Flash 0731, released July 31, 2026, with the API model ID `deepseek-v4-flash`, a 1M-token context window, text input and output modalities, and reasoning support. It lists a published input price of $0.14 and o BenchLM flags the model as tracked but not yet publicly ranked, noting no category has enough eligible evidence for a comparative rank and that provisional rows remain visible separately from verified ones. Maximum output length, knowledge cutoff, and several other spec fields are marked "Not sourced yet," so those att

Ollama Cloud

CoverageAnalysis

A third-party compute analysis from Executive Mind (Jul 9, 2026) directly examines Ollama Cloud's deepseek-v4-flash tier and explains how Ollama weights usage by GPU compute difficulty rather than raw tokens. It reports that DeepSeek V4 Flash operates with 13B active parameters and consumes roughly 73% less compute tha The same analysis frames active parameters per token and hidden reasoning-token overhead as the dominant cost drivers on Ollama Cloud, concluding that a 13B-active MoE like V4 Flash costs a small fraction of a 49B-active MoE for the same 1,000-token response. While the figures come from a single analyst blog without di

Azure

CoverageBenchmark

DeepSeek shipped the official DeepSeek-V4-Flash-0731 release on 31 July 2026, superseding the April preview with the same architecture and initial price. According to the article's source material, the model is a Mixture-of-Experts design in the open-weight V4 family with 284 billion total parameters and 13 billion act The piece details efficiency gains at long context: at one million tokens V4-Flash is said to use roughly 10% of the single-token inference FLOPs and 7% of the KV cache of DeepSeek V3.2, extending the same family of improvements credited to V4-Pro. Distribution is via Hugging Face under the MIT License, the DeepSeek AP

Videos about deepseek-v4-flash

More models around deepseek-v4-flash