Sulat.com
AI models
Deep Infra logo

Model details

DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is a large mixture-of-experts model built around a 552B-parameter architecture that activates roughly 8B parameters for input and 16B for output, giving it the efficiency profile typical of modern MoE designs while keeping the active compute budget low. This sparse activation pattern is what lets the model deliver Flash-tier responsiveness without paying the full inference cost of its total parameter count, which positions it for high-throughput applications such as coding assistants, retrieval-heavy agents, and long-document analysis. The design reflects DeepSeek's broader strategy of pushing capability-per-dollar rather than chasing the largest dense model, making V4.1 Flash a practical workhorse rather than a flagship reasoning giant.

Released under MIT-licensed open weights, DeepSeek V4.1 Flash lowers the barrier for self-hosting and downstream fine-tuning compared with proprietary peers, and the model's intended use centers on everyday production workloads where latency and cost dominate over raw frontier reasoning. Its open-weight status also encourages community optimization efforts, as evidenced by early discussion threads on enthusiast hardware like NVIDIA's DGX Spark / GB10 platform exploring local deployment. In practice, the model fits teams that want a capable general-purpose assistant with image understanding and tool use, without committing to the expense or closed-source constraints of competing Flash-tier offerings.

Deep Infradeepseek-ai/DeepSeek-V4.1-Flashdeepseek-flash

Quick Info

Powered by
Provider
Deep Infra
Model key
deepseek-ai/DeepSeek-V4.1-Flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$1.20

Limits

Output tokens
384,000 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare DeepSeek V4.1 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4.1 Flash

Deep Infra

Coverage

DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552B-parameter Mixture-of-Experts model and the smallest member of a new architecture family, with weights published on Hugging Face under the MIT license. According to the DeepSeek-authored paper indexed on alphaXiv, it uses a Causal Encoder-Decoder desig The same paper documents the release's KV-cache compression work: a Compressed Sparse Attention 2 mechanism with cross-layer KV reuse plus FP4 KV caching, supported by a deployment optimization called SWA Bounded Replay, which together reduce the global KV footprint to roughly 890 bytes per token — about one quarter of

Deep Infra

CoverageBenchmark

Coursiv's write-up, drawing on DeepSeek's official model card, frames DeepSeek-V4.1-Flash as a generational replacement for V4 Flash rather than a point update, with the old V4 Flash and V4 Flash Vision names temporarily routed to the new model so existing code keeps running. It reports that from 12:00 Beijing time (04 The blog's spec table contrasts V4 Flash (284B backbone, 13B active, MoE decoder) with V4.1 Flash (552B backbone, 8B active reading / 16B generating, causal encoder-decoder with 20 encoder + 20 decoder layers, a separate 196B-parameter Engram conditional memory, native vision trained in from pre-training, ~890-byte glo

Deep Infra

Coverage

Requesty's blog documents that at 06:10 UTC on 10 September 2026 DeepSeek posted a six-tweet thread introducing DeepSeek-V4.1-Flash, and reports community-driven social traction including 21,110 likes, 3,311 bookmarks, and 2.78 million impressions on the announcement within ten hours, plus 70 X and Reddit posts naming The same piece covers community-reported on-disk weight composition from r/LocalLLaMA — roughly 552B in the main model plus ~197B of optional Engram parameters and ~14B of speculative-decoding weights, for a total near 748B on disk, a relevant caveat for anyone planning to self-host rather than use a hosted endpoint. I

Deep Infra

CoverageBenchmark

Independent tracking site AI Release Tracker logs DeepSeek-V4.1-Flash as a new release dated 10 September 2026, arriving 28 days after DeepSeek-V4-Pro-0813, and provides a broad benchmark sweep that includes coding, terminal, cybersecurity, and business-workflow evaluations. Its leaderboard positioning shows DeepSeek-V Concrete published scores include NL2Repo-Bench 65.4 (best tracked), DeepSWE v1.1 74.2 (2nd, behind Muse Spark 1.3's 75.4), Terminal-Bench 3.0 30% (best tracked), Terminal-Bench 2.1 90.6 (best tracked), Terminal-Bench 4.0 31.2 (8th), CyberGym 88.1 (best tracked), and AutomationBench 54.8 (best tracked), corroborating t

Videos about DeepSeek V4.1 Flash

More models around DeepSeek V4.1 Flash