Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
CoreWeave logo

Model details

DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is a large mixture-of-experts model built around a 552B-parameter architecture that activates roughly 8B parameters for input and 16B for output, giving it the efficiency profile typical of modern MoE designs while keeping the active compute budget low. This sparse activation pattern is what lets the model deliver Flash-tier responsiveness without paying the full inference cost of its total parameter count, which positions it for high-throughput applications such as coding assistants, retrieval-heavy agents, and long-document analysis. The design reflects DeepSeek's broader strategy of pushing capability-per-dollar rather than chasing the largest dense model, making V4.1 Flash a practical workhorse rather than a flagship reasoning giant.

Released under MIT-licensed open weights, DeepSeek V4.1 Flash lowers the barrier for self-hosting and downstream fine-tuning compared with proprietary peers, and the model's intended use centers on everyday production workloads where latency and cost dominate over raw frontier reasoning. Its open-weight status also encourages community optimization efforts, as evidenced by early discussion threads on enthusiast hardware like NVIDIA's DGX Spark / GB10 platform exploring local deployment. In practice, the model fits teams that want a capable general-purpose assistant with image understanding and tool use, without committing to the expense or closed-source constraints of competing Flash-tier offerings.

CoreWeavedeepseek-ai/DeepSeek-V4.1-Flashdeepseek

Quick Info

Powered by
Provider
CoreWeave
Model key
deepseek-ai/DeepSeek-V4.1-Flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$0.65

Limits

Output tokens
1,048,576 tokens
Context window
1,048,576 tokens

Latest news about DeepSeek V4.1 Flash

Baseten

Coverage

The DeepSeek-V4.1-Flash paper abstract on alphaXiv introduces the model as a multimodal Mixture-of-Experts with 552B backbone parameters and up to 1M-token context. It uses a Causal Encoder-Decoder (CED) architecture that activates only 8B parameters per token during prefill and 16B during decode, a design explicitly a To compress KV cache footprint, V4.1 Flash combines cross-layer reuse in Compressed Sparse Attention 2 (CSA2) with FP4 KV caching in E2M1, reducing the global KV cache (always in HBM) to roughly 890 bytes per token, about one-quarter of DeepSeek-V4-Flash. A deployment optimization called SWA Bounded Replay further cuts

CoreWeave

CoverageBenchmark

DeepSeek V4.1 Flash launched on September 10-11, 2026 as a 552B-parameter MoE model with 8B active parameters per input token and 16B during output generation, designed around agent workloads with a one-million-token context window and roughly 890 bytes/token of KV-cache footprint—about one quarter of V4 Flash's global For API access, the stable model identifier is now `deepseek-flash`; the legacy `deepseek-v4-flash` and `deepseek-v4-flash-vision-exp` IDs temporarily route to V4.1 Flash for compatibility, and from September 14, 2026 the `deepseek-v4-pro` endpoint also serves V4.1 Flash at Flash pricing until a future V4.1 Pro launch.

Deep Infra

CoverageBenchmark

Coursiv's write-up, drawing on DeepSeek's official model card, frames DeepSeek-V4.1-Flash as a generational replacement for V4 Flash rather than a point update, with the old V4 Flash and V4 Flash Vision names temporarily routed to the new model so existing code keeps running. It reports that from 12:00 Beijing time (04 The blog's spec table contrasts V4 Flash (284B backbone, 13B active, MoE decoder) with V4.1 Flash (552B backbone, 8B active reading / 16B generating, causal encoder-decoder with 20 encoder + 20 decoder layers, a separate 196B-parameter Engram conditional memory, native vision trained in from pre-training, ~890-byte glo

Hugging Face

CoverageBenchmark

A third-party Flowtivity blog post dated September 10, 2026 summarizes the DeepSeek V4.1 Flash release for an Australian business audience, restating figures from the official model card: a 552B-parameter MoE backbone that activates 8B parameters per token during prefill and 16B during decoding, a 1M-token context wind The same post notes DeepSeek's stated rationale for the release, paraphrasing the company as saying V4.1 Flash has comprehensively surpassed V4 Pro across performance, cost, speed, and task completion time, alongside a DeepSeek-cited figure of $0.30 per million input tokens at peak (pricing that is not independently ve

CoreWeave

CoverageBenchmark

An independent community editorial frames DeepSeek V4.1 Flash as the production target to test first for new high-volume deployments, noting DeepSeek treats it as the current Flash target, routes older Flash aliases through compatibility logic, and claims V4.1 Flash surpasses V4 Pro on performance, cost, speed, and tot The deployment profile covers a 1M-token context window, 384K-token maximum output, and peak/off-peak pricing tiers, while the editorial explicitly flags evidence-maturity limits: the September 14 V4 Pro route switch is still scheduled, and same-version independent benchmark results for V4.1 Flash are not yet as comple

Deep Infra

Coverage

Requesty's blog documents that at 06:10 UTC on 10 September 2026 DeepSeek posted a six-tweet thread introducing DeepSeek-V4.1-Flash, and reports community-driven social traction including 21,110 likes, 3,311 bookmarks, and 2.78 million impressions on the announcement within ten hours, plus 70 X and Reddit posts naming The same piece covers community-reported on-disk weight composition from r/LocalLLaMA — roughly 552B in the main model plus ~197B of optional Engram parameters and ~14B of speculative-decoding weights, for a total near 748B on disk, a relevant caveat for anyone planning to self-host rather than use a hosted endpoint. I

Deep Infra

CoverageBenchmark

Independent tracking site AI Release Tracker logs DeepSeek-V4.1-Flash as a new release dated 10 September 2026, arriving 28 days after DeepSeek-V4-Pro-0813, and provides a broad benchmark sweep that includes coding, terminal, cybersecurity, and business-workflow evaluations. Its leaderboard positioning shows DeepSeek-V Concrete published scores include NL2Repo-Bench 65.4 (best tracked), DeepSWE v1.1 74.2 (2nd, behind Muse Spark 1.3's 75.4), Terminal-Bench 3.0 30% (best tracked), Terminal-Bench 2.1 90.6 (best tracked), Terminal-Bench 4.0 31.2 (8th), CyberGym 88.1 (best tracked), and AutomationBench 54.8 (best tracked), corroborating t

Hugging Face

Coverage

DeepSeek officially released DeepSeek-V4.1-Flash on 2026-09-10, according to its first-party API documentation changelog. The model is described as the smallest entry in a new architecture family, offering native multimodal visual understanding and designed for a higher capability ceiling, faster inference, higher thro The same changelog entry documents API-level changes accompanying the V4.1-Flash release: V4.1 Flash is invoked via the new model name 'deepseek-flash' on the DeepSeek API with native multimodal support, while the previous-generation V4 Flash and V4 Flash Vision Exp have been retired (their old names are temporarily ro

Videos about DeepSeek V4.1 Flash

More models around DeepSeek V4.1 Flash