Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
302.AI logo

Model details

DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is positioned as the smallest member of a new architecture family, built around an asymmetric Causal Encoder–Decoder design inside a 552B-parameter Mixture-of-Experts framework. Only 8B parameters activate on the input side while 16B activate on the output side, an arrangement DeepSeek says is engineered for greater capability, faster inference, and higher throughput than the previous generation. The model is launched with native visual understanding, allowing it to take in image inputs alongside text and return text outputs, which broadens its utility for tasks that combine documents, screenshots, or diagrams with natural language prompts.

DeepSeek attributes the model's reported gains to fresh pretraining methods combined with larger-scale reinforcement-learning post-training, with benchmark results said to surpass flagship models including DeepSeek-V4-Pro. A second efficiency story comes from the memory side: the KV cache has been compressed to roughly one quarter of the previous generation's HBM footprint and one eighth of its SSD footprint, a meaningful reduction for agent-style workloads where cache-hit reads typically dominate cost. Together, these architectural and training choices aim at practitioners who want strong multimodal reasoning at lower per-query cost, particularly in long-context and high-throughput agent pipelines.

302.AIdeepseek-flashdeepseek-flash

Quick Info

Powered by
Provider
302.AI
Model key
deepseek-flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.60

Limits

Output tokens
384,000 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare DeepSeek V4.1 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4.1 Flash

302.AI

CoverageBenchmark

BenchLM.ai's model page for DeepSeek V4.1 Flash aggregates published benchmark, pricing, and speed data, listing 22 published rows with verified coverage in agentic (8/8), coding (6/6), knowledge (3/3), math (1/1), and multimodal (3/3) categories. The page reports API prices of $0.30 per million input tokens and $1.20 The page notes that DeepSeek V4.1 Flash is currently not eligible for a comparative public rank, with capability field-median at 56.3, and that reasoning and multilingual categories are not measured. Verified evidence sources are separated from provisional rows, and category scores are listed as pending across several

302.AI

CoverageBenchmark

A third-party eesel AI analysis published 2026-09-11 reports that DeepSeek shipped V4.1 Flash on 2026-09-10 as a new architecture design rather than a tweak of the prior V4 Flash, calling it 'the smallest model in our new architecture family, with native visual understanding.' The piece highlights the cleanup: V4 Flash The article also surfaces community concern from Hacker News about API customers being silently migrated to a different model without a deprecation window, which has practical implications for users pinning model versions in production. DeepSeek's counterpoint noted is that it is a small lab realistically able to host

302.AI

CoverageBenchmark

Coursiv's blog post dated September 10, 2026 covers the release of DeepSeek-V4.1-Flash and contrasts it with the prior DeepSeek V4 Flash. Compared to V4 Flash's 284B backbone and 13B active parameters per token, V4.1 Flash is a 552B-parameter Causal Encoder-Decoder with 20 encoder plus 20 decoder layers, activating 8B The blog also reports a routing change on DeepSeek's platform: from 12:00 Beijing time on September 14, 2026 (04:00 UTC), requests to deepseek-v4-pro are routed to V4.1 Flash and billed at Flash rates until a V4.1 Pro is released. DeepSeek's stated rationale is that V4.1 Flash surpassed V4 Pro across performance, cost,

302.AI

CoverageBenchmark

Flowtivity's analysis describes DeepSeek V4.1 Flash as a multimodal Mixture-of-Experts model released 2026-09-10 with 552B backbone parameters, 8B active per token during prefill and 16B during decoding, a 1M-token context window with up to 384K output tokens, and native image and text processing. Weights are released The article reports pricing at $0.30 per million input tokens at peak and claims V4.1 Flash beats GPT-5.6 Sol and Claude Opus-5.0 on four of five hard agentic benchmarks per the official model card: DeepSWE v1.1 74.2 versus Sol's 73.0, AutomationBench 54.8 versus 45.8, Agents' Last Exam 31.8 versus 26.7, and CyberGym 8

302.AI

Coverage

An alphaXiv-hosted paper dated September 10, 2026 introduces DeepSeek-V4.1-Flash as a multimodal Mixture-of-Experts model with 552B backbone parameters and support for contexts of up to one million tokens. The model uses a Causal Encoder-Decoder (CED) architecture that activates 16B parameters per token during decode b The paper documents DeepSeek-V4.1-Flash's KV cache compression design, combining cross-layer KV cache reuse in Compressed Sparse Attention 2 (CSA2) with FP4 KV caching. This reduces the global KV cache footprint (always in HBM) to 890 bytes per token, roughly 1/4 of the corresponding footprint of DeepSeek-V4-Flash. A d

302.AI

Coverage

DeepSeek's official API changelog dated 2026-09-10 announces the release of DeepSeek-V4.1-Flash, described as the smallest model in DeepSeek's new architecture family with native multimodal visual understanding. The release is positioned for a higher capability ceiling, faster inference, and higher throughput, scaling The changelog also documents API-level changes accompanying the release. V4.1 Flash is available on the DeepSeek API with native multimodal support and is invoked using the model name "deepseek-flash". The previous-generation V4 Flash and V4 Flash Vision Exp models have been retired, with their legacy names temporarily

Videos about DeepSeek V4.1 Flash

More models around DeepSeek V4.1 Flash