Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
CoralBricks logo

Model details

DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is a multimodal Mixture-of-Experts model from DeepSeek-AI that reads text and images and writes text autoregressively, with open weights released under an MIT license on the DeepSeek-AI Hugging Face organization. Its backbone spans 552B parameters, yet the model activates only a small slice of them per token, roughly 8B during prefill and 16B during decode, an efficiency pattern tailored to input-heavy agentic workloads where prompt processing dominates cost. The release is documented in a companion technical report and is positioned for both commercial and non-commercial use.

The defining technical bet is a Causal Encoder-Decoder layout: a 40-layer Transformer split into a 20-layer causal encoder and a 20-layer decoder, where the decoder's global KV cache is projected from the final encoder hidden states instead of being maintained layer by layer. A Sliding-Window-Attention Bounded Replay scheme rebuilds missing SWA KV states from the most recent tokens, letting the model push KV cache compression further without giving up long-range behavior. Combined with continuously controllable reasoning effort and a context that scales to roughly one million tokens, the practical fit is long-conversation assistants, document and image analysis pipelines, and tool-using agents that need sustained reasoning over very large inputs.

CoralBricksdeepseek-v4.1-flash-fastdeepseek-flash

Quick Info

Powered by
Provider
CoralBricks
Model key
deepseek-v4.1-flash-fast
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$1.20

Limits

Output tokens
131,072 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare DeepSeek V4.1 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4.1 Flash

CoralBricks

CoverageBenchmark

A third-party breakdown of the V4.1 Flash release frames it as a generational replacement for V4 Flash rather than a point update. Compared to V4 Flash, backbone parameters grow from 284B to 552B while active parameters per token fall from 13B to 8B for input and 16B for output, and the architecture shifts from a Mixture-of-Experts decoder to a Causal Encoder-Decoder with 20 encoder plus 20 decoder layers and a 196B-parameter Engram conditional memory. Vision moves from a separate experimental model to native training from pre-training start, and global KV cache per token drops to roughly a quarter of V4 Flash at 890 bytes. Terminal-Bench 2.1 rises from 82.7 to 90.6, Terminal-Bench 4.0 from 7.0 to 31.2, and DeepSWE v1.1 from 54.4 to 74.2, all at MIT license with a 1M-token context. The piece documents DeepSeek's decision to route all deepseek-v4-pro traffic to V4.1 Flash from 12:00 Beijing time on September 14, 2026 (04:00 UTC), billing at Flash rates until a V4.1 Pro ships, which drops cache-miss input cost by about 77 percent and output cost by about 70 percent for Pro users. V4 Flash and experimental V4 Flash Vision are retired, with their old model names temporarily routing to V4.1 Flash so existing code keeps running. DeepSeek's stated reason is that V4.1 Flash surpassed V4 Pro on performance, cost, speed, and total time, supported by its own agentic benchmark numbers.

CoralBricks

CoverageBenchmark

A third-party analysis of the official DeepSeek V4.1 Flash model card reports the model beats GPT-5.6 Sol and Claude Opus-5.0 on four of five hardest agentic benchmarks: DeepSWE v1.1 (74.2 vs 73.0), AutomationBench (54.8 vs 45.8), Agents' Last Exam (31.8 vs 26.7), and CyberGym (88.1 vs 84.5). The architecture is a 552B-parameter Mixture-of-Experts that activates only 8B parameters per token during prefill, ships with a 1M-token context window, reads images natively, is MIT-licensed, and costs $0.30 per million input tokens at peak. DeepSeek is retiring the roughly four-times-more-expensive V4 Pro on September 14, 2026, stating that V4.1 Flash has comprehensively surpassed it. The article frames V4.1 Flash as a major step for agentic coding workflows, with full benchmark tables, pricing math, and KV-cache economics analysis tied to Australian teams. It also notes V4 Pro is being retired and that a V4.1 Pro is coming. The piece serves as derivative analysis of the official model card rather than primary reporting, and confirms the September 10, 2026 release date, the multimodal MoE architecture, and the head-to-head results against frontier competitors cited above.

CoralBricks

CoverageBenchmark

An aggregator page ranks DeepSeek-V4.1-Flash 15th on the LLM Stats composite score, with capability tiers placing it in the S tier for Tool Calling (4 of 202), A tier for Coding (7 of 275) and Reasoning (18 of 371), and B tier for Vision (31 of 215) and Math (47 of 329). Tracked scores include Codeforces 3471 at rank 1, GPQA Diamond 0.91 at rank 23, and Terminal-Bench 2.1 0.91 at rank 1, with methodology noting maximum reasoning effort (reasoning_effort=100), temperature 1.0, and top_p 0.95. Blended price is listed at $0.24 per million tokens at a composite LLM Stats Score of 51.2. The page situates DeepSeek-V4.1-Flash against peers like Gemma 4 E4B ($0.024, score 13.8), GPT OSS 120B ($0.043, 28.7), DeepSeek-V4-Flash-0731 ($0.066, 43.8), and Claude Opus 5.5 ($4.76, 60.3) on a log-scale price-versus-score plot. The data was last refreshed on Sun Oct 04 2026. Per the page, DeepSeek-V4.1-Flash scores are sourced from the model's scorecard, paper, or official blog posts rather than re-evaluated independently.

CoralBricks

Coverage

DeepSeek officially released DeepSeek-V4.1-Flash on 2026-09-10 as the smallest model in a new architecture family with native multimodal visual understanding, designed for a higher capability ceiling, faster inference, and higher throughput. The official benchmark list includes GPQA Diamond 90.9, HLE 36.8 (39.1 text-only), Codeforces 3471, MathArena Apex 65.6, Terminal-Bench 2.1 90.6, Terminal-Bench 3.0 30.0, Terminal-Bench 4.0 31.2, DeepSWE v1.1 74.2, ProgramBench 20.3, NL2Repo-Bench 65.4, CyberGym 88.1, SEC-Bench Pro 62.8, ExploitGym 15.3, HLE w/tools 63.9, Automation-Bench 54.8, Agents' Last Exam 31.8, Chartography w/tools 78.9, BabyVision w/tools 89.6, and ZeroBench-main w/tools 49.0. The model name deepseek-flash now calls the V4.1 Flash endpoint. The changelog also retires the previous-generation V4 Flash and V4 Flash Vision Exp, with old model names temporarily routed to V4.1 Flash for compatibility. In response to demand, DeepSeek will continue providing API services for V4 Pro after September 14, 2026, with unchanged billing. API prices have been reduced with the V4.1 Flash release, with details on the Models & Pricing page. This is the first-party changelog entry announcing the model, its benchmark suite, and the associated API surface changes.

Videos about DeepSeek V4.1 Flash

More models around DeepSeek V4.1 Flash