Model details
DeepSeek V4.1 Flash
DeepSeek V4.1 Flash is a lightweight member of the DeepSeek family designed as a generational replacement for the prior V4 Flash line, with DeepSeek reporting that it surpasses the V4 Pro flagship on performance, cost, speed, and overall throughput. The model is intended as a fast default for production traffic rather than a small specialty variant, which is why existing V4 Pro requests are being routed to it at Flash rates until a future V4.1 Pro appears. Practitioners tend to treat this tier as the practical workhorse for high-volume assistants and pipelines where latency and cost dominate over absolute recall depth, and the Coursiv coverage frames it as DeepSeek's recommended everyday engine for agent-style work rather than knowledge-heavy recall.
The version-4.1 release marks a clear step up from the experimental V4 Flash Vision generation, retiring those endpoints and folding their behavior into V4.1 Flash so existing integrations keep running. In independent discussion threads, users highlight that early benchmarks place it roughly on par with competing fast-tier models while delivering noticeably better cost economics, with cache-miss input costs down around 77 percent and output costs down around 70 percent compared with V4 Pro once the routing switch takes effect. The open-weights status makes V4.1 Flash attractive for self-hosting and fine-tuning scenarios, and its positioning as a fast reasoning and tool-calling model suits teams that want a single, economical backbone for chat, structured output, and lightweight agent loops rather than separate models per task.
Quick Info
Powered by- Provider
- Fireworks AI
- Model key
- accounts/fireworks/models/deepseek-v4p1-flash
- Release date
- Sep 10, 2026
- Last updated
- Sep 10, 2026
- Knowledge cutoff
- 2025-05
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.22
- Output token cost
- $0.66
Limits
- Output tokens
- 384,000 tokens
- Context window
- 1,000,000 tokens
OpenCode
Model variants
Transparent token rates
Compare DeepSeek V4.1 Flash pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about DeepSeek V4.1 Flash
Fireworks AI
A Coursiv blog dated 2026-09-10 provides a side-by-side comparison of DeepSeek V4 Flash versus DeepSeek V4.1 Flash, sourced from DeepSeek's official model card. Key architectural shifts include backbone parameters growing from 284B to 552B, active parameters dropping from 13B to 8B on input and 16B on output, the move The article highlights that from 12:00 Beijing time on September 14, 2026 (04:00 UTC), all requests to 'deepseek-v4-pro' are redirected to V4.1 Flash and billed at Flash rates until a V4.1 Pro ships, yielding approximately 77% lower cache-miss input cost and 70% lower output cost for existing Pro users. DeepSeek's stat
Fireworks AI
DeepSeek released V4.1 Flash on 10 September 2026 as a 552B-parameter open-weights Mixture-of-Experts model that activates just 8B parameters during prefill and 16B during decoding, with a 1M-token context window and native image input, according to Flowtivity's benchmark write-up citing the official model card. The MI Flowtivity reports that the official model card shows V4.1 Flash beating GPT-5.6 Sol and Claude Opus-5.0 on four of five agentic benchmarks: DeepSWE v1.1 (74.2 vs 73.0), AutomationBench (54.8 vs 45.8), Agent's Last Exam (31.8 vs 26.7), and CyberGym (88.1 vs 84.5). DeepSeek is so confident in the model that it announced
Fireworks AI
Neowin reports that DeepSeek released DeepSeek-V4.1 Flash on September 10, 2026, calling it smarter, faster, and more efficient as the smallest model in its new architecture family that now natively supports visual understanding. The article confirms a 552 billion parameter Mixture of Experts design with a new Causal E The piece reproduces DeepSeek's full benchmark comparison table against V4-Pro 0813, V4-Flash 0731, GLM 5.3, Kimi K3, GPT 5.6-Sol, and Claude Opus 5, with V4.1 Flash showing e.g. GPQA Diamond 90.9, HLE 36.8 (39.1 text-only), Codeforces 3471, Terminal-Bench 2.1 90.6, Terminal-Bench 3.0 30.0, DeepSWE v1.1 74.2, and NL2Re
Fireworks AI
Local AI Zone's technical deep-dive describes DeepSeek V4.1 Flash as a fundamental architectural departure using a Causal Encoder-Decoder (CED) design with 552B backbone parameters, only 8B active during prefill and 16B active during decode, and projected global KV cache that the analysis claims reduces KV cache requir The analysis details additional architectural components including Compressed Sparse Attention 2 with Full/Reindex/Reuse modes, FP4 KV cache compression using E2M1 format, SWA Bounded Replay for sliding window attention reconstruction, Single-Pass mHC residual connections, Engram conditional memory at 196B parameters,
Fireworks AI
DeepSeek launched DeepSeek V4.1 Flash on September 10, 2026 as a multimodal model released under an open MIT license, extending the Flash tier with vision capabilities alongside the earlier V4 Pro and Flash releases, according to Emergent's coverage of the announcement. The article frames V4.1 Flash as maintaining the The MIT licensing framing lowers barriers for commercial and research adoption and positions V4.1 Flash to compete with proprietary multimodal systems while preserving DeepSeek's open-source posture, per the article. Emergent notes that the architecture is reported to balance throughput with multimodal performance, tho
Fireworks AI
DeepSeek officially released DeepSeek-V4.1-Flash on September 10, 2026, as the smallest model in its new architecture family with native multimodal visual understanding, according to the DeepSeek API changelog. The release notes state the new architecture targets higher capability ceilings, faster inference, higher thr The changelog additionally documents multimodal capability additions, expanded agent and tool-use benchmarks (including HLE with tools at 63.9, Automation-Bench 54.8, Chartography with tools 78.9, BabyVision with tools 89.6, and ZeroBench-main with tools 49.0), and confirms that DeepSeek V4 Pro API service will continu
Fireworks AI
DeepSeek V4.1 Flash was officially released on September 10, 2026, according to its Baidu Baike entry, which names the model explicitly. It is the smallest member of DeepSeek's new architecture series and ships with native multimodal visual understanding. DeepSeek simultaneously launched the model on the DeepSeek API u The page documents V4.1 Flash as a 552B-parameter Mixture-of-Experts model using a Causal-Encoder-Decoder architecture with asymmetric activations — 8B for input and 16B for output — yielding lower cost than comparably sized models. It extends context length from 4K to 1M tokens, reduces HBM requirements to one-quarter
Videos about DeepSeek V4.1 Flash
More models around DeepSeek V4.1 Flash
This exact model name is also listed by 47 other providers.