Sulat.com
AI models
CrossModel logo

Model details

DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is framed by its developers around the theme of "Pushing the Limits of KV Cache Compression," signaling that long-context efficiency is the central design goal rather than a raw scale story. Per the official model card, it is a multimodal Mixture-of-Experts model with 552B backbone parameters and native support for contexts of up to the cataloged API limit, processing images and text as input and emitting text autoregressively. That combination makes it most naturally suited to retrieval-heavy, document-rich, and agentic tasks where very large prompts dominate the cost profile and where image understanding can be mixed into the same conversation.

The architectural emphasis on cache compression shows up in the choice of a Causal Encoder-Decoder layout: a 40-layer Transformer split into a 20-layer causal encoder and a 20-layer decoder, with the decoder's global KV cache projected from the final encoder hidden states instead of being maintained per decoder layer. This design activates only about 8B parameters during prefill and roughly 16B during decode, yielding meaningful cost savings on input-bound workloads. The model is released under an MIT license, which lowers the barrier for self-hosting, fine-tuning, and integration into custom pipelines for teams that want long-context multimodal reasoning without surrendering ownership of their stack.

CrossModeldeepseek/deepseek-v4.1-flashdeepseek-flash

Quick Info

Powered by
Provider
CrossModel
Model key
deepseek/deepseek-v4.1-flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.27
Output token cost
$1.08

Limits

Output tokens
384,000 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare DeepSeek V4.1 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4.1 Flash

NanoGPT

Official sourceAnnouncement

DeepSeek announced DeepSeek-V4.1-Flash on September 9–10, 2026, introducing the smallest model in a new architecture family with native visual understanding. The model is a 552B-parameter mixture-of-experts that uses a new Causal Encoder–Decoder design, activating just 8B parameters for input and 16B for output. Accord V4.1-Flash went live on the DeepSeek API as `deepseek-flash` with native multimodal support and 1M-token context (384K max output), while V4-Flash and V4-Flash-Vision-Exp were retired and their old IDs temporarily aliased to the new model. DeepSeek also disclosed that starting 04:00 UTC on September 14, 2026, every `de

NanoGPT

CoverageBenchmark

Coursiv's write-up, drawn from DeepSeek's official model card, compares DeepSeek V4 Flash and V4.1 Flash side by side. Backbone parameters roughly doubled from 284B to 552B, while active parameters per token shifted from 13B to 8B on input and 16B on output, and the architecture changed from a MoE decoder to a 20-plus- The article also notes that from 04:00 UTC on September 14, 2026, every `deepseek-v4-pro` request is silently routed to V4.1 Flash at Flash rates until a V4.1 Pro ships, with DeepSeek citing about a 77% drop in cache-miss input cost and roughly 70% lower output cost versus V4 Pro. V4 Flash and V4 Flash Vision were reti

NanoGPT

CoverageBenchmark

BenchLM's profile for DeepSeek V4.1 Flash, dated September 10, 2026, confirms the model's core specifications and pricing: API model ID `deepseek-flash`, 1M-token context, open-weight release, and DeepSeek-published pricing of $0.30 per million input tokens and $1.20 per million output tokens, with cached input at $0.0 The profile is tracked but not yet publicly ranked because no category has enough eligible evidence for a comparative placement, and independent runtime speed has not been measured. Knowledge-cutoff and maximum-output fields are explicitly listed as not sourced yet. The page acts as a neutral aggregator tying each spec

NanoGPT

Coverage

Requesty's coverage walks through DeepSeek's own description of V4.1-Flash: a 552B-parameter MoE described as the smallest in a new architecture family, MIT-licensed on Hugging Face, with a Causal Encoder–Decoder design that activates 8B parameters on input and 16B on output — the asymmetric compute split being the hea The piece flags a sharp lifecycle change: from 04:00 UTC on September 14, 2026, pinned `deepseek-v4-pro` model IDs will start serving V4.1-Flash at Flash rates until V4.1-Pro launches, a redirect that production teams need to plan around. Community notes from r/LocalLLaMA estimate on-disk footprint at roughly 552B main

NanoGPT

CoverageBenchmark

Benchmark tracker AI Release Tracker logged DeepSeek-V4.1-Flash as released on September 10, 2026, 28 days after DeepSeek-V4-Pro-0813, and compiled published scores across NL2Repo-Bench, Terminal-Bench 3.0/2.1/4.0, CyberGym, AutomationBench, Chartography, and additional suites. Notable figures include 65.4% on NL2Repo- The tracker frames V4.1-Flash as a leader on coding, agentic terminal, and cybersecurity benchmarks relative to the models it tracks, while its Terminal-Bench 4.0 placement shows there is still headroom on the newest, recalibrated agentic terminal test where Claude Mythos 5.1 leads at 60.9%. Together with DeepSeek's ow

Videos about DeepSeek V4.1 Flash

More models around DeepSeek V4.1 Flash