Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

DeepSeek V4.1 Flash (Consensus Protocol)

DeepSeek-V4.1-Flash is positioned as the smallest and most efficient member of a new architecture family, designed for greater capability, faster inference, higher throughput, and clean scaling toward larger variants. The model uses a 552B-parameter Mixture-of-Experts design with a new Causal Encoder–Decoder layout, activating just 8B parameters for input processing and 16B for output generation, an asymmetric split that DeepSeek says delivers more intelligence per compute unit while keeping serving costs low. Native visual understanding is built into the architecture rather than bolted on, so image inputs can be handled alongside text in the same workflow.

Training combines new pretraining methods with larger-scale reinforcement-learning post-training, and DeepSeek reports that benchmark results place V4.1-Flash ahead of its flagship DeepSeek-V4-Pro on agentic evaluations, a notable claim for a model explicitly aimed at efficiency. The KV cache footprint has been sharply reduced compared with the previous generation, requiring only a quarter of the HBM and an eighth of the SSD storage, which is particularly valuable for agent workloads where cache-hit charges dominate cost. V4-Flash and V4-Flash-Vision-Exp have been retired, with traffic temporarily routed to V4.1-Flash for compatibility.

LLM Gatewayconsensusprotocol/deepseek-v4.1-flashdeepseek-flash

Quick Info

Powered by
Provider
LLM Gateway
Model key
consensusprotocol/deepseek-v4.1-flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$0.60

Limits

Output tokens
384,000 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare DeepSeek V4.1 Flash (Consensus Protocol) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4.1 Flash (Consensus Protocol)

LLM Gateway

CoverageBenchmark

A third-party review published 2026-09-10 characterizes DeepSeek-V4.1-Flash as the newest Flash model from DeepSeek, designed for higher capability ceiling, faster inference, higher throughput, and scaling to larger models, with native multimodal visual understanding integrated into the new architecture rather than add The same review notes that published coding results improve sharply versus the previous V4 Flash 0731 release, with Terminal-Bench 2.1 rising from 82.7 to 90.6 and DeepSWE from 54.4 to 74.2. It reports DeepSeek's Flash pricing effective 2026-09-10 at $0.003 per million cache-hit input tokens, $0.15 per million uncached

LLM Gateway

CoverageAnalysis

A third-party technical analysis published on 2026-09-10 describes DeepSeek V4.1 Flash as introducing a Causal Encoder-Decoder (CED) architecture with a 552B-parameter backbone, activating roughly 8B parameters during prefill and 16B during decode, and claims a roughly 4× KV cache reduction versus V4 Flash and 437× ver The same analysis reports that V4.1 Flash outperforms DeepSeek V4 Pro on all measured benchmarks despite roughly one-third the total parameters and one-sixth the active parameters, framing this as the first time a Flash-tier model fully replaces a Pro-tier model in the DeepSeek lineup. It positions the gains as an effi

LLM Gateway

CoverageAnalysis

A Substack analysis dated 2026-09-15 traces DeepSeek's KV cache compression trajectory, noting that DeepSeek-V4 (April 2026) introduced CSA and HCA techniques that reduced KV cache storage by roughly 4× and 128× respectively, and built them atop sparse attention with a top-k of 512 indexer to further cut compute during The same article positions V4.1 as a further step beyond V4, achieving approximately 3.9× additional KV cache compression on top of an already compressed state, extending the KV-reduction storyline that DeepSeek has pursued over the last three years. The analysis discusses these compression techniques in the context of

LLM Gateway

Coverage

DeepSeek's official API changelog for 2026-09-10 documents the release of DeepSeek-V4.1-Flash as the smallest model in the company's new architecture family, with native multimodal visual understanding, a design intended for higher capability, faster inference, higher throughput, and scaling to larger models. The chang The same changelog entry confirms that DeepSeek V4.1 Flash is available on the DeepSeek API with native multimodal support, and that the new canonical model name is 'deepseek-flash', which routes to V4.1 Flash, while the previous V4 Flash and V4 Flash Vision Exp have been retired (their legacy names are temporarily ali

Videos about DeepSeek V4.1 Flash (Consensus Protocol)

More models around DeepSeek V4.1 Flash (Consensus Protocol)