Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

DeepSeek V4.1 Flash (Alibaba Cloud)

DeepSeek V4.1 Flash is positioned by DeepSeek as the smallest member of a new architecture family, built around an asymmetric Mixture-of-Experts design that the creator describes as a Causal Encoder–Decoder. The full model is reported at 552B parameters, with only 8B active during input processing and 16B active during output generation, allowing a single deployment to deliver dense-tier reasoning quality while keeping per-request compute modest. DeepSeek attributes its capabilities to new pretraining methods combined with larger-scale reinforcement-learning post-training, and the company claims benchmark results that land ahead of its own flagship DeepSeek-V4-Pro. The release also ships with native visual understanding, making the model multimodal on the input side while remaining text-only on output, a fit for image-grounded agentic and analytical workflows.

A defining practical strength of V4.1 Flash is its drastically reduced memory and storage footprint relative to the previous generation. DeepSeek reports that the model's KV cache requires roughly one-quarter of the HBM and one-eighth of the SSD storage of its predecessor, which translates directly into lower cache-hit costs for long-running agents and higher throughput under tight memory budgets. The release includes a dedicated agentic benchmark comparison and a publicly hosted technical report on Hugging Face, signalling an emphasis on reproducible evaluation and on serving as an efficient workhorse for tool-using, multi-step tasks. Together, the compact active-parameter profile and the compressed cache make V4.1 Flash well suited to latency-sensitive production deployments where both cost per call and sustained throughput matter more than headline parameter count.

LLM Gatewayalibaba/deepseek-v4.1-flashdeepseek-flash

Quick Info

Powered by
Provider
LLM Gateway
Model key
alibaba/deepseek-v4.1-flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$1.20

Limits

Output tokens
393,216 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare DeepSeek V4.1 Flash (Alibaba Cloud) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4.1 Flash (Alibaba Cloud)

LLM Gateway

CoverageBenchmark

DeepSeek-V4.1-Flash shipped September 10, 2026, and from 12:00 Beijing time (04:00 UTC) on September 14, 2026, every request to deepseek-v4-pro is routed to V4.1 Flash and billed at Flash rates until a V4.1 Pro is released. V4 Flash and the experimental V4 Flash Vision are retired, with legacy names temporarily aliasin Compared with V4 Flash, V4.1 Flash nearly doubles backbone size (284B → 552B) while dropping active parameters per input token from 13B to 8B, and adds 196B Engram conditional memory plus a Causal Encoder-Decoder layout (20 encoder + 20 decoder layers) versus a pure MoE decoder. Terminal-Bench 4.0 jumps from 7.0 to 31.

LLM Gateway

Coverage

DeepSeek announced V4.1-Flash on 2026-09-10 via a thread on X and published the weights on Hugging Face under the MIT licence, according to tech-press coverage. From 04:00 UTC on 2026-09-14, every request to V4-Pro will be routed to V4.1-Flash at Flash-tier pricing until a V4.1-Pro ships, effectively retiring V4-Pro as The article also surfaces DeepSeek's self-reported benchmark table: DeepSWE v1.1 74.2 (Claude Opus 5 74.0, GPT-5.6 Sol 73.0), CyberGym 88.1 (claimed best of listed rivals), but gaps on Humanity's Last Exam at 36.8 versus Opus 5's 56.3, ProgramBench 20.3 against Opus 5's 37.0, and trailing results on Terminal-Bench 3.0.

LLM Gateway

Coverage

DeepSeek released V4.1-Flash at 04:00 UTC on September 10, 2026, alongside a 50-page technical report on Hugging Face, as the debut of a distinct V4.1 architecture lineage rather than an incremental V4 update. The model is a 552-billion-parameter multimodal design that reduces the active key-value cache memory required The technical article attributes the cache-memory reduction to a causal encoder-decoder split combined with 4-bit cache quantization. It also notes that starting September 14, 2026 at 04:00 UTC, all API traffic directed to the deepseek-v4-pro endpoint will be automatically rerouted to V4.1-Flash at V4.1-Flash rates, ma

LLM Gateway

CoverageBenchmark

DeepSeek V4.1 Flash was released September 10, 2026 as a multimodal MoE model served via the deepseek-flash API endpoint, with a 552B-parameter backbone, 8B active parameters during prefill and 16B during decode, a 1M-token context window, up to 384K output tokens, and native image understanding. The model supports thi The release changes input/output cost economics by separating prefill from decode and compressing the KV cache for long agent traces, with teams on retired V4 Flash aliases advised to check routing before comparing old and new runs. DeepSeek is positioning V4.1 Flash as the temporary destination for V4 Pro traffic whil

LLM Gateway

CoverageAnalysis

DeepSeek released DeepSeek-V4.1-Flash on September 10, 2026 as a generational departure from the prior V4 Flash line, introducing a Causal Encoder-Decoder (CED) architecture with 552B backbone parameters and only 8B active during prefill and 16B during decode, reducing KV cache requirements roughly 4× versus V4 Flash. Per the model's own published benchmarks, V4.1 Flash outperforms the larger V4 Pro (1.6T/49B active) across measured evaluations — a first for a Flash-tier model in the DeepSeek lineup. The page frames V4.1 Flash as an inflection point where efficiency gains come from architectural innovation rather than parameter scal

LLM Gateway

CoverageBenchmark

AI Release Tracker records DeepSeek-V4.1-Flash as released on Thursday, September 10, 2026, 28 days after DeepSeek-V4-Pro-0813, with benchmark coverage spanning NL2Repo-Bench, Terminal-Bench 3.0, Terminal-Bench 2.1, CyberGym, AutomationBench, BullshitBench v2, and others. The model is listed as available across multipl Per-provider rates sourced from OpenRouter on September 23, 2026 show DeepInfra at $0.14 input / $0.42 output per 1M tokens (fp8) as the lowest priced tier and DeepSeek's own listing at $0.15 input / $0.60 output per 1M tokens, with most providers clustering around $0.30 input / $1.20 output per 1M tokens. The page not

LLM Gateway

Coverage

DeepSeek officially released DeepSeek-V4.1-Flash on 2026-09-10 as the smallest model in a new architecture family with native multimodal visual understanding, designed for higher capability ceiling, faster inference, higher throughput, and scaling to larger models. Reported benchmarks include GPQA Diamond 90.9, HLE 36. API changes ship alongside the release: DeepSeek V4.1 Flash is now available on the DeepSeek API with native multimodal support, callable via the model name deepseek-flash. The previous-generation deepseek-v4-flash and deepseek-v4-flash-vision-exp endpoints have been retired but are temporarily routed to V4.1 Flash for

Videos about DeepSeek V4.1 Flash (Alibaba Cloud)

More models around DeepSeek V4.1 Flash (Alibaba Cloud)