Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Eden AI logo

Model details

DeepSeek V4.1 Flash (FlexAI)

The model overview is being prepared.

Eden AIflexai/DeepSeek-V4.1-Flashdeepseek-flash

Quick Info

Powered by
Provider
Eden AI
Model key
flexai/DeepSeek-V4.1-Flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.14
Output token cost
$0.42

Limits

Output tokens
384,000 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare DeepSeek V4.1 Flash (FlexAI) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4.1 Flash (FlexAI)

Eden AI

CoverageBenchmark

DeepSeek V4.1 Flash integrates native multimodal visual understanding directly into its architecture, replacing the earlier V4 Flash text model paired with a separate Vision Exp path. The API family offers a 1M-token context with up to 384K maximum output. DeepSeek reports strong benchmark results including 90.9 on GPQA Diamond, a 3,471 Codeforces rating, 90.6 on Terminal-Bench 2.1, 74.2% on DeepSWE v1.1, 88.1 on CyberGym, and 65.4 on NL2Repo-Bench. Compared with the previous V4 Flash 0731 release, coding results improve sharply: Terminal-Bench 2.1 rises from 82.7 to 90.6, and DeepSWE from 54.4 to 74.2. From September 10, 2026, DeepSeek's Flash pricing is $0.003 per million cache-hit input tokens, $0.15 per million uncached input tokens, and $0.60 per million output tokens off-peak, with peak pricing at double those rates. V4 Pro requests are being routed to V4.1 Flash and billed at Flash pricing until V4.1 Pro arrives.

Eden AI

CoverageAnalysis

DeepSeek V4.1 Flash, released September 10, 2026, introduces a Causal Encoder-Decoder (CED) architecture that reduces KV cache requirements by approximately 4× relative to V4 Flash. The model uses a 552B-parameter backbone with only 8B active during prefill and 16B during decode, employing Compressed Sparse Attention 2 (CSA2) with Full, Reindex, and Reuse modes, plus FP4 KV cache compression in E2M1 format and SWA Bounded Replay for sliding window reconstruction. Additional innovations include Single-Pass mHC residual connections and Engram conditional memory at 196B parameters. The model achieves 552B backbone parameters with sparse MoE activation, outperforming V4 Pro (1.6T/49B active) across measured benchmarks despite having one-third the total parameters and one-sixth the active parameters. This marks the first time a Flash-tier model in the DeepSeek lineup entirely replaces a Pro-tier model, driven by architectural innovation rather than parameter scaling. Native multimodal vision is integrated via DeepSeek-ViT with jointly trained visual and text embeddings, and the model is MIT-licensed with weights available on HuggingFace.

Eden AI

CoverageBenchmark

DeepSeek V4.1 Flash is a 552B-parameter sparse mixture-of-experts model that activates roughly 8B parameters for input and 16B for output, shipping with MIT-licensed open weights. Peak pricing is $0.30/$1.20 per 1M tokens, dropping to $0.15/$0.60 off-peak, undercutting competitors like Gemini 3.8 Flash and GLM-5.3-Flash while sitting roughly 30× below GPT-6 Astra and Claude Fable 5.1. The model includes native image input as a core capability. The headline change versus the previous generation is efficiency: DeepSeek shrank the KV cache to roughly a quarter of V4 Flash's size, reducing memory footprint for agentic and high-volume coding workloads. V4.1 Flash phases out V4 Pro 0813 on cost, speed, and performance, positioning the Flash tier as the efficient high-throughput option rather than a research-grade heavyweight. The model targets coding, software agents, and long-context reasoning workloads.

Videos about DeepSeek V4.1 Flash (FlexAI)

More models around DeepSeek V4.1 Flash (FlexAI)