Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Eden AI logo

Model details

DeepSeek V4.1 Flash (Deep Infra)

DeepSeek V4.1 Flash is positioned as a compact, efficiency-focused member of a new DeepSeek architecture family, and is described as the first Flash-tier variant to ship with native vision and multimodal understanding. Its design uses a Mixture of Experts setup with 552B total parameters, activating only 8B parameters for input and 16B for output, built on what the source calls a Causal Encoder Decoder architecture. Open weights are published on Hugging Face, making the model usable outside of hosted APIs and attractive for teams that want to self-host or fine-tune.

Practical strengths emphasized for the model include frontier-leaning intelligence at modest active parameter counts, a long context window of 1M tokens, and a maximum output of the cataloged API limit tokens, with both thinking and non-thinking modes plus tool calling and structured JSON output. The source claims it surpasses DeepSeek V4 Pro on performance, cost, speed, and task completion time, and reports a score of 88.1 on CyberGym, while a heavily compressed KV cache of roughly 890 bytes per token is highlighted as a way to lower cache hit costs in agent workloads. It is served in FP8 on Deep Infra with prompt caching, making it well suited to coding agents, long-context retrieval, high-throughput batch processing, and multimodal assistants.

Eden AIdeepinfra/deepseek-ai/DeepSeek-V4.1-Flashdeepseek-flash

Quick Info

Powered by
Provider
Eden AI
Model key
deepinfra/deepseek-ai/DeepSeek-V4.1-Flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$0.60

Limits

Output tokens
384,000 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare DeepSeek V4.1 Flash (Deep Infra) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4.1 Flash (Deep Infra)

Eden AI

CoverageAnalysis

A long-form technical analysis of DeepSeek V4.1 Flash describes a Causal Encoder-Decoder (CED) design with 552B backbone parameters, 8B active during prefill and 16B during decode, plus a 196B Engram conditional memory. The piece claims a roughly 4× KV-cache reduction versus V4 Flash and highlights components such as C The same analysis states that DeepSeek V4.1 Flash outperforms V4 Pro (1.6T with 49B active) across all measured benchmarks despite using about one-third the total parameters and one-sixth the active parameters, calling it the first time a Flash-tier model fully replaces a Pro-tier model in the DeepSeek lineup. All arch

Eden AI

Coverage

DeepSeek officially released DeepSeek-V4.1-Flash on 2026-09-10 as the smallest model in a new architecture family, adding native multimodal visual understanding. The changelog notes the design targets a higher capability ceiling, faster inference, higher throughput, and scaling to larger models, and lists first-party b The DeepSeek API now exposes the model as `deepseek-flash` with native multimodal support, while legacy names `deepseek-v4-flash` and `deepseek-v4-flash-vision-exp` are temporarily routed to V4.1 Flash. The V4 Flash and V4 Flash Vision Exp models have been retired, V4 Pro API service is extended past September 14, 2026

Videos about DeepSeek V4.1 Flash (Deep Infra)

More models around DeepSeek V4.1 Flash (Deep Infra)