Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

DeepSeek V4.1 Flash (DeepSeek)

DeepSeek V4.1 Flash represents a mid-cycle architectural refresh of the Flash family, arriving roughly forty days after the prior base release and bringing what DeepSeek describes as a new model architecture designed for stronger capability, faster speed, and lower cost. Unlike the earlier V4 Flash Vision-Exp variant, which attached a visual encoder and Aligner externally to a text-only base, V4.1 Flash integrates text and image input natively, making it the first model in the lineup built from the ground up for multimodal use rather than retrofitting vision onto a language model.

In practical terms, V4.1 Flash is positioned as a lightweight, responsive option for developers who want multimodal understanding without paying for a flagship-tier model, and the beta artifact was circulated under the internal name deepseek-v4.1-flash-expires-on-0910 ahead of its wider rollout. DeepSeek has not yet published a formal technical report or benchmark sheet for the model, so its real-world fit is best understood through DeepSeek's own framing: a faster, cheaper, natively multimodal Flash-tier model aimed at general assistant, document understanding, and image-aware workflows where quick iteration matters more than top-of-class reasoning depth.

LLM Gatewaydeepseek/deepseek-v4.1-flashdeepseek-flash

Quick Info

Powered by
Provider
LLM Gateway
Model key
deepseek/deepseek-v4.1-flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.60

Limits

Output tokens
393,216 tokens
Context window
1,050,000 tokens

Transparent token rates

Compare DeepSeek V4.1 Flash (DeepSeek) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4.1 Flash (DeepSeek)

LLM Gateway

Coverage

DeepSeek V4.1 Flash is a 552B-parameter model whose benchmark performance rivals the much larger 1.6T-parameter V4 Pro, with Artificial Analysis assigning it an index score of 40 and DeepSeek itself announcing V4 Pro would be taken offline with requests routed accordingly. The core architectural innovation is dramatic Additional features include an upgraded Single-Pass mHC for training stability, Engram conditional memory module integration for static knowledge retrieval, and DSpark speculative decoding. The API exposes a "reasoning effort" parameter (b=25 to 100) with low/high/max settings to trade output verbosity for task perform

LLM Gateway

CoverageBenchmark

DeepSeek released DeepSeek V4.1 Flash on September 10, 2026 as a multimodal, open-weight model published under the MIT licence on Hugging Face and served via the deepseek-flash API endpoint. The model uses a 552-billion-parameter mixture-of-experts backbone with an asymmetric compute design, activating roughly 8B param The release positions V4.1 Flash as a practical choice for high-volume or input-heavy coding, tool-use, and long-context agent workflows, with DeepSeek cutting hosted API rates relative to the outgoing V4 Pro tariff. Thinking mode is enabled by default but can be disabled. According to the supplied page excerpt, the ma

LLM Gateway

CoverageBenchmark

DeepSeek Agent's guide details V4.1 Flash's architecture as a 552B-parameter MoE with 8B active parameters per input token and 16B during output generation, a 1M-token context for long sessions, and roughly one-quarter of V4 Flash's global KV-cache footprint at 890 bytes/token. The model introduces a Causal Encoder–Dec Official benchmark gains cited include Terminal-Bench 2.1 rising from 82.7 to 90.6, Terminal-Bench 3.0 from 7.6 to 30.0, Terminal-Bench 4.0 from 7.0 to 31.2, DeepSWE v1.1 from 54.4 to 74.2, and NL2Repo-Bench from 54.2 to 64.0, with evaluations run across Claude Code, Codex, OpenCode, Pi, mini-SWE, DSH Minimal, DSH Stan

LLM Gateway

CoverageBenchmark

VentureBeat reported on 10 September 2026 that DeepSeek launched DeepSeek-V4.1-Flash with a 552-billion-parameter MoE backbone, native vision, a 1-million-token context window, and an architecture designed to make repeatedly reading large contexts cheaper, shipping open weights under an MIT license on Hugging Face. Dur DeepSeek's peak pricing window runs Monday through Friday from 01:00–04:00 UTC and 06:00–10:00 UTC, making workload scheduling a cost-control lever for schedulable agent jobs that repeatedly reread repositories, tool definitions, system instructions, or conversation history. The article cites VentureBeat Pulse Research

LLM Gateway

CoverageBenchmark

DeepSeek released V4.1 Flash on 10 September 2026 as a 552-billion-parameter Mixture-of-Experts model that activates only 8B parameters per token during prefill, ships under an MIT license, and supports native image inputs alongside a 1-million-token context window. According to the official model card cited in the art DeepSeek has announced it will retire the roughly four-times-more-expensive V4 Pro on 14 September 2026 because V4.1 Flash "has comprehensively surpassed V4 Pro across all key metrics, including performance, cost, speed, and task completion time," signaling a strategic consolidation around the Flash tier. The Flowtivit

LLM Gateway

CoverageBenchmark

DeepSeek V4.1 Flash launched September 10, 2026 as the smallest member of DeepSeek's new architecture family, designed for higher capability, faster inference, higher throughput, and scaling to larger models. The release introduces native multimodal visual understanding into the Flash line rather than pairing a text mo From September 10, 2026, DeepSeek's published Flash pricing is $0.003 per million cache-hit input tokens, $0.15 per million uncached input tokens, and $0.60 per million output tokens off-peak, with peak pricing at double those rates. The article notes V4 Pro requests will be routed to V4.1 Flash and billed at Flash pri

LLM Gateway

Coverage

DeepSeek is rolling out V4.1-Flash as the smallest model in a new architecture family, adding native visual input and positioning it as an automatic replacement for customers on the higher-tier V4-Pro API. According to the article, beginning at 12:00 Beijing time on September 14, requests sent to the deepseek-v4-pro en DeepSeek's own testing reportedly puts V4.1-Flash ahead of V4-Pro on performance, cost, speed, and total runtime, though the article flags these as company-reported results that may not generalize across every customer workload. Both V4-Flash and V4-Pro were introduced with million-token context windows, framing V4.1-F

LLM Gateway

CoverageRelease Notes

DeepSeek released DeepSeek-V4.1-Flash on September 10, 2026 as an open-weight mixture-of-experts model with 552 billion total parameters, activating roughly 8 billion per token during prefill and 16 billion during decode, published under the MIT licence on Hugging Face. The model supports a context window of up to one From 04:00 UTC on September 14, 2026, DeepSeek began automatically routing all API requests naming deepseek-v4-pro to V4.1-Flash instead, billed at Flash's cheaper rate, until a successor V4.1-Pro model ships — effectively retiring V4-Pro as a distinct served model rather than issuing a conventional side-by-side upgrad

Videos about DeepSeek V4.1 Flash (DeepSeek)

More models around DeepSeek V4.1 Flash (DeepSeek)