Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Melious logo

Model details

DeepSeek V4.1 Flash

DeepSeek V4.1 Flash introduces a new Causal Encoder–Decoder architecture that powers the smallest variant in DeepSeek's refreshed model family. The model activates only 8B parameters for input and 16B for output within a 552B-parameter MoE design, an asymmetric setup that aims to deliver flagship-level capability while keeping inference cost lower. DeepSeek pairs this architecture with new pre-training methods and a larger-scale reinforcement learning post-training stage, reporting benchmark results that exceed its own V4-Pro flagship across reasoning, coding, and agent tasks.

In practice, V4.1 Flash targets developers who need strong reasoning, tool use, and visual understanding in a single open-weights deployment. Reported scores include 90.9 on GPQA Diamond, 36.8 on HLE text-only, 3471 on Codeforces, and 90.6 on Terminal-Bench 2.1, alongside agent-oriented results such as 74.2 on DeepSWE v1.1, 88.1 on CyberGym, 62.8 on SEC-Bench Pro, and 78.9 on Chartography with tools. Native multimodal support and a redesigned KV cache that needs roughly one-quarter the HBM and one-eighth the SSD storage of the previous generation make it well suited to long-context, cache-heavy agent workloads where serving efficiency matters as much as raw accuracy.

Meliousdeepseek-v4.1-flashdeepseek-flash

Quick Info

Powered by
Provider
Melious
Model key
deepseek-v4.1-flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.23184
Output token cost
$1.1592

Limits

Output tokens
384,000 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare DeepSeek V4.1 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4.1 Flash

Melious

Coverage

mem0's 2026-09-15 technical analysis details V4.1 Flash as a 552B MoE trained on 45 trillion multimodal tokens with a Causal Encoder-Decoder architecture, 8B active parameters for input and 16B for output, and a 1M-token context window. The post highlights a KV cache compressed to 890 bytes per token, a claimed 437x re The same analysis cites off-peak input pricing of $0.15 per million tokens ($0.003 cached) and reports V4.1 Flash beating Claude Opus 5 and GPT-5.6 Sol on Terminal-Bench 2.1 (90.6 vs 89.1 and 88.8) and DeepSWE v1.1 (74.2 vs 74.0 and 73.0). The piece frames long context as working memory and positions external memory la

Melious

CoverageBenchmark

Flowtivity's 2026-09-10 breakdown describes DeepSeek V4.1 Flash as a 552B-parameter Mixture-of-Experts model that activates 8B parameters per token during prefill and 16B during decoding, with a 1M-token context window and native image input. According to the article, V4.1 Flash beats OpenAI's GPT-5.6 Sol and Anthropic The same post reports peak input pricing of $0.30 per million tokens, an MIT license release, and DeepSeek's stated intent to retire the roughly four-times-more-expensive V4 Pro on 2026-09-14 because V4.1 Flash "comprehensively surpassed V4 Pro across all key metrics." Architecture and benchmark figures are sourced fro

Melious

CoverageBenchmark

The BenchLM aggregator page tracks DeepSeek V4.1 Flash with a release date of September 10, 2026, open-weight licensing, reasoning capability, and a 1M token context window, listing the API model id as 'deepseek-flash'. It surfaces 22 published benchmark rows split across categories with verified coverage in Agentic (8 The page also restates pricing at $0.30 input and $1.20 output per million tokens, with a $0.006 cached input rate and a $0.75 blended rate, against field medians of $1 input and $0.06 cached, and notes that independent runtime speed has not been measured and time-to-first-token is not measured. Maximum output length a

Melious

Coverage

A Requesty analysis published 2026-09-10 corroborates DeepSeek's release of DeepSeek-V4.1-Flash and adds technical detail sourced from the official release note and Hugging Face weights (MIT licensed): a 552B-parameter mixture-of-experts model with an asymmetric causal encoder-decoder design that activates 8B parameter Requesty also highlights a community parameter-count discussion from r/LocalLLaMA, where a safetensors reading showed roughly 552B in the main model plus about 197B of optional engram parameters and 14B of speculative decoding weights, for a total nearer 748B on disk, a consideration relevant to self-hosters. On the li

Melious

Coverage

DeepSeek officially released DeepSeek-V4.1-Flash on 2026-09-10 as the smallest model in a new architecture family with native multimodal visual understanding, per the DeepSeek API change log. The release entry lists a 19-benchmark score block including GPQA Diamond 90.9, HLE 36.8, Codeforces rating 3471, Terminal-Bench The same change log notes that API prices were reduced in step with the V4.1 Flash release and that DeepSeek will continue serving V4 Pro after September 14, 2026 with billing unchanged. The model ships with native multimodal support on the DeepSeek API, positioning V4.1 Flash as a higher-throughput, higher-capability-

Videos about DeepSeek V4.1 Flash

More models around DeepSeek V4.1 Flash