Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

DeepSeek V4.1 Flash (luminal)

The model overview is being prepared.

LLM Gatewayluminal/deepseek-v4.1-flashdeepseek-flash

Quick Info

Powered by
Provider
LLM Gateway
Model key
luminal/deepseek-v4.1-flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.135
Output token cost
$0.54

Limits

Output tokens
384,000 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare DeepSeek V4.1 Flash (luminal) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4.1 Flash (luminal)

LLM Gateway

CoverageBenchmark

DeepSeek V4.1 Flash, released September 10, 2026, is the newest Flash model designed for higher capability ceiling, faster inference, and native multimodal visual understanding integrated into a single architecture. The API offers a 1M-token context with up to 384K maximum output, and DeepSeek reports strong benchmark results including 90.9 on GPQA Diamond, 3,471 Codeforces rating, 90.6 on Terminal-Bench 2.1, 74.2% on DeepSWE v1.1, 88.1 on CyberGym, and 65.4 on NL2Repo-Bench. Compared to V4 Flash 0731, coding results improve sharply: Terminal-Bench 2.1 rises from 82.7 to 90.6 and DeepSWE from 54.4 to 74.2. Flash pricing from September 10, 2026 is $0.003 per million cache-hit input tokens, $0.15 per million uncached input, and $0.60 per million output off-peak, with peak rates double. DeepSeek also routes V4 Pro requests to V4.1 Flash at Flash pricing until V4.1 Pro ships.

LLM Gateway

CoverageAnalysis

DeepSeek V4.1 Flash represents a fundamental architectural shift via a Causal Encoder-Decoder design that cuts KV cache requirements approximately 4× relative to V4 Flash and 437× relative to DeepSeek V1. The model uses 552B backbone parameters with only 8B active during prefill and 16B during decode, plus an additional 196B-parameter Engram conditional memory, delivering major cost efficiency for agentic workloads. Key innovations include Compressed Sparse Attention 2 with three modes, FP4 KV cache compression in E2M1 format, and DSpark speculative decoding. Released September 10, 2026 with MIT-licensed weights, V4.1 Flash outperforms V4 Pro (1.6T/49B active) across measured benchmarks despite having one-third the total parameters and one-sixth the active parameters, marking the first time a Flash-tier model entirely replaces a Pro-tier model. Native multimodal vision is integrated via DeepSeek-ViT, and the analysis covers the full pre-training and post-training pipeline verified against official HuggingFace model card and DeepSeek documentation.

LLM Gateway

CoverageBenchmark

DeepSeek-V4.1-Flash ranks 15th on the LLM Stats composite score, earning top-2% standing in Tool Calling (4 of 202) and top-10% in Coding (7 of 275), with average-tier results in Reasoning, Vision, and Math. The model achieves a CodeForces rating of 3471, GPQA Diamond score of 0.91, and Terminal-Bench 2.1 score of 0.91 at maximum reasoning effort. Its blended cost efficiency places it at $0.24 per million tokens with an LLM Stats Score of 51.2. Cost-efficiency positioning shows DeepSeek-V4.1-Flash at $0.24 with a score of 51.2, positioned between GPT OSS 120B at $0.043/28.7 and Claude Opus 5.5 at $4.76/60.3. The model significantly outperforms its predecessor DeepSeek-V4-Flash-0731, which sits at $0.066 with a score of 43.8, demonstrating the architecture's value improvement. Benchmark methodology uses maximum reasoning effort with temperature 1.0 and top-p 0.95.

LLM Gateway

Coverage

DeepSeek V4.1 Flash officially launched on September 10, 2026, as the smallest model in DeepSeek's new architecture family, featuring native multimodal visual understanding. It is a 552B-parameter mixture-of-experts model with a Causal-Encoder-Decoder architecture using asymmetric activation of 8B input parameters and 16B output parameters, reducing costs relative to same-size peers. The model ships on the DeepSeek API under the name deepseek-flash with 1M-token context, and DeepSeek has open-sourced weights on Hugging Face alongside a technical report. Performance-wise, V4.1 Flash surpasses DeepSeek V4 Pro and rivals flagship models like GLM 5.3 and Kimi-K3 on benchmarks such as the Agentic Benchmark, while cutting HBM requirements to one-fourth and SSD needs to one-eighth versus V4 Flash. The release also introduced Peak-Valley Pricing effective 12:00 Beijing time on September 10, with off-peak rates at half the peak price. After September 14, 2026, all deepseek-v4-pro requests route to V4.1 Flash at Flash pricing until V4.1 Pro arrives.

Videos about DeepSeek V4.1 Flash (luminal)

More models around DeepSeek V4.1 Flash (luminal)