Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

DeepSeek V4.1 Flash (SCX.ai)

DeepSeek V4.1 Flash is positioned as the smallest member of a new architecture family, built to combine stronger reasoning with lower inference cost. It uses a 552-billion-parameter mixture-of-experts design paired with a new causal encoder-decoder structure, activating only 8 billion parameters on input and 16 billion on output. This asymmetric setup is the central efficiency idea: most tokens pass through the lighter input path, while the heavier output path engages only where generation demands more capacity, letting the model scale toward larger siblings without paying the full parameter cost on every request.

The release pairs that architecture with new pretraining methods and broader reinforcement-learning post-training, and DeepSeek reports benchmark results that land ahead of its own flagship V4-Pro model. Operationally, V4.1 Flash also shrinks the KV cache to roughly a quarter of the previous generation's HBM usage and one-eighth of its SSD footprint, a meaningful reduction for agent workloads where cached context dominates spend. The model ships with native multimodal support and is intended for fast, high-throughput deployments that still need strong agentic and general reasoning quality, fitting builders who want frontier-adjacent capability without flagship-class serving cost.

LLM Gatewayscx-ai-gp/deepseek-v4.1-flashdeepseek-flash

Quick Info

Powered by
Provider
LLM Gateway
Model key
scx-ai-gp/deepseek-v4.1-flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.60

Limits

Output tokens
384,000 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare DeepSeek V4.1 Flash (SCX.ai) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4.1 Flash (SCX.ai)

LLM Gateway

CoverageBenchmark

A reference write-up describes DeepSeek V4.1 Flash as a non-routine refresh of V4 Flash, with changes to architecture, expanded native multimodal support, reduced long-context inference storage, and a new API price schedule. Headline specifications include a 552-billion-parameter Mixture-of-Experts backbone, 8B active The source notes DeepSeek is positioning V4.1 Flash as the temporary destination for V4 Pro traffic while preparing a future V4.1 Pro, and that legacy V4 Flash aliases should be checked for routing and billing behavior before comparing old and new runs. Documented features include thinking and non-thinking modes, JSON

LLM Gateway

CoverageAnalysis

Released September 10, 2026, DeepSeek V4.1 Flash introduces a Causal Encoder-Decoder (CED) architecture that reduces KV cache requirements by roughly 4× relative to V4 Flash and 437× relative to DeepSeek V1, according to a research-grade technical analysis published September 11, 2026. The model carries 552B backbone p The analysis reports that V4.1 Flash outperforms V4 Pro (1.6T parameters, 49B active) across all measured benchmarks despite having one-third the total parameters and one-sixth the active parameters, marking what the source calls the first time a Flash-tier model entirely replaces a Pro-tier model in the DeepSeek lineu

LLM Gateway

CoverageBenchmark

An aggregator page tracks DeepSeek-V4.1-Flash benchmarks and pricing, listing a blended price of $0.24 per million tokens with an LLM Stats Score of 51.8 and an overall rank of 13 on its composite leaderboard. The page benchmarks the model at rank 1 on CodeForces with a score of 3471.00/3000 using maximum reasoning eff The cost-efficiency view contrasts DeepSeek-V4.1-Flash against siblings and competitors, placing it above DeepSeek-V4-Flash-0731 ($0.066/M, score 44.7) and below GPT-6 Astra ($11.9/M, score 59.6), while Gemma 4 E4B and GPT OSS 120B appear at lower price points. Performance-by-conversation-depth data is presented for ho

Videos about DeepSeek V4.1 Flash (SCX.ai)

More models around DeepSeek V4.1 Flash (SCX.ai)