Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Requesty logo

Model details

DeepSeek V4.1 Flash (EU)

DeepSeek V4.1 Flash is positioned by its creator as the smallest member of a new architecture family, built around a Causal Encoder–Decoder design rather than a conventional decoder-only stack. The total parameter count reaches 552 billion in a Mixture-of-Experts layout, but only 8 billion parameters activate on input and 16 billion on output, letting the model serve lighter requests cheaply while still routing harder prompts through a larger expert pool. DeepSeek pairs this sparse design with new pretraining methods and a larger-scale reinforcement-learning post-training stage, which the announcement credits for agentic benchmark results that sit ahead of its own flagship-tier V4 Pro. Native visual understanding is built into the model rather than bolted on, so image and text inputs can be handled in a single forward pass.

The practical payoff that DeepSeek highlights is inference efficiency rather than raw scale. Compared with the previous generation, V4.1 Flash needs roughly one quarter of the HBM and one eighth of the SSD storage for its KV cache, which matters disproportionately for long-running agent loops where cache-hit tokens dominate the bill. Third-party testing described in coverage of the launch reports that V4.1 Flash beats V4 Pro on performance, cost, speed, and end-to-end task latency, prompting DeepSeek to redirect V4-Pro traffic to V4.1 Flash from mid-September onward. A technical report is published on Hugging Face alongside the weights, making it well suited to teams that want an open-weight multimodal model with a small active-parameter footprint and low cache overhead for production agentic workloads.

Requestydeepseek-v4.1-flash@eudeepseek-flash

Quick Info

Powered by
Provider
Requesty
Model key
deepseek-v4.1-flash@eu
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.50
Output token cost
$1.50

Limits

Output tokens
393,216 tokens
Context window
1,048,576 tokens

Latest news about DeepSeek V4.1 Flash (EU)

No articles yet. Fetch the latest news to show it here.

Videos about DeepSeek V4.1 Flash (EU)

More models around DeepSeek V4.1 Flash (EU)