Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Eden AI logo

Model details

DeepSeek V4.1 Flash (Nebius)

DeepSeek V4.1 Flash marks the debut of the company's Causal Encoder-Decoder (CED) architecture, a sparse mixture-of-experts design with a 552B-parameter backbone that activates only 8B parameters on input and 16B on output. This asymmetric split is intended to keep per-token compute low relative to the model's total size, and the architecture bakes in image understanding from the start: visual and text embeddings are trained jointly during pre-training rather than bolted on afterward, giving the model native multimodal capability in a single stack.

DeepSeek describes V4.1 Flash as the cost-efficient tier of the V4.1 family, tuned for coding, terminal, and computer-use agents as well as long-horizon tasks that have to run to completion across many steps. A new pretraining recipe combined with larger-scale reinforcement-learning post-training is credited with pushing benchmark results ahead of the flagship V4-Pro, and a compressed KV cache cuts HBM usage to about a quarter of the previous Flash generation while trimming SSD storage needs to roughly an eighth, materially lowering the cache-hit costs that dominate agentic workloads.

Eden AInebius/deepseek-ai/DeepSeek-V4.1-Flashdeepseek-flash

Quick Info

Powered by
Provider
Eden AI
Model key
nebius/deepseek-ai/DeepSeek-V4.1-Flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$1.20

Limits

Output tokens
384,000 tokens
Context window
1,048,000 tokens

Transparent token rates

Compare DeepSeek V4.1 Flash (Nebius) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4.1 Flash (Nebius)

Eden AI

CoverageBenchmark

Coursiv's comparison confirms that DeepSeek V4.1 Flash is a generational replacement for V4 Flash rather than a point update: the backbone nearly doubles to 552B parameters while active parameters per token drop from 13B to 8B on input and 16B on output, the architecture shifts from a Mixture-of-Experts decoder to a Ca The article documents DeepSeek's September 10, 2026 release and the resulting pricing change: from 04:00 UTC on September 14, 2026, all requests to the deepseek-v4-pro endpoint are routed to V4.1 Flash and billed at Flash rates until a V4.1 Pro ships, yielding roughly a 77 percent drop in cache-miss input cost and abou

Eden AI

Coverage

The Next Web reports that DeepSeek released V4.1-Flash on September 10, 2026, publishing the weights on Hugging Face under the MIT license and announcing in an X thread that the model is positioned as the smallest member of a new causal encoder-decoder architecture family. It uses 552B total parameters with 8B active p The article documents the simultaneous retirement of V4-Pro: from 04:00 UTC on September 14, 2026, all deepseek-v4-pro requests are routed to V4.1-Flash and billed at Flash rates until a V4.1-Pro is released. DeepSeek's published benchmark table shows V4.1-Flash scoring 74.2 on DeepSWE v1.1 (ahead of Claude Opus 5's 74

Eden AI

CoverageBenchmark

BuildFastWithAI's review confirms DeepSeek V4.1 Flash is the newest Flash-series model with native multimodal visual understanding integrated into the architecture, a 1M-token context window, and up to 384K maximum output tokens, targeting coding, software agents, long-context reasoning, and high-throughput API workloa The review documents DeepSeek's official Flash pricing effective September 10, 2026: $0.003 per million cache-hit input tokens, $0.15 per million uncached input tokens, and $0.60 per million output tokens off-peak, with peak pricing at double those rates. DeepSeek also confirms that V4-Pro traffic will be routed to V4.

Eden AI

CoverageAnalysis

A research-grade technical analysis published on September 10, 2026, documents that DeepSeek V4.1 Flash introduces a Causal Encoder-Decoder (CED) architecture, replacing the pure decoder MoE design of V4 Flash with 20 encoder plus 20 decoder layers. The model uses a 552B-parameter backbone with asymmetric activation of The deep dive details that V4.1 Flash cuts global KV cache requirements to roughly 890 bytes per token—about a quarter of V4 Flash's footprint—via Compressed Sparse Attention 2 modes and FP4 (E2M1) cache compression, while supporting 1M-token context and native multimodal vision via jointly trained DeepSeek-ViT embeddi

Eden AI

CoverageAnalysis

DeepSeek V4.1-Flash, released September 10, 2026 with MIT-licensed weights, is engineered around the observation that long-horizon agents are input-heavy workloads where prefill compute and KV-cache memory—not raw intelligence—are the primary cost bottleneck. The model splits its 40 Transformer layers into a 20-layer c Key verified specifications include a 552B MoE backbone across 40 layers, 890 bytes per token KV cache (one-quarter of V4-Flash), $0.003 per million cached input tokens at off-peak pricing, and a 1M token context via YaRN ×16. The architecture combines Compressed Sparse Attention 2 (CSA2) with three static attention mo

Eden AI

Coverage

DeepSeek released DeepSeek-V4.1 Flash on September 10, 2026, describing it as the smallest model in a new architecture family that now natively supports visual understanding. The model uses a 552 billion parameter Mixture of Experts design with a new Causal Encoder-Decoder architecture, activating just 8 billion parame The article includes a comprehensive benchmark comparison table showing V4.1-Flash's performance against V4-Pro 0813, V4-Flash 0731, GLM 5.3, Kimi K3, GPT 5.6-Sol, and Claude Opus 5 across evaluations including GPQA Diamond (90.9), Terminal-Bench 2.1 (90.6), Terminal-Bench 4.0 (31.2), DeepSWE v1.1 (74.2), Codeforces ra

Videos about DeepSeek V4.1 Flash (Nebius)

More models around DeepSeek V4.1 Flash (Nebius)