Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Inco logo

Model details

DeepSeek V4.1 Flash

The model overview is temporarily unavailable.

Incodeepseek-v4.1-flash:fastdeepseek-flash

Quick Info

Powered by
Provider
Inco
Model key
deepseek-v4.1-flash:fast
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.60
Output token cost
$2.40

Limits

Output tokens
384,000 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare DeepSeek V4.1 Flash pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4.1 Flash

Inco

CoverageBenchmark

Coursiv's analysis frames V4.1 Flash as a generational replacement rather than a point update, comparing it against V4 Flash using DeepSeek's official model card. The backbone grew from 284B to 552B parameters, while active parameters per token dropped from 13B to 8B for input reading and 16B for output generation. The Notable benchmark jumps include Terminal-Bench 2.1 rising from 82.7 to 90.6, Terminal-Bench 4.0 from 7.0 to 31.2, and DeepSWE v1.1 from 54.4 to 74.2. Coursiv also reports that from September 14, 2026 at 04:00 UTC, every request to "deepseek-v4-pro" is routed to V4.1 Flash and billed at Flash rates, yielding approximate

Inco

CoverageBenchmark

BenchLM.ai tracks DeepSeek V4.1 Flash as an open-weight reasoning model released September 10, 2026, with a 1M token context window and text+image input modalities. The canonical API model ID is documented as "deepseek-flash", consistent with DeepSeek's official changelog. The tracker lists 22 sourced benchmark rows sp Independent third-party tracker pricing is reported at $0.30 per million input tokens and $1.20 per million output tokens, with cached input at $0.006 and a blended rate around $0.75, plus a measured throughput of approximately 212 tok/s and a first-token latency of about 10.62 seconds. These figures are sourced back t

Inco

Coverage

DeepSeek officially released DeepSeek-V4.1-Flash on September 10, 2026, positioning it as the smallest model in a new architecture family with native multimodal visual understanding. The model is designed for a higher capability ceiling, faster inference, higher throughput, and scaling to larger models. DeepSeek publis The release introduced notable API changes: the model is now callable via the name "deepseek-flash", while the previous-generation V4 Flash and V4 Flash Vision Exp have been retired, with "deepseek-v4-flash" and "deepseek-v4-flash-vision-exp" temporarily routed to V4.1 Flash for backward compatibility. DeepSeek also an

Videos about DeepSeek V4.1 Flash

More models around DeepSeek V4.1 Flash