Sulat.com
AI models
Ollama Cloud logo

Model details

DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is positioned as a generational successor to the earlier V4 Flash line, including the experimental V4 Flash Vision variant, with prior model names temporarily rerouted to the new release so existing integrations keep running. The release carries a different focus than a typical incremental update: DeepSeek's own published numbers describe V4.1 Flash as surpassing the previous V4 Pro flagship on performance, cost, speed, and overall throughput, with the strongest evidence appearing in agent-oriented workloads while knowledge-recall benchmarks present a more mixed picture. Community discussion in the NVIDIA DGX Spark / GB10 user forum already treated the model as imminent in early September 2026, with an experimental API identifier suggesting pre-release availability tied to the Flash family.

The practical effect of the release is a notable shift in the inference cost curve for applications previously built on V4 Pro. Beginning September 14, 2026, traffic to the V4 Pro endpoint is rerouted to V4.1 Flash and billed at Flash rates, yielding roughly a 77 percent drop in cache-miss input pricing and around a 70 percent drop in output pricing until a dedicated V4.1 Pro is introduced. Combined with the open-weights posture and broad capability surface, V4.1 Flash is shaped to fit production agent pipelines, tool-using assistants, and other throughput-sensitive deployments where the agent-work gains outweigh the softer showing on knowledge-recall tasks.

Ollama Clouddeepseek-v4.1-flashdeepseek-flash

Quick Info

Powered by
Provider
Ollama Cloud
Model key
deepseek-v4.1-flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Limits

Output tokens
384,000 tokens
Context window
1,048,576 tokens

Latest news about DeepSeek V4.1 Flash

Ollama Cloud

Official sourceAnnouncement

The Hacker News thread (417 points, 217 comments) reproduces DeepSeek's official announcement text: V4.1 Flash is targeted for official release around 2026-09-10 Beijing Time, has "comprehensively surpassed" V4 Pro across performance, cost, speed, and task completion time, and will see all `deepseek-v4-pro` requests ro Beyond the release quotes, the thread collects independent user reports of pathologies in the prior V4 Flash web build (language-following drift, thinking chains in the wrong language) and coding-harness issues such as infinite loops, hallucinations, premature termination, and invalid tool calls. These are anecdotal co

Ollama Cloud

CoverageBenchmark

eesel AI's piece confirms the 2026-09-10 release and quotes DeepSeek's framing that V4.1 Flash is "the smallest model in our new architecture family, with native visual understanding," then lays out the consolidation: V4 Flash, V4 Flash Vision Exp, and eventually V4 Pro all collapse onto V4.1 Flash. It highlights the 2 The post is otherwise a vendor-flavored summary aimed at customer-support automation buyers, with thinner first-party technical detail than DeepSeek's change log. It treats architecture, benchmarks, and pricing at a higher level and adds eesel's product framing rather than independent measurements of V4.1 Flash itself.

Ollama Cloud

CoverageBenchmark

BenchLM's tracker page lists DeepSeek V4.1 Flash as released 2026-09-10 with 1M context, open weights, reasoning support, and 22 sourced benchmark rows covering agentic (8/8 verified), coding (6/6), knowledge (3/3), math (1/1), and multimodal (3/3) categories, with reasoning, multilingual, and instruction-following sti The page also surfaces coverage gaps and decision caveats: no public comparative rank yet because eligible evidence is thin, independent runtime speed not measured, time-to-first-token not measured, maximum output length tracked separately, and knowledge cutoff not sourced. Maximum output length is explicitly flagged a

Ollama Cloud

CoverageBenchmark

Coursiv's comparison table attributes to DeepSeek's official model card a generational shift from V4 Flash to V4.1 Flash: backbone parameters roughly double from 284B to 552B, active parameters per token move from 13B to 8B on input and 16B on output, the architecture changes from a MoE decoder to a causal encoder-deco The post also reports agentic benchmark jumps DeepSeek attributes to the model card (Terminal-Bench 2.1 82.7 → 90.6, Terminal-Bench 4.0 7.0 → 31.2, DeepSWE v1.1 54.4 → 74.2), a KV cache roughly four times larger per token on V4 Flash than on V4.1 Flash (890 bytes), and an effective price drop of about 77% on cache-miss

Ollama Cloud

Coverage

Requesty's analysis treats the V4.1 Flash launch as both a model release and a lifecycle event, flagging that from 2026-09-14 the pinned `deepseek-v4-pro` model ID will silently serve a different model. It reproduces architectural claims from DeepSeek: a 552B-parameter MoE backbone, an asymmetric causal encoder-decoder The piece adds community context: launch-tweet engagement (~21,110 likes, ~2.78M impressions), a r/LocalLLaMA reading of the safetensors putting on-disk weight nearer ~748B (552B main plus ~197B engram and ~14B speculative decoding weights), and pricing reference points of $0.30 input / $1.20 output per million tokens.

Ollama Cloud

Coverage

DeepSeek's own API change log dated 2026-09-10 announces DeepSeek-V4.1-Flash as the smallest model in a new architecture family with native multimodal visual understanding, and documents a full benchmark block including GPQA Diamond 90.9, HLE 36.8 (39.1 on the pure-text subset), Codeforces 3471, Terminal-Bench 2.1 90.6 The same entry formalizes the V4 Pro lifecycle change: after 12:00 Beijing Time on 2026-09-14, all requests to `deepseek-v4-pro` will be routed to V4.1 Flash and billed at Flash pricing until a future V4.1 Pro ships, on DeepSeek's claim that V4.1 Flash has surpassed V4 Pro across performance, cost, speed, and total tim

Videos about DeepSeek V4.1 Flash

More models around DeepSeek V4.1 Flash