Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
SCNet Token Plan logo

Model details

DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is positioned as the smallest member of a new architecture family designed for higher capability ceilings, faster inference, greater throughput, and clean scaling to larger models. It ships as an open-weight release under an MIT license, with a Mixture-of-Experts design totaling 763.2 billion parameters trained on 45.0 trillion tokens, giving it a 59x tokens-to-parameters ratio. Native multimodal visual understanding is built directly into the architecture, so the same checkpoint can accept both text and image inputs and is intended to be a unified foundation for reasoning, coding, and agent workflows rather than a vision add-on.

In practice, the model leans into agentic and coding tasks. It ranks first on Codeforces-style competitive programming with a rating of 3471 and first on Terminal-Bench 2.1 at 0.91 in the cataloged API limit context, while also leading the BabyVision early visual reasoning benchmark at 0.90. Reported reasoning and knowledge scores include GPQA Diamond at 0.91, MathArena Apex at 65.6, and HLE at 36.8, rising to 63.9 when tools are enabled on the pure-text subset. Strong agent results such as DeepSWE v1.1 at 74.2, NL2Repo-Bench at 65.4, CyberGym at 88.1, and SEC-Bench Pro at 62.8, together with top-tier vision performance, make DeepSeek V4.1 Flash a practical fit for long-horizon coding agents, tool-using assistants, and multimodal pipelines that benefit from a self-hostable open-weight checkpoint.

SCNet Token PlanDeepSeek-V4.1-Flashdeepseek-flash

Quick Info

Powered by
Provider
SCNet Token Plan
Model key
DeepSeek-V4.1-Flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
384,000 tokens
Context window
1,000,000 tokens

Latest news about DeepSeek V4.1 Flash

AMD

CoverageBenchmark

BenchLM's aggregator profile lists DeepSeek V4.1 Flash pricing at $0.30 per million input tokens and $1.20 per million output tokens, with a $0.006 cached input rate and a blended rate of $0.75. Reported speed is 207 tokens per second with a first-token latency of 10.83 seconds, and the context window is listed at 1,00 Coverage tracks 22 published benchmark rows across Agentic (8/8 verified), Coding (6/6 verified), Knowledge (3/3 verified, category score 57.8), Math (1/1 verified), and Multimodal (3/3 verified), while Reasoning, Multilingual, and Instruction Following are listed as not measured. The model currently sits without an el

Melious

Coverage

mem0's 2026-09-15 technical analysis details V4.1 Flash as a 552B MoE trained on 45 trillion multimodal tokens with a Causal Encoder-Decoder architecture, 8B active parameters for input and 16B for output, and a 1M-token context window. The post highlights a KV cache compressed to 890 bytes per token, a claimed 437x re The same analysis cites off-peak input pricing of $0.15 per million tokens ($0.003 cached) and reports V4.1 Flash beating Claude Opus 5 and GPT-5.6 Sol on Terminal-Bench 2.1 (90.6 vs 89.1 and 88.8) and DeepSWE v1.1 (74.2 vs 74.0 and 73.0). The piece frames long context as working memory and positions external memory la

EmpirioLabs AI

Coverage

Modelspectra's research brief frames DeepSeek-V4.1-Flash as a serving-first architecture with three mechanisms: a causal encoder-decoder splitting 40 layers into 20 encoder and 20 decoder layers to roughly halve prefill FLOPs, CSA2 attention sharing KV across layers, and an Engram conditional memory of 196B parameters. The brief cites DeepSWE v1.1 at 74.2 versus Claude Opus 5 at 74.0 and GPT-5.6 Sol at 73.0, with a 1M-token context under open weights. It positions the model as evidence that cost-per-useful-output now rivals peak-score as the relevant frontier axis, while flagging that hosted list price and serving specs were not fina

OpenCode Go

CoverageBenchmark

DeepSeek shipped V4.1 Flash on 10 September 2026, positioning it as "the smallest model in our new architecture family, with native visual understanding." According to the eesel AI write-up, the release is part of a broader cleanup in which three prior DeepSeek endpoints are being collapsed into one: the old deepseek-v The eesel blog frames the technical story as a redesign rather than an incremental tweak: a new architecture with native multimodal visual input aimed at a higher capability ceiling, faster inference, higher throughput, and easier scaling to larger models, with DeepSeek claiming V4.1 Flash now beats the V4 Pro flagship

GreenPT

CoverageBenchmark

VentureBeat's launch coverage highlights V4.1 Flash's pricing structure aimed at agent economics. During off-peak hours DeepSeek prices the model at $0.003 per million input tokens on a cache hit, $0.15 per million on a cache miss, and $0.60 per million output tokens, with peak rates double those figures. DeepSeek's pe The model ships with a 552B-parameter mixture-of-experts backbone, native vision, a 1M-token context window, and MIT-licensed open weights on Hugging Face for commercial use. VentureBeat frames the release as particularly relevant to enterprises running coding agents and other long-running workflows against repositorie

Vivgrid

CoverageBenchmark

Coursiv's analysis characterizes V4.1 Flash as a generational replacement rather than a point update, comparing it against the retired V4 Flash and V4 Flash Vision. The comparison table shows backbone parameters growing from 284B to 552B, active parameters dropping from 13B to 8B for input reading and 16B for output ge The article highlights key benchmark gains: Terminal-Bench 2.1 improved from 82.7 to 90.6, Terminal-Bench 4.0 jumped from 7.0 to 31.2, and DeepSWE v1.1 rose from 54.4 to 74.2, all while maintaining a 1M token context window and MIT license. From 12:00 Beijing time on September 14, 2026, requests to deepseek-v4-pro were

Videos about DeepSeek V4.1 Flash

More models around DeepSeek V4.1 Flash