Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

DeepSeek V4.1 Flash (Runware)

DeepSeek V4.1 Flash is positioned by its creator as a general-purpose, high-throughput model in the DeepSeek family, succeeding the earlier V4 Flash line. Built for workflows where latency and cost matter, it has been reported to sustain processing speeds up to 427 tokens per second, making it well suited to time-sensitive tasks such as game development, 3D modeling, and complex coding pipelines. A notable shift from its predecessor is that vision is now native to the Flash tier rather than served through a separate vision model, simplifying requests that combine text with JPEG, PNG, GIF, or WebP imagery.

The V4.1-Flash weights are published by DeepSeek alongside other families in its catalog, reflecting a continued emphasis on open distribution for the Flash tier. In practice, the model fits teams that need long-context, high-volume routine work or multimodal input, while heavier V4 Pro remains the option for workloads where measured quality gains justify a higher rate. Reviewers note occasional overthinking tendencies in specific scenarios, suggesting prompt tuning or pairing with higher-tier reasoning when complex deliberation is required. The transition consolidated prior hosted identifiers under the new V4.1-Flash serving path, signaling a streamlined API surface for developers.

LLM Gatewayrunware/deepseek-v4.1-flashdeepseek-flash

Quick Info

Powered by
Provider
LLM Gateway
Model key
runware/deepseek-v4.1-flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.15
Output token cost
$0.60

Limits

Output tokens
384,000 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare DeepSeek V4.1 Flash (Runware) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4.1 Flash (Runware)

LLM Gateway

CoverageBenchmark

DeepSeek-V4.1-Flash launched on September 10, 2026, and from 04:00 UTC on September 14 every request to deepseek-v4-pro is routed to V4.1 Flash and billed at Flash rates until a V4.1 Pro ships. Anyone building on Pro gets the switch automatically, with cache-miss input cost dropping about 77 percent and output cost abo The model is a generational replacement rather than a point update: backbone parameters grew from 284B to 552B while active parameters fell from 13B to 8B (input) and 16B (output), the architecture moved from a Mixture-of-Experts decoder to a Causal Encoder-Decoder (20 encoder plus 20 decoder layers), vision is now nat

LLM Gateway

CoverageBenchmark

A third-party review published September 10, 2026 characterizes DeepSeek-V4.1-Flash as the newest Flash model that changes the role of the Flash tier, integrating native multimodal visual understanding directly into the new architecture rather than pairing a text model with a separate vision path. The model offers a 1M The review reports strong benchmark gains over the previous V4 Flash 0731 release, with Terminal-Bench 2.1 rising from 82.7 to 90.6 and DeepSWE from 54.4 to 74.2. From September 10, 2026, DeepSeek's Flash pricing schedule is $0.003 per million cache-hit input tokens, $0.15 per million uncached input tokens, and $0.60 p

LLM Gateway

CoverageAnalysis

A technical deep dive published September 11, 2026 analyzes DeepSeek V4.1 Flash's Causal Encoder-Decoder (CED) architecture, which reportedly reduces KV cache requirements by approximately 4x relative to V4 Flash and 437x relative to DeepSeek V1. The model features 552B backbone parameters with only 8B active during pr The analysis notes that V4.1 Flash outperforms V4 Pro (1.6T/49B active) across all measured benchmarks despite having roughly one-third the total parameters and one-sixth the active parameters, marking the first instance where a Flash-tier model entirely replaces a Pro-tier model in the DeepSeek lineup. Native multimod

LLM Gateway

Coverage

DeepSeek officially released DeepSeek-V4.1-Flash on September 10, 2026, as the smallest model in its new architecture family with native multimodal visual understanding. According to the official changelog, the model is designed for higher capability ceiling, faster inference, higher throughput, and scaling to larger m The release published a comprehensive benchmark suite including GPQA Diamond 90.9, Codeforces rating 3471, Terminal-Bench 2.1 90.6, DeepSWE v1.1 74.2, CyberGym 88.1, and NL2Repo-Bench 65.4, positioning V4.1 Flash strongly across coding, agentic, cybersecurity, and reasoning evaluations. API prices were reduced with the

LLM Gateway

Coverage

Chinese AI startup DeepSeek launched DeepSeek-V4.1-Flash on September 10, 2026, calling it the smallest model in its new architecture family. The company said the new model is designed for greater capability, faster inference, higher throughput and scaling to larger models, in a statement reported by Reuters from Beiji The release comes as DeepSeek is reported to be preparing for an initial public offering on Shanghai's tech-focused STAR Market. The Reuters dispatch is brief on technical detail and does not document architecture, benchmark numbers, or hosting variants, serving primarily as a launch confirmation for the V4.1 Flash fam

LLM Gateway

CoverageBenchmark

DeepSeek-V4.1-Flash is ranked 13th on the LLM Stats Score composite, with a blended price around $0.24 per million tokens versus higher-ranked Claude Opus 5.5 at $4.76 and Gemma 4 E4B at $0.024. It scores 51.4 on the composite, slotting between DeepSeek-V4-Flash-0731 at 43.8 and Claude Opus 5.5 at 59.7 on the cost-effi On individual benchmarks sourced from the model's scorecard, paper, or official blog posts, DeepSeek-V4.1-Flash ranks 1st on CodeForces with a score of 3471 (maximum reasoning effort, temperature 1.0, top-p 0.95) and ranks 23rd on GPQA Diamond with a score of 0.91 at Pass@1. On Terminal-Bench 2.1 it ranks 1st with a Pa

LLM Gateway

CoverageLeaks

DeepSeek released V4.1 Flash for real on September 10, 2026, across its app, web interface and API, where it answers to the model name deepseek-flash. The model spent September 8 and 9 inside DeepSeek's live API under the ID deepseek-v4.1-flash-expires-on-0910 before the public release. DeepSeek also published a techni The release is positioned as larger than the beta implied, with the new Causal Encoder-Decoder architecture and the publicly released MIT weights (rather than a private beta) marking the key change in status. Native multimodality and the new conditional-memory mechanism are part of the model design, and the public weig

Videos about DeepSeek V4.1 Flash (Runware)

More models around DeepSeek V4.1 Flash (Runware)