Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

DeepSeek V4.1 Flash (NovitaAI)

DeepSeek V4.1 Flash is positioned as the smallest member of a new DeepSeek architecture family that introduces native visual understanding, making it a multimodal reasoning model suited for tasks that combine text and image inputs. DeepSeek designed the model for greater capability, faster inference, and higher throughput than prior generations, explicitly aiming to scale the same approach to larger models. The model is described as a particularly cost-conscious option, owing to a substantially reduced KV cache footprint that uses roughly a quarter of the high-bandwidth memory and an eighth of the SSD storage of the previous generation.

Under the hood, V4.1 Flash uses a 552 billion-parameter MoE design paired with a new Causal Encoder-Decoder architecture, activating around 8 billion parameters for input and 16 billion for output. DeepSeek attributes its performance to fresh pre-training methods combined with larger-scale reinforcement-learning post-training, and the company claims benchmark results that surpass several flagship models including its own V4-Pro. This combination of efficient sparse activation, multimodal grounding, and aggressive post-training makes the model a practical fit for high-volume agentic workloads, code generation, and visual question answering where inference cost and latency matter as much as raw reasoning quality.

LLM Gatewaynovita/deepseek-v4.1-flashdeepseek-flash

Quick Info

Powered by
Provider
LLM Gateway
Model key
novita/deepseek-v4.1-flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$1.20

Limits

Output tokens
393,216 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare DeepSeek V4.1 Flash (NovitaAI) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4.1 Flash (NovitaAI)

LLM Gateway

CoverageBenchmark

A third-party review from September 10, 2026 characterizes DeepSeek-V4.1-Flash as the new Flash-series model with native multimodal visual understanding integrated into the architecture rather than split off as a separate Vision Exp variant. It highlights a 1M-token context window with up to 384K maximum output tokens, The same review states that DeepSeek's official Flash pricing from September 10, 2026 is $0.003 per million cache-hit input tokens, $0.15 per million uncached input tokens, and $0.60 per million output tokens off-peak, with peak pricing at double those rates. It also notes DeepSeek's announcement that V4 Pro requests w

LLM Gateway

CoverageAnalysis

A research-grade third-party technical analysis dated 2026-09-11 details DeepSeek-V4.1-Flash's architecture as a Causal Encoder-Decoder (CED) design with a 552B-parameter backbone that activates only 8B parameters during prefill and 16B during decode, alongside a 196B-parameter Engram conditional memory. The article re The same analysis states that DeepSeek V4.1 Flash outperforms V4 Pro (1.6T total, 49B active) on DeepSeek's published benchmarks despite having about one-third the total parameters and one-sixth the active parameters, marking what the author calls the first time a Flash-tier model entirely replaces a Pro-tier model. Al

LLM Gateway

Coverage

DeepSeek's official change log dated 2026-09-10 announces the release of DeepSeek-V4.1-Flash, describing it as the smallest model in a new architecture family with native multimodal visual understanding. The post lists first-party benchmark results including GPQA Diamond 90.9, HLE 36.8 (39.1 on the pure-text subset), C The same DeepSeek change log describes the API changes accompanying the launch: DeepSeek V4.1 Flash is served via the first-party endpoint name `deepseek-flash`, and the previous-generation V4 Flash and V4 Flash Vision Exp have been retired, with `deepseek-v4-flash` and `deepseek-v4-flash-vision-exp` temporarily routed

LLM Gateway

CoverageBenchmark

A third-party engineering breakdown describes DeepSeek-V4.1-Flash as a 552B-parameter Asymmetric Causal Encoder-Decoder Mixture-of-Experts model that activates only 8B parameters during prefill and 16B during decoding, departing from a monolithic decoder-only transformer. The article cites 1M-token native context, 384K The same article frames the architecture as splitting execution into an 8B prefill path that ingests 1,000,000-token contexts with about 114 ms latency and a 16B autoregressive decode path that produces roughly 122 tokens per second while compressing KV cache memory overhead by 75%. The author presents this as a delibe

LLM Gateway

Coverage

DeepSeek released V4.1 Flash on September 10, 2026, positioning it as the smallest member of a new architecture family that now natively supports visual understanding. The model is designed for greater capability, faster inference, higher throughput, and scaling to larger models, and DeepSeek calls it smarter, faster, DeepSeek-V4.1 Flash uses a 552 billion parameter Mixture of Experts design built on a new Causal Encoder-Decoder architecture, with only 8 billion active parameters for input and 16 billion for output. DeepSeek combined new pre-training methods with larger-scale reinforcement-learning post-training, and the published b

Videos about DeepSeek V4.1 Flash (NovitaAI)

More models around DeepSeek V4.1 Flash (NovitaAI)