Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

DeepSeek V4.1 Flash (DeepInfra)

DeepSeek V4.1 Flash is a 552-billion-parameter Mixture-of-Experts model built on a Causal Encoder Decoder design that activates only 8 billion parameters on input and 16 billion on output. This sparse activation pattern lets a very large model deliver strong reasoning while keeping the per-token compute budget closer to a small model. The architecture also marks the first Flash-tier release in the new DeepSeek family line to include native vision and multimodal understanding, so a single model can handle text and images together rather than relying on a separate vision encoder pipeline.

The model is positioned for agent and long-context workloads where cache efficiency dominates cost. Its KV cache is reported at roughly 890 bytes per token, which is dramatically smaller than earlier DeepSeek generations and helps repeated-prompt and tool-loop scenarios stay affordable. DeepSeek V4.1 Flash is also said to outperform the larger DeepSeek V4 Pro on performance, cost, speed, and task completion time while reaching a CyberGym score of 88.1. That combination of a one-million-token context window, very large maximum output, structured JSON output, tool calling, and open-weight availability makes the model a practical fit for coding agents, retrieval-heavy assistants, high-throughput batch processing, and other applications that need frontier-class reasoning without paying frontier-class prices.

LLM Gatewaydeepinfra/deepseek-v4.1-flashdeepseek-flash

Quick Info

Powered by
Provider
LLM Gateway
Model key
deepinfra/deepseek-v4.1-flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$0.60

Limits

Output tokens
393,216 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare DeepSeek V4.1 Flash (DeepInfra) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4.1 Flash (DeepInfra)

LLM Gateway

CoverageBenchmark

DeepSeek released DeepSeek-V4.1-Flash on September 10, 2026, as a generational replacement for V4 Flash rather than a point update. The new model doubles the backbone to 552B parameters while cutting active parameters to 8B for input and 16B for output, moves to a 40-layer Causal Encoder-Decoder architecture (20 encode On Terminal-Bench 4.0 the score jumps from 7.0 (V4 Flash) to 31.2, and Terminal-Bench 2.1 rises from 82.7 to 90.6, with DeepSWE v1.1 climbing from 54.4 to 74.2. DeepSeek also announced that from 12:00 Beijing time on September 14, 2026 (04:00 UTC), every request to deepseek-v4-pro is routed to V4.1 Flash at Flash rates

LLM Gateway

CoverageBenchmark

DeepSeek V4.1 Flash ships as a multimodal Mixture-of-Experts model with a 552B-parameter backbone, 8B active parameters during prefill and 16B during decode, a 1M-token context window, up to 384K tokens of output, and native text-plus-image inputs producing autoregressive text. The architecture is a 20-layer causal enc The release centers on a redesigned KV-cache strategy that compresses global per-token cache to roughly a quarter of the previous Flash generation, a key change for long-context agent workloads. The API supports JSON output, tool calls, the Responses API, the Anthropic-compatible endpoint at https://api.deepseek.com/an

Videos about DeepSeek V4.1 Flash (DeepInfra)

More models around DeepSeek V4.1 Flash (DeepInfra)