Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

DeepSeek V4.1 Flash (Fireworks AI)

DeepSeek V4.1 Flash introduces a new Causal Encoder-Decoder architecture within the broader DeepSeek family, designed for greater capability and faster inference than its predecessors. It is a 552-billion-parameter mixture-of-experts model that activates only 8 billion parameters for input processing and 16 billion for output generation. Combined with new pretraining methods and larger-scale reinforcement learning post-training, the model is positioned as the smallest entry in a new architecture family that scales toward larger flagship variants, with DeepSeek reporting benchmark performance ahead of its own top-tier DeepSeek-V4-Pro model.

The architecture's efficiency focus translates into practical advantages for agentic workloads, particularly around memory and cost. V4.1 Flash requires roughly one-quarter of the high-bandwidth memory and one-eighth of the SSD storage for its KV cache compared with the previous generation, lowering agent memory costs substantially. The release also notes native multimodal support, and compatibility aliases were introduced so that retiring V4-Flash and V4-Flash-Vision-Exp endpoint identifiers temporarily route to the new V4.1 Flash model. As an open-weights release from DeepSeek, it fits well for teams building reasoning-heavy agents, tool-using assistants, and high-throughput services that benefit from the asymmetric parameter design and reduced cache footprint.

LLM Gatewayfireworks/deepseek-v4.1-flashdeepseek-flash

Quick Info

Powered by
Provider
LLM Gateway
Model key
fireworks/deepseek-v4.1-flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.22
Output token cost
$0.66

Limits

Output tokens
393,216 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare DeepSeek V4.1 Flash (Fireworks AI) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4.1 Flash (Fireworks AI)

LLM Gateway

CoverageBenchmark

VentureBeat reports DeepSeek launched DeepSeek-V4.1-Flash on September 10, 2026, describing it as a 552-billion-parameter mixture-of-experts model with native vision, a 1-million-token context window, and a Causal Encoder-Decoder architecture designed to make repeatedly reading large contexts cheaper. The MIT-licensed The piece frames the $0.003 cached-input rate as a significant shift for agent workloads that repeatedly reread the same repository, tool definitions, system instructions, or conversation history, noting that DeepSeek itself acknowledges cache-hit charges can dominate agent costs. It benchmarks V4.1 Flash against front

LLM Gateway

CoverageBenchmark

Coursiv's coverage of the DeepSeek V4.1 Flash release documents the September 10, 2026 launch and the subsequent routing change effective 12:00 Beijing time on September 14, 2026, after which every request to deepseek-v4-pro is redirected to V4.1 Flash and billed at Flash rates until a future V4.1 Pro ships. The articl A spec table sourced from DeepSeek's official model card compares V4 Flash to V4.1 Flash: backbone parameters grew from 284B to 552B while active parameters dropped from 13B to 8B for input and sit at 16B for output, the architecture shifted from Mixture-of-Experts decoder to a Causal Encoder-Decoder with 20 encoder an

LLM Gateway

Coverage

DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with 552B backbone parameters and 196B additional Engram parameters, supporting a 1M-token context window. It activates 8B parameters per token during prefill and 16B during decode, with a global KV cache footprint of 890 bytes per token — roughly one-quarter The architecture introduces a Causal Encoder-Decoder split: a 20-layer causal encoder paired with a 20-layer decoder, where the decoder derives its global KV from the final encoder hidden state via per-layer projection weights, nearly halving prefill compute. Sliding-window attention with a 128-token window uses Decode

LLM Gateway

CoverageAnalysis

Released September 10, 2026, DeepSeek-V4.1-Flash introduces a Causal Encoder-Decoder (CED) architecture that reduces KV cache requirements by roughly 4x relative to V4 Flash and 437x relative to DeepSeek V1. The model achieves its 552B backbone parameter count with only 8B active during prefill and 16B during decode, a The architecture combines CED encoder-decoder with projected global KV cache, Compressed Sparse Attention 2 (CSA2) with three static attention modes (Full, Reindex, Reuse), FP4 KV cache compression using E2M1 format, SWA Bounded Replay, Single-Pass mHC residual connections, Engram conditional memory (196B parameters),

LLM Gateway

CoverageBenchmark

The AI Release Tracker entry for DeepSeek-V4.1-Flash records the model as released by DeepSeek on September 10, 2026, 28 days after DeepSeek-V4-Pro-0813. It aggregates third-party provider pricing sourced from openrouter.ai on September 15, 2026, listing Fireworks at $0.22 input and $0.66 output per 1M tokens at a 1M c On benchmarks, the tracker reports V4.1 Flash scoring 65.4% on NL2Repo-Bench, described as the best published NL2Repo score of all tracked models, alongside other coding and agent evaluations. The page contextualizes Fireworks as one of multiple inference hosts serving the same DeepSeek-released weights at varying prec

LLM Gateway

CoverageBenchmark

DeepSeek-V4.1-Flash ranks 13th on the LLM Stats composite leaderboard, with capability tiers placing it in the top 2% for Tool Calling (2 of 194) and the top 10% for Coding (5 of 267), while landing at 18 of 363 in Reasoning, 31 of 208 in Vision, and 43 of 327 in Math. The model posts a 3,471 Codeforces rating, a 0.91 Specific benchmark methodology cited includes maximum reasoning effort at temperature 1.0 and top-p 0.95, with sources traced to the model's official Hugging Face scorecard and paper. The leaderboard positions DeepSeek-V4.1-Flash above Gemma 4 E4B and below GPT-6 Astra on the LLM Stats Score versus blended price chart,

LLM Gateway

Coverage

DeepSeek officially released DeepSeek-V4.1-Flash on September 10, 2026 via its API change log. The model is described as the smallest member of a new architecture family with native multimodal visual understanding, designed for a higher capability ceiling, faster inference, higher throughput, and scaling to larger mode The change log lists the model's official benchmark results, including GPQA Diamond 90.9, Codeforces Rating 3471, MathArena Apex 65.6, Terminal-Bench 2.1 90.6, Terminal-Bench 3.0 30.0, Terminal-Bench 4.0 31.2, DeepSWE v1.1 74.2, ProgramBench 20.3, NL2Repo-Bench 65.4, CyberGym 88.1, SEC-Bench Pro 62.8, ExploitGym 15.3,

Videos about DeepSeek V4.1 Flash (Fireworks AI)

More models around DeepSeek V4.1 Flash (Fireworks AI)