Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
LLM Gateway logo

Model details

DeepSeek V4.1 Flash (Together AI)

DeepSeek V4.1 Flash is a multimodal Mixture-of-Experts model built around KV cache compression, pairing a 552B-parameter backbone with selective activation so that only 8B parameters fire during prefill and 16B during decode. This narrow activation strategy is the through-line of the design: by projecting the decoder's global KV cache from the final encoder hidden states rather than maintaining per-layer caches, and by reconstructing sliding-window attention states through bounded replay, the model is engineered to keep memory and compute flat as contexts stretch toward one million tokens. The result is a long-context model that natively ingests images and text and emits text autoregressively, intended for input-heavy agentic loops where serving economics matter as much as raw capability.

In practice, the model is positioned as a balanced workhorse in the Flash tier: it accepts multimodal input, supports reasoning and tool calling, and is released as open weights under the MIT license, making it well suited to teams that want to self-host or fine-tune a vision-capable reasoning model without paying flagship-tier prices. Compared with the larger V4 Pro sibling it trades some intelligence index headroom for markedly higher output throughput, while sitting above the base V4 Flash line on aggregate benchmarks, reflecting the payoff from the cache-compression work documented in the accompanying technical report. For practitioners building retrieval-heavy agents, document analysis pipelines, or any workload that benefits from a million-token window at modest per-token cost, V4.1 Flash offers a practical middle ground between capability and serving efficiency.

LLM Gatewaytogether-ai/deepseek-v4.1-flashdeepseek-flash

Quick Info

Powered by
Provider
LLM Gateway
Model key
together-ai/deepseek-v4.1-flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$1.20

Limits

Output tokens
393,216 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare DeepSeek V4.1 Flash (Together AI) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4.1 Flash (Together AI)

LLM Gateway

Coverage

DeepSeek released V4.1-Flash, a 552-billion-parameter multimodal Mixture-of-Experts model built around a Causal Encoder-Decoder architecture with four-bit cache quantization that cuts the active key-value cache HBM requirement for long-running agents to roughly one-quarter, and persistent SSD cache to roughly one-eight From September 14, 2026 at 04:00 UTC, all API traffic directed to the deepseek-v4-pro endpoint will be automatically rerouted to V4.1-Flash at V4.1-Flash rates, making the migration mandatory for every developer currently calling that endpoint. The article frames the KV-cache reduction as a structural rather than margi

LLM Gateway

CoverageBenchmark

DeepSeek-V4.1-Flash is positioned as the newest Flash model and the smallest member of a new architecture family, designed for a higher capability ceiling, faster inference, higher throughput, and scaling to larger models rather than as a cheap sibling to a Pro model. The release brings native multimodal visual underst Pricing on DeepSeek's Flash schedule from September 10, 2026 is $0.003 per million cache-hit input tokens, $0.15 per million uncached input tokens, and $0.60 per million output tokens off-peak, with peak pricing at double those rates. DeepSeek also confirmed that V4 Pro requests will be routed to V4.1 Flash and billed

LLM Gateway

Coverage

DeepSeek AI released DeepSeek-V4.1-Flash on September 10, 2026, targeting the KV cache bottleneck that long-horizon agents create through repeated prefills and million-token contexts. The model is a multimodal Mixture-of-Experts with a 552B-parameter backbone, an additional 196B Engram parameters, and a 1M-token contex The architecture introduces a Causal Encoder-Decoder split that nearly halves prefill compute: the 40-layer backbone is divided into a 20-layer causal encoder and a 20-layer decoder, with per-layer projection weights deriving the decoder's global KV from the final encoder hidden state in a YOCO-inspired design. A 128-t

LLM Gateway

CoverageAnalysis

A research-grade technical deep dive confirms DeepSeek V4.1 Flash released on September 10, 2026, as a Causal Encoder-Decoder (CED) model that cuts KV cache requirements approximately 4x relative to V4 Flash and 437x relative to DeepSeek V1. With 552B backbone parameters and only 8B active during prefill and 16B during The deep dive enumerates the full architectural stack: CED encoder-decoder with projected global KV cache; Compressed Sparse Attention 2 (CSA2) with three static attention modes (Full, Reindex, Reuse); FP4 KV cache compression using E2M1 format; SWA Bounded Replay for sliding window attention reconstruction; Single-Pas

LLM Gateway

CoveragePreview

DeepSeek V4.1 Flash is available through the official DeepSeek API under the model name deepseek-flash, with the September 10 release adding the new model behind the Flash name and lowering listed prices. The legacy identifiers deepseek-v4-flash and deepseek-v4-flash-vision-exp are temporarily accepted and routed to V4 The post warns developers that a stable request string does not prove the underlying model stayed the same, and recommends running the same acceptance checks before and after the September 14 boundary — including structured outputs, tool loops, and billing records — rather than relying solely on chat prompts. It also c

LLM Gateway

CoverageLeaks

DeepSeek V4.1 Flash graduated from a two-day experiment to a real release on September 10, 2026, now answering to the model name deepseek-flash across DeepSeek's app, web interface, and API. DeepSeek published a technical report titled "Pushing the Limits of KV Cache Compression" and released MIT-licensed weights for a The post situates V4.1 Flash alongside other recent listings — including Orca: OrcaCyber Zero 1.0, Orca: OrcaVerify Text 1.0, GPT-6 Astra, Gemini 3.8 Flash, Qwen3.8 Max, Claude Fable 5.1, Qwen3.8 Flash, GLM 5.3 Flash, DeepSeek V4 Flash Vision (Exp), GLM 5.3, Qwen3.8 27B, DeepSeek V4 Pro 0813, Grok 4.6, Muse Spark 1.2,

LLM Gateway

Coverage

DeepSeek V4.1 Flash is a 552-billion-parameter MoE model using a new Causal-Encoder-Decoder asymmetric structure that activates only 8B parameters during input and 16B during output, paired with a new pretraining method and larger-scale reinforcement-learning post-training that the source says lets it exceed flagships The official API exposes V4.1 Flash under the model name deepseek-flash, with support for thinking and non-thinking modes, tool calls, JSON output, the Responses API, and OpenAI- and Anthropic-compatible interface formats. New peak-hour pricing of $0.006 per million tokens for cache-hit input, $0.30 for cache-miss inpu

Videos about DeepSeek V4.1 Flash (Together AI)

More models around DeepSeek V4.1 Flash (Together AI)