Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Eden AI logo

Model details

DeepSeek V4.1 Flash (Together AI)

DeepSeek V4.1 Flash is an open-weight multimodal model from DeepSeek designed to handle very long contexts while keeping inference efficient. It is built as a 552B-parameter Mixture-of-Experts backbone that activates only 8B parameters per token during prefill and 16B during decode, a configuration aimed at cutting the cost of input-heavy and agent-style workloads. The architecture uses a 40-layer Causal Encoder-Decoder layout in which a 20-layer encoder feeds a 20-layer decoder whose global key-value cache is projected from the final encoder states rather than rebuilt at every decoder layer. Native multimodal input lets the model accept both images and text and produce text autoregressively, while techniques such as SWA Bounded Replay help reconstruct missing sliding-window attention states for long sequences.

Practically, DeepSeek V4.1 Flash targets teams that need a reasoning-capable model with a million-token context window for tasks like document analysis, multi-turn agent loops, and tool-driven workflows. Its hybrid-attention design and KV-cache compression focus translate into competitive throughput, with reported output speeds around the mid-hundreds of tokens per second alongside a measured intelligence index that positions it ahead of other DeepSeek V4 Flash variants on command-line leaderboards. The combination of open weights, vision support, and strong long-context reasoning makes it a flexible foundation for experimentation and deployment where both cost efficiency and analytical depth matter.

Eden AItogether_ai/deepseek-ai/DeepSeek-V4.1-Flashdeepseek-flash

Quick Info

Powered by
Provider
Eden AI
Model key
together_ai/deepseek-ai/DeepSeek-V4.1-Flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.30
Output token cost
$1.20

Limits

Output tokens
384,000 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare DeepSeek V4.1 Flash (Together AI) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4.1 Flash (Together AI)

Eden AI

CoveragePreview

This developer blog, checked September 10, 2026, documents DeepSeek V4.1 Flash as available through the official DeepSeek API under the recommended model name `deepseek-flash`, with legacy aliases `deepseek-v4-flash` and `deepseek-v4-flash-vision-exp` temporarily served by the new release. It announces a September 14, The post also cautions that vendor-reported benchmarks do not prove real-world agent performance and advises retaining a baseline of structured-output, tool-loop, and billing tests before the September 14 boundary so the same checks can be rerun afterward. It explicitly notes it does not establish that every third-part

Eden AI

CoverageLeaks

This post characterizes DeepSeek V4.1 Flash as a formally released (rather than experimental) model, with DeepSeek publishing a technical report titled "Pushing the Limits of KV Cache Compression" and MIT-licensed weights for a 552B-parameter multimodal mixture-of-experts model built on a new Causal Encoder-Decoder arc The page is hosted by OrcaRouter, a model router/aggregator, so its model-listing rows and pricing entries reflect that vendor's view rather than DeepSeek's official API pricing, and the excerpt does not name Together AI or address Eden AI routing. Architecture and shard-count claims are valuable technical context but

Eden AI

Coverage

Reuters reports that Chinese AI startup DeepSeek launched DeepSeek-V4.1-Flash on September 10, 2026, describing it as the smallest model in its new architecture family. According to DeepSeek's own statement cited in the wire, the model is designed for greater capability, faster inference, higher throughput, and scaling This Reuters article is concise and authoritative on the launch event itself, but the supplied excerpt does not include benchmark numbers, API specifics, or any mention of Together AI as a hosting partner. For developer-facing technical detail such as API model names, routing changes, and capability lists, other candid

Eden AI

Coverage

DeepSeek officially released DeepSeek V4.1 Flash on September 10, 2026, as the smallest model in a new architectural series that pairs higher capability with faster inference, greater throughput, and lower cost, per AIHub's Chinese-language summary. The model is a 552B-parameter MoE with an asymmetric Causal-Encoder-De According to AIHub, the official API exposes a 1M token context with up to 384K tokens of output and supports thinking and non-thinking modes, tool calling, JSON output, the Responses API, and OpenAI- and Anthropic-compatible interfaces; the API model name is deepseek-flash, while the legacy aliases deepseek-v4-flash a

Eden AI

Coverage

DeepSeek released DeepSeek-V4.1 Flash on September 10, 2026, as the smallest model in a new architecture family that now natively supports visual understanding. According to Neowin's coverage, the model is a 552 billion parameter Mixture of Experts design using a new Causal Encoder-Decoder architecture that activates o The Neowin article publishes a detailed benchmark table comparing DeepSeek V4.1-Flash against V4-Pro 0813, V4-Flash 0731, GLM 5.3, Kimi K3, GPT 5.6-Sol, and Claude Opus 5 across GPQA Diamond, HLE, Codeforces Rating, MathArena Apex, Terminal-Bench 2.1/3.0/4.0, DeepSWE v1.1, ProgramBench, NL2Repo-Bench, CyberGym, SEC-Ben

Videos about DeepSeek V4.1 Flash (Together AI)

More models around DeepSeek V4.1 Flash (Together AI)