Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Eden AI logo

Model details

DeepSeek V4.1 Flash (Greenference)

The model overview is being prepared.

Eden AIgreenference/deepseek-v4.1-flashdeepseek-flash

Quick Info

Powered by
Provider
Eden AI
Model key
greenference/deepseek-v4.1-flash
Release date
Sep 10, 2026
Last updated
Sep 10, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.0015
Output token cost
$0.09

Limits

Output tokens
384,000 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare DeepSeek V4.1 Flash (Greenference) pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4.1 Flash (Greenference)

Eden AI

Coverage

DeepSeek released V4.1-Flash on September 10, 2026 at 04:00 UTC, a 552-billion-parameter multimodal model built around four interlocking architectural techniques that reduce active KV cache memory to one-quarter of V4-Flash levels. The release went live alongside a 50-page technical report on Hugging Face, marking the debut of DeepSeek's V4.1 model series as a distinct architecture lineage. For developers running production agent workloads, cache hit charges on agentic tasks routinely account for the majority of inference spending. The structural change cuts per-token agent memory to 890 bytes and reduces persistent SSD cache storage to one-eighth of V4-Flash requirements. Starting September 14, 2026 at 04:00 UTC, all API traffic directed to the deepseek-v4-pro endpoint will be automatically rerouted to V4.1-Flash at V4.1-Flash rates, making the transition mandatory for every developer currently calling that endpoint. This architectural innovation rather than parameter scaling represents a significant efficiency inflection point for long-running agentic deployments.

Eden AI

CoverageAnalysis

DeepSeek V4.1 Flash represents a fundamental architectural departure from the V4 generation, introducing a Causal Encoder-Decoder design that reduces KV cache requirements by approximately 4× relative to V4 Flash and 437× relative to DeepSeek V1. The model achieves 552B backbone parameters with only 8B active during prefill and 16B during decode, enabling substantial cost efficiency for agentic workloads. Key techniques include Compressed Sparse Attention 2 with three static attention modes, FP4 KV cache compression in E2M1 format, and native multimodal vision via DeepSeek-ViT. The model outperforms V4 Pro (1.6T/49B active) across all measured benchmarks despite having one-third the total parameters and one-sixth the active parameters, marking the first instance where a Flash tier model entirely replaces a Pro tier model. Additional features include Engram conditional memory at 196B parameters, DSpark speculative decoding, and Single-Pass mHC residual connections. All architectural claims are verified against official DeepSeek documentation and the HuggingFace model card published September 10, 2026, under MIT license.

Eden AI

Coverage

DeepSeek V4.1 Flash is a 552B-parameter MoE model employing a Causal-Encoder-Decoder architecture with asymmetric activation: only 8B during input and 16B during output, resulting in lower costs than known models of the same size. It is the smallest model in the new architecture series and features native multimodal visual understanding, surpassing DeepSeek V4 Pro and outperforming models like GLM5.3 and Kimi-K3 on Agentic Benchmarks. Compared with V4 Flash, the model reduces HBM requirements to one-fourth and SSD requirements to one-eighth while extending context length from 4K to 1M. DeepSeek has open-sourced V4.1 Flash on Hugging Face and released a technical report alongside the launch. The model adopted a new pretraining approach with larger-scale reinforcement learning post-training, and new pricing with Peak-Valley Pricing took effect at 12:00 on September 10, 2026. The older V4 Flash and V4 Flash Vision Exp versions were taken offline, with their names temporarily routed to V4.1 Flash, and V4 Pro requests will be routed to V4.1 Flash after September 14, 2026.

Eden AI

Coverage

DeepSeek V4.1 Flash officially launched on September 10, 2026, as a multimodal model released under an open MIT license for commercial and research use. The release extends the Flash family with native vision capabilities, enabling developers to process both text and image inputs within a single model for document analysis and visual question answering. The model became generally available through DeepSeek's standard API endpoints, representing a point update focused on multimodal integration within the V4 family cycle rather than a ground-up architectural redesign. This addition positions DeepSeek to compete with proprietary multimodal systems while maintaining its commitment to open-source AI development. V4.1 Flash maintains the efficiency characteristics expected from the Flash designation while adding cross-modal reasoning, jointly processing visual and textual information increasingly essential for enterprise workflows. The open MIT license lowers barriers for enterprise adoption, allowing organizations to freely deploy the model across commercial applications without restrictive licensing terms.

Videos about DeepSeek V4.1 Flash (Greenference)

More models around DeepSeek V4.1 Flash (Greenference)