Eden AI
DeepSeek released V4.1-Flash on September 10, 2026 at 04:00 UTC, a 552-billion-parameter multimodal model built around four interlocking architectural techniques that reduce active KV cache memory to one-quarter of V4-Flash levels. The release went live alongside a 50-page technical report on Hugging Face, marking the debut of DeepSeek's V4.1 model series as a distinct architecture lineage. For developers running production agent workloads, cache hit charges on agentic tasks routinely account for the majority of inference spending. The structural change cuts per-token agent memory to 890 bytes and reduces persistent SSD cache storage to one-eighth of V4-Flash requirements. Starting September 14, 2026 at 04:00 UTC, all API traffic directed to the deepseek-v4-pro endpoint will be automatically rerouted to V4.1-Flash at V4.1-Flash rates, making the transition mandatory for every developer currently calling that endpoint. This architectural innovation rather than parameter scaling represents a significant efficiency inflection point for long-running agentic deployments.