CoralBricks
A third-party breakdown of the V4.1 Flash release frames it as a generational replacement for V4 Flash rather than a point update. Compared to V4 Flash, backbone parameters grow from 284B to 552B while active parameters per token fall from 13B to 8B for input and 16B for output, and the architecture shifts from a Mixture-of-Experts decoder to a Causal Encoder-Decoder with 20 encoder plus 20 decoder layers and a 196B-parameter Engram conditional memory. Vision moves from a separate experimental model to native training from pre-training start, and global KV cache per token drops to roughly a quarter of V4 Flash at 890 bytes. Terminal-Bench 2.1 rises from 82.7 to 90.6, Terminal-Bench 4.0 from 7.0 to 31.2, and DeepSWE v1.1 from 54.4 to 74.2, all at MIT license with a 1M-token context. The piece documents DeepSeek's decision to route all deepseek-v4-pro traffic to V4.1 Flash from 12:00 Beijing time on September 14, 2026 (04:00 UTC), billing at Flash rates until a V4.1 Pro ships, which drops cache-miss input cost by about 77 percent and output cost by about 70 percent for Pro users. V4 Flash and experimental V4 Flash Vision are retired, with their old model names temporarily routing to V4.1 Flash so existing code keeps running. DeepSeek's stated reason is that V4.1 Flash surpassed V4 Pro on performance, cost, speed, and total time, supported by its own agentic benchmark numbers.