LLM Gateway
DeepSeek-V4.1-Flash launched on September 10, 2026, and from 04:00 UTC on September 14 every request to deepseek-v4-pro is routed to V4.1 Flash and billed at Flash rates until a V4.1 Pro ships. Anyone building on Pro gets the switch automatically, with cache-miss input cost dropping about 77 percent and output cost abo The model is a generational replacement rather than a point update: backbone parameters grew from 284B to 552B while active parameters fell from 13B to 8B (input) and 16B (output), the architecture moved from a Mixture-of-Experts decoder to a Causal Encoder-Decoder (20 encoder plus 20 decoder layers), vision is now nat