LLM Gateway
DeepSeek V4.1 Flash, released September 10, 2026, is the newest Flash model designed for higher capability ceiling, faster inference, and native multimodal visual understanding integrated into a single architecture. The API offers a 1M-token context with up to 384K maximum output, and DeepSeek reports strong benchmark results including 90.9 on GPQA Diamond, 3,471 Codeforces rating, 90.6 on Terminal-Bench 2.1, 74.2% on DeepSWE v1.1, 88.1 on CyberGym, and 65.4 on NL2Repo-Bench. Compared to V4 Flash 0731, coding results improve sharply: Terminal-Bench 2.1 rises from 82.7 to 90.6 and DeepSWE from 54.4 to 74.2. Flash pricing from September 10, 2026 is $0.003 per million cache-hit input tokens, $0.15 per million uncached input, and $0.60 per million output off-peak, with peak rates double. DeepSeek also routes V4 Pro requests to V4.1 Flash at Flash pricing until V4.1 Pro ships.