DeepSeek V4 Pro serves as the flagship reasoning model in DeepSeek's V4 generation, positioned ahead of the smaller V4-Flash sibling. DeepSeek's official changelog describes the GA release of V4-Pro as a significant upgrade to agent capabilities, with particularly strong improvements in production environments, and notes that both V4-Pro and V4-Flash expose three thinking effort levels (low, high, and max) so users can tune reasoning depth to task complexity. The V4-Pro API also gained native support for the OpenAI Responses API format adapted for Codex, signaling tighter integration with agent and coding workflows.
Benchmark evidence from the V4-Pro GA update highlights the model's agent and coding focus. DeepSeek reports scores of 42.7 on HLE (without tools) and 60.0 with tools, 87.9 on Terminal Bench 2.1, 61.5 on NL2Repo, 83.3 on Cybergym, 62.7 on DeepSWE, 74.1 on Toolathlon-Verified, 25.7 on Agents' Last Exam, 31.8 on AutomationBench (Public), 71.1 on DSBench-FullStack, and 67.2 on DSBench-Hard, reflecting a balanced profile across terminal-driven, full-stack, and tool-mediated agent tasks. DeepSeek subsequently announced that V4-Pro API service would continue past its initial retirement date in response to user demand, underscoring the model's continued role in DeepSeek's lineup as newer V4.1-Flash variants benchmark against it as the prior flagship reference.