Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Nvidia logo

Model details

DeepSeek V4 Pro

DeepSeek V4 Pro serves as the flagship reasoning model in DeepSeek's V4 generation, positioned ahead of the smaller V4-Flash sibling. DeepSeek's official changelog describes the GA release of V4-Pro as a significant upgrade to agent capabilities, with particularly strong improvements in production environments, and notes that both V4-Pro and V4-Flash expose three thinking effort levels (low, high, and max) so users can tune reasoning depth to task complexity. The V4-Pro API also gained native support for the OpenAI Responses API format adapted for Codex, signaling tighter integration with agent and coding workflows.

Benchmark evidence from the V4-Pro GA update highlights the model's agent and coding focus. DeepSeek reports scores of 42.7 on HLE (without tools) and 60.0 with tools, 87.9 on Terminal Bench 2.1, 61.5 on NL2Repo, 83.3 on Cybergym, 62.7 on DeepSWE, 74.1 on Toolathlon-Verified, 25.7 on Agents' Last Exam, 31.8 on AutomationBench (Public), 71.1 on DSBench-FullStack, and 67.2 on DSBench-Hard, reflecting a balanced profile across terminal-driven, full-stack, and tool-mediated agent tasks. DeepSeek subsequently announced that V4-Pro API service would continue past its initial retirement date in response to user demand, underscoring the model's continued role in DeepSeek's lineup as newer V4.1-Flash variants benchmark against it as the prior flagship reference.

Nvidiadeepseek-ai/deepseek-v4-prodeepseek-thinkingdeprecated

Quick Info

Powered by
Provider
Nvidia
Model key
deepseek-ai/deepseek-v4-pro
Release date
Apr 24, 2026
Last updated
Apr 24, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.435
Output token cost
$0.87

Limits

Output tokens
393,216 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare DeepSeek V4 Pro pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4 Pro

Nebius Token Factory

Coverage

DeepSeek officially released its V4-Pro model (general availability build designated DeepSeek-V4-Pro-0813) on August 13, 2026, after a preview period that began in April, making the model available across DeepSeek's app, web interface, and API. According to the supplied QZ coverage citing DeepSeek's own disclosures, th The same QZ report notes that a price increase for the V4 model family takes effect at 16:00 UTC on August 16, 2026, with V4-Pro output tokens rising to $3.96 per million at peak hours from the prior flat rate of $0.87 per million, and that DeepSeek is introducing peak and off-peak billing with off-peak rates at half t

Together AI

Coverage

DeepSeek has launched DeepSeek-V4-Pro into general availability under an MIT license, making it one of the largest open-weight AI systems released. The Mixture-of-Experts model carries 1.6 trillion total parameters with 49 billion active per token and supports a 1 million-token context window for long-form tasks. The GA rollout integrates V4 Pro across DeepSeek's web platform, mobile app, and API, adding Responses API support and Codex integration aimed at agentic workflows. Hybrid attention mechanisms are cited as an architectural improvement for handling both short queries and sprawling long-context inputs. V4 Pro follows an April 24, 2026 preview, with a lighter V4 Flash variant (284B/13B active) also available.

Nvidia

CoverageBenchmark

SemiAnalysis published a Day 0 to Day 43 inference performance analysis of DeepSeek V4's 1.6T-parameter model on June 9, 2026, measuring results across GB300 NVL72, Huawei Ascend 950DT, MI355X, and B200 hardware. The study, produced with the open-source InferenceX engineering team, tracks iterative performance improvem The piece includes a detailed first look at DeepSeek V4 inference on Huawei Ascend 950DT, breaking down compute and communication overlap and the compute streams used to optimize performance, and frames the model as co-designed in part for Huawei Ascend. Most optimizations measured during the study were merged into the

Together AI

CoveragePreview

DeepSeek's official API docs announce the V4 Preview release, making DeepSeek-V4-Pro live and open-sourced. The model uses a Mixture-of-Experts design with 1.6 trillion total parameters and 49 billion active per token, positioned as rivaling top closed-source systems. DeepSeek also introduced V4-Flash at 284B total / 13B active for faster, cheaper inference. The Preview notes highlight agentic coding benchmarks where V4-Pro achieves open-source SOTA, plus leading knowledge and reasoning results among open models. Architectural changes include token-wise compression and DeepSeek Sparse Attention, with 1M context set as the default. The API exposes both models with Thinking and Non-Thinking modes, while deepseek-chat and deepseek-reasoner are slated for retirement on July 24, 2026.

SiliconFlow

Coverage

According to the DeepSeek API Change Log (2026-09-10), DeepSeek announced that, in response to user demand, it will continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged. The change log entry was filed alongside the release of DeepSeek-V4.1-Flash, the The notice affects V4 Pro specifically: despite an earlier plan to retire V4 Pro, DeepSeek states the model will remain available on the DeepSeek API past the September 14, 2026 cutoff, with no announced change to pricing or usage terms, and promises further notice if conditions change. V4.1-Flash replaces the retired

Videos about DeepSeek V4 Pro

More models around DeepSeek V4 Pro