Sulat.com
AI models
UnoRouter logo

Model details

DeepSeek V4 Pro

DeepSeek V4 Pro is a Mixture-of-Experts language model from the DeepSeek-V4 preview series, pairing 1.6 trillion total parameters with 49 billion activated per token. Its defining architectural choice is a hybrid attention stack that fuses Compressed Sparse Attention (CSA) with Heavily Compressed Attention (HCA), targeting long-context efficiency rather than brute-force scaling. In a one-the cataloged API limit setting, the model reportedly requires only about 27% of the single-token inference FLOPs and roughly 10% of the KV cache that the prior DeepSeek-V3.2 needs, a meaningful jump in compute and memory efficiency for very long inputs.

Post-training follows a two-stage pipeline: independent domain-expert cultivation through supervised fine-tuning combined with GRPO, then unified model consolidation via on-policy distillation. The released weights are openly available under an MIT-style license and are packaged for commercial deployment, making the model attractive for teams that want to self-host a reasoning-oriented LLM tuned for the cataloged API limit workloads. Practical fits include document analysis, code reasoning across large repositories, agentic pipelines that need tool calling and structured outputs, and any application where reducing inference cost at extreme context lengths matters more than maximizing raw single-prompt throughput.

UnoRouterdeepseek-v4-prodeepseek-thinking

Quick Info

Powered by
Provider
UnoRouter
Model key
deepseek-v4-pro
Release date
Apr 24, 2026
Last updated
Apr 24, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.8999
Output token cost
$1.7999

Limits

Output tokens
384,000 tokens
Context window
1,000,000 tokens

Latest news about DeepSeek V4 Pro

Videos about DeepSeek V4 Pro

More models around DeepSeek V4 Pro