Currently listed through these providers:
Model details
DeepSeek-V3
DeepSeek-V3 is a large-scale open-source language model developed by DeepSeek-AI, built on a Mixture-of-Experts architecture with 671B total parameters and 37B activated per token. This sparse activation design allows the model to maintain broad capacity while keeping per-token compute manageable, making it well suited for general-purpose text generation, reasoning, and knowledge-intensive tasks. The model is positioned as a successor in the DeepSeek family, continuing architectural ideas that were validated in DeepSeek-V2 and refining them for stronger overall capability.
The architecture centers on Multi-head Latent Attention paired with the DeepSeekMoE framework, and it introduces an auxiliary-loss-free load balancing strategy alongside a multi-token prediction training objective. DeepSeek-V3 was pre-trained on 14.8 trillion diverse and high-quality tokens and then refined through Supervised Fine-Tuning and Reinforcement Learning stages, with the creators reporting only 2.788M H800 GPU hours for full training and a notably stable run with no irrecoverable loss spikes or rollbacks. The technical report describes DeepSeek-V3 as outperforming other open-source models and performing comparably to leading closed-source systems, suggesting it is a strong fit for teams seeking open-weight frontier-quality language modeling.
Quick Info
Powered by- Provider
- Together AI
- Model key
- deepseek-ai/DeepSeek-V3
- Release date
- Dec 26, 2024
- Last updated
- May 29, 2025
- Knowledge cutoff
- 2024-07
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.25
- Output token cost
- $1.25
Limits
- Output tokens
- 131,072 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare DeepSeek-V3 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about DeepSeek-V3
No articles yet. Fetch the latest news to show it here.