Currently listed through these providers:
Model details
DeepSeek V3
DeepSeek V3 is an open-weight transformer-based large language model that continues the architectural lineage established by DeepSeek-V2, carrying over the DeepSeekMoE mixture-of-experts design for the feed-forward layers while introducing auxiliary-loss-free load balancing as the main structural refinement over its predecessor. By removing the need for an explicit auxiliary loss to keep experts balanced, the approach aims to simplify training and preserve the model's representational capacity, making the Mixture-of-Experts configuration easier to scale without the gradient interference that traditional balancing penalties can introduce.
The model also adopts Multi-Head Latent Attention for efficient autoregressive inference, jointly compressing keys and values into a latent space to shrink the key-value cache that otherwise dominates memory usage during long-context generation. Positional information is preserved through RoPE applied to the compressed representations, supplemented by an additional projection matrix that carries a rotation key, and queries are compressed in parallel to minimize the memory footprint before being expanded back to full dimensionality at the attention output. Together, these design choices reflect an emphasis on economical training and inference, positioning DeepSeek V3 as a practical open-weight option for developers who want MoE-scale capacity without paying the full memory cost of dense attention.
Quick Info
Powered by- Provider
- Alibaba (China)
- Model key
- deepseek-v3
- Release date
- Dec 1, 2024
- Last updated
- Dec 1, 2024
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.287
- Output token cost
- $1.147
Limits
- Output tokens
- 8,192 tokens
- Context window
- 65,536 tokens
Transparent token rates
Compare DeepSeek V3 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about DeepSeek V3
No articles yet. Fetch the latest news to show it here.