Currently listed through these providers:
Model details
DeepSeek R1 Distill Llama 70B
DeepSeek R1 Distill Llama 70B is a distilled large language model built upon the Llama-3.3-70B-Instruct architecture, engineered to bring frontier-level reasoning capabilities to a more compact and efficient form. With 70 billion parameters, it serves as a distilled version of the larger DeepSeek R1 series, designed to handle reasoning-heavy workloads in math, coding, and analytical problem-solving. The model targets developers and enterprises seeking high accuracy on complex tasks without the computational overhead of larger frontier models.
The model achieves its performance through advanced knowledge distillation, learning from outputs generated by the full DeepSeek R1 system. This process transfers sophisticated reasoning patterns into the smaller Llama backbone, resulting in competitive benchmark results: 70% pass@1 on AIME 2024, 94.5% on MATH-500, and a CodeForces rating of 1633. The architecture combines the efficiency of the 70B Llama foundation with DeepSeek's reasoning training, enabling the model to perform multi-step reasoning, tool calling, and temperature-controlled generation. Available through multiple inference providers including DeepInfra, OpenRouter, and NVIDIA NIM, the model offers practical deployment flexibility for production applications requiring fast, accurate reasoning at scale.
Quick Info
Powered by- Provider
- Alibaba (China)
- Model key
- deepseek-r1-distill-llama-70b
- Release date
- Jan 1, 2025
- Last updated
- Jan 1, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.287
- Output token cost
- $0.861
Limits
- Output tokens
- 16,384 tokens
- Context window
- 32,768 tokens
Transparent token rates
Compare DeepSeek R1 Distill Llama 70B pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about DeepSeek R1 Distill Llama 70B
No articles yet. Fetch the latest news to show it here.