Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
NovitaAI logo

Model details

DeepSeek R1 Distill LLama 70B

DeepSeek R1 Distill Llama 70B is a distilled reasoning model built on the Llama 3.3 70B architecture, designed to bring advanced chain-of-thought capabilities to a more compact and efficient footprint. Rather than training from scratch, the model inherits the reasoning blueprint and outputs generated by the larger DeepSeek R1, distilling that intelligence into the established Llama framework. This makes it well-suited for tasks requiring methodical problem-solving, mathematical rigor, and multi-step logical reasoning.

The training approach relies on using DeepSeek R1's own outputs as supervision signals to fine-tune the Llama 70B base—a process commonly called knowledge distillation. This technique yields a model that, according to benchmark results, outperforms the original Llama 70B on mathematical and factual precision, achieving 94.5% on MATH-500 and 86.7% on the AIME 2024 competition exam. Released under the MIT license with full open-weight access, the model strikes a practical balance between capability and deployability. Its reasoning strengths, combined with the ability to run on accessible infrastructure, make it a strong candidate for developers and enterprises looking to integrate sophisticated AI reasoning into real-world applications.

NovitaAIdeepseek/deepseek-r1-distill-llama-70bdeepseek-thinking

Quick Info

Powered by
Provider
NovitaAI
Model key
deepseek/deepseek-r1-distill-llama-70b
Release date
Jan 27, 2025
Last updated
Jan 27, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.80
Output token cost
$0.80

Limits

Output tokens
8,192 tokens
Context window
8,192 tokens

Transparent token rates

Compare DeepSeek R1 Distill LLama 70B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek R1 Distill LLama 70B

NovitaAI

CoverageBenchmark

R1 Distill Llama 70B pricing: $0.70/M input, $0.80/M output. Compare with 10 similar models, see benchmarks, and find the cheapest provider.

NovitaAI

CoverageBenchmark

DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). $0.70 per million input tokens, $0.80 per million output tokens. 131,072 token context window, maximum output of 16,384 tokens

Videos about DeepSeek R1 Distill LLama 70B

More models around DeepSeek R1 Distill LLama 70B