Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Azure Cognitive Services logo

Model details

DeepSeek-R1

DeepSeek-R1 is a reasoning-focused language model built to handle complex multi-step problems in math, code, and logical deduction. The model draws from a lineage of 671 billion parameters, with 37 billion parameters active during inference, positioning it as a substantial architecture designed for deep reasoning chains rather than quick surface-level responses. Rather than relying purely on supervised fine-tuning, DeepSeek-R1 was shaped through large-scale reinforcement learning, a design choice that lets the model explore and develop reasoning strategies organically. This approach enables the model to generate visible chain-of-thought processes and deliver answers that reflect genuine problem decomposition, making it particularly suited for tasks where accuracy depends on sustained logical effort.

The model traces its origins to DeepSeek-R1-Zero, an earlier variant trained purely through reinforcement learning without any supervised fine-tuning preamble. DeepSeek-R1 builds on that foundation by adding cold-start data and multi-stage training pipelines before the RL phase, which helped address readability and language-mixing challenges seen in the zero variant. From this base, the team distilled six smaller models ranging from 1.5B to 70B parameters using Qwen and Llama architectures, and the 32B and 70B versions proved competitive with OpenAI-o1-mini on reasoning benchmarks. The open-source release under MIT licensing has made DeepSeek-R1 especially attractive to the research community, enabling local deployment on consumer hardware like RTX 4090 and Apple M3 Max. Even as the ecosystem has grown to include newer versions like DeepSeek-R1-0528 with improved benchmarks and capabilities like function calling and JSON output, the original R1 remains a widely deployed open-weight option for developers seeking strong reasoning without proprietary constraints.

Azure Cognitive Servicesdeepseek-r1deepseek-thinkingdeprecated

Quick Info

Powered by
Provider
Azure Cognitive Services
Model key
deepseek-r1
Release date
Jan 20, 2025
Last updated
Jan 20, 2025
Knowledge cutoff
2024-07
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.35
Output token cost
$5.40

Limits

Output tokens
163,840 tokens
Context window
163,840 tokens

Transparent token rates

Compare DeepSeek-R1 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek-R1

Vercel AI Gateway

CoverageBenchmark

A peer-reviewed meta-analysis published in the Journal of Big Data (Volume 13, article 26, 2026; published 19 December 2025) directly benchmarks DeepSeek-R1 against GPT-4 Turbo, Gemini Ultra, Qwen, and LLaMA 3.1 using standardized tasks including MMLU, HumanEval, FLORES-200, and TyDiQA. The hybrid meta-analysis aggrega Reported results for DeepSeek-R1 include 80.2 ± 1.5% on HumanEval and 78.5 ± 1.8% on MMLU, compared with ChatGPT-4 Turbo's 86.5 ± 1.9% on HumanEval, with the gap falling within observed heterogeneity (I² = 14.6%). The study concludes R1 demonstrates strong coding and multilingual efficiency, trails GPT-4 Turbo in reaso

Videos about DeepSeek-R1

More models around DeepSeek-R1