Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

DeepSeek-R1

DeepSeek-R1 is a reasoning-focused language model designed to tackle complex problem-solving tasks, particularly in mathematics, coding, and multi-step logical reasoning. The model was built on a foundation of 671 billion total parameters, with 37 billion active during each inference pass, giving it the capacity to develop and express extended reasoning chains. Unlike many contemporary models that rely primarily on supervised fine-tuning, DeepSeek-R1 emerged from a novel training lineage that began with DeepSeek-R1-Zero—a version trained entirely through large-scale reinforcement learning without any supervised fine-tuning step. This pure RL approach allowed the model to naturally develop sophisticated reasoning behaviors, though it sometimes struggled with readability and language consistency. The main DeepSeek-R1 model refined this foundation by incorporating cold-start data before its reinforcement learning phase, multi-stage training pipelines, and rejection sampling techniques to address those quality issues.

The model's development story reflects a deliberate push toward democratizing access to frontier-level reasoning capabilities. DeepSeek's team open-sourced both the full model weights and its technical report under the MIT license, enabling the research community to study, fine-tune, and build upon the work. A notable outcome of this openness was the creation of six distilled variants—ranging from 1.5B to 70B parameters—built on Qwen and Llama architectures, with the larger distilled models reportedly reaching performance levels comparable to OpenAI's o1-mini. An upgrade released in May 2025 further improved benchmark performance while reducing hallucinations and adding practical features like function calling and JSON output support. The combination of competitive reasoning performance, open-source accessibility, and significantly lower cost positioning makes DeepSeek-R1 particularly compelling for developers and organizations seeking to integrate advanced AI reasoning into applications without relying on proprietary API dependencies.

Vercel AI Gatewaydeepseek/deepseek-r1deepseek-thinking

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
deepseek/deepseek-r1
Release date
Jan 20, 2025
Last updated
May 29, 2025
Knowledge cutoff
2024-07
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.35
Output token cost
$5.40

Limits

Output tokens
32,768 tokens
Context window
128,000 tokens

Transparent token rates

Compare DeepSeek-R1 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek-R1

Vercel AI Gateway

CoverageComparison

Leading AI models include GPT-4o, OpenAI o1, OpenAI o3-mini, Claude 3.7 Sonnet, Gemini 2.5 Pro, and DeepSeek-R1.

Vercel AI Gateway

CoverageBenchmark

Run DeepSeek R1 locally on RTX 4090 or M3 Max. Detailed benchmarks, quantization comparisons, token/s performance metrics, and setup guide for consumer GPUs.

Vercel AI Gateway

CoverageBenchmark

A peer-reviewed meta-analysis published in the Journal of Big Data (Volume 13, article 26, 2026; published 19 December 2025) directly benchmarks DeepSeek-R1 against GPT-4 Turbo, Gemini Ultra, Qwen, and LLaMA 3.1 using standardized tasks including MMLU, HumanEval, FLORES-200, and TyDiQA. The hybrid meta-analysis aggrega Reported results for DeepSeek-R1 include 80.2 ± 1.5% on HumanEval and 78.5 ± 1.8% on MMLU, compared with ChatGPT-4 Turbo's 86.5 ± 1.9% on HumanEval, with the gap falling within observed heterogeneity (I² = 14.6%). The study concludes R1 demonstrates strong coding and multilingual efficiency, trails GPT-4 Turbo in reaso

Videos about DeepSeek-R1

More models around DeepSeek-R1