Sulat.com
AI models
Get 10-25% off
Get 10-25% off from Qwen
Alibaba (China) logo

Model details

DeepSeek R1 0528

DeepSeek R1 0528 is a large-scale reasoning model built on a 671-billion-parameter Mixture-of-Experts architecture that activates 37 billion parameters per forward pass during inference. Rather than training on curated human-written reasoning chains, DeepSeek applied reinforcement learning directly to the base DeepSeek-V3 model weights, allowing behaviors such as self-verification, self-reflection, and extended chain-of-thought reasoning to emerge organically. The model's design prioritizes transparent, open reasoning tokens rather than the locked outputs common in proprietary competitors, and it ships under the permissive MIT License for commercial use. This architectural philosophy supports complex, multi-step problem-solving without requiring developers to coax out extended thinking processes.

The May 2025 update to R1 brought measurable gains in reasoning depth, pushing AIME 2025 accuracy from approximately 70% to 87.5% while increasing average tokens-per-question from 12K to 23K for deeper deliberation. Benchmarks across mathematical reasoning (MMLU-Redux at 93.4, GPQA-Diamond at 81.0), live coding (LiveCodeBench at 73.3), and competition mathematics (AIME 2025 at 87.5) position the model near leading closed systems. Notably, DeepSeek distilled the chain-of-thought patterns from this model to train smaller variants—including the DeepSeek-R1-0528-Qwen3-8B—which achieved state-of-the-art results among open-source models on AIME 2024. The model now supports system prompts, enhanced function calling, and JSON output, making it practical for applications ranging from advanced mathematical problem-solving to vibe coding workflows.

Alibaba (China)deepseek-r1-0528deepseek-thinking

Quick Info

Powered by
Provider
Alibaba (China)
Model key
deepseek-r1-0528
Release date
May 28, 2025
Last updated
May 28, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.574
Output token cost
$2.294

Limits

Output tokens
16,384 tokens
Context window
131,072 tokens

Transparent token rates

Compare DeepSeek R1 0528 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek R1 0528

Cortecs

CoverageAnalysis

A third-party deep dive analyzes DeepSeek-R1-0528 as more than a cosmetic patch, detailing that it retains the Mixture-of-Experts architecture scaled up with 128K context (extendable via RoPE scaling) while delivering large mathematical-reasoning improvements. The post highlights the AIME 2025 jump from 70% to 87.5% as The piece contextualizes R1-0528 within DeepSeek's open-source trajectory and notes improvements relevant to vibe coding and function calling. It frames the release as a stealth incremental update that nonetheless shifts the competitive balance against closed frontier models in mathematical and multi-step reasoning wor

Videos about DeepSeek R1 0528

More models around DeepSeek R1 0528