Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
NovitaAI logo

Model details

DeepSeek R1 0528

DeepSeek R1 0528 represents an evolution within the DeepSeek R1 family, built upon the DeepSeek-V3 architecture and enhanced with FP8 quantization for efficient complex computation. The model employs a mixture-of-experts architecture with multi-headed latent attention and multi-token prediction capabilities, enabling it to tackle involved reasoning tasks by routing specialized capacity only to the experts needed for each problem. This design allows a 671B parameter model to keep only 37B parameters active per forward pass, balancing capability against computational efficiency. The architecture is deliberately shaped for strength in mathematical reasoning, code generation, and multi-step logical problem solving.

Unlike models trained purely through supervised fine-tuning, DeepSeek R1 0528 puts reinforcement learning at the heart of its development, letting the model improve through trial and error rather than relying solely on pre-labeled data. This approach yields self-improving reasoning capabilities and contributes to the model's notably reduced hallucination rates compared to earlier iterations. The May 2025 release brought measurable benchmark improvements, enhanced front-end performance, JSON output support, and native function calling—making the model practical for developers building structured workflows and tool-augmented applications. Its performance now approaches leading closed models on reasoning and math tasks, while fully open weights let the research community inspect, fine-tune, and build upon the model.

NovitaAIdeepseek/deepseek-r1-0528deepseek-thinking

Quick Info

Powered by
Provider
NovitaAI
Model key
deepseek/deepseek-r1-0528
Release date
May 28, 2025
Last updated
May 28, 2025
Knowledge cutoff
2024-07
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.70
Output token cost
$2.50

Limits

Output tokens
32,768 tokens
Context window
163,840 tokens

Transparent token rates

Compare DeepSeek R1 0528 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek R1 0528

Cortecs

CoverageAnalysis

A third-party deep dive analyzes DeepSeek-R1-0528 as more than a cosmetic patch, detailing that it retains the Mixture-of-Experts architecture scaled up with 128K context (extendable via RoPE scaling) while delivering large mathematical-reasoning improvements. The post highlights the AIME 2025 jump from 70% to 87.5% as The piece contextualizes R1-0528 within DeepSeek's open-source trajectory and notes improvements relevant to vibe coding and function calling. It frames the release as a stealth incremental update that nonetheless shifts the competitive balance against closed frontier models in mathematical and multi-step reasoning wor

Videos about DeepSeek R1 0528

More models around DeepSeek R1 0528