Model details
DeepSeek R1 0528 Qwen3 8B
DeepSeek R1 0528 Qwen3 8B is a compact reasoning powerhouse built by distilling the chain-of-thought capabilities of DeepSeek-R1-0528 into the Qwen3 8B Base model. The goal was to pack state-of-the-art reasoning performance into a lightweight, efficient form factor that can run on modest hardware without sacrificing the depth-of-thought that made the parent model famous. The model is specifically optimized for tasks that demand multi-step logical chains, mathematical rigor, and programming problem-solving, making it especially useful for developers and researchers who need strong analytical capabilities but lack access to massive GPU clusters.
The lineage here is key: this model was created by continuing the post-training process on Qwen3 8B using Chain-of-Thought demonstrations harvested from the larger DeepSeek-R1-0528. That larger sibling had already pushed reasoning and inference to the brink of flagship models like O3 and Gemini 2.5 Pro through increased compute and algorithmic refinement during post-training. By distilling those capabilities down, the 8B variant achieves roughly a 10 percentage point improvement over standard Qwen3 8B while actually matching the reasoning scores of the 235-billion-parameter Qwen3-235B-thinking giant on the AIME 2024 benchmark. It sets a new state-of-the-art among open-source models of comparable size, offering a compelling balance of capability, efficiency, and local deployment flexibility.
Quick Info
Powered by- Provider
- NovitaAI
- Model key
- deepseek/deepseek-r1-0528-qwen3-8b
- Release date
- May 29, 2025
- Last updated
- May 29, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.06
- Output token cost
- $0.09
Limits
- Output tokens
- 32,000 tokens
- Context window
- 128,000 tokens
Latest news about DeepSeek R1 0528 Qwen3 8B
No articles yet. Fetch the latest news to show it here.