Currently listed through these providers:
Model details
DeepSeek R1 Distill Qwen 1.5B
DeepSeek R1 Distill Qwen 1.5B is a compact, efficient language model built upon the Qwen 2.5 architecture. It is specifically designed to bring advanced reasoning capabilities to smaller, resource-constrained hardware environments. By focusing on a streamlined parameter count, the model serves as a practical solution for developers who need to perform complex logical and mathematical tasks without the heavy computational overhead typically associated with larger, dense language models.
The model is developed through a distillation process that transfers the knowledge and reasoning patterns of the larger DeepSeek R1 model into a more portable 1.5B parameter framework. This lineage allows the model to maintain strong performance in specialized benchmarks, effectively inheriting the reasoning strengths of its larger counterpart. Its design makes it highly suitable for local deployment and edge computing, where it can be further optimized through quantization methods to balance memory footprint with inference speed and reasoning accuracy.
Quick Info
Powered by- Provider
- Alibaba (China)
- Model key
- deepseek-r1-distill-qwen-1-5b
- Release date
- Jan 1, 2025
- Last updated
- Jan 1, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 16,384 tokens
- Context window
- 32,768 tokens