Sulat.com
AI models
Get 10-25% off
Get 10-25% off from Qwen
Alibaba (China) logo

Model details

DeepSeek R1 Distill Qwen 14B

DeepSeek R1 Distill Qwen 14B is a 14-billion parameter model distilled from DeepSeek R1, built on a Qwen base architecture so that chain-of-thought patterns from the larger reasoning parent transfer into a more efficient package. The distillation approach aims to bring reasoning-grade behavior to a mid-sized footprint, sitting between lightweight models and much larger 70B-class variants. Released in January 2025 under the MIT license, the model is positioned for developers who want stronger logical and mathematical performance without the hardware demands of flagship reasoning systems.

In benchmark evaluations, the model reaches 69.7% on AIME 2024 and 93.9% on MATH-500, illustrating effective knowledge transfer in the 14B range, with additional reporting of 74% on MMLU-Pro and 48% on GPQA Diamond alongside 56% on AIME 2025. It operates within a roughly 32K-to-33K token context window with up to 16K tokens of output, supporting long-form chain-of-thought responses. Practically, the model fits resource-constrained deployments and applications that need moderate reasoning, tool use, and structured output without the latency or cost profile of larger alternatives.

Alibaba (China)deepseek-r1-distill-qwen-14bqwen

Quick Info

Powered by
Provider
Alibaba (China)
Model key
deepseek-r1-distill-qwen-14b
Release date
Jan 1, 2025
Last updated
Jan 1, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.144
Output token cost
$0.431

Limits

Output tokens
16,384 tokens
Context window
32,768 tokens

Latest news about DeepSeek R1 Distill Qwen 14B

Alibaba (China)

CoverageBenchmark

R1 Distill Qwen 14B pricing: $0.15/M input, $0.15/M output. Compare with 10 similar models, see benchmarks, and find the cheapest provider.

Videos about DeepSeek R1 Distill Qwen 14B

More models around DeepSeek R1 Distill Qwen 14B