Alibaba (China)
R1 Distill Qwen 14B pricing: $0.15/M input, $0.15/M output. Compare with 10 similar models, see benchmarks, and find the cheapest provider.
Model details
DeepSeek R1 Distill Qwen 14B is a 14-billion parameter model distilled from DeepSeek R1, built on a Qwen base architecture so that chain-of-thought patterns from the larger reasoning parent transfer into a more efficient package. The distillation approach aims to bring reasoning-grade behavior to a mid-sized footprint, sitting between lightweight models and much larger 70B-class variants. Released in January 2025 under the MIT license, the model is positioned for developers who want stronger logical and mathematical performance without the hardware demands of flagship reasoning systems.
In benchmark evaluations, the model reaches 69.7% on AIME 2024 and 93.9% on MATH-500, illustrating effective knowledge transfer in the 14B range, with additional reporting of 74% on MMLU-Pro and 48% on GPQA Diamond alongside 56% on AIME 2025. It operates within a roughly 32K-to-33K token context window with up to 16K tokens of output, supporting long-form chain-of-thought responses. Practically, the model fits resource-constrained deployments and applications that need moderate reasoning, tool use, and structured output without the latency or cost profile of larger alternatives.
Alibaba (China)
R1 Distill Qwen 14B pricing: $0.15/M input, $0.15/M output. Compare with 10 similar models, see benchmarks, and find the cheapest provider.