SiliconFlow (China)
Run DeepSeek R1 locally on RTX 4090 or M3 Max. Detailed benchmarks, quantization comparisons, token/s performance metrics, and setup guide for consumer GPUs.
Model details
DeepSeek-R1 represents a significant shift in model design, focusing on native reasoning capabilities rather than relying solely on traditional supervised fine-tuning. The architecture is built to handle complex, multi-step problem solving, with a design intent that prioritizes deep logical processing for tasks like mathematics and programming. By moving beyond standard training paradigms, the model is engineered to perform at a level comparable to leading industry benchmarks, offering a robust framework for users who require high-fidelity reasoning and structured output in their technical workflows.
The model's lineage is rooted in large-scale reinforcement learning, which allows it to develop sophisticated reasoning behaviors autonomously. While early iterations like DeepSeek-R1-Zero demonstrated the potential of this approach, they faced challenges with readability and repetition. The final version addresses these through refined training methods, resulting in a more stable and readable output. This advancement has also enabled the creation of smaller, distilled versions that maintain high performance, making the technology more accessible for diverse applications and empowering developers to leverage these reasoning strengths in a wider range of environments.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
SiliconFlow (China)
Run DeepSeek R1 locally on RTX 4090 or M3 Max. Detailed benchmarks, quantization comparisons, token/s performance metrics, and setup guide for consumer GPUs.
SiliconFlow (China)
A peer-reviewed benchmark published in Clinics (Volume 81, 2026; DOI 10.1016/j.clinsp.2026.101021) evaluated DeepSeek-R1 on 321 text-based USMLE-style questions, where it achieved 92.5% overall accuracy (95% CI 89.1%–94.9%), significantly outperforming three OpenAI models — GPT-4 Omni, OpenAI o3-mini, and OpenAI o1 pro DeepSeek-R1 also surpassed reported average human examinee performance across all three USMLE steps and showed markedly stronger reasoning on discordant questions, achieving 82.8% accuracy versus 14.1%–28.1% for the OpenAI models (p < 0.0001). Inter-model consensus between DeepSeek-R1 and OpenAI o1 pro reached 94.9% ac