D.Run (China)
Run DeepSeek R1 locally on RTX 4090 or M3 Max. Detailed benchmarks, quantization comparisons, token/s performance metrics, and setup guide for consumer GPUs.
Model details
DeepSeek R1 belongs to a first generation of "thinking" or reasoning-oriented language models from DeepSeek-AI, built to tackle multi-step inference, logic, and analytical problem solving rather than simple text generation. The model was introduced in the DeepSeek-R1 paper alongside DeepSeek-R1-Zero, an RL-only variant that demonstrated that strong reasoning behaviors could emerge purely from large-scale reinforcement learning, though it suffered from readability and language-mixing problems. DeepSeek R1 itself addresses those issues through a multi-stage pipeline that adds a cold-start data phase before reinforcement learning, producing cleaner outputs while preserving the emergent reasoning strengths. The same release also open-sourced six distilled dense checkpoints at 1.5B, 7B, 8B, 14B, 32B, and 70B parameters, based on Qwen and Llama base models, giving developers a spectrum of sizes for different deployment scenarios.
On reasoning benchmarks the DeepSeek-R1 paper reports performance comparable to OpenAI's o1-1217 model, positioning R1 as a credible open-weights alternative to closed frontier systems for analytical workloads. Because the weights are openly available, third-party guides have shown that R1 can be run locally on consumer hardware such as NVIDIA RTX 4090 and Apple M3 Max systems, with quantization choices trading off memory footprint against tokens-per-second throughput. That combination of strong reported reasoning accuracy, flexible deployment, and a family of distilled sizes makes DeepSeek R1 a practical fit for research experimentation, cost-sensitive production reasoning tasks, and on-device or private-cloud use cases where sending prompts to a proprietary API is not desirable.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
D.Run (China)
Run DeepSeek R1 locally on RTX 4090 or M3 Max. Detailed benchmarks, quantization comparisons, token/s performance metrics, and setup guide for consumer GPUs.