Abacus
Run DeepSeek R1 locally on RTX 4090 or M3 Max. Detailed benchmarks, quantization comparisons, token/s performance metrics, and setup guide for consumer GPUs.
Model details
DeepSeek R1 belongs to a new generation of AI systems often described as "thinking models," designed to break down complex problems, weigh alternatives, and produce answers through visible chains of reasoning rather than single-shot responses. By exposing its internal deliberation, the model aims to tackle challenges in logical inference, mathematics, and multi-step analysis where conventional text generators tend to falter. Its release framed it as a competitive open alternative to leading proprietary reasoning systems, bringing strategic, human-like problem solving into a fully open-weights package that developers can study, fine-tune, or run locally on capable hardware ranging from high-end consumer GPUs to workstations.
The model's lineage centers on a two-stage development path described in the DeepSeek team's technical paper. The first variant, DeepSeek-R1-Zero, was trained through large-scale reinforcement learning without a supervised fine-tuning step, which allowed chain-of-thought behavior and other reasoning patterns to emerge naturally. To fix readability and language-mixing issues, the flagship DeepSeek-R1 added cold-start data, reasoning-oriented reinforcement learning, rejection sampling, and a second supervised fine-tuning pass before a final reinforcement learning stage tuned for general scenarios. The team also open-sourced six dense distilled models ranging from 1.5B to 70B parameters based on Qwen and Llama, using 800k samples from R1 to bring reasoning ability to smaller footprints. Independent summaries report benchmark parity with OpenAI-o1-1217, including 79.8% Pass@1 on AIME 2024 and 97.3% on MATH-500, alongside strong Codeforces ratings, making the model a flexible option for coding assistants, tutoring, research support, and long-context analysis in both hosted and locally quantized deployments.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Abacus
Run DeepSeek R1 locally on RTX 4090 or M3 Max. Detailed benchmarks, quantization comparisons, token/s performance metrics, and setup guide for consumer GPUs.
Abacus
DeepSeek’s New R1–0528: Performance Analysis and Benchmark Comparisons TL;DR: The newest DeepSeek R1 model is the most powerful among open-weight models, approaching performance of the leading …
Abacus
The official DeepSeek API changelog records two August 2026 updates relevant to the DeepSeek lineup that surrounds R1. On August 21, 2026 DeepSeek released the experimental multimodal model DeepSeek-V4-Flash-Vision-Exp (model name `deepseek-v4-flash-vision-exp`), with vendor-reported scores including Terminal Bench 2.1 On August 13, 2026 DeepSeek rolled out the GA of DeepSeek-V4-Pro on the APP, Web and API (model name `deepseek-v4-pro`), reporting gains in agent benchmarks (HLE 42.7/60.0, Terminal Bench 2.1 87.9, NL2Repo 61.5, Cybergym 83.3, DeepSWE 62.7, Toolathlon-Verified 74.1, Agents' Last Exam 25.7, AutomationBench Public 31.8,