DeepSeek-R1 was developed as part of DeepSeek's first-generation reasoning line, alongside the sibling variant DeepSeek-R1-Zero. R1-Zero was trained using large-scale reinforcement learning directly on the base model without any preliminary supervised fine-tuning, which let reasoning behaviors emerge naturally but produced outputs with readability and language-mixing issues. R1 addresses those limitations through a multi-stage pipeline that starts with a cold-start phase, continues with reasoning-oriented reinforcement learning, then uses rejection sampling and supervised fine-tuning before a final reinforcement stage aimed at broader scenarios. The release also open-sourced six dense distilled models ranging from 1.5B to 70B parameters, built on Qwen and Llama bases, so smaller deployments can inherit part of R1's reasoning ability.
The hosted variant served through the OpenRouter router is a 671B-parameter mixture-of-experts architecture that activates around 37B parameters per inference pass, making it a heavyweight reasoning system rather than a lightweight chat model. On math, code, and general reasoning benchmarks reported by the DeepSeek team, R1 reaches performance comparable to OpenAI's o1-1217, and the May 28th revision tightens that comparison further while exposing its full reasoning traces for inspection. For practitioners it fits well in workflows that need chain-of-thought inspection, tool integration, and structured outputs on long contexts, and its open weights combined with permissive licensing make it attractive for teams that want to study or self-host a frontier-class reasoning model without paying closed-API premiums.