DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang et al. · Nature volume 645, pages 633-638 (2025) · 2025
Abstract
Shows that extended reasoning can be learned from reinforcement learning against verifiable answers alone, with no supervised examples of reasoning steps. Left to optimise, the model lengthens its own chains of thought and develops self-correction unprompted. A second version adds a small supervised warm start to fix readability and language mixing, and the behaviour distils into much smaller models.
Why it matters
The result that reframed reasoning as something trained rather than prompted. That the behaviour emerges from outcome-only rewards — nobody demonstrated the reasoning — is the part worth sitting with.
https://arxiv.org/abs/2501.12948