DeepSeek R1 Architecture: Reinforcement Learning Without Supervised Fine-Tuning
DeepSeek R1 demonstrates that pure reinforcement learning incentivizes complex reasoning behaviors, dramatically reducing training costs and democratizing open-weights AI.