# Liang Wenfeng's $294K DeepSeek-R1 RL Breakthrough Reached Nature _How Liang Wenfeng's DeepSeek-R1 used Group Relative Policy Optimization and pure reinforcement learning to produce emergent reasoning capabilities for $294,000 in training compute, and why the paper reached the cover of Nature._ **Published:** 2026-06-30 **Source:** https://www.startuphub.ai/ai-news/ai-figures/2026/figure-liang-wenfeng-deepseek-r1-technical-contribution-2026-06-30 --- **A $294,000 training run with no human-labelled reasoning data produced DeepSeek-R1, the paper that subsequently reached the cover of Nature. Liang Wenfeng, co-founder and CEO of DeepSeek, is listed as the corresponding author. The paper, published on arXiv in January 2025 under the title [DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning](https://arxiv.org/abs/2501.12948), demonstrated that a language model trained on pure reinforcement learning signals, without any supervised fine-tuning on human-curated reasoning chains, could produce world-class mathematical and coding reasoning.** Quick context Liang Wenfeng co-founded DeepSeek as the AI research arm of High-Flyer Capital, a Chinese quantitative hedge fund. DeepSeek's earlier model, DeepSeek-V3, was [reported to have a $5.6M training budget](/ai-news/ai-figures/2026/figure-liang-wenfeng-deepseek-financial-breakdown-2026-06-09) and served as the base on which the R1 work was built. This piece focuses on the specific technical architecture of R1 and what made it significant enough for *Nature*'s editors. ## Training without a teacher: how GRPO replaced supervised fine-tuning The standard large language model training recipe in 2024 followed three stages: pre-train a base model on large text corpora; fine-tune it on human-curated examples with step-by-step reasoning (supervised fine-tuning, or SFT); then apply reinforcement learning, typically Proximal Policy Optimization (PPO), to align outputs further. The SFT stage for reasoning required expensive chains of thought labelled by humans or distilled from a more capable teacher model. DeepSeek-R1-Zero, the experimental variant described in the paper, removed the SFT stage entirely. Starting from DeepSeek-V3-Base, the team applied reinforcement learning directly using a custom algorithm, Group Relative Policy Optimization (GRPO). GRPO's structural advantage over PPO is that it does not require a separate evaluator model of the same scale as the policy model. Instead, it samples a group of responses and uses relative performance within that group as the reward signal. As the [paper states](https://arxiv.org/abs/2501.12948), GRPO "directly estimates the baseline from the group scores," cutting the compute footprint of the RL stage substantially. The reward signal itself was minimal: positive feedback for a correct final answer on math or coding problems, and for adherence to the required output format (thinking wrapped in tags, final answer in a separate tag). No partial credit, no reasoning-quality scoring, no human feedback at any stage. The cost implication is direct. DeepSeek-V3's training was [reported by 36kr](https://eu.36kr.com/en/p/3471471431046531) to have run to approximately $5.6 million; the R1 training run cost $294,000, per the same reporting citing the DeepSeek team. The savings came from removing the SFT stage and eliminating the need for a matching-scale evaluator in the RL phase. ![](https://quickchart.io/chart?width=600&height=300&c=%7B%22type%22%3A%22bar%22%2C%22data%22%3A%7B%22labels%22%3A%5B%22DeepSeek-V3%20%28Dec%202024%29%22%2C%22DeepSeek-R1%20%28Jan%202025%29%22%5D%2C%22datasets%22%3A%5B%7B%22label%22%3A%22Training%20Compute%20Cost%20%28USD%29%22%2C%22data%22%3A%5B5600000%2C294000%5D%2C%22backgroundColor%22%3A%5B%22%236366f1%22%2C%22%2310b981%22%5D%7D%5D%7D%2C%22options%22%3A%7B%22plugins%22%3A%7B%22title%22%3A%7B%22display%22%3Atrue%2C%22text%22%3A%22DeepSeek%20Training%20Costs%3A%20V3%20vs%20R1%22%7D%7D%2C%22scales%22%3A%7B%22y%22%3A%7B%22beginAtZero%22%3Atrue%2C%22title%22%3A%7B%22display%22%3Atrue%2C%22text%22%3A%22USD%22%7D%7D%7D%7D%7D) DeepSeek training compute costs: V3 ($5.6M) versus R1 ($294K). Source: [36kr](https://eu.36kr.com/en/p/3471471431046531); [DeepSeek-R1 paper](https://arxiv.org/abs/2501.12948). ## The 'aha moment' and what Nature's editors recognised The most notable result in the paper was not the cost reduction. It was the behaviour the model developed without being instructed to. DeepSeek-R1-Zero, during RL training, began allocating more reasoning tokens to harder problems and started explicitly noting within its chain of thought when an earlier step was incorrect and correcting it mid-chain. The paper documents the training checkpoint where this first occurs as an "aha moment." The four emergent behaviours, none of them specified in the reward function, were: self-reflection, self-verification, dynamic strategy adaptation, and active exploration of alternative solution approaches. The significance of this result for *Nature*'s editors was its bearing on a theoretical question that had been debated since 2023: whether reinforcement learning alone, applied to a base language model with no reasoning scaffolding, could induce the kind of structured self-correction that had previously required explicit chain-of-thought supervision. [36kr reported](https://eu.36kr.com/en/p/3471061555959170) that the DeepSeek-R1 paper reached the cover of *Nature*, with Liang Wenfeng as the corresponding author; the team subsequently addressed public questions about reproducibility, per the same reporting. ![](https://quickchart.io/chart?width=600&height=300&c=%7B%22type%22%3A%22doughnut%22%2C%22data%22%3A%7B%22labels%22%3A%5B%22Self-reflection%22%2C%22Self-verification%22%2C%22Dynamic%20strategy%22%2C%22Exploration%22%5D%2C%22datasets%22%3A%5B%7B%22data%22%3A%5B25%2C25%2C25%2C25%5D%2C%22backgroundColor%22%3A%5B%22%236366f1%22%2C%22%2310b981%22%2C%22%23f59e0b%22%2C%22%23ef4444%22%5D%7D%5D%7D%2C%22options%22%3A%7B%22plugins%22%3A%7B%22title%22%3A%7B%22display%22%3Atrue%2C%22text%22%3A%22Emergent%20RL%20Behaviors%20in%20DeepSeek-R1-Zero%22%7D%2C%22legend%22%3A%7B%22position%22%3A%22bottom%22%7D%7D%7D%7D) Four emergent behaviors identified in DeepSeek-R1-Zero during RL training; equal segments reflect the paper's qualitative treatment, not a severity ranking. Source: [DeepSeek-R1 arXiv paper](https://arxiv.org/abs/2501.12948). ## Distilled models and the open-weight downstream effect Alongside the R1 model, the January 2025 release included a suite of smaller open-weight models built via knowledge distillation: R1-Distill-Qwen-7B, R1-Distill-Qwen-14B, R1-Distill-Qwen-32B, and R1-Distill-Llama-70B. These were not trained with RL from scratch; instead, R1's reasoning traces were used as training data for smaller base models, enabling them to carry the chain-of-thought format into more deployable sizes. The open-weight release continued DeepSeek's practice with its earlier models. Within weeks of publication, GRPO was incorporated into open-source training frameworks and researchers at other labs began publishing variants of the RL-only training approach. The [contrast with the closed-lab model](/ai-news/ai-figures/2026/figure-ilya-sutskever-ssi-vs-openai-safety-strategy-2026-06-24) championed by labs like Safe Superintelligence drew attention in the research community, where the rapid replication of DeepSeek-R1's approach served as a practical test of the paper's claims. ![](https://quickchart.io/chart?width=600&height=300&c=%7B%22type%22%3A%22horizontalBar%22%2C%22data%22%3A%7B%22labels%22%3A%5B%22R1-Distill-Qwen-7B%22%2C%22R1-Distill-Qwen-14B%22%2C%22R1-Distill-Qwen-32B%22%2C%22R1-Distill-Llama-70B%22%5D%2C%22datasets%22%3A%5B%7B%22label%22%3A%22Parameters%20%28billions%29%22%2C%22data%22%3A%5B7%2C14%2C32%2C70%5D%2C%22backgroundColor%22%3A%5B%22%2310b981%22%2C%22%2310b981%22%2C%22%236366f1%22%2C%22%236366f1%22%5D%7D%5D%7D%2C%22options%22%3A%7B%22plugins%22%3A%7B%22title%22%3A%7B%22display%22%3Atrue%2C%22text%22%3A%22DeepSeek-R1%20Distilled%20Open-Weight%20Models%20%28Jan%202025%29%22%7D%7D%2C%22scales%22%3A%7B%22x%22%3A%7B%22beginAtZero%22%3Atrue%2C%22title%22%3A%7B%22display%22%3Atrue%2C%22text%22%3A%22Parameters%20%28B%29%22%7D%7D%7D%7D%7D) DeepSeek-R1 distilled open-weight models by parameter count, released January 2025. Source: [DeepSeek-R1 arXiv paper](https://arxiv.org/abs/2501.12948). ## What it means DeepSeek-R1 is notable for two results that are easy to conflate. The first is cost: a $294,000 training run producing reasoning capabilities that matched models costing far more. The second is the scientific result: demonstrating that pure RL, without curated reasoning data, produces self-correcting behaviour in language models, and that GRPO offers a computationally lighter path than PPO for that RL stage. The first fact made headlines in early 2025; the second is what *Nature* published. Both represent Liang Wenfeng's specific contribution as the corresponding author and technical leader of the project. The subsequent replication by other labs, and the incorporation of GRPO into open-source training pipelines, is the standard measure of a research result that holds. ## Sources - ![](https://www.google.com/s2/favicons?domain=arxiv.org&sz=32) [arXiv: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning](https://arxiv.org/abs/2501.12948) - ![](https://www.google.com/s2/favicons?domain=36kr.com&sz=32) [36kr: DeepSeek on Nature's Cover -- Liang Wenfeng Leads Team to Address Doubts, R1 Training Costs $294,000](https://eu.36kr.com/en/p/3471471431046531) - ![](https://www.google.com/s2/favicons?domain=36kr.com&sz=32) [36kr: DeepSeek-R1 Paper Makes It to Nature's Cover, Corresponding Author: LIANG Wenfeng](https://eu.36kr.com/en/p/3471061555959170) - ![](https://www.google.com/s2/favicons?domain=en.wikipedia.org&sz=32) [Wikipedia: DeepSeek (biographical background)](https://en.wikipedia.org/wiki/DeepSeek) Editorial standards: every claim is sourced. Tips: editor@startuphub.ai --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.