一龍馬/AI 情報站讀懂消息背後的脈絡
星期日
搜尋

Sebastian Raschka 發布從零實作推理模型的第六集,主題是 RLVR 與 GRPO

中文摘要

內容涵蓋獎勵設計、GRPO 與 PPO 差異、KL 項、取樣與損失實作,並以 MATH 資料做訓練與評估。從貼文可確認的是教學大綱,實作細節需看影片本身。

一龍馬判讀

對想理解 DeepSeek-R1 類訓練流程的工程師與學生是實用教材,但只看大綱無法判斷程式正確性。

原文節錄

Sebastian Raschka · @rasbt

An introduction (and implementation) of Reinforcement Learning…

取得全文 · 不代表內容已獨立查證

查看原文
完整收錄文字與來源

Reasoning from scratch, round number 6! An introduction (and implementation) of Reinforcement Learning with Verifiable Rewards (RLVR) and Group Relative Policy Optimization (GRPO). 00:00 Introduction 01:54 What makes a reasoning model different? 04:25 Reasoning traces and model capability 08:29 Accuracy and format rewards 11:34 Aha moments and DeepSeek-R1 training 14:41 Reasoning effort and answer length 18:38 RLHF and RLVR 23:04 GRPO vs. PPO 26:40 GRPO explained with a cooking analogy 31:43 The KL term and simplified GRPO 35:04 Loading the pretrained model 36:07 Loading the MATH training data 39:26 Sampling model responses 46:30 Computing verifiable rewards 49:55 Computing advantages 51:54 Token and sequence log probabilities 55:29 Implementing sequence log probabilities 57:37 Fixing the inference-mode error 1:02:24 Computing the GRPO loss 1:04:37 Putting the GRPO step together 1:09:19 The GRPO training loop 1:12:57 Training settings, logging, and checkpoints 1:17:24 Running training and inspecting outputs 1:19:28 Loading and evaluating checkpoints 1:22:33 MATH-500 results and training stability 1:24:05 Memory requirements and next steps

收錄日期
2026-10-04
來源
Nitter RSS(公開貼文)
抓取時間
2026/10/04 08:14(台北)