跳到主要内容
步芽

【极致中配】2025年春季 斯坦福大学 CS224R 深度强化学习 Deep Reinforcement Learning

斯坦福CS224R系统讲授深度强化学习,从模仿学习、策略梯度、Q学习到离线RL、面向大模型的RL及机器人应用等前沿方向。

难度
难度 5/5研究生级课程,需扎实的机器学习、概率与深度学习基础,涵盖前沿RL方向
适合人群
有ML基础、想系统学习深度强化学习的研究生及研究者
前置要求
机器学习(监督学习、梯度下降等基础)、深度学习(神经网络与反向传播)、概率论与统计(马尔可夫决策过程、期望)、Python编程(熟悉PyTorch等框架)
课程规模
19 · 1.3 万播放

主题覆盖

模仿学习策略梯度Actor-CriticQ学习离线强化学习奖励学习大模型RL模型驱动RL多任务与元RL探索策略分层RL机器人强化学习

课程大纲(19 讲)

  1. P1 · Lecture 1: Class Intro32 分钟
  2. P2 · Lecture 2: Imitation Learning39 分钟
  3. P3 · Lecture 3: Policy Gradients37 分钟
  4. P4 · Lecture 4: Actor-Critic Methods38 分钟
  5. P5 · Lecture 5: Off-Policy Actor Critic42 分钟
  6. P6 · Lecture 6: Q-Learning42 分钟
  7. P7 · Lecture 7: Offline RL39 分钟
  8. P8 · Lecture 8: Reward Learning37 分钟
  9. P9 · Lecture 9: RL for LLMs40 分钟
  10. P10 · Lecture 10: RL for LLM Reasoning41 分钟
  11. P11 · Lecture 11: Model-Based RL40 分钟
  12. P12 · Lecture 12: Multi-Task RL38 分钟
  13. P13 · Lecture 13: Meta RL38 分钟
  14. P14 · Lecture 14: Exploration39 分钟
  15. P15 · Lecture 15: Hierarchical RL and IL37 分钟
  16. P16 · Lecture 16: RL for Robots36 分钟
  17. P17 · Lecture 17: Advancing Robot Intelligence31 分钟
  18. P18 · Lecture 18: Frontiers40 分钟
  19. P19 · Tutorial Session: Review of Q-Learning29 分钟

本课程卡由 AI 生成,可能存在误差,欢迎反馈。