跳到主要内容
步芽

【极致中配】哈佛大学 CS2881R 人工智能安全|2025年秋季

哈佛2025秋季AI安全研讨课,涵盖LLM安全训练、鲁棒性、模型规范、递归自我提升、欺骗行为、机制可解释性及AI经济社会影响等前沿议题。

难度
难度 4/5前沿AI安全研究生研讨课,需机器学习与大语言模型基础
适合人群
关注AI安全与对齐的研究生、研究者及从业者
前置要求
机器学习基础(了解监督学习、神经网络)、大语言模型原理(熟悉LLM训练与推理)、深度学习(Transformer等架构)、概率与统计(有助于理解评估与实验)
课程规模
15 · 1427播放

主题覆盖

LLM安全训练鲁棒性模型规范策略合规递归自我提升欺骗行为AI经济影响机制可解释性心理健康与情感依赖AI未来展望GDPval学生项目展示

课程大纲(15 讲)

  1. P1 · Lecture 185 分钟
  2. P2 · Lecture 2- Modern LLM training and safety training81 分钟
  3. P3 · Lecture 3: Robustness56 分钟
  4. P4 · Lecture 4 Model Specs74 分钟
  5. P5 · Lecture 5: Experiment on Policy compliance9 分钟
  6. P6 · Lecture 6: Recursive Self Improvement85 分钟
  7. P7 · Lecture 7: Lab vs Field: Guest lecture by Joel Becker70 分钟
  8. P8 · Lecture 8: Scheming5 分钟
  9. P9 · Lecture 8 Experiment: Under Pressure20 分钟
  10. P10 · Lecture 9: Economic Impacts of AI95 分钟
  11. P11 · Lecture 10: Mechanistic Intepretability95 分钟
  12. P12 · Lecture 11: Mental Health and Emotional Attachment45 分钟
  13. P13 · When the earth dies (codex clip)1 分钟
  14. P14 · Lecture 12: AI in 2035 and GDPval68 分钟
  15. P15 · Oral presentations of student projects48 分钟

本课程卡由 AI 生成,可能存在误差,欢迎反馈。