【极致中配】哈佛大学 CS2881R 人工智能安全|2025年秋季
哈佛2025秋季AI安全研讨课,涵盖LLM安全训练、鲁棒性、模型规范、递归自我提升、欺骗行为、机制可解释性及AI经济社会影响等前沿议题。
- 难度
- 难度 4/5 — 前沿AI安全研究生研讨课,需机器学习与大语言模型基础
- 适合人群
- 关注AI安全与对齐的研究生、研究者及从业者
- 前置要求
- 机器学习基础(了解监督学习、神经网络)、大语言模型原理(熟悉LLM训练与推理)、深度学习(Transformer等架构)、概率与统计(有助于理解评估与实验)
- 课程规模
- 15 讲 · 1427播放
主题覆盖
LLM安全训练鲁棒性模型规范策略合规递归自我提升欺骗行为AI经济影响机制可解释性心理健康与情感依赖AI未来展望GDPval学生项目展示
课程大纲(15 讲)
- P1 · Lecture 185 分钟
- P2 · Lecture 2- Modern LLM training and safety training81 分钟
- P3 · Lecture 3: Robustness56 分钟
- P4 · Lecture 4 Model Specs74 分钟
- P5 · Lecture 5: Experiment on Policy compliance9 分钟
- P6 · Lecture 6: Recursive Self Improvement85 分钟
- P7 · Lecture 7: Lab vs Field: Guest lecture by Joel Becker70 分钟
- P8 · Lecture 8: Scheming5 分钟
- P9 · Lecture 8 Experiment: Under Pressure20 分钟
- P10 · Lecture 9: Economic Impacts of AI95 分钟
- P11 · Lecture 10: Mechanistic Intepretability95 分钟
- P12 · Lecture 11: Mental Health and Emotional Attachment45 分钟
- P13 · When the earth dies (codex clip)1 分钟
- P14 · Lecture 12: AI in 2035 and GDPval68 分钟
- P15 · Oral presentations of student projects48 分钟
本课程卡由 AI 生成,可能存在误差,欢迎反馈。