The Information 播客第4期:与 Stuart Russell 聊 AI 对齐问题
Stuart Russell 上播客聊对齐,讲清楚人类反馈为啥会奖励错误行为,还聊了关机开关怎么设计,视频带时间戳可以挑着看。
The Information 的 AI Deep Dive 第4期邀请 UC Berkeley 计算机科学教授 Stuart Russell 与 rocketalignment 对谈。节目从对齐问题讲起,解释 AI 如何习得目标,以及为什么人类反馈可能奖励错误行为。对谈涵盖第7分18秒起的失齐表现、第26分05秒起的人类反馈能否修复对齐,以及第58分06秒的 assistance games 和关机开关设计。
🚀 AI Deep Dive Episode 4: Can We Solve AI’s Alignment Problem?
We want AI to follow instructions. But what if our instructions are the problem?
@rocketalignment and UC Berkeley Computer Science Professor Stuart Russell explore how AI learns its goals, why human feedback can reward the wrong behavior and what it would take to prove a powerful system is safe.
00:00 – Intro 00:43 – The alignment problem 07:18 – How AI misalignment shows up 16:07 – Training AI to imitate humans 26:05 – Can human feedback fix alignment? 37:47 – When humans become the obstacle 40:04 – Did AI take a wrong turn? 44:24 – AI & existential risk 51:09 – AI labs & safety evidence 58:06 – Assistance games & the off switch