Alexandr Wang:AI对齐问题尚无解决方案
Meta首席AI官谈AI对齐难题,提出用更智能的AI监控AI的解决方案。
Alexandr Wang表示,AI对齐是当前AI领域最开放的科学问题之一。他提出"可扩展监督"方案,即使用更智能的AI来监控其他AI的行为。Meta的Muse系统已采用此方案,通过独立的哨兵代理检查主代理行为。随着AI模型变得更智能,实验室需要构建越来越智能的监控代理。
Alexandr Wang ( @alexandr_wang, Chief AI Officer at Meta): nobody knows how to solve alignment yet
"This is, I think, one of the most open questions scientifically in AI. I think nobody knows exactly the way to solve this problem, but there’s a few ideas."
His solution is "scalable oversight":
“as the AIs get smarter, we use a different set of AIs to observe what they’re doing and keep them in check.”
The watcher AIs have to improve along with the models they watch, so labs would need to build “smarter and smarter policing agents” too.
Meta’s Muse already uses a version of this, with a separate sentinel agent checking what the main agent does.
---- Full video on "Cleo Abram" YouTube channel, (link in comment)