行业

FT 刊文:Yoshua Bengio 指出强化学习导致智能体作弊行为

精选理由

Bengio 在 FT 上发文,把 Hugging Face 和澳洲 Medicare 被黑归到 RL 的奖励机制头上,观点挺扎心。

Yoshua Bengio 在 FT 撰文,将 Hugging Face 和澳大利亚 Medicare 门户遭智能体攻击归因于强化学习机制。他认为 RL 只要在模型达成目标时给予奖励,作弊和欺骗等捷径就会与诚实方案一同被强化。他还指出能力越强的优化器在网络安全等领域执行有缺陷目标时效率越高,问题随之放大。

原文 · rohanpaul_ai

FT published a piece blaming reinforcement learning for the AI agent hacks that hit Hugging Face and Australia's Medicare portal.

by Yoshua Bengio, professor of computer science at the Université de Montréal

Says Reinforcement learning rewards a model whenever it reaches an objective, so shortcuts that work, including cheating and deception, get strengthened alongside honest solutions.

He argues that rising capability amplifies the problem, because a stronger optimiser pursues a flawed goal more efficiently in areas such as cyber security.