Marcus质疑OpenAI安全报告可靠性

adding to my point that we just don’t have m the science here that we need

精选理由

Marcus和Tracy共同质疑OpenAI安全报告的充分性,指出其评估方法不透明且可能低估风险。

AI 摘要

GaryMarcus支持TylerTracy观点,认为OpenAI的安全报告无法证明能有效防止Astra模型规避安全系统。OpenAI使用难以从外部评估的内部评估方法,未详细说明人工升级/响应流程,且未充分测试Astra的攻击和规避能力。OpenAI过去曾低估此类风险,Marcus对其当前安全措施缺乏信心。

原文 · Gary Marcus

adding to my point that we just don’t have m the science here that we need

adding to my point that we just don’t have m the science here that we need Tyler Tracy @tylertracy321 I don't think the public can conclude from OpenAI's safety report that, if Astra were trying to subvert its safety systems to, say, start a rogue deployment, OpenAI could sufficiently prevent it. This report suggests that Astra could be stealthy enough to do a bunch of bad things internally without anyone ever being alerted in time. But more importantly, I feel uncertain about this because OpenAI doesn't provide nearly enough evidence of this for us to accurately assess their safety measures. They use internal evals that are very hard to assess from the outside, and they don't explain their human escalation/response process in enough detail for us to understand if it is sufficient. It also seems to me that they didn't do much to elicit Astra's attack and stealth capabilities, so they might be underestimating the risk. They are relying heavily on the fact that they have measured the alignment of their model to be good enough that it won't try to evade them and that their security is good enough. But OpenAI has underestimated this in the past, and I don't feel confident, based on their public comms, that they have done better this time. 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 0 👀 730 ⚡

Marcus质疑OpenAI安全报告可靠性 · AITOP · AI热报