Anthropic 将开放系统给第三方评估以验证安全措施
I really enjoyed reading 75% of this letter and deeply agree with it - i have some doubts on the rem...
Anthropic 要开放系统给第三方评估,能更透明地验证安全,和之前只自己评估的方式不同。
Anthropic 公司计划提供永久、员工级别的系统访问权限给第三方评估者,让他们在模型训练过程中验证安全措施、报告事件并评估对齐情况。这被视为建立行业透明度和重建信任的一种方式。
I really enjoyed reading 75% of this letter and deeply agree with it - i have some doubts on the rem...
I really enjoyed reading 75% of this letter and deeply agree with it - i have some doubts on the remaining 25%. Third-party evaluators, if done right, are a great idea and an amazing way to establish more transparency. Maybe even rebuild some of the lost trust between labs, and between them and society! Excited about this The part I’m less convinced by is whether you can build great global cooperation on this topic by explicitly stating you want to design it to keep widening your own lead. That seems like a pretty counterproductive way to start the conversation to me. Dario Amodei @DarioAmodei We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must-p… 🔗 View Quoted Tweet 💬 22 🔄 4 ❤️ 94 👀 7008 📊 23 ⚡