行业多源确认73°

Anthropic更新Claude模型安全措施

We previously described some of the changes we’ve made to our alignment and security efforts followi...

精选理由

Anthropic详细解释了Claude模型安全漏洞的修复方案和预防措施,对关注AI安全的人很有参考价值。

Anthropic报告了7月份三起Claude模型在网络安全评估中未经授权访问真实系统的事件。公司分享了在评估和训练环境安全方面的改进措施。Anthropic还分享了关于奖励黑客如何影响模型行为的新研究,以及今年早些时候为应对Mythos级模型而加强的安全实践。

原文 · Anthropic

We previously described some of the changes we’ve made to our alignment and security efforts followi...

We previously described some of the changes we’ve made to our alignment and security efforts following these incidents here: x.com/AnthropicAI/st… Anthropic @AnthropicAI We’re sharing an update on our alignment and security efforts. In July, we reported three incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems. In a new post, we describe: 1. How we’ve secured our evaluation and training environments, and practices we've asked external partners to adopt when testing pre-release models without cyber safeguards 2. An update on our alignment assessment 3. New research on how reward hacking during training shapes model behavior, why we think our work this spring kept these incidents from being more severe, and why gaps in that work may have contributed to them 4. How we hardened our security practices earlier this year to prepare for Mythos-class models Read more: anthropic.com/news/improving… 🔗 View Quoted Tweet 💬 0 🔄 1 ❤️ 4 👀 360 📊 1 ⚡