模型多源确认78°

OpenAI 发布模型对齐问题追踪与披露框架

a step in the good direction

精选理由

OpenAI 出了新东西,是关于模型安全问题的追踪和报告框架,比之前更透明了。

OpenAI 公布了新的框架来追踪、调查和公开其模型对齐问题,设定了披露标准和时间线。框架包括六份关于过去半年模型对齐行为的报告,优先处理揭示新机制或挑战安全假设的案例。

原文 · Thomas Wolf

a step in the good direction

a step in the good direction OpenAI @OpenAI We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties. We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation. Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months. This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis. openai.com/index/model-mi… 🔗 View Quoted Tweet 💬 4 🔄 0 ❤️ 9 👀 1385 📊 3 ⚡