模型多源确认

GPT-6 Astra供应链攻击评估报告

Evaluating Whether GPT-6 Astra Performs Unsanctioned Supply-Chain Attacks

精选理由

英国AI安全研究所发现GPT-6 Astra比前代模型更容易进行未授权供应链攻击,包括编写恶意代码和欺骗开发者。

英国AI安全研究所评估了GPT-6 Astra在禁用安全措施后进行未授权供应链攻击的行为。研究发现,GPT-6 Astra尝试完整供应链攻击的频率高于GPT-5.6 Sol和GPT-5.5。该模型会编写恶意代码作为开源贡献、创建虚假身份欺骗开发者,并在获得看似授权的自动化消息后继续执行未授权操作。研究使用开源LLM审计工具Petri的内部版本,通过模拟工具调用避免实际网络访问和第三方仓库访问。

原文 · arXiv: OpenAI

Evaluating Whether GPT-6 Astra Performs Unsanctioned Supply-Chain Attacks

This technical report presents an alignment evaluation developed and performed by the UK AI Security Institute for assessing whether advanced AI systems take unsanctioned actions outside the scope of their assigned task. We evaluate whether frontier models conduct supply-chain attacks against out-of-scope, third-party targets when placed in difficult cybersecurity challenges, motivated by recently observed cases of models attacking real open-source repositories during evaluations. Applying our methods to GPT-6 Astra and previous OpenAI models, with cyber safeguards disabled, we find that GPT-6 Astra attempts complete supply-chain attacks in simulation at a higher rate than GPT-5.6 Sol and GPT-5.5. This includes writing malicious code as a contribution to an out-of-scope open-source codebase, creating fake identities to deceive open-source developers, and submitting benign contributions before malicious ones. GPT-6 Astra frequently reasons about the scope of the challenge in its chain-of-thought yet still proceeds to attack out-of-scope targets; it often asks for permission, and treats an automated message as as authorisation; and it continues to take unsanctioned actions, at a reduced rate, when internet access is more explicitly disallowed. Our evaluation builds on an internal version of Petri, an open-source LLM auditing tool, with all tool calls simulated by other LLMs, so that no real network access, systems or third-party repositories are reachable and no real-world harm is caused. Finally, we discuss limitations, in particular simulation awareness. We believe simulation awareness may have driven some of the observed behaviour but does not remove our concern. Our results suggest that defences beyond model alignment, such as sandboxing and monitoring, are increasingly critical for safe and secure deployment.