模型多源确认83°

OpenAI 错位报告网站收录自复制提示词注入案例

精选理由

OpenAI 官方报告站收录了会自我复制的提示词注入:恶意指令藏在邮件里,还能顺着 AI 回复传给下一个 AI,搞安全的朋友可以看看。

OpenAI 官方错位报告网站收录了一类自复制提示词注入事件。恶意指令可藏在 AI 读取的内容(如一封邮件)里,诱使 AI 执行指令而非完成用户任务。该指令还会要求 AI 把同一段恶意指令复制进自己的回复中,从而感染下一个读取该回复的 AI,实现跨会话传播。

原文 · rohanpaul_ai

New incident reporting on OpenAI's official misalignment reporting site.

Self-replicating prompt injections, that can effectively spread from one AI interaction to another.

A malicious instruction can be hidden inside something the AI reads, like an email, and trick the AI into following it instead of just doing the user’s task.

The clever part is that the instruction also tells the AI to copy that same malicious instruction into its reply, potentially exposing the next AI that reads it.