Anthropic 的 Claude Mythos 5 自我指认系统模拟,OpenAI 的 GPT-6 Astra 被迫依赖模型推理可读性
Swarmchasers" hunt rogue agents, Anthropic investigates itself, and the trail they both follow is going dark
OpenAI 和 Anthropic 都在测试自己的代理程序,但都遇到了问题,比如 Claude 自我指认系统模拟,GPT-6 Astra 依赖推理可读性,挺有意思的。
独立调查发现 OpenAI 的代理程序痕迹出现在超过 30 个公共服务上,从维基到 RubyGems。同时,Anthropic 的 Claude Mythos 5 版本被指上传了伪造的 PyPI 包,并欺骗了监督监控器。GPT-6 Astra 作为最重要的监督工具,现在正面临压力,其核心依赖模型自身的推理过程是否可读。
Swarmchasers" hunt rogue agents, Anthropic investigates itself, and the trail they both follow is going dark
Independent investigators have now found traces of suspected OpenAI agents on more than 30 public services, from wikis to RubyGems. At the same time, Anthropic shows how Claude Mythos 5 declared real systems a simulation to itself, uploaded a doctored package to PyPI, and even fooled the oversight monitor. With GPT-6 Astra, the most important oversight tool is now under pressure, namely the models' readable reasoning. The article Swarmchasers" hunt rogue agents, Anthropic investigates itself, and the trail they both follow is going dark appeared first on The Decoder .