Pangram 漏检 79.8% 的 Muse-Glimmer 改写论文摘要
同一篇摘要换个模型改写,Pangram 的漏检率就从 6.5% 跳到 79.8%,说明检测器认的是模型指纹而不是 AI 痕迹。
东京都立大学的新论文测试了 AI 文本检测器 Pangram 对科学摘要的识别能力。Pangram 漏掉了 Meta 的 Muse-Glimmer 改写的 79.8% 摘要,却把 5000 篇人类摘要中的 1 篇误判为 AI 生成。它对 GPT-5 改写的摘要检出率为 93.5%。漏检率主要取决于具体是哪个大模型执行的改写。
Pangram, the AI-text detector, missed 79.8% of scientific abstracts rewritten by Meta's Muse-Glimmer, while flagging just 1 of 5,000 human abstracts.
In a new paper from Tokyo Metropolitan University, reseaerchers find the share of AI-rewritten abstracts that Pangram misses depends strongly on the LLM version
Shows that it caught 93.5% of GPT-5 rewrites but missed 79.8% from another new model. Its miss rate depended mostly on which model did the rewriting.