论文

IdeaLens:从写作思路而非文字判断内容是否由 AI 生成

IdeaLens: Detecting AI Ideas in Long-form Writing

精选理由

一群研究者做了个叫 IdeaLens 的检测器,专看想法是人的还是 AI 的。有人照 AI 详细大纲写文,Pangram 照样误报 92%,它只误报 7%,挺有意思的思路。

arXiv 论文提出 IdeaLens 检测器,判断文档的想法来自人类还是 AI,而不是判断文字是谁写的。方法是把文档表示成大纲,每条包含语篇角色和对内容的改写描述,尽量去除表面文字信息。模型用 100 万篇 FineWeb 文档训练,标签来自 Pangram 检测器的 silver 标注。在受控实验中,随着人类计划越来越详细,IdeaLens 的 AI 标记率从 95% 降到 7%,而 Pangram 4 仍标记 92%;在 50 篇人类按 AI 计划写的故事数据集上,IdeaLens 标记了 68%,Pangram 4 只有 8%。在 19 个现有检测基准上,IdeaLens 在低误报率下保持较高检出率,且跨领域、格式和语言表现稳定。

原文 · arXiv cs.AI

IdeaLens: Detecting AI Ideas in Long-form Writing

While modern AI detectors identify who wrote the words, emerging policies on AI use increasingly hinge on a different question: who came up with the ideas? We introduce IdeaLens, a detector that identifies whether a document's ideas came from a human or AI (idea provenance), regardless of who wrote its words. To focus IdeaLens on ideas rather than prose, we represent documents as outlines: lists of items that each pair a discourse role with a brief, paraphrased description of the content, minimizing word-level overlap with the raw text. We train IdeaLens on 1M FineWeb documents with silver labels from Pangram, a prose provenance detector. Since the outlines are largely stripped of surface-level information, the labels must be fit mainly through the ideas. In a controlled study, IdeaLens's AI flag rate drops from 95% to 7% as models write from increasingly detailed human plans, while Pangram 4 still flags 92%; from AI-derived plans, IdeaLens stays above 96%. Conversely, on a new dataset of 50 stories that human authors wrote from AI-generated plans, IdeaLens flags 68% of the stories as AI, compared to 8% for Pangram 4. On a comprehensive suite of 19 existing detection benchmarks, we show that IdeaLens maintains strong detection rates at low false positive rates, suggesting that ideas themselves provide a powerful discriminative signal, and its performance holds across domains, formats, and languages. Finally, we examine 90K predictions from IdeaLens to characterize systematic differences between human and AI ideation. We release our models and labeled datasets to facilitate future research on idea provenance detection.