论文

Insight Provenance:判断科研评审洞见来自人类还是 LLM

Who Wrote It Is Not Enough: Detecting Who Contributed the Insight

精选理由

研究者做了个数据集专门判断评审点子是人写的还是 AI 写的,发现 AI 想法大多离不开论文本身,人会带外部知识进来,挺有意思的角度。

一项 arXiv 研究提出 Insight Provenance 任务,判定评审洞见出自人类、LLM 还是混合贡献。团队基于 4,057 篇论文和 12,660 条人类评审构建 InsightProv-v0 数据集,用 GPT-4o、Gemini、DeepSeek 模拟不同 LLM 参与度,并在句子级标注来源。实验发现模型常利用语言和文本署名捷径,去偏评估后性能明显下降,因此提出两阶段对抗框架抑制捷径信号。分析显示 AI 洞见多停留在通用或论文自带信息,而人类洞见更多引入外部知识和独立判断。

原文 · arXiv: DeepSeek

Who Wrote It Is Not Enough: Detecting Who Contributed the Insight

As LLMs increasingly assist scientific writing and peer review, detecting who wrote the text is no longer sufficient: we need to determine who contributed the underlying insight. We introduce Insight Provenance, the task of identifying whether a review insight originates from a human, an LLM, or their hybrid contribution. We construct InsightProv-v0 from 4,057 scientific papers and 12,660 human reviews, simulating different levels of LLM involvement with GPT-4o, Gemini, and DeepSeek and annotating provenance at the sentence level. We show that strong performance on raw data can be misleading, as models exploit linguistic and textual-authorship shortcuts that degrade substantially under progressively debiased evaluation. We therefore propose a two-stage adversarial framework that suppresses shortcut signals while preserving provenance-relevant information. Beyond detection, extensive analyses reveal what makes intellectual authorship identifiable: paper grounding and neighboring review context provide complementary provenance signals, while human, hybrid, and AI insights systematically differ in their information sources and failure modes. Most strikingly, AI insights predominantly remain close to generic or paper-provided information, whereas human insights more often introduce external knowledge and independent judgment. These findings suggest that while wording can be rewritten by an LLM, the provenance of an idea leaves a deeper and more persistent signal.