CASPER:面向审讯对话的属性特定可控摘要框架
Controlled Attribute-Specific Summarization of Interrogative Dialogues
论文提出了 CASPER 框架来摘要审讯对话,还配套 6000 条标注的 MINDSum 数据集,做法是让警官、督察等多角色打分迭代,ROUGE 和 BERTScore 都超了基线。
论文提出 CASPER,一种基于思维链的属性特定提示框架,用于生成审讯者与受访者交互的高质量摘要。作者构建了 MINDSum 数据集,在 MIND 语料基础上扩展出 6000 条话语对,标注事件细节、事实陈述、人物描述和填充语。框架引入 RoleEval 分层评估机制,由警官、督察、高级督察三种角色按预设标准迭代评估摘要质量。实验显示 CASPER 在 ROUGE 和 BERTScore 两类指标上均超过标准摘要模型,人类评估也确认其与专家推理一致。
Controlled Attribute-Specific Summarization of Interrogative Dialogues
Effective summarization of interrogative dialogues is a critical task in forensic and investigative settings, requiring high factual accuracy, coherence, and attribute-specific relevance. In this work, we introduce CASPER, a novel Chain-of-Thought Attribute-Specific Prompting for Evaluative Summarization framework that leverages structured prompting and iterative refinement to generate high-quality summaries of interrogator-witness interactions. We construct MINDSum, a dataset extending the MIND corpus, comprising 6,000 utterance pairs annotated with event details, factual statements, character descriptions, and fillers. CASPER employs RoleEval, a hierarchical evaluation mechanism where multiple roles (officer, inspector, senior inspector) iteratively assess summaries based on predefined criteria. By integrating entity extraction and structured feedback loops, CASPER significantly improves factual consistency and contextual completeness compared to existing baselines. Experimental results demonstrate that our framework outperforms standard summarization models on both lexical (ROUGE) and semantic (BERTScore) metrics, while human evaluation confirms its alignment with expert reasoning. Our findings underscore the potential of controlled summarization in high-stakes domains, paving the way for AI-driven forensic intelligence.