论文

Jev 用作放射报告事实性评估器:RadEvalX 相关性达 0.573

Can Jev Judge Radiology Reports? Evaluating a System One Model for Clinical Factuality

精选理由

医院里 AI 写放射报告容易漏病灶或编造发现,这篇论文用 Jev 做低成本裁判,成本不到 3 美分评 100 对报告,做医疗 AI 评估的可以看看。

arXiv 论文研究了 Jev(一个 System One 决策模型)作为放射报告事实性评估器的效果,判断 AI 生成的报告与医生参考报告是否一致。双向比对配置在 RadEvalX 上达到 0.573 的 Kendall 相关性,在 RadEvalExpert 上达到 0.398,超过同等条件下的开放自然语言推理评估器。每条陈述只问一个支持性问题时,判断输入 token 减少 43-45%,专家一致性相近。按文档标价计算,每 100 对报告的判断成本低于 3 美分,另在受控错误测试中检测假阴性错误的 AUROC 达 0.977。

原文 · arXiv cs.AI

Can Jev Judge Radiology Reports? Evaluating a System One Model for Clinical Factuality

An AI-generated radiology report can resemble a physician's report while omitting an abnormality, adding an unsupported finding, or reversing its presence. Measuring these factual differences is essential for evaluating report generators. We study Jev, a System One decision model, as a simple, low-cost judge of agreement with physician-written reference reports. Our evaluator checks whether each statement is supported by the other report and combines these judgments in both directions to capture unsupported claims and omissions. A single-question configuration reaches Kendall correlations of 0.573 on RadEvalX and 0.398 on RadEvalExpert with expert error counts, outperforming an open natural language inference judge under matched decomposition and aggregation. One support question per statement retains similar expert agreement to seven while using 43-45% fewer judgment input tokens. At the documented API price, judgments cost under three cents per hundred report pairs, excluding local decomposition. In a separate controlled-error test, Jev detects false negation with an AUROC of 0.977. Local RadMatch achieves stronger agreement on clinically significant errors in both expert datasets and on total errors in the shared RadEvalExpert subset. Finding-count and error-scope analyses show that benchmark agreement reflects report size and error definitions as well as medical error detection. These results support Jev as a practical judgment component for measuring factual differences in generated radiology reports and identify where more elaborate evaluation remains valuable.