论文

AI安全基准测试的心理学审计

Searching for "Harmful Refusal": A Psychometric Audit of an AI Safety Benchmark

精选理由

这篇论文揭示了AI安全基准测试的测量问题,特别是HarmBench数据集无法准确衡量单一安全属性。

研究人员对HELM Safety基准中的HarmBench数据集进行了心理测量分析。该研究使用多维项目反应理论模型发现HarmBench并非测量单一属性。差异项目功能分析显示,不同开发者的模型在相同拒绝能力得分上表现不同。研究指出,任何将多个数据集和项目平均的安全分数都可能掩盖饱和现象并混淆行为。

原文 · arXiv cs.AI

Searching for "Harmful Refusal": A Psychometric Audit of an AI Safety Benchmark

Safety benchmarks typically report one overall score for a suite of datasets, each of which may target one or more safety-related attributes, so models with similar overall scores can have very different attribute profiles. Comparing models is more tractable at the level of individual attributes, yet it is often unclear whether even a single dataset's scores isolate any single attribute. One plausible candidate for such an attribute is harmful refusal, a model's tendency to refuse dangerous or policy-violating prompts. We examine whether it constitutes a single, measurable attribute in HELM Safety. Using a construct validity framework that stipulates that an attribute must exist before a test can measure it, we start with HELM Safety's four datasets that might plausibly target harmful refusal, but find that three are saturated. We subject the remaining dataset, HarmBench, to two psychometric tests to determine if a single attribute like harmful refusal could stand behind its score. First, multidimensional item response theory modeling strongly suggests that HarmBench does not measure a singular attribute. Second, a differential item functioning analysis finds items where models from different developers with the same refusal ability score differently. These flags largely disappear under scope-specific matching, a pattern consistent with aggregation effects but not sufficient to rule out domain-specific developer differences. Zooming out, HarmBench collapses distinct harm behaviors into one score, and the overall HELM safety aggregate further collapses HarmBench and scores from other datasets into a single top-line number. Any safety score that averages over datasets and items can hide saturation and conflate behaviors this way. We argue that a score should earn its single-attribute reading before models are compared with it.