论文

Systemic Risk Index:基于 EU AI Act 行为准则的开源风险评估面板

An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice

精选理由

有人把 EU AI Act 那套系统性风险分类做成了开源评估面板,18 个模型最坏情况评分比平均分低 14–37 分,证据可逐条追溯。

Systemic Risk Index 是一个开源评估流水线和可视化面板,将 19 个公开基准按 EU GPAI Code of Practice 定义的 4 类系统性风险(CBRN、网络攻击、有害操纵、失控)组织起来。它用保留危害性的扰动和模拟部署场景对模型进行评估,用户可在面板上切换平均值与最坏情况聚合,并追溯每项风险评分的基准证据。在 18 个模型的测试中,最坏情况聚合使评分下降 14 到 37 分。LLM 裁判与人类评分者的一致性达到 κ=0.78–0.82,盲审中 83% 的样本扰动保留了原始危害。

原文 · arXiv cs.AI

An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice

Claims about AI safety reach audiences well beyond the AI community, yet many rely on opaque evidence or static assessments, when supporting evidence is accessible at all. We present the Systemic Risk Index, an open evaluation pipeline and dashboard built to make empirical evidence more transparent and traceable to the public. Our work organizes 19 public benchmarks into four systemic-risk categories defined by the EU GPAI Code of Practice---CBRN, cyber offense, harmful manipulation, and loss of control---and evaluates models using harm-preserving perturbations and simulated deployment contexts. The interactive dashboard lets users alternate between average and worst-case aggregation, vary how model capability affects the aggregate score, and trace each risk rating to its benchmark evidence. Across 18 models, scores fall by 14 to 37 points under worst-case aggregation, highlighting information that can be hidden by an average assessment of model risk. LLM judges show agreement with human graders comparable to human--human agreement ($κ= 0.78\text{--}0.82$), and a blind audit finds that $83\%$ of sampled transformations preserve the original harm. In a survey ($N = 21$), most participants report that scores are easy to understand and that the dashboard encouraged them to view model evaluations under different settings