CausalArena 基准测试框架发布,用于评估因果发现模型
CausalArena: Benchmarking Causal Discovery in the Foundation Model Era
这个新基准测试框架能更全面地评估因果发现模型的性能,避免单一基准测试的局限性。
CausalArena 是一个统一的因果发现基准测试框架,它通过合成数据、语义操作数据集和公式化数据集来测试模型。实验显示,不同基准测试下的模型排名会发生显著变化,表明在基础模型时代评估因果发现能力面临挑战。
CausalArena: Benchmarking Causal Discovery in the Foundation Model Era
Causal discovery aims to uncover causal structures from data and is fundamental to scientific reasoning and intervention-based decision making. Its evaluation relies heavily on structural causal models (SCMs), which specify a causal graph together with the mechanisms that generate data, yet existing studies differ substantially in graph families, mechanisms, and evaluation protocols. The emergence of causal discovery foundation models (CDFMs) further complicates evaluation: performance may reflect not only causal discovery ability, but also overlap between pretraining environments and test SCMs, making results on fixed synthetic benchmarks difficult to interpret. We introduce CausalArena, a unified and evolvable benchmark for causal discovery under a common protocol. Synthetic SCMs supply controlled breadth over structures and mechanisms; semantic operational SCMs provide human-auditable, semantically grounded environments beyond standard synthetic generators; and formula-grounded SCMs test discovery under explicit scientific mechanisms. Public real-world datasets provide an additional external-validity check. Experiments across classical, neural, and pretrained methods reveal substantial ranking shifts across SCM families and protocols, showing that strong performance in one benchmark regime does not reliably transfer to others. These results highlight benchmark diversity and pretraining--evaluation overlap as central challenges for evaluating causal discovery in the foundation model era.