论文

SPADE 让组合搜索式因果发现可扩展至 1600 变量

Mapping and Advancing the Scalability-Accuracy Frontier of Nonlinear Causal Discovery

精选理由

研究因果发现的朋友可以看看:SPADE 把组合搜索做到了 1600 变量级别,还顺手把四类主流方法的优劣捋了一遍,省你自己跑实验。

论文系统比较了可微结构学习、摊销结构学习、score-matching 和组合搜索四类非线性因果发现方法的准确率与运行时间取舍,指出各自瓶颈。作者据此提出 SPADE,一种基于样条的评分计算方案,只编译一次充分统计量并在组合搜索中复用。在有界入度条件下,其高斯变体把算法复杂度从 O(nd^3) 降到 O(nd^2+d^3)。实验显示 SPADE 可在数秒内求解 100 变量、16 万样本的问题,数分钟内求解 1600 变量、2500 样本的问题,同时保持较高结构准确率。

原文 · arXiv cs.AI

Mapping and Advancing the Scalability-Accuracy Frontier of Nonlinear Causal Discovery

Scalable nonlinear causal discovery requires methods that combine flexible mechanism estimators with efficient search over large graph spaces. Several algorithmic families have been proposed to address this challenge, yet their accuracy-runtime trade-offs remain poorly understood. We empirically compare the four major approaches: differentiable structure learning, amortized structure learning, score-matching, and combinatorial search. Our results reveal complementary bottlenecks: differentiable and amortized methods scale well but exhibit an accuracy gap, score-matching methods can be accurate in low dimensions but degrade quickly for increasing feature sizes, and combinatorial methods remain accurate but are slowed by repeated and redundant local scoring. Motivated by this bottleneck, we develop SPADE, a spline-based score-evaluation scheme that compiles sufficient statistics once and reuses them throughout combinatorial search. Under bounded indegree, its Gaussian variant reduces algorithmic complexity from O(nd^3) to O(nd^2+d^3). Empirically, SPADE shifts the observed scalability-accuracy frontier by orders of magnitude: it solves 100-variable problems with 160K samples in seconds and 1600-variable problems with 2.5K samples in minutes, while retaining high structural accuracy across synthetic and real-world benchmarks. These results reveal a substantial shift in the practical scale of combinatorial search and highlight the importance of evaluating scalable causal-discovery methods along the full accuracy-runtime frontier.