论文

论文:数据归因估计结果不一致,根源在于反事实规范差异

Which Influence Are We Estimating? The Role of Counterfactual Specifications in Data Attribution

精选理由

做数据归因或模型调试的朋友可以看看,它解释了为什么不同影响力估计方法排序对不上,还给出对齐行为规范的具体做法。

这篇 arXiv 论文研究训练数据影响力估计方法为何给出互不兼容的排序。作者指出核心问题不是近似误差,而是规范不匹配:影响力取决于被归因的行为、对训练样本的干预方式、以及连接干预与模型响应的反事实训练过程。论文将影响力形式化为反事实估计量,并通过噪声标签检测与 LLM 归因实验证明,行为代理(如 query loss、logit、margin)的选择显著影响归因质量。行为对齐的规范能找到基于默认损失或相似度的规范遗漏的目标特定训练样本。

原文 · arXiv cs.LG

Which Influence Are We Estimating? The Role of Counterfactual Specifications in Data Attribution

Estimating the influence of training examples on model behavior is essential for data debugging, valuation, and attribution. Existing influence estimators often produce incompatible rankings, which are commonly ascribed to approximation error. We argue that a more fundamental source of disagreement is specification mismatch: influence depends on the behavior being attributed, the intervention applied to each training example, and the counterfactual training process that maps the intervention to a model response. These choices are especially important when the target behavior requires a tractable surrogate, such as query loss, a logit, or a margin. We formalize influence as a counterfactual estimand, distinguish specification mismatch across estimands from approximation error in estimating a fixed estimand, and organize representative estimators by their implied specifications. We further derive a local decomposition that exposes how behavior signals, training signals, and counterfactual parameter responses interact. Controlled experiments show that exact estimands under different specifications can induce different rankings, whereas approximation error grows as perturbations move farther from their linearization points. Experiments on noisy label detection and LLM attribution show that specification choices significantly affect attribution quality, especially for the choice of behavior surrogate. Behavior-aligned specifications can identify target-specific training examples obscured by default loss-based or similarity-based specifications. These results establish specification analysis as a necessary first step for interpreting and comparing data influence estimators.