SR4-Fit:可解释规则学习框架,准确率媲美黑盒模型
SR4-Fit: A Unified Interpretable Rule-Based Machine Learning Framework for Informative and Trustworthy Decision-Making
一篇解决黑盒模型解释不可靠问题的论文,SR4-Fit 用稀疏正则生成规则集,在 14 个数据集上跑赢了 RuleFit 和决策树。
SR4-Fit 是一种内在可解释的规则学习算法,支持分类和回归任务,能生成紧凑且稳定的规则集。在美国人口普查局 American Community Survey 数据上,它预测美国众议院选举结果并揭示黑盒模型遗漏的人口统计交互。在 14 个基准数据集(6 个分类、8 个回归)上,SR4-Fit 在准确率、稳定性和紧凑性方面优于 RuleFit 和决策树,预测能力与黑盒模型相当。
SR4-Fit: A Unified Interpretable Rule-Based Machine Learning Framework for Informative and Trustworthy Decision-Making
In many high-stakes applications, machine learning is dominated by black-box models that require post hoc explanations to justify their predictions. These explanations are often unreliable because they do not reflect the model's actual computations, limiting accountability and trust. A natural alternative is to use models that are interpretable by design. However, existing rule-based approaches, such as RuleFit and decision trees, while transparent, often lack stability and predictive strength, reinforcing a perceived trade-off between traditional performance measures and model understandability. To address this, we propose Sparse Relaxed Regularized Regression Rule-Fit (SR4-Fit), an intrinsically interpretable algorithm for both classification and regression that produces compact and stable rule sets without sacrificing performance. Using demographic data from the U.S. Census Bureau's American Community Survey, SR4-Fit predicts U.S. House election outcomes with high accuracy and interpretability while uncovering demographic interactions missed by black-box models. We further validate SR4-Fit across fourteen benchmark datasets (six classification and eight regression), where it outperforms existing rule-based methods, including RuleFit and decision trees in terms of accuracy, stability, and compactness while remaining competitive with black-box models in predictivity. These results demonstrate that interpretability and predictive reliability need not be mutually exclusive, offering a practical and transparent alternative for high-stakes decision-making.