表格基础模型架构与稀疏先验对齐研究
Architecture Alignment With Sparse Priors in Tabular Foundation Models
TabPFN v2在处理无关特征时比TabDPT更稳定,研究揭示了架构如何影响特征选择能力。
研究比较了TabDPT和TabPFN v2两种表格基础模型在不相关特征抑制能力上的差异。TabDPT在面对空特征时预测性能下降更严重,而TabPFN v2保持相对稳定。通过在相同稀疏到密集线性先验下训练简化模型,研究发现稀疏预测需要上下文相关的特征门控,密集预测仅需均匀特征加权。交替轴模型在稀疏任务上更接近贝叶斯最优预测器。
Architecture Alignment With Sparse Priors in Tabular Foundation Models
Tabular foundation models (TFMs) are increasingly popular because they deliver strong predictions on new datasets through in-context learning, without task-specific training or extensive tuning. Yet released TFMs differ simultaneously in their pretraining priors, architectures, and objectives, obscuring their respective inductive biases. We therefore examine one concrete capability: irrelevant-feature suppression. Across synthetic tasks and real-world datasets, adding null features causes substantially greater predictive degradation in the row-token model TabDPT, whereas the cell-token alternating-axis model TabPFN v2 and other TFMs remain comparatively stable. This gap motivates us to ask whether architecture contributes to irrelevant-feature suppression. Because released TFMs remain confounded by other design choices, we train streamlined row-token and alternating-axis transformers under identical sparse-to-dense linear priors. Exact Bayes analysis shows that sparse prediction requires context-dependent feature gating, whereas the dense endpoint requires only uniform feature weighting. Consistent with this distinction, the alternating-axis model is substantially closer to the Bayesian optimal predictor on sparse tasks, while the architecture gap becomes negligible on dense tasks; almost all of the sparse gap arises from linear coefficient-estimation error. Finally, in both the controlled model and frozen TabPFN v2, we examine the effect of interventions on the feature-attention outputs on the linear coefficients, finding evidence of task-dependent selective routing of computation through feature-indexed pathways. Together, these results support architecture-prior alignment: preserving an addressable feature axis provides an inductive bias for task-adaptive relevance inference. Code is available at https://github.com/Tianqi-Zhao/ArchitecturePriorTFMs.