Marformer:预测缺失数据分布的Transformer模型
Marformer: A Transformer for Predicting Missing Data Distributions
Marformer直接预测条件边际分布,无需建模完整联合分布,在缺失数据处理任务上表现优异且速度快。
Marformer是一种新型Transformer模型,专为预测缺失数据的条件边际分布而设计。该模型在三个合成数据领域(贝叶斯网络、离散多元高斯结构和结构化注释数据)上进行了评估,性能匹配或超越传统缺失数据处理方法。在真实注释数据集上,Marformer在最大训练规模时优于评估基线,且速度显著快于生成式基线模型。
Marformer: A Transformer for Predicting Missing Data Distributions
Real decisions are made under incomplete information. If we observe only some of the random variables we need, we can predict the others. The \textbf{conditional marginals} over the missing variables are the key ingredient for computing Bayes risk and Value of Information (VOI), the expected gain from acquiring one more observation before deciding. We present the Marformer, a Transformer trained to directly predict conditional marginals given any set of observed values. Like BERT, which is trained to predict missing words from context, the Marformer constructs a hidden-vector representation for each distribution $p(X_i)$ and iteratively refines it through attention to other distributions $p(X_j)$. Unlike generative approaches, the Marformer does not model the full joint distribution, requires no domain knowledge of the data-generating process, and makes all predictions in a single forward pass. We evaluate across three synthetic domains with missing data---Bayesian networks, discretized multivariate Gaussians, and structured annotation data. The Marformer can match or outperform classical missing-data methods, even when those methods are given the true model family and prior that generated the synthetic data. We also evaluate on a real annotation dataset, where the Marformer outperforms the evaluated baselines at the largest training size. In both cases, the Marformer is substantially faster than the evaluated generative baselines.