论文

CDMD跨数据集混合类型扩散模型

CDMD: A Cross-Dataset Mixed-Type Diffusion Model for Tabular Data

精选理由

CDMD让一个模型搞定多种表格数据,比单独训练多个模型参数更少效果更好,代码已开源。

CDMD是一种表格数据扩散模型,可在不同架构和数值/分类特征数量的异构数据集上联合训练。该模型在七个真实数据集上实现了最高的平均生成质量,参数量远低于单独训练的模型集合。在337个数据集预训练后,CDMD在目标数据有限和适应周期有限的情况下仍能提升未见数据集的生成效果。

原文 · arXiv cs.LG

CDMD: A Cross-Dataset Mixed-Type Diffusion Model for Tabular Data

Generative models for tabular data are typically trained separately for each dataset, limiting knowledge transfer and requiring the storage of many specialized models. In this paper, we introduce CDMD, a tabular diffusion model trained jointly across heterogeneous datasets with different schemas and variable numbers of numerical and categorical features. Unlike existing cross-dataset tabular diffusion models that operate in continuous representation spaces, CDMD defines diffusion directly over the mixed-type feature space and is trained end-to-end. To accommodate heterogeneous categorical domains, we introduce a schema-restricted reverse-process parameterization for masked diffusion models, in which the output space dynamically adapts to each feature's vocabulary. We then compose numerical and categorical feature-level diffusion processes into a schema-dependent row-level process. A shared schema-aware Transformer denoiser captures dependencies between features and parameterizes the reverse process across varying schemas. On seven real-world datasets, a single jointly trained CDMD achieves the highest average generation quality among strong single-dataset and cross-dataset baselines, while using substantially fewer total parameters than the collection of separately trained models. Furthermore, pre-training on a corpus of 337 datasets improves generation on previously unseen datasets under both limited target data and limited adaptation epochs. These results demonstrate the potential of direct mixed-type diffusion for shared and transferable tabular data generation. Our code is available at https://github.com/ketatam/cdmd.