论文

持久表征学习让像素空间 Drifting Models 一步生成质量大幅提升

Learning Discriminative Geometry for Drifting Models

精选理由

这篇论文解释了 Drifting Models 为什么在像素空间效果差,还提出从像素直接学判别表征,FID 降了 82-95%,做生成模型方向的可以看看。

arXiv 论文(编号 2610.04703)指出 Drifting Models 在像素空间构建 drifting field 时生成质量差,根源在于表征的判别几何影响核密度估计(KDE)的样本加权。作者提出 persistent representation learning,在生成器跨批次演化过程中持续学习更具判别性的表征几何,无需预训练编码器。论文还证明了 KDE ratio loss 与 drift regression loss 在匹配条件下的当前步梯度等价性。该方法在多个数据集上将 FID 相比原像素空间 Drifting Models 降低约 82-95%,配合速度裁剪可进一步收益。

原文 · arXiv cs.LG

Learning Discriminative Geometry for Drifting Models

Recently proposed Drifting Models shift iterative distribution refinement from inference to training, enabling effective one-step generation. However, their performance on complex image datasets depends strongly on the representation used to construct the drifting field: pixel-space drifting performs poorly, whereas pretrained feature spaces substantially improve sample quality for reasons that remain unclear. We trace this gap to the discriminative geometry of the representation, which determines sample weighting in kernel density estimation (KDE) and, consequently drift. We introduce persistent representation learning, which continuously learns a more discriminative representation geometry as the generator evolves across batches. We further establish a current-step gradient equivalence between the KDE ratio loss and drift regression loss under matched conditions, connecting density-ratio-based generator optimization to empirical drifting and motivating direct control of the drifting velocity. Across multiple datasets, our method learns effective discriminative representations directly from pixels and reduces FID by approximately $82-95\%$ over the original pixel-space Drifting Models, without pretrained encoders. Adapting pretrained representations and applying velocity clipping provide further gains.