扩散模型记忆化可通过能量盆地几何与循环去噪提前检测
Early Signatures of Memorization in Diffusion Models via Basin Geometry and Cyclic Denoising
生成出复制图之前就能查出扩散模型背了哪些训练数据,CelebA 和 Stable Diffusion v1.4 上的数字都给出来了,做模型审计的可以看看。
扩散模型在训练早期泛化、后期会复现训练样本,但常规检测要等一次性生成出近似复制才报警。论文提出潜在记忆化概念,用 score divergence 和 basin volume 发现训练样本周围会先形成局部盆地,其出现时点与记忆化时间同样遵循 O(n) 缩放。通过循环去噪方法,在一个一次性复制率仅 0.1% 的 CelebA 检查点上,500 次循环将记忆化比例提升到 30% 以上。在 Stable Diffusion v1.4 上,条件与无条件生成的 score divergence 差值区分记忆化与非记忆化提示词,AUC 达 0.944、1% FPR 下 TPR 为 0.866。
Early Signatures of Memorization in Diffusion Models via Basin Geometry and Cyclic Denoising
Diffusion models generalize early in training and later reproduce individual training samples. Standard tests detect memorization only once one-shot generation produces near-copies, leaving a released model unaudited until its outputs fail. We show that memorization is encoded in the geometry of the learned energy landscape before it appears in generated samples, a state we call latent memorization. Using score divergence and basin volume, we find that localized basins form around training samples and separate them from held-out samples before the first memorized sample appears, with an onset that follows the same $O(n)$ scaling as the memorization time. We probe these basins with cyclic denoising, which repeatedly applies partial noising and denoising. Under the exact empirical score, we prove that cycling started near an isolated training sample recovers it and returns to it over any finite number of cycles with high probability. In trained models, cycling recovers training images from CelebA and CIFAR-10 checkpoints whose one-shot samples contain no copies, and at a CelebA checkpoint with 0.1% one-shot copies, 500 cycles raise the memorized fraction above 30%. Cycling also reveals degenerate attractors that match no single training image and fade as training proceeds, so residence in a basin does not by itself imply memorization. These findings hold on a Gaussian mixture, CelebA, and CIFAR-10 across optimizers, architectures, noise schedules, and training-set sizes, and extend to off-the-shelf Stable Diffusion v1.4, where the cycled conditional-unconditional divergence gap separates memorized from non-memorized prompts with an AUC of 0.944 and a TPR of 0.866 at 1% FPR. More broadly, what a diffusion model has memorized is a property of the geometry and stability of its learned distribution, and assessing it requires examining this structure rather than generated outputs alone.