论文

VSDD:用 Stackelberg 博弈让离散扩散模型学会改写已生成 token

Learning to Re-Draft: A Variational Stackelberg Game for Discrete Diffusion

精选理由

这篇论文教离散扩散模型边生成边改错,分子生成有效性比均匀扩散和掩码扩散都高,做生成模型研究的可以看看博弈式训练的思路。

论文提出 Variational Stackelberg Discrete Diffusion(VSDD),用于学习语义感知的前向腐蚀过程。VSDD 把训练建模为领导者-追随者博弈:领导者基于去噪器的 token 嵌入定义马尔可夫腐蚀过程,追随者固定腐蚀过程优化变分去噪目标。领导者根据去噪器从腐蚀中学习后的改进幅度来奖励腐蚀,而非腐蚀重建难度。在分子、文本和歌单生成三项任务上,VSDD 将分子有效性大幅超越均匀扩散和掩码扩散,文本困惑度低于均匀扩散并与掩码扩散保持竞争力,离线歌单推荐指标也有可观提升。

原文 · arXiv cs.AI

Learning to Re-Draft: A Variational Stackelberg Game for Discrete Diffusion

Discrete diffusion models offer the ability to re-draft, revisiting and correcting earlier tokens throughout generation. This capability depends on the forward corruption process that defines what the denoiser learns to correct. Masked diffusion models fix tokens once they are unmasked, while uniform diffusion permits revisions but relies on uniformly random token substitutions. We instead learn which substitutions are most useful for training the denoiser to re-draft. We introduce Variational Stackelberg Discrete Diffusion (VSDD), a framework for learning a semantically aware corruption process. VSDD formulates training as a leader-follower game: the leader defines a Markovian corruption process parameterized by the denoiser's token embeddings, while the follower optimizes a variational denoising objective with the corruption process held fixed. The leader rewards corruptions based on how much the denoiser improves after learning from them, rather than on how easily the current denoiser can reconstruct them. We measure this improvement under a fixed reference corruption process, approximate the follower's response with a one-step gradient update, and optimize the leader using a score-function estimator. We evaluate VSDD across molecular, text, and playlist generation. VSDD substantially improves molecular validity over uniform and masked diffusion, reduces text perplexity relative to uniform diffusion while remaining competitive with masked diffusion, and achieves sizable improvements in offline playlist recommendation metrics.