研究发现单步生成模型把去噪计算转移到网络深度中
Depth as Time in One-Step Generative Models
一篇挺有意思的分析论文:单步生成模型里,扩散的多步去噪没消失,而是挪到了网络各层里,还能借此把 MeanFlow SiT-L/2 压缩 16.6 倍参数。
论文提出"depth as time"的经验观察:多步扩散在采样步之间完成的去噪计算,会以分层形式出现在单步生成模型的一次前向传播里,可通过模型自身输出头解码中间层来恢复。研究发现这种深度方向计算取决于流图所训练的传输任务,MeanFlow 在更短传输间隔的探针中同时呈现去噪与再加噪,而没有时间索引传输任务的 drifting 模型则没有该现象。基于此,作者将分层计算显式当作流来训练单个时间条件块,把 MeanFlow SiT-L/2 模型的参数压缩 16.6 倍。
Depth as Time in One-Step Generative Models
The recent wave of one-step generative models, which compress the multi-step trajectory of diffusion via either distillation or learned flow maps, has reached an inflection point where they can generate high-quality images. Here, we ask a natural question that follows from these advances: what happens to the denoising trajectory of multi-step diffusion when generation is compressed into a single forward pass? We offer an empirical observation we call \textit{depth as time}: the denoising computation that multi-step diffusion performs across sampling steps appears to unfold across the depth of a single forward pass, and can be recovered by decoding intermediate layers with the model's own output head. Most interestingly, we show that this depthwise computation depends on the transport task a flow map is trained to solve. The most surprising case is MeanFlow, where probing shorter transport intervals reveals both denoising and renoising within a single network evaluation. In contrast, generators trained without a time-indexed transport task, such as drifting models, do not exhibit the same depthwise denoising. Consequently, we show that models that exhibit the depthwise denoising phenomenon are more compressible across the layerwise computation: a MeanFlow \texttt{SiT-L/2} model can be compressed by $16.6\times$ in parameters into a single time-conditioned block. We offer an explanation for this denoise-then-renoise behavior and show that, when we treat the layerwise computation explicitly as a flow, a single time-conditioned block can be trained to denoise across layers, compressing a MeanFlow \texttt{SiT-L/2} model by $16.6\times$ in parameters. Together, these results suggest that the temporal computation of diffusion is not eliminated by one-step generation, but reorganized across network depth.