Looped Diffusion Transformer研究发布
Looped Diffusion Transformer
OpenAI团队发布Looped-DiT,260M参数模型性能超越6.5倍大模型,计算效率提升4.9倍。
研究人员提出Looped Diffusion Transformer (Looped-DiT),通过在去噪步骤中重复运行共享Transformer块来扩展计算深度。260M参数的循环模型在多个文本到图像基准测试中超越了参数量6.5倍更大的模型,推理计算需求降低4.9倍。该模型结合深度监督和自调制注意力机制稳定循环特征更新,展现出类似潜在推理的行为。
Looped Diffusion Transformer
Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the parameter count fixed. This looped computation enables iterative refinement of internal representations without explicit reasoning tokens. However, naive looping fails to consistently improve image quality. We trace this problem to weak supervision across intermediate loops and unregulated attention updates that progressively erode local information. To overcome these challenges, we propose Looped Diffusion Transformer (Looped-DiT), which combines deep supervision across intermediate loops with self-modulating attention to stabilize looped feature updates. Under matched-parameter and matched-compute settings, Looped-DiT consistently outperforms non-looped baselines. Notably, a 260M-parameter looped model can surpass a model 6.5x larger across multiple text-to-image benchmarks while requiring 4.9x lower inference compute. Beyond this performance gain, we find that looped computation can offer a more effective form of iterative computation for diffusion models, with increasing loop depth yielding larger gains than adding more denoising steps under a fixed inference budget. Furthermore, deeper loops can progressively correct mistakes made in earlier loops, exhibiting behaviors suggestive of latent reasoning. Together, these results show that looped computation offers a promising way to scale visual generation models.