论文

Level-of-Token Diffusion:按需分配算力的多分辨率图像视频生成框架

Level-of-Token Diffusion

精选理由

一篇很巧的论文:让扩散模型在需要细节的地方用细 token、别处用粗 token,预训练模型不用重训就能提速,还能接智能体的规划布局。

Level-of-Token(LoT)Diffusion 针对扩散模型对每个区域分配同等计算量的问题,把画面细节的空间分布转成显式的多分辨率 token 布局:细节密集处用细 token,其余区域用粗 token。通过 patch 级非对称 flow 参数化和多分辨率 token 嵌入,可直接适配预训练的扩散 Transformer,在每个去噪步骤保留全分辨率 flow 预测,同时只处理更少的 token。论文展示了基于语义掩码、边界框、纹理方差、景深线索以及智能体规划的布局来源,图像和视频生成均获得由 token 预算决定的可观加速。项目页面已公开。

原文 · arXiv cs.AI

Level-of-Token Diffusion

Image and video diffusion models allocate equal computation to every region, even when the intended scene calls for varying levels of detail. The spatial distribution of detail can often be anticipated before generation, indicating where computation can be reduced. We introduce Level-of-Token (LoT) Diffusion, a framework that turns this knowledge into an explicit multiresolution token layout (Level-of-Token layout) for adaptive and efficient generation. Tokens represent rectangular patches of varying sizes and shapes, allocating finer tokens where detail is needed and coarser tokens elsewhere. We adapt pretrained diffusion transformers to LoT layouts through a patch-wise asymmetric flow parametrization and embeddings for multiresolution tokens, preserving full-resolution flow prediction at every denoising step while processing only a reduced token sequence. LoT Diffusion enables layout-adaptive generation while preserving pretrained generative priors. We demonstrate LoT with layouts derived from semantic masks, bounding boxes, texture variance, and depth-of-field cues, as well as agentic plans. Across image and video generation, LoT offers favorable quality-efficiency tradeoffs, with significant speedups determined by the layout's token budget. Our project website is at https://georgenakayama.github.io/lotdiffusion/.