LoGo:局部+全局奖励提升长视频 3D 一致性
LoGo: Local-Global Rewards for Consistent Long-Horizon Video Generation
做视频生成方向的可以看看,LoGo 用局部加全局两种奖励解决相机移动时物体乱飘的问题,还附带了新基准 TrajectoryBench。
论文提出 LoGo 方法,针对相机控制视频生成中的 3D 不一致问题,将局部奖励与全局奖励结合用于后训练。局部奖励提供细粒度的信用分配,减少物体漂移和伪影;全局奖励负责保持相机轨迹跟随和画质。该方法在 DL3DV 和 TrajectoryBench 两个基准上测试,后者是为长时程复杂相机控制新提出的评测集。实验覆盖三个基础模型,均显示一致改进。
LoGo: Local-Global Rewards for Consistent Long-Horizon Video Generation
Camera-controlled video models are rapidly advancing toward long generation horizons and complex camera control. A key failure mode is 3D inconsistency: as the camera moves, objects lose permanence and scene structures shift. Existing post-training techniques, which assign a single scalar reward to the entire generation, are poorly suited to correcting these inconsistencies over long horizons. We introduce LoGo, which blends global and spatially localized rewards for camera-controlled video models. The local reward provides fine-grained credit assignment, which substantially improves 3D consistency, while the global reward preserves camera following and video quality. Across three base models, LoGo shows a clear advantage on DL3DV and TrajectoryBench, a new benchmark for long-horizon, complex-camera-control generation that current evaluations lack. LoGo effectively reduces local object shifts, artifacts, and global scene changes, illustrating the importance of credit assignment in post-training video models. Project website: https://ziqi-ma.github.io/logo-website/