WeLike2Party实现多人图像动画
WeLike2Party! In-Context Motion Transfer for Multi-Human Image Animation
WeLike2Party让多人图像动画更逼真,解决了遮挡问题,还有新数据集和基准测试。
WeLike2Party是一种多人类人动画框架,通过直接上下文视频条件实现,无需显式姿态或网格提取。研究团队构建了MotionTwin数据集,包含14.4K跨身份视频对,总计84.3小时逼真视频。引入参考非对称RoPE条件和身份绑定监督,解决多人交互时的遮挡问题。在MotionTwin-Bench基准测试中,该方法在视觉保真度和身份-运动绑定方面优于现有最先进方法。
WeLike2Party! In-Context Motion Transfer for Multi-Human Image Animation
Human image animation aims to transfer motion from a driving video to subjects in a reference image. Despite remarkable progress in video generation, achieving high-fidelity animation of multiple interacting subjects remains a challenge. Many existing approaches rely on explicit motion representations such as 2D skeletons or parametric body meshes and struggle to preserve identity-motion binding under inter-person occlusion. To address this limitation, we propose WeLike2Party, a multi-human animation framework built on direct in-context video conditioning without explicit pose or mesh extraction at inference. We further introduce Reference Asymmetric RoPE Conditioning to preserve fine-grained appearance details, and Identity Binding Supervision to associate each reference identity with its intended motion trajectory. To support cross-identity training, we construct MotionTwin, a large-scale synthetic dataset comprising 14.4K cross-identity video pairs with shared subject and camera motions, totaling 84.3 hours of photorealistic video. We additionally present MotionTwin-Bench, a cross-identity benchmark specifically designed to evaluate subject-level visual fidelity and identity-motion binding. Extensive experiments on MotionTwin-Bench and real-world videos demonstrate that WeLike2Party outperforms recent state-of-the-art methods in subject-level visual fidelity, identity-motion binding, and overall perceptual quality, particularly in multi-person interactions with substantial occlusion.