CMA-OT模型通过分层专家监督提升舞蹈配乐生成效果
CMA-OT: Hierarchical Expert Supervision for Dance-to-Music Generation
这个模型用分层专家监督解决了舞蹈配乐生成中的语义不匹配问题,效果比现有方法更好。
CMA-OT是一种用于舞蹈配乐生成的新方法,它通过引入外部音乐专家来提供分层监督,解决了现有方法中舞蹈动作与音乐生成的语义不匹配问题。该方法在两个数据集上进行了实验,结果显示其在节奏同步性和感知质量方面达到了SOTA水平。
CMA-OT: Hierarchical Expert Supervision for Dance-to-Music Generation
Dance-to-music (D2M) generation aims to synthesize music that is rhythmically and stylistically aligned with dance videos. A key challenge arises from the semantic mismatch between sparse dance cues, such as rhythm and style, and the dense information required for music composition, including structure, instrumentation, and expressive dynamics. Existing methods typically rely on these sparse cues and supervise only the final audio output, resulting in poorly learned music representations and generated music with limited musicality and structural coherence. To address these issues, we propose Curriculum-guided Multi-scale representation Alignment with scale-aware Optimal Transport (CMA-OT), a novel paradigm that leverages an external music expert to provide hierarchical supervision for the generator's latent features, bridging the semantic gap and enhancing representation learning. To effectively incorporate hierarchical supervision, we introduce a curriculum-guided multi-scale learning strategy that progressively transfers musical knowledge from the expert to the music generator, enabling stable and effective representation learning. Moreover, to accommodate the semantic and structural variations across different expert scales and achieve fine-grained alignment under temporal mismatch, we propose a scale-aware optimal transport alignment mechanism, which models soft correspondences between hierarchical expert representations and the generator's latent features. Extensive experiments on two datasets demonstrate that CMA-OT achieves state-of-the-art performance in rhythmic synchronization, perceptual quality, and overall music generation.