论文

SynCo:通过对比交互残差学习跨模态协同

SynCo: Learning Cross-Modal Synergy by Contrasting Interaction Residuals

精选理由

SynCo方法能直接解决多模态学习中协同信息训练不足的问题,可作为插件集成到现有框架中。

SynCo是一种新的多模态对比学习方法,专注于捕获模态间的协同信息。该方法通过线性预测器从单模态特征预测融合表示,并对交互残差进行专门的对比监督。在Trifeature基准测试中,SynCo比基线方法提升5.98%的协同捕获性能,在MultiBench、DARai和MM-IMDb等真实世界基准上也表现优异。

原文 · arXiv cs.LG

SynCo: Learning Cross-Modal Synergy by Contrasting Interaction Residuals

Multimodal contrastive learning is a dominant paradigm for learning transferable representations from unlabeled data, but standard objectives primarily capture information that is redundant between modalities. Partial Information Decomposition (PID) shows that task-relevant information in multimodal data decomposes into three components: redundancy shared between modalities, uniqueness specific to each modality, and synergy available only from their joint observation. Recent frameworks extend contrastive learning to capture all three components, yet synergy remains undertrained in practice. We propose SynCo (Synergy Contrastive Learning), a method that directly addresses synergy undertraining through dedicated supervision on an interaction residual. SynCo fits a linear projector to predict the fused representation from independently computed unimodal features, and the resulting interaction residual, which removes the linearly unimodal-predictable component, receives dedicated contrastive supervision at negligible computational cost. On the controlled Trifeature benchmark, SynCo achieves state-of-the-art synergy capture with a $+5.98\%$ gain over the baseline, and on real-world benchmarks from MultiBench, DARai, and MM-IMDb, SynCo consistently outperforms or matches prior methods across diverse modality combinations and task types. The method operates as a plug-in to existing contrastive multimodal frameworks without modifying the underlying fusion architecture and can further improve synergy capture when combined with other methods.