LSC-DPO:动态控制学习信号的偏好优化新方法
LSC-DPO: Learning-Signal-Controlled Direct Preference Optimization
训练偏好模型的老问题:DPO 训到后面梯度就没劲了。这篇把 sigmoid 因子当成可调信号来动态控制,三个基准上都比 DPO 强,搞对齐训练的可以看看实现思路。
论文提出 LSC-DPO,从损失函数几何视角分析 DPO,将 sigmoid 因子视为刻画目标局部敏感度的学习信号。该方法动态调节学习信号使其保持在目标区间,并给出对数空间下的稳定追踪条件。在 AlpacaEval 2、MT-Bench 和 Anthropic-HH 三个基准上,LSC-DPO 的成绩稳定优于 DPO 及其他强偏好优化基线。作者还发现不同系数初始化会导致瞬态学习信号轨迹不同,据此推导出信号预算补偿规则,显著降低了不同初始化之间的性能波动。
LSC-DPO: Learning-Signal-Controlled Direct Preference Optimization
Direct Preference Optimization (DPO) has become a standard reward-model-free approach for aligning language models with preference data. However, as the scaled preference margin grows during training, the logistic DPO loss becomes progressively less sensitive to further changes. We study DPO from a loss-level geometric perspective and identify the sigmoid factor as a learning signal that characterizes the local sensitivity of the objective. Based on this view, we propose Learning-Signal-Controlled Direct Preference Optimization (LSC-DPO), which dynamically regulates the learning signal near a target regime. A log-space analysis establishes conditions for stable tracking of the target learning-signal regime. Experiments on AlpacaEval 2, MT-Bench, and Anthropic-HH show that LSC-DPO consistently improves over DPO and strong preference-optimization baselines. We further find that different coefficient initializations induce distinct transient learning-signal trajectories even when their later signal levels become similar. Based on this observation, we derive a signal-budget compensation rule that adjusts the target learning signal to compensate for these transient differences. The resulting compensation substantially reduces performance variation across coefficient initializations.