核理论解释 Subliminal Learning:幽灵输出为何能传递任务能力
Why Ghost Outputs Teach: A Kernel-Based Understanding of Subliminal Learning
学生模型从看似无关的输出里学会没见过的任务,这篇论文用核函数把背后的数学原理讲透了。
arXiv 论文《Why Ghost Outputs Teach: A Kernel-Based Understanding of Subliminal Learning》为 Subliminal Learning 现象给出首个基于学习动力学的数学解释。作者推导出 chained cross-task kernel,通过共享骨干表征将教师模型的 ghost-output 监督信号与学生任务预测变化联系起来。理论证明共享初始化下传递算子严格半正定(PSD),学生无需任务标签即可对齐教师任务目标;ghost-output 维度则构成决定特征传递的秩瓶颈。框架还解释了为何高熵随机噪声比结构化数据更利于 Subliminal 传递——随机输入是最大化跨任务核重叠的宽带探针。在标准 ghost-output 实验设置中,三项理论预测均获得验证。
Why Ghost Outputs Teach: A Kernel-Based Understanding of Subliminal Learning
Subliminal Learning (SL) is a recently identified phenomenon in which a student model acquires downstream task capabilities by matching seemingly unrelated auxiliary outputs from a teacher, despite never observing task labels, task-specific outputs, or the original training data. While recent studies have identified where subliminal signals may reside, the optimization mechanism underlying this phenomenon remains poorly understood. In this work, we provide a mechanistic understanding of SL through the lens of learning dynamics. Specifically, we derive a chained cross-task kernel that explicitly links ghost-output supervision to changes in task predictions through shared backbone representations. Our unified analytical framework provides a rigorous mathematical explanation for three central empirical puzzles in SL: (i) under shared initialization, the transfer operator forms a strictly Positive Semi-Definite (PSD) structure, guaranteeing that ghost-output optimization aligns the student with the teacher's true task objective without explicit label exposure; (ii) the ghost-output dimensionality acts as an explicit rank bottleneck governing the transfer of task-relevant features; and (iii) synthetic, high-entropy inputs function as broadband probes that maximize cross-task kernel overlap, explaining why random noise consistently outperforms structured data for subliminal transfer. Experiments on the canonical ghost-output setting validate all three theoretical predictions, providing the first learning-dynamics-based theoretical explanation of how ghost-output supervision gives rise to subliminal learning.