论文:RL 后训练可轻易绕过蒸馏防御
Distillation Defenses Easily Break After Reinforcement Learning
说蒸馏防御没用的人来了:只要蒸馏后再跑一轮 RL,很多防御就失效。做模型安全或 API 防护的值得看看。
一篇 arXiv 论文指出,现有蒸馏防御通常只在蒸馏完成后立即评估,假设攻击者不再继续训练。作者提出更现实的威胁模型应在蒸馏后加入强化学习(RL)阶段,并展示一些在蒸馏后看似有效的防御会在 RL 后失效。实验表明,攻击者用现有 API 易于获取的数据即可从闭源模型窃取推理能力,效果等同于提取完整隐藏推理轨迹的复杂攻击。论文结论是,任何泄露足够信息以重建近似推理轨迹的防御都可能无效,并讨论了批级蒸馏防御的可行方向。
Distillation Defenses Easily Break After Reinforcement Learning
Distillation attacks copy the reasoning capabilities of closed-source large language models, allowing bad actors to replicate state-of-the-art performance at low cost. Attackers systematically collect a large volume of frontier model reasoning traces and then train (i.e., "distill") their own models on these traces. Existing defenses against distillation attacks are typically evaluated immediately after distillation, implicitly assuming attackers do not train their models any further. In this paper, we argue that a more realistic threat model includes further training with reinforcement learning after distillation. A misspecified threat model can give a false sense of security -- some defenses that seem effective after distillation can be broken after subsequent reinforcement learning. Practically, reinforcement learning lowers the bar for a distillation attack to be effective. We show that simple attacks can steal reasoning capabilities from existing closed-source language models using data easily obtainable from current APIs, yielding reasoning improvements equivalent to more sophisticated attacks that extract the full hidden traces. Results indicate that any distillation defense that leaks sufficient information to reconstruct approximate reasoning traces is likely ineffective. We conclude by discussing broader implications and batch-level distillation defenses which could be more effective.