论文

解码器LLM在隐私设置下权重绑定研究

Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?

精选理由

这篇论文发现,在隐私保护训练中,不绑定权重比传统设计效果更好,准确率更高且内存使用更少。

研究团队使用GPT2和DistilGPT2模型,探索了差分私有训练(DP-SGD)中权重绑定的作用。实验显示,不绑定权重的模型在SST-2、QNLI和QQP基准上准确率最高提升4.74个百分点。不绑定权重还能实现内存效率更高的ghost clipping技术,内存使用降低60%以上。

原文 · arXiv cs.LG

Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?

Differentially Private Stochastic Gradient Descent (DP-SGD) is a leading approach for privacy-preserving fine-tuning of large language models (LLMs). Many decoder-only LLMs employ weight tying between input and output embeddings, a design choice originally introduced for parameter efficiency and improved language modeling performance in the non-private setting. However, the impact of weight tying under differentially private training remains largely unexplored. In this work, we investigate the role of weight tying in the DP setting using GPT2 and DistilGPT2 as representative decoder-only architectures. Interestingly, we find that untied embeddings consistently outperform weight-tied models under DP-SGD, achieving gains of up to 4.74% points in accuracy on SST-2, QNLI, and QQP. Beyond improved utility, untying embeddings enables the use of memory-efficient ghost clipping for DP-SGD. By contrast, weight tying introduces shared-parameter interactions that complicate standard ghost norm computation and largely negate its computational advantages. As a result, untied models achieve over 60% lower memory usage while preserving the benefits of ghost clipping. Our results indicate that untied embeddings provide a more effective and scalable design for differentially private training of decoder-only LLMs and highlight the need to revisit standard LLM architectural choices in the privacy-preserving setting.