论文

Local Support Learning解决大模型灾难性遗忘

Local Support Learning

精选理由

这篇论文提出LSL框架,让大模型在训练新任务时不会忘记旧能力,特别适合需要持续学习的场景。

Local Support Learning (LSL)是一种新框架,专为解决大模型灾难性遗忘问题设计。该研究将遗忘视为权重矩阵输入空间中的几何问题,提出无需访问先前数据即可保留模型先验能力的方法。LSL结合标准权重适配器和门控函数,基于高斯混合模型(GMM)实现数据路由,在7B参数LLM上验证有效。

原文 · arXiv cs.AI

Local Support Learning

We explore catastrophic forgetting in the context of large pre-trained models. By considering forgetting as a geometric problem in the input space of each weight matrix, we uncover a natural retention objective under which updates produced by gradient-based optimizers are suboptimal. Following this observation, we propose Local Support Learning (LSL), a general-purpose framework that augments gradient-based training for retention of prior capabilities without access to prior data. During a new learning phase, LSL pairs two components with distinct roles: a standard weight adapter, trained as usual to minimize the loss, and a gating function that enables the adapter only on input activations from its own training distribution, making the update local to that distribution. The key challenge is that this gate must route data from all learning phases while training only on data from the current one. We address this with a gate based on a Gaussian Mixture Model (GMM), whose likelihood decays rapidly away from its training data, giving it a natural tendency to stay closed on data from prior phases. We show that this post-training approach can resolve forgetting in LLMs of up to 7 billion parameters, retaining both pretrained and finetuned capabilities across multiple training phases, while being efficient in memory and compute, robust to hyperparameter choice, and showing scaling potential.