论文

READ 方法让 LoRA 适配器只读不写,多技能组合更稳定

New LoRA Skills Should Read but Never Write

精选理由

一篇讲 LoRA 技能怎么叠加不互相干扰的论文,思路是把新适配器设计成只读旧技能的输入,合并进基座权重还零推理开销,SuperGLUE 比现有方法高了 20 多分。

arXiv 论文提出 READ(Read-only Expansion of Adapter Deltas),解决多个独立训练的 LoRA 适配器难以合并的问题。方法将每个适配器改写成保持更新完全一致的平衡规范形式,并让新旧技能之间的耦合只单向增长:新技能可以读取旧技能的输入子空间,但不能写入其输出子空间。每次追加技能时唯一可训练对象是新技能在耦合矩阵中的一行,组合后的更新直接折叠进基础权重,推理时无额外开销。在 4 个基准套件和 2 个模型家族上逐一追加技能的评测中,READ 在 SuperGLUE 上超过最强已发表基线 20 分以上,在领域套件上超过 7 分以上。

原文 · arXiv cs.LG

New LoRA Skills Should Read but Never Write

Low-rank adapters (LoRA) make it cheap to fine-tune a large language model once per task, but combining several independently trained adapters into one model remains difficult: merging the updates in weight space causes interference, retraining on all task data is expensive, and routing between separate adapters gives up the goal of a single combined model. We trace the difficulty to two choices that every composition method makes implicitly. A LoRA update admits infinitely many equivalent factorizations; the choice among them is invisible while an adapter serves alone, but it determines what a learned interaction between adapters can see. A coupling between an old skill and a new one can likewise point in either direction, and the direction decides whether the old skills keep computing what they computed before. We introduce READ (Read-only Expansion of Adapter Deltas), which fixes both choices: each adapter is rewritten into a balanced canonical form that preserves its update exactly, and the coupling grows in one direction only, so a new skill can read the input subspaces of old skills but cannot write into their output subspaces. The only trainable object at each append is the new skill's row of the coupling matrix, and the composed update folds into the base weights with no inference cost, routing, or task-specific rules. We evaluate READ across four benchmark suites and two model families, adding skills one at a time. Across several families, READ improves every suite average over the strongest published baselines built from the same adapters---by more than twenty points on SuperGLUE and more than seven points on the domain suite---and nearly all complete addition sequences end above every direct baseline. Factor coordinates and coupling direction, which a lone adapter never exposes, are what decide whether composed skills survive.