论文

百万参数排序模型做个性化生成,打分延迟降低四个数量级

Rethinking Personalized Generation: Test-Time Alignment via Factorized Ranking Models

精选理由

这几个作者发现个性化生成其实卡在匹配而不是生成能力上,用不到0.4%参数的小 MLP 就把十亿级 reward model 全线打过了,思路很值得琢磨。

论文提出用测试时对齐来解决 LLM 个性化生成问题,核心做法是把个性化视为候选匹配问题。作者训练百万参数规模的 MLP 排序模型,直接复用基础生成器的内部嵌入,替代十亿参数级 reward model。在覆盖三种个性化场景的九个数据集上,该模型在全部数据集上超过十亿参数通用 reward model,参数量不到后者的 0.4%,打分延迟低四个数量级。排序模型还能引导生成过程,减少生成 N 个候选的开销。

原文 · arXiv cs.LG

Rethinking Personalized Generation: Test-Time Alignment via Factorized Ranking Models

Aligning large language models (LLMs) to diverse user preferences is fundamentally hindered by standard alignment paradigms that optimize for monolithic users. In this work, empirical studies are first used to reveal the existence of a massive, untapped performance headroom for personalized generation through test-time alignment. We demonstrate that personalized generation is uniquely suited for test-time scaling methods like Best-of-N (BoN) because it can be viewed primarily as a candidate matching problem rather than a generator capability bottleneck. While reward models could in principle exploit this headroom, they are poorly calibrated for personalization, and their billion-parameter scale makes scoring large candidate pools prohibitively expensive. To overcome this limitation, we propose a parameter-efficient framework utilizing million-parameter scale multi-layer perceptron (MLP) ranking models. Our personalized ranking model directly reuses the internal embeddings of the base generator with minimal overhead. By scaling train-time data to provide fine-grained personalized preferences, this million-parameter ranking model accurately scores large candidate pools and can seamlessly guide generation to reduce the cost of materializing N candidates. Extensive experiments on nine datasets spanning three personalized generation settings show that our personalized ranking model effectively exploits the discovered headroom, outperforming billion-parameter generalist reward models on every dataset, with under 0.4% of their parameters and four orders of magnitude lower scoring latency.