论文

论文提出 Nash decoding:用博弈论方法改写文本生成

Nash Equilibrium Text: A Game-Theoretic Decoding Framework for Text Generation

精选理由

有人把文本生成当成博弈来解,用小模型算出 Nash 均衡,问答成绩超过大 18 倍的自回归模型,思路挺新颖。

arXiv 论文《Nash Equilibrium Text》把文本改写建模成一个博弈:token 位置是玩家,词表是动作,效用函数是语言模型的对数条件概率。论文证明随着序列长度增加,Nash 均衡的似然可以指数级高于自回归输出。作者提出 Nash decoding 算法,在给定条件联合概率下可在 O(1/ε) 时间内逼近 ε-Nash 均衡。在 CLAPNQ、PubMedQA、CoQA 三个问答基准上,用掩码语言模型得到的均衡结果在 F1 和 ROUGE 上超过最大 18 倍规模的自回归模型,且无需微调或重训练。

原文 · arXiv cs.AI

Nash Equilibrium Text: A Game-Theoretic Decoding Framework for Text Generation

Text revision has become an integral component of large language models. This paper formulates revision such that it admits a Nash equilibrium: Token positions are players, vocabulary items are actions, and each player's utility is the language model's log conditional probability. We motivate the revision by showing that Nash equilibria can have exponentially higher likelihood than autoregressive outputs as the sequence length grows. We further propose Nash decoding, an algorithm that reaches an $\varepsilon$-Nash equilibrium in $O(1/\varepsilon)$ time given access to the joint probability of tokens conditioned on a prompt. In practice, we run Nash decoding using conditional probability estimates from large language models and evaluate the resulting equilibria on question-answering benchmarks. On CLAPNQ, PubMedQA, and CoQA, Nash equilibria obtained from masked language models achieve higher F1 and ROUGE scores than autoregressive models up to $18\times$ larger, without any fine-tuning or retraining, at the cost of additional test-time computation.