论文

研究:特定起始 token 提示可让 base 模型接近 RL 训练后的推理表现

Base Models Can Reason By Taking a Cue From Training Data

精选理由

不用 RL 训练,只在开头加个 "Okay" 就能让 Olmo-3-7B 数学成绩翻近一倍,这篇论文把原因挖得很透,做模型训练的值得看看。

arXiv 论文研究训练数据如何让 base 模型的起始 token 与后续推理行为形成关联。固定起始 token 线索后,Olmo-3-7B 在 MATH-500 上的 pass@1 准确率从 42% 升到 78%,Qwen3-14B 从 72% 升到 87%,接近各自经 RL 训练的版本。通过因果数据干预,作者能把 "chicken" 这类任意词改造成有效推理线索,让 "Think duck duck goose" 达到 "Think step by step" 的效果。研究还发现不同线索对应训练集中不同文档类型,并影响模型的安全拒绝行为。

原文 · arXiv cs.AI

Base Models Can Reason By Taking a Cue From Training Data

In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows. First, we demonstrate that fixing particular starting token cues makes a base model's performance competitive with that of its reinforcement learning (RL)-trained counterparts on math and coding. For instance, the cue ".\n\nOkay" raises Olmo-3-7B's MATH-500 pass@1 accuracy from 42% to 78%, while "Alright," raises Qwen3-14B's from 72% to 87%. Second, RL makes these cues more likely, while fixing them recovers much of its performance gain over the base model. Third, we trace the reasoning effects of token cues to the training data. We perform causal data interventions to turn an arbitrary word, such as "chicken", into an effective reasoning cue, or remove an existing cue's effect. A similar edit makes the prompt instruction "Think duck duck goose" as effective as "Think step by step" at eliciting reasoning. We also find that the hidden state representations induced by different cues correlate with different document types from the training set. Finally, we extend our study of token cues with a case study in language model safety, finding that different cues elicit distinct refusal and compliance behaviors that correspond to different types of training data.