Forking: Sudden Overfitting Under Replay
Deep学习又发现新现象'分叉',比Grokking和Double Descent还奇怪,AutoResearch agents的技巧可能带来意外副作用。
该研究在NanoGPT模型中发现了名为'分叉'的泛化失败现象。在数据重放条件下,具有过编码n-gram记忆分支的模型在训练和验证损失间出现急剧分离,形状类似叉子。研究团队在标准NanoGPT设置中复现了这一现象,并在DeepSeek风格的Engram模型中观察到相同效果。低频上下文对损失差距贡献最大,而更大的训练预算和密集数据表会抑制这种现象。
This paper studies forking, a generalization failure discovered in NanoGPT autoresearch. Under data replay, models with an over-encoding n-gram memory branch show a sharp separation of training and validation loss at epoch boundaries, resembling the shape of forks. We study this phenomenon in a controlled vanilla NanoGPT setting and reproduce it in a DeepSeek-style model with Engram. Mechanistically, repeated updates sharpen the continuations observed in training while suppressing the probability of unseen continuations, whose loss grows with each pass. The n-gram module creates weakly interacting context-specific subspaces, amplifying this effect. Low-frequency contexts contribute most of the gap, whereas larger training budgets and heavily crowded tables suppress it. We also observe forking in short-budget, heavily repeated SFT and RL-like regimes. The contributions of this paper are twofold: (1) Forking reveals yet another curious phenomenon in deep learning, in addition to grokking and double descent. (2) Forking is an unexpected and unpleasant by-product of tricks proposed by autoresearch agents. While these agents produce an enormous number of results that seem useful, we should always be careful with their results.