QAM 提出二次精度检查点合并方法,误差达 O(h^3) 最优界
QAM: Quadratic-Accurate Checkpoint Merging via Sequential Consistency
一篇理论味很浓的 checkpoint 合并论文:证明信息论下界,给出达到 O(h^3) 最优误差的合并系数,还在 SmolLM3-3B 上做了实测对比。
论文 QAM 研究如何用训练轨迹上保存的 checkpoint 精确重构串行训练的终点。在局部转移模型下,两个 checkpoint 索引矩条件刻画了所有与串行参考二阶一致的凸合并方案,并证明了 GD 历史无法达到 o(h^3) 误差的信息论下界。QAM 方法达到匹配的 O(h^3) 一致误差界,其显式系数在所有固定二次目标上精确匹配串行 GD 参考。在 SmolLM3-3B 和 OpenEuroLLM-Prelude-9B 两条 Adam checkpoint 轨迹、每模型 3 个窗口和 3 种配置、共 15 个任务的实验中,QAM 在长窗口场景下优于 Warmup-Stable and Merge(WSM),短窗口表现则好坏参半。GSM8K 诊断还显示局部一致性并不完全决定下游成绩。
QAM: Quadratic-Accurate Checkpoint Merging via Sequential Consistency
Saved checkpoints record states along a training trajectory, but generally do not determine the updates at states that would be visited under a different schedule. We study how accurately these checkpoints can reconstruct the endpoint of a sequential reference with prescribed update strengths. Under a common local transition model, two checkpoint-index moment conditions characterize all convex merges that agree with this reference through second order. We then prove an information limit that for nondegenerate profiles, no algorithm using only a fixed-length gradient-descent (GD) history with step size $h$ can achieve $o(h^3)$ endpoint error uniformly over a fixed class of smooth, strongly convex losses. The lower bound follows from two losses with identical GD checkpoint histories but sequential reference endpoints separated by $Ω(h^3)$. \textbf{Quadratic-Accurate Merging} (QAM) achieves a matching uniform $O(h^3)$ endpoint error bound. Its explicit coefficients also define the unique profile-dependent merge that exactly matches the sequential GD reference across all fixed quadratic objectives. Across two public Adam checkpoint trajectories (SmolLM3-3B and OpenEuroLLM-Prelude-9B), three windows and three profiles per model, and 15 tasks, QAM shows mixed results for short windows and broader advantages over \textbf{Warmup-Stable and Merge} (WSM) for longer windows. Matched-moment GSM8K diagnostics further show that local consistency alone does not fully determine downstream scores. These results characterize the reconstruction limits of saved histories, provide a coefficient rule that attains the optimal rate, and assess its practical utility.