论文

TaH2 自适应循环 Transformer 提升 AIME 测试时扩展效率

Improving Test-Time Scaling with Adaptive Looped Transformers

精选理由

清华团队提出 TaH2,让循环 Transformer 只对需要的 token 多迭代,AIME 上同等算力多拿 3.4 分,代码已开源。

论文研究了循环 Transformer 在测试时扩展中的表现,提出 TaH2 方法,通过 lookahead depth supervision 联合训练骨干网络和迭代决策器,只在有益的 token 上分配额外迭代。在 AIME 基准上,TaH2 将 accuracy-compute slope 提升 53%(2.74 vs 1.79),在相同测试时算力下比非循环基线峰值准确率高约 3.4 点。当最大迭代深度从 2 增至 8 时,TaH2 对基线的领先从 +2.8 点扩大到 +3.9 点,而已有循环模型在此过程中趋于停滞。

原文 · arXiv cs.LG

Improving Test-Time Scaling with Adaptive Looped Transformers

Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs. However, to the best of our knowledge, whether looping improves test-time scaling as outputs grow longer remains underexplored. Through post-training looped transformers, we study the accuracy-compute slope, measured as the accuracy gain per doubling of test-time decoding FLOPs. We find that existing looped transformers often yield steeper slopes than their non-looped baseline, yet underperform it at matched compute. While fixed-depth looping spends extra iterations on every token, our analysis shows that many tokens do not benefit from extra iterations. We therefore propose TaH2, which enables the model to focus extra iterations on the tokens that benefit from looping. It jointly post-trains the backbone and an iteration decider through lookahead depth supervision, which uses online labels indicating whether further iteration improves the prediction. TaH2 improves both the efficiency and attainable accuracy of test-time scaling. On challenging AIME benchmarks, TaH2 improves the accuracy-compute slope by 53% (2.74 vs. 1.79) over the non-looped baseline, exceeding the baseline's peak accuracy by about 3.4 points at matched test-time compute. As the maximum iteration depth increases, existing looped models largely plateau, while TaH2's gain over the non-looped baseline continues to grow from +2.8 points at depth 2 to +3.9 points at depth 8. Our code is available at https://github.com/thu-nics/TaH.