论文73°

近似值迭代在自我博弈中的有效性研究

The Surprising Effectiveness of Approximate Value Iteration in Self-Play

精选理由

这篇论文发现近似值迭代(AVI)在多个棋类游戏中比 AlphaZero 更准确且成本更低,挑战了 MCTS 主导地位。

研究人员在 Connect Four、Hex(7x7) 等中等规模游戏中测试了近似值迭代(AVI)方法。实验表明,AVI 学习的价值函数比 AlphaZero 更准确,且单步前瞻贪心策略与基于蒙特卡洛树搜索(MCTS)的策略竞争性相当。AVI 在训练和推理成本上显著低于 MCTS 方法,在 Othello 和 Go(9x9) 上也展现出稳定训练能力。

原文 · arXiv cs.AI

The Surprising Effectiveness of Approximate Value Iteration in Self-Play

Combining search with function approximation has driven major advances in game-playing programs, making self-play algorithms more competitive than ever. Still, the computational overhead of the most popular methods, based on Monte Carlo Tree Search (MCTS), can be substantial. In this work, we investigate whether simpler methods remain competitive in non-trivial, moderately sized games such as Connect Four, Hex(7x7) and synthetic games. We train a minimal self-play implementation of Approximate Value Iteration (AVI) and use ground-truth oracles for exact evaluation. Contrary to expectations, our results demonstrate the surprising effectiveness of AVI: it learns more accurate value functions than those learned by AlphaZero, while its one-step-lookahead greedy policies remain competitive with MCTS-based policies at substantially lower training and inference costs. Preliminary experiments on Othello and Go(9x9) show that AVI trains stably on larger games and learns effective value functions. These findings suggest that the success of MCTS-based methods may have eclipsed simpler approaches that have become increasingly practical with modern deep-learning tools.