清华论文:AI代理通过比赛回放可登顶排行榜
清华团队展示了AI如何通过学习比赛回放提升游戏表现,详细回放反馈比单纯胜负记录效果好得多。
清华大学新论文发现,AI代理通过学习比赛回放可超越人类排行榜,但在规则复杂的游戏中表现停滞。研究团队构建了包含12款游戏的AAArena平台,使用1,920个历史人类程序作为对手。在吃豆人游戏中,使用详细回放的代理达到第1名,而仅使用胜负反馈的代理仅排名第11名。
New Tsinghua paper finds that AI agents improving game bots from match replays can top human leaderboards, but mostly stall on games with complex rules.
Getting AI to learn a winning game strategy from a limited number of matches is still hard, especially against changing rivals.
They built AAArena from 12 games in Tsinghua's yearly bot-building contest, with 1,920 archived human programs as rivals. A coding agent, with its model weights unchanged, reads the rules, picks opponents, studies replays, and rewrites its bot within a match budget.
Detailed replays beat win/loss-only feedback in all 3 games tested. With replays, a Pacman bot reached rank 1, versus rank 11 without them.
Tripling the match budget did not push any of 4 stuck bots to rank 1.
– arxiv. org/abs/2610.12341
Title: "Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent Competition"