Kaggle 推出 Game Arena:用对弈游戏评估大模型策略能力
Game Arena: Strategic LLM Evaluation in Competitive Environments
Kaggle 做了个让模型下棋、打扑克、玩狼人杀互相打分的评测场,比静态跑分更能看出策略水平,可以直接看各家模型的对局结果。
Kaggle 发布 Game Arena,一个通过竞技游戏评估大语言模型的开放平台,模型之间进行正面交锋的对局,避免静态基准的性能饱和问题。首批提供 Chess、Poker 和 Werewolf 三个游戏环境,分别对应完全信息、不完全信息和多人博弈场景。该平台用于系统研究模型的策略规划、适应能力和不确定性下的稳健性,并基于大规模真实对局结果提供可复现的评估。
Game Arena: Strategic LLM Evaluation in Competitive Environments
We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena enables models to play head-to-head matchups in structured environments where the gameplay strength naturally increases as models evolve, preventing performance saturation. This technical report details the infrastructure behind Game Arena and describes the three pilot game environments: Chess, Poker, and Werewolf. These environments span perfect information, imperfect information, and multiplayer game settings, enabling a systematic study of models' strategic planning, adaptation, and robustness under uncertainty. For each game, we provide a detailed description of the environment, evaluation metrics, and results from running full competitions across models. Through robust infrastructure and large-scale ground-truth based evaluation, Game Arena ensures reproducibility, transparency and generalizability to new games and variants over time.