论文

MA-WAM:多智能体世界动作模型实现测试时规划

MA-WAM: Multi-Agent World-Action Model for Test-Time Planning

精选理由

多智能体协作一直难在联合动作的相互影响上,这篇提出了首个给多智能体流策略做测试时规划的世界模型,30 个设置平均提升 22% 还只慢 12 毫秒,做 MARL 的可以看看。

MA-WAM 是一个测试时规划框架,让冻结的多智能体流策略对候选联合动作的未来结果进行评估。据论文所述,这是首个面向多智能体流策略的测试时世界模型规划器,通过跨智能体依赖关系预测联合动作后果并高效评分。在 MAMuJoCo、SMAC 和 MPE 共 30 个离线 MARL 设置上,MA-WAM 相比直接执行平均相对提升 22.0%,相比均匀动作选择提升 25.6%。在 A100 GPU 的标准评测协议下,每次规划仅增加 12.1 ms 开销,占生成与评分总时间的 2.5%。

原文 · arXiv cs.AI

MA-WAM: Multi-Agent World-Action Model for Test-Time Planning

Multi-agent cooperative tasks require different agents to execute a joint action simultaneously, and each agent's action affects both the observations and responses of the other agents. Hence, a world model is needed to predict the team return resulting from the joint actions of all agents. A naive extension directly applies a single-agent world model to each agent's action when predicting the team return step by step. However, such an extension fails to capture the dependencies among the simultaneous actions of multiple agents. We propose Multi-Agent World-Action Model (MA-WAM), a test-time planning framework that enables a frozen multi-agent flow policy to evaluate futures of candidate joint actions. To our knowledge, MA-WAM is the first test-time world-model planner for multi-agent flow policies. MA-WAM predicts the consequences of each joint action according to cross-agent dependencies and enables efficient candidate scoring. Across 30 offline multi-agent reinforcement learning (MARL) settings on MAMuJoCo, SMAC, and MPE, MA-WAM achieves mean relative gains of 22.0% over direct execution and 25.6% over uniform action selection. Under the standard evaluation protocol on an A100 GPU, MA-WAM adds 12.1 ms, accounting for 2.5% of the measured generation-and-scoring time.