论文

MADBench评测多智能体辩论安全性

MADBench: Benchmarking the Security of Multi-Agent Debate

精选理由

MADBench首次系统评估多智能体辩论安全性,发现三个恶意智能体只能影响28.3%的正确答案。

研究人员发布MADBench基准,评估多智能体辩论(MAD)的安全性。该基准包含356个源任务和3,958个测试案例,涵盖6种攻击类型。结果显示,MAD不一定能提升大语言模型推理能力,在问答任务中可减轻攻击对准确率的影响,但在工作空间任务中会放大未授权读写行为。

原文 · arXiv cs.AI

MADBench: Benchmarking the Security of Multi-Agent Debate

Multi-agent debate (MAD) can improve large language model (LLM) reasoning by allowing multiple agents to exchange and critique their answers to the same task. However, the interactions that enable agents to correct mistakes can also spread adversarial errors and steer the agents toward an incorrect answer. Although some efforts have been made to examine particular attack types on MAD, systematic evaluation of MAD under diverse attacks remains limited. A central question is whether debate mitigates adversarial influence or amplifies it. In this paper, we present MADBench, a benchmark for evaluating the security of MAD. We organize attacks into a layered taxonomy following the MAD workflow, incorporating both established attacks and new strategies tailored to debate. We evaluate six attack families over 356 source tasks and 3,958 test cases, examining their effects on the final answer and the propagation of adversarial influence. Our results show that, under attacks, MAD does not necessarily improve LLM reasoning. Compared with a single-agent baseline, MAD can mitigate attacks on answer accuracy in question-answering tasks while amplifying unauthorized reads or writes in both question-answering and workspace tasks. Moreover, even when three out of five agents collude, the attack changes the final answer from correct to wrong on only 28.30\% of tasks answered correctly without attack, while only 3.26\% of initially correct honest agents switch to wrong answers during debate.