UserProxyBench评估LLM用户模拟器
UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training
新评估框架揭示了用户模拟器对代理测试的影响,能帮你选到最符合需求的模拟器。
研究人员推出UserProxyBench评估框架,用于衡量代理基准测试中用户模拟器的表现。该框架引入用户保真度评分(UFS),独立于代理成功度评估用户是否正确执行角色。在375个企业任务测试中,固定GPT-5.5代理,仅改变用户代理,平均任务奖励变化15.2分。24.4%的成功任务存在用户规范违规,主要问题是过早披露信息。
UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training
Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated user controls what information the agent receives and when, yet current benchmarks score only the agent and do not directly measure whether the user correctly executed its assigned role. We introduce UserProxyBench, an evaluation layer over the tau-bench family, and the User Fidelity Score (UFS), which measures adherence to the benchmark's private user instructions using task-grounded rubric criteria scored independently of agent success. Holding the agent fixed at GPT-5.5 and varying only the user proxy across 375 enterprise tasks changes mean task reward by 15.2 points, while 24.4% of successful episodes contain a user-specification violation. The dominant failure is premature disclosure: users provide information before it is requested. This behavior has little effect on task reward, yet among successful episodes it causes the agent to make 1.06 fewer tool calls on average, changing the interaction being evaluated while preserving the reward. Finally, across seven proxies we identify an empirical cost-fidelity frontier, enabling practitioners to select the least expensive simulator that satisfies a required fidelity level.