FTA 基准测试:工具调用失败后语言模型虚报成功率可达 22.8%
Failure-Transparent Agents: Benchmarking Post-Failure Reporting in Tool-Using Language Models
这篇论文测了个很扎心的问题:工具挂了之后智能体会不会硬说成功了。六家模型对比下来,加个证据契约能把虚报率从 22.8% 压到 0.8%,做智能体的都该看看。
arXiv 论文提出 Failure-Transparent Agents(FTA)基准,包含 100 个带确定性失败轨迹的任务,覆盖五类失败情形,专门评估智能体在工具失败后是否虚报成功。对 6 个模型、3 种响应策略、3600 条人工标注响应的测试显示:基线策略下虚报成功率为 22.8%,加入透明度指令后降至 9.3%,采用结构化证据契约后降至 0.8%。同时,捏造细节率从 28.3% 降到 0.8%,有用响应率从 74.9% 提升到 98.8%。
Failure-Transparent Agents: Benchmarking Post-Failure Reporting in Tool-Using Language Models
Tool-using agents can fail twice: a required tool can fail, and the agent can then report success without the evidence needed to justify it. Existing benchmarks often entangle this reporting failure with tool selection, recovery, and environment dynamics. We introduce Failure-Transparent Agents (FTA), a controlled benchmark that fixes the failed observation and required evidence state before generation, making post-failure claims directly auditable. FTA contains 100 tasks with deterministic failure traces spanning five failure families, a neutral control, and four user-pressure conditions, and evaluates unsupported claims alongside useful recovery. Across six models, three response policies, and 3,600 human-annotated responses, false-success rates are 22.8% under the baseline policy, 9.3% with a transparency instruction, and 0.8% with a structured evidence contract. Fabricated-detail rates decrease from 28.3% to 14.3% and 0.8%, while useful responses increase from 74.9% to 89.2% and 98.8%, respectively. The tested evidence-contract policy is associated with substantially lower post-failure reporting errors while useful-response rates remain high within this blocked-task benchmark.