研究提出 CoVer 方法:LLM 无法自行判定代理验证器的检查边界
Not Self-Decidable: LLMs Cannot Draw the Boundary of What an Agent Verifier Can Check
论文测了 4 个模型、2.2 万条标签,发现 LLM 在监管规则上判断不可靠,还给出 CoVer 这个修复方案,做合规验证的值得看看。
arXiv 论文《Not Self-Decidable》研究了代理验证器如何决定哪些规则可以由固定检查验证、哪些需要人工评审。团队在包括 EU AI Act、FINRA 指南和一个已部署信贷代理在内的 6 个语料库上收集约 22,000 条标签,涉及 3 家实验室的 4 个模型,发现各模型在监管文本上判断崩溃且错误方向相反。论文提出 CoVer(corroborate-then-verify)方法,只有当为某谓词合成的检查能通过干预测试时才采纳该谓词。对照实验显示,与参考评审模型达成一致无法提供认证:一致率从 30% 升到 77%,但真正可判定的比例没有变化。
Not Self-Decidable: LLMs Cannot Draw the Boundary of What an Agent Verifier Can Check
A verifier for an agent faces rules of two kinds: the ones a fixed check can settle and the ones that require a judge. A team that derives its own checks fixes that split up front. Where the requirements come from outside, as in finance, healthcare and law, the agent enforces rules it did not write, so the split falls to runtime, recurring for every predicate of every rule on every action at a rate no reviewer can audit. Every escalation scheme assumes a model can make that decision itself, that it is self-decidable. Across six corpora, including the EU AI Act, FINRA guidance and a deployed credit agent, we collect roughly 22,000 labels from four models built by three labs. They agree almost perfectly where the answer is obvious and collapse on regulatory text; their errors run in opposite directions, so no model can be trusted as the conservative choice; and on the deployed agent's own rule-set they err together, over-claiming that a fixed check will do, the direction that never gets escalated. We introduce CoVer (corroborate-then-verify), which treats unanimity as a nomination, admitting a predicate only when the check synthesized for it survives intervention, reading fields the agent cannot write and holding under deterministic rewording. That gate rejects most of what corroboration wrongly admits, at a cost in coverage we report rather than tune away. The obvious alternative, agreement with a reference judge, certifies nothing: it climbs from 30% to 77% across calibration bands while the genuinely decidable share does not move, because a judge drawn from the population under indictment ratifies the blind spot it shares. Self-decidability is not a capability to elicit from a model but a boundary the verifier must construct.