法律大模型引用权威与实际决策一致性研究
Cited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought Faithfulness
发现法律大模型只是'假装'遵循引用的法律权威,实际判决对法律依据变化不敏感,这对法律AI应用是个警示。
研究测试了7个开源法律大模型(8B-70B参数)在4个基准测试中的表现。当要求模型解释判决依据时,模型能正确命名相关法律权威的比例达66.7%-100%。但当实际替换法律权威时,判决随之改变的比例却低得多:CaseHOLD基准为0.0%-21.7%,ECHR和SCOTUS基准为30.0%-76.7%,ContractNLI基准为43.3%-50.0%。专门优化的法律推理模型也无法缩小这一差距。
Cited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought Faithfulness
Large language models increasingly justify legal decisions by naming the statute or precedent behind a verdict, treated as evidence that the decision follows from it. We test this directly: holding case facts fixed, we substitute the named legal authority for an unrelated one and decode a model's evolving verdict from its hidden states. Across seven open-weight models (8B-70B) and four benchmarks spanning judicial and contractual reasoning, when explicitly required to justify a verdict by naming the governing authority, models name the correct one in 66.7%-100% of generations, while the verdict changing when the authority changes is far less consistent: 0.0%-21.7% on CaseHOLD, 30.0%-76.7% on ECHR and SCOTUS, and 43.3%-50.0% on ContractNLI. Neither scale nor a purpose-built legal-reasoning model (a best-effort LoRA reproduction; Section 6) closes this gap. A red-teaming evaluation on five core models finds compliance with an adversarial instruction hidden in the case facts (73.3%-96.4%) exceeds verdict-swap sensitivity by a wide margin, holding without exception across model rankings. Naming a legal authority is thus a poor proxy for a verdict's dependence on it, while the same verdict remains separately vulnerable to adversarial manipulation. Both findings replicate across checks ruling out prompt-wording noise and confounded sampling, and bear directly on the use of generated legal explanations as compliance or audit artefacts.