1,042 题哥伦比亚法律基准:LLM 引用法条近半数有错
Do Large Language Models Know Colombian Law? A Reliability Benchmark for the Colombian Legal System
哥伦比亚法律考了 15 个模型,Gemini 3.1 Pro 选择题 0.905 但写答案正确率不到 0.45,引用的法条一半是编的,做法律 AI 的都该看看。
研究团队构建了一个由专家审核的哥伦比亚法律基准,包含 1,042 道题、覆盖十个法律领域和三种题型。15 个模型的选择题准确率从 0.577 到 Gemini 3.1 Pro 的 0.905 不等,但自由文本回答的事实正确率没有任何模型超过 0.45。研究还发现答案相关性与正确性呈负相关(Spearman rho = -0.46),模型听起来答得对但实际经常出错。独立 LLM 评审和人类专家盲评都复现了排序(rho >= 0.88),并发现模型引用的规范约一半是错的或根本不存在。
Do Large Language Models Know Colombian Law? A Reliability Benchmark for the Colombian Legal System
Large language models (LLMs) are increasingly used to support legal practice, education, and research, yet their reliability in national legal systems outside the United States remains largely undocumented. We introduce an expert-validated benchmark for evaluating LLM reliability on the Colombian legal system. The benchmark comprises 1,042 items spanning ten areas of law and three question formats (closed multiple-choice, semi-open, and open-ended IRAC), built through a human-in-the-loop pipeline with multi-stage expert review. We evaluate 15 contemporary proprietary and open-weight models with format-appropriate metrics. Accuracy on closed questions ranges widely, from 0.905 (Gemini 3.1 Pro) to 0.577, but on free-text legal answers factual correctness never exceeds 0.45 (on a 0-1 scale) for any model. We find a dissociation between answer relevancy and correctness (Spearman rho = -0.46): models reliably sound responsive while frequently being wrong, a pattern of particular concern for non-expert users. Closed-question accuracy and free-text correctness are strongly rank-correlated (rho = 0.94), so cheap multiple-choice screening predicts model ranking but overstates absolute reliability. An independent rubric-based LLM judge and blind human expert scoring both reproduce the free-text ranking (rho >= 0.88). The judge further reveals that only about half of the norms models cite are correct; the rest are wrong or non-existent. Reliability varies systematically by legal area and follows an inverted-U across question complexity. Our results indicate that current LLMs require expert supervision for Colombian legal tasks, and that grounding answers in authoritative sources is a promising path to higher reliability. We release the benchmark construction pipeline to support reproducible evaluation.