Harvey LAB-AA v1.1 加入幻觉检查,Grok 4.7 以 9.4% 登顶法律基准
法律基准加了幻觉门槛后排名大洗牌:Muse Spark 本来领先,过滤幻觉后掉到第二,GPT-6 Astra 几乎不缩水。看各家法律任务幻觉差距挺有意思。
Artificial Analysis 与 Harvey 联合更新 Legal Agent Benchmark 评分方法,推出 LAB-AA v1.1,新增幻觉检查,只有满足全部评分标准且不含重大幻觉的任务才计入新指标 Hallucination-Gated All-Pass Rate。Grok 4.7(xhigh)以 9.4% 排名第一,Muse Spark 1.3(max)和 GPT-6 Astra(max)分别为 8.9% 和 8.6%。Muse Spark 1.3 在不加幻觉门槛时以 26.7% 领先,但其三分之二的通过结果含重大幻觉;GPT-6 Astra 在 120 个任务中平均每次任务仅 0.03 个重大幻觉,Gemini 3.8 Flash(high)则高达 13.96 个。成本方面,榜首的 Grok 4.7 每任务约 9.50 美元,不到最贵的 Claude Fable 5.1(约 21.70 美元)的一半。
Today we are announcing Harvey LAB-AA v1.1 in collaboration with Harvey. This updates our scoring methodology for the Legal Agent Benchmark (LAB) to add a hallucination check and require correct responses to not include material misstatements. LAB-AA v1.1's new headline metric, Hallucination-Gated All-Pass Rate, only credits a task when the deliverables satisfy every rubric criterion and contain no material hallucinations.
We define a material hallucination as one that would mislead a reader on a substantive point, such as a wrong contractually required date, while a minor hallucination is a real error that is unlikely to meaningfully affect the legal interpretation of a deliverable. Minor hallucinations are reported separately and do not affect the headline score.
Grok 4.7 (xhigh) leads at 9.4% Hallucination-Gated All-Pass Rate, and >60% of otherwise passing results across the models tested at launch contain a material hallucination.
This is the first step in enhancing the methodology for Harvey LAB-AA. In future updates, we’re working with Harvey to better account for the full set of factors lawyers value, including usability features like style and tone.
Key takeaways:
➤ Top of the leaderboard: Grok 4.7 (xhigh) from @SpaceXAI leads at 9.4% Hallucination-Gated All-Pass Rate, narrowly ahead of Muse Spark 1.3 (max) from @AIatMeta at 8.9% and GPT-6 Astra (max) from @OpenAI at 8.6%
➤ Hallucinations reshape the leaderboard: without the hallucination gate, Muse Spark 1.3 (max) would lead clearly with a 26.7% all-pass rate, but two thirds of those passes contain one or more material hallucinations, resulting in a Hallucination-Gated All-Pass Rate of 8.9%. GPT-6 Astra (max) retains almost all of its passes after the hallucination gate (8.9% to 8.6%) and moves from joint 10th to 3rd
➤ The GPT-6 model family is the most grounded: GPT-6 Astra (max) averages 0.03 material hallucinations per task (4 across all 120 tasks) and GPT-6 Sol (max) averages 0.07. When comparing six checker models on a 20-task subset, GPT-6 Astra had zero material hallucinations under every checker, including Claude Opus 5.5 (high). By comparison, Gemini 3.8 Flash (high) averages 13.96 material hallucinations per task.
➤ Top-scoring models aren’t the most expensive: Grok 4.7 (xhigh) takes the top spot on the leaderboard at a cost of ~$9.50 per task, under half the cost of Claude Fable 5.1 (max with fallback) at ~$21.70, the most expensive model. Muse Spark 1.3 (max) represents strong performance for its cost, landing in second for ~$4.20 per task
Thank you to @nikogrupen and @ItsJulioPereyra from @harvey for their work on Harvey LAB and collaboration.