论文

CLEAR-Med 双智能体框架用 SQL 锚定临床数据问答,准确率较 ChatGPT 基线提升 54.4 个百分点

Large Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and Validation

精选理由

医生朋友做临床数据分析的可以看看这个 CLEAR-Med 框架,让模型先写 SQL 再独立校验,准确率从 12% 拉到 66%,还能弃答不乱猜。

论文提出 CLEAR-Med,一个双智能体框架,用 Invocation Agent 把问题翻译成可执行 SQL,再由独立调用的 Validation Agent 做校验、限次修复或弃答。实验基于 21 家医院的 532 条去标识新生儿缺氧缺血性脑病记录,含约 1300 个变量。在 25 个查询的开发基准上重复 5 次共 125 次回答中,Invocation Agent 答对 83 次(66.4%),无 SQL 锚定的 ChatGPT 基线仅答对 15 次(12.0%),配对提升 54.4 个百分点。数值结论始终与已执行 SQL 关联,无法解决的案例可保守弃答。

原文 · arXiv cs.AI

Large Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and Validation

Objective: To develop and characterize CLEAR-Med, a dual-agent framework for natural-language analysis of structured clinical data that separates SQL-based invocation from independent validation. Methods: CLEAR-Med uses one agent to translate a question into executable Structured Query Language (SQL), retain the executed query and database result, and produce a draft. Deterministic checks and a separately invoked cross-provider Validation Agent then accept the draft, request one bounded repair, or abstain. We formalized the system as a bounded selective pipeline and evaluated CLEAR-Med's configuration and scalability, and the Invocation Agent's accuracy and consistency on a 25-query development benchmark, using a harmonized 21-site neonatal hypoxic-ischemic encephalopathy table containing 532 de-identified infant records and approximately 1,300 variables. Results: CLEAR-Med completed all six nominal scalability configurations, including 500x1300. Across 25 development-benchmark queries repeated five times, the Invocation Agent answered 83 of 125 responses correctly (66.4%; query-cluster bootstrap 95% CI, 48.0-83.2%), compared with 15 of 125 (12.0%; 95% CI, 3.2-22.4%) for the ungrounded ChatGPT baseline, a paired improvement of 54.4 percentage points (95% CI, 36.8-72.0%). Conclusion: CLEAR-Med provides a general architecture for traceable analysis of structured clinical data: numerical claims remain linked to executed SQL, and unresolved cases can fail closed. The reported experiments characterize CLEAR-Med's configuration and scalability and the Invocation Agent's accuracy, while the formal analysis establishes the encoded-property guarantee of the complete control flow; a prospective full-pipeline evaluation of the validation and abstention stages is the next stage of this work.