六款开源模型自动提取间作研究数据,零样本提示效果最好
Can Generative AI Automate Data Extraction for Meta-Analysis? A Case Study on Intercropping Research
农业 meta 分析要人工翻文献提数据,太累。这篇用六款开源模型测了三种提取方法,零样本最靠谱,但 F1 才 0.577,离能用还有距离。
一项研究评估了用 LLM 自动化 meta 分析数据提取的可行性,测试对象是间作(intercropping)文献。研究对比了直接零样本提示、分阶段工作流和多智能体系统三种方法,共涉及六款开源权重模型。结果显示直接零样本提示效果最稳定,相似度调整 F1 平均达到 0.577,但没有一种方法接近完全准确。在下游统计分析中,大多数模型与方法组合能恢复预测变量和结果变量之间关系的方向,但无法准确估计其大小。
Can Generative AI Automate Data Extraction for Meta-Analysis? A Case Study on Intercropping Research
Meta-analysis is the synthesis of information from multiple sources to arrive at an overarching conclusion. There is a large need for meta-analysis in agricultural research to synthesize what is known and analyze overarching patterns. Extracting data from published literature is, however, labor-intensive, time-consuming, and tedious, and is impeded by a lack of standardization in research design, units of measurement, and terminology. These challenges are particularly evident in the domain of crop species mixtures, also called intercropping. With the growing capabilities of LLMs, many recent attempts have focused on building systems and tools to automate data collection, yet rigorous assessment against human-labeled ground truth is often missing. In this research, we evaluate three LLM-based approaches---direct zero-shot prompting, a staged workflow, and a multi-agent system---with six open-weight models to extract data from the intercropping literature. The results are evaluated against the manually curated ground truth and through a downstream statistical analysis. Overall, direct zero-shot prompting is the strongest and most consistent approach, achieving the highest mean similarity-adjusted F1 of 0.577, although none of the approaches is close to fully accurate. In the downstream analysis, most model--approach combinations recover the direction of the relationship between the predictor and outcome variables, but do not estimate its magnitude accurately.