论文精选

BioEVAL 基准测试:前沿 LLM 生物工程答题最高达 90% 准确率

精选理由

一篇生物工程版的 LLM 考试报告,前沿模型选择题能考到 90%,但一到实验设计类题就露怯,各子领域强弱差异也列得很清楚。

多机构联合发起的 BioEVAL 基准用于评估 LLM 在生物工程各子领域的实验推理能力。多家云端模型在 MCQ 基准上整体准确率超过 85%,领先模型达到 90%。实验类问题对多数模型比概念类问题更难。能力分布不均:Diagnostics、Biosensing、Neuroengineering、Immunoengineering 等领域表现较高,而 Drug Delivery、Genetics、Systems and Synthetic Biology、Bioimaging 等领域多数模型得分偏低。

原文 · Tanishq Abraham (论文推介)

Found an interesting paper as a former biomedical engineer...

"We assembled BioEVAL (BioEngineering Validation of AI and LLMs), a global, multi-institutional initiative designed to assess experimental reasoning capability across bioengineering (BE) subfields."

"A central finding is that current frontier LLMs already demonstrate substantial BE domain knowledge. Several cloud-scale models achieved greater than 85% overall accuracy on the MCQ benchmark, with the leading model reaching 90% accuracy."

"experimental questions were consistently more difficult for most models than conceptual questions."

"Subfield-resolved performance further showed that LLM capability is not uniformly distributed across BE. Some areas, including Diagnostics, Biosensing, and Bioelectronics, Neuroengineering and Neurobiology, and Immunoengineering, showed consistently high MCQ performance across many models. In contrast, Bioimaging, Biophotonics, and Optics, Biomaterials and Biomolecules, Genetics, Systems and Synthetic Biology, and Drug Delivery, Therapy, and Nanomedicine showed lower performance across most models"

link: https://t.co/i8f9oWqCNM