BioPhys-Bridge 基准数据集发布,用于评估跨学科生物物理研究中的科学推理能力
BioPhys-Bridge: A Benchmark for Interdisciplinary Scientific Reasoning in Physics-Grounded Biological Research
朋友发了这个新基准数据集,专门用来测试模型在生物物理领域的科学推理能力,比之前的基准更复杂,能看出不同模型在这个专业领域的真实水平。
为解决语言模型在分析生物物理领域跨学科研究文献时的挑战,研究人员推出了 BioPhys-Bridge 基准数据集。该数据集包含 500 个案例,用于评估模型在基于证据进行科学推理的能力。初步评估显示,DeepSeek-V4-Flash 在证据标识符 F1 分数上表现最佳(0.360),其次是 Qwen3.7-Max(0.316)和 GPT-4o-mini(0.294)。
BioPhys-Bridge: A Benchmark for Interdisciplinary Scientific Reasoning in Physics-Grounded Biological Research
Language models face unique challenges in analyzing interdisciplinary scientific research literature. In biophysics research, faithful answers require grounding observed data in source evidence, interpreting it through a quantitative physics model, and linking it to a biological mechanism. To address this challenge, we introduce BioPhys-Bridge, a novel benchmark dataset for evidence-grounded scientific reasoning over biophysical literature. Each case contains evidence blocks, stable evidence IDs, quantitative values, units, equations, assumptions, mechanisms, and next decisions as grounding targets for question answering (QA) and retrieval-augmented generation (RAG). The initial release contains 500 cases, 1,517 agent-facing tasks, and covers six biological domains and nine physical model families, including three sparse families reserved for future expansion. We enforce strict quality gates for all cases in schema, evidence-integrity, quantitative-grounding, source-license, duplicate, unit-normalization, with domain expert review and annotation for 81 cases. Preliminary evaluations show that DeepSeek-V4-Flash obtain the highest evidence-ID $F_1$ score (0.360), followed by Qwen3.7-Max (0.316) and GPT-4o-mini (0.294). BioPhys-Bridge is an interdisciplinary benchmark for evaluating attribution, faithfulness, hallucination reduction, and biological experiment design with complex, multi-step scientific reasoning. Future works will increase the size and complexity of the dataset and perform comprehensive evaluations. Code and data are available in the GitHub repository and on Hugging Face.