ScienceClaw:衡量科学智能体持续自我进化的新基准
ScienceClaw: Benchmarking Continual Self-Evolution of AI-for-Science Agents Across the Natural and Social Sciences
一群研究者做了 ScienceClaw,专门测科研智能体能不能边干活边自我改进,覆盖 23 个学科,代码已开源,做智能体进化的可以看看。
ScienceClaw 是一个研究论文提出的基准框架,用于评估 AI-for-Science 智能体的持续自我进化能力。它将自我进化形式化为固定参数的程序级更新,统一任务求解、科学验证与程序更新三个环节。ScienceClaw-Eval 覆盖 23 个学科,通过序列任务流和独立重置评估来测量科学正确性、进化增益、保留率、跨数据集迁移和进化成本。框架只在源任务重放复现修复、且独立科学任务有提升时才保留更新,代码已在 GitHub 开源。
ScienceClaw: Benchmarking Continual Self-Evolution of AI-for-Science Agents Across the Natural and Social Sciences
Large language model agents are accelerating scientific automation, yet verified executions rarely become persistent program-level improvements, and existing evaluations do not examine this process across sequential tasks in both the natural and social sciences. We formalize ScienceClaw as fixed-parameter program self-evolution that unifies task solving, scientific verification, and program updates. ScienceClaw-Eval spans 23 disciplines and measures scientific correctness, evolutionary gain, retention, cross-dataset transfer, and evolution cost through sequential streams and independent reset evaluation. Our framework repairs executable workflows through multi-turn interaction, converts re-execution-verified failure--success trajectories into linked Skill and Operator candidates, and retains an update only when source-task replay reproduces the repair and independent scientific tasks improve. Code is available at https://github.com/beita6969/ScienceClaw.