AI Night-Scientist:用强化学习提升大模型科研创意能力
Reinforcing Agentic Creativity in Scientific Ideation with Night Science
这篇论文教模型“跳出常规”做科研构思,方向多了近三成,原创性涨 66 分,做科研的人可以看看怎么让 LLM 帮你想点子。
arXiv 论文提出 AI Night-Scientist,一个基于强化学习的智能体框架,用于解决 LLM 在开放式科学构思中输出同质化的问题。该方法将创造力建模为行动、过程、结果三个维度,并用 GRPO 训练模型学会何时偏离可预测推理。相比基座模型,生成的研究方向扩展 27.8%,贡献类型扩展 14.9%,预测引用影响力最多提升 32.0 个百分点,原创性提升 66.2 分。实验显示仅调高解码温度无法复现这些收益,关键在于指定创造力类型的语义引导。
Reinforcing Agentic Creativity in Scientific Ideation with Night Science
Large language models (LLMs) excel at structured, verifiable tasks, but their low-entropy bias can produce homogeneous and predictable outputs, limiting their utility for open-ended scientific ideation. Effective discovery, however, spans a broader creative spectrum: from structured day science to loosely structured, serendipitous night science that reaches ideas beyond those typically considered. We introduce AI Night-Scientist, an agentic framework that uses reinforcement learning to teach models when and how to depart from predictable reasoning. Grounded in cognitive science, we model creativity along three axes: action (what to do and how creatively), process (when to explore versus exploit), and outcome (the novelty and usefulness of the resulting idea). We use these axes to train models with GRPO, exposing them to varying degrees and forms of creativity throughout training. This produces substantially more diverse scientific proposals, expanding the range of research directions by 27.8% and contribution types by 14.9% over the base model. It also improves predicted citation impact by up to 32.0 percentage points and originality by 66.2 points. These gains cannot be reproduced by simply increasing decoding temperature; instead, we find that semantic guidance specifying what kind of creativity to pursue is critical. Overall, our results suggest that creativity is a learnable, multi-level ability that can be shaped to help researchers reach ideas beyond those typically explored by LLMs.