SciExam for ENSO 基准:AI 智能体能自主构建气候模型吗
SciExam for ENSO: Can AI Agents Build Climate Models?
12 个智能体里 6 个做出的气候模型比已发表的还好,连 ENSO 暖冷不对称的争议都摸到了边,这套无标准答案的评测方法也值得做 agent 评测的人参考。
arXiv 论文提出 SciExam for ENSO 基准,让语言模型智能体在 6 小时内基于真实观测数据构建 ENSO(厄尔尼诺-南方涛动)的低阶随机气候模型。评测不依赖已知答案或 LLM 评审,而是用隐藏评分器检验模型能否复现 ENSO 统计特征、还原未观测变量并预测留存年份。在 12 个智能体系统中,6 个的模型得分超过已发表的模型,主要优势在重构和预测环节。其中最强的模型结构分别与 ENSO 暖冷不对称性的两种竞争解释之一相容,而这一开放学术争议并未在任务中提及。
SciExam for ENSO: Can AI Agents Build Climate Models?
Language-model agents are increasingly asked to carry out open-ended scientific research, yet their results are usually graded against a known answer, a rubric, or a language-model reviewer, none of which can tell whether a new scientific model is valid. The AI Science Exam for El Nino-Southern Oscillation (SciExam for ENSO) is a benchmark in which agents build low-order stochastic models of ENSO, the dominant mode of interannual climate variability, from real observations. Within a six-hour budget, agents process the observations, write their own diagnostics, which are then frozen, and develop a model using only these diagnostics as feedback. Hidden graders then test whether the model reproduces ENSO's statistics, recovers unobserved variables, and forecasts held-out years, and score a published model in the same way. Across twelve agent systems, six produce models that score higher than the published model, mainly through better reconstruction and forecasting. The simplified forms of the stronger models are each compatible with one of the two competing explanations of ENSO's warm-cold asymmetry, an open debate that the task never mentions. Controlled runs of the top system under varied information suggest that its scores do not come from recalling the dated observational record and that the information it receives shapes how it builds its model. SciExam for ENSO can thus evaluate agent research where no answer is known, and the results suggest that agents can already build competitive models whose structures bear on questions that scientists still debate.