论文精选

Amazon 论文:用 LLM 预测训练实验收益

精选理由

Amazon 这篇论文说,把历史实验记录喂给 LLM,它就能提前猜出哪个训练改动有用,OLMo3-100M 上相关性冲到 0.892,比硬堆推理努力管用。

Amazon 发表论文,将 LLM 用作研究世界模型,在消耗 GPU 前预测训练改动是否有效。实验基于 9 个设置下的 2653 条真实实验记录,覆盖预训练到推理。加入同一设置的过往记录后,5 个设置的平均排序相关性从 0.506 升至 0.774。在 OLMo3-100M 设置上,低推理努力加记录得 0.892,最高努力无记录仅 0.648。

原文 · rohanpaul_ai

New Amazon paper shows that an LLM can act as a research world model, predicting whether a training change will help before anyone spends GPU time on it.

AI research agents can propose experiments far faster than teams can afford to run them. Choosing what gets GPU time means guessing outcomes in advance.

They used an LLM as a research world model that predicts an experiment's gain before it runs. They tested it on 2,653 real experiment records from 9 setups, from pretraining to inference.

Past records from the same setup raised average ranking correlation with actual results from 0.506 to 0.774 across 5 setups. On the OLMo3-100M setup, adding records at low reasoning effort scored 0.892, while max effort without records reached 0.648.

If you run research agents, log every experiment, including failures, and feed those records to whatever model picks the next run.