论文

LLMs高估历史温度研究

Who Warmed the Archives? LLMs Overestimate Historical Warmth

精选理由

六种主流LLM在历史气候分析中普遍存在'暖化偏差',相关性指标不足以评估其可靠性。

研究比较了六种大语言模型在德国五世纪文本中提取温度指数的表现。所有LLM模型均表现出随年代增长的正向偏差,斜率在+0.13至+0.34/世纪之间。Gemini 2.5 Flash相关性最佳(r=0.32),但误差是最佳词汇方法的两倍。移除文本中的日期标记后偏差趋势不变,表明模型倾向于将历史数据误判为现代数据。

原文 · arXiv: DeepSeek

Who Warmed the Archives? LLMs Overestimate Historical Warmth

Historical archives are an under-used source for extending the instrumental climate record backward in time, and LLMs offer a way to extract the indices climatologists derive by hand. Beyond measuring how well systems extract this signal, we check whether their errors are safe to use for cross-century comparison, since a good correlation score does not rule out systematic, era-linked bias. Comparing lexical baselines, fine-tuned historical transformers, and LLM prompting on the Pfister temperature index across five centuries of German text, lexical methods beat every fine-tuned transformer we test, including one pretrained from scratch on historical German (r=-0.016). All six LLMs we test (Gemini 2.5 Flash, GPT-5-mini, DeepSeek v4 Flash, Claude Sonnet 4.6, Qwen3.7-Plus, Kimi-K2.6-Fast) show a warm bias that grows with calendar year, with the same sign in every model (slopes +0.13 to +0.34/century, p<0.01). The effect is modest in size (r-squared approx equal to 0.01 to 0.05) but consistent across six independently developed models. The best-correlated of the six, Gemini 2.5 Flash, matches the best lexical correlation (r=0.32) at double the error. An ablation stripping explicit dates and calendar-era markers from the quotes leaves this trend essentially unchanged, favoring an anachronistic present-day prior over the model correctly inferring the quote's era. Correlation alone is thus insufficient for vetting an LLM as a historical-climate-index oracle.