英国气象局将强化学习集成到全球天气预报模型,显著降低多个纬度带的预报误差。
英国气象局统一模型通过分布式模型-智能体耦合实现在线强化学习。研究者在70个垂直模型层级上应用DDPG智能体,对位势温度进行有界修正。在10个 nudged 训练预报中,学习策略在6小时预报中使 Z500 平均绝对误差在四个纬度带降低,热带地区降幅达45.8%和40.8%。MSLP 误差在三个纬度带减少,0-30°N 区域最大降幅27.3%。
Online Reinforcement Learning in the Met Office Unified Model through Distributed Model-Agent Coupling
Machine-learnt corrections can complement numerical weather prediction only if they adapt to the evolving model state while preserving dynamical consistency and numerical stability. To test this within a global forecasting model, we couple the Met Office (UKMO) Unified Model (UM) with distributed RL agents through rank-local tensors. A DDPG actor shares weights across the 70 vertical model levels of each atmospheric column and applies bounded potential-temperature corrections to the model tendencies. Across ten nudged training forecasts, nudging calculations towards the UKMO operational analysis provides an immediate counterfactual target. The frozen policy is then evaluated in a non-nudged forecast for inference. The coupled workflow successfully completes training and remains numerically stable in the evaluated case. Relative to a matched native UM forecast at +6 h, the learnt policy reduces Z$_{500}$ MAE in four of six latitude bands, including reductions of 45.8% and 40.8% in the northern and southern tropics. MSLP error too decreases in three bands, with a maximum reduction of 27.3% at 0-30°N. This single-case experiment demonstrates significant promise and feasibility of distributed online learning followed by non-nudged inference, laying the groundwork for RL-based bias correction and parametrisations within operational systems.