论文

高能物理 ML 扩展定律论文:在 ATLAS JetSet2 上系统比较架构

How to scale your HEP ML models: A recipe for robust architecture comparisons at scale

精选理由

做物理机器学习的朋友可以看看,论文给出在 ATLAS JetSet2 上跑扩展定律对比架构的具体配方,连学习率和 batch size 一起优化了。

论文提出一套在高能物理任务中推导扩展定律的系统流程,先在玩具问题上验证完整扩展轨迹,再应用于约 110 亿 jet 的 ATLAS JetSet2 数据集上的多任务 transformer,覆盖计算受限与数据受限两种场景。作者首次联合预测了早停下的最优模型规模、训练时长、学习率和 batch size。在计算最优扩展下,模型与数据集规模呈近似相等的 sqrt(C) 依赖,辅助目标在同等计算预算下降低了 jet 分类主损失。论文还指出数据集规模低于阈值时,损失几乎不携带高计算量扩展的信息。

原文 · arXiv cs.LG

How to scale your HEP ML models: A recipe for robust architecture comparisons at scale

Much of the recent progress in machine learning domains such as language models has come from scaling laws that predict performance as a function of training effort. In high-energy physics (HEP) similar behavior has now been observed. To aid further study, we present a systematic procedure to derive robust scaling laws and compare design choices on the relevant budget axes for HEP tasks. We first validate the full scaling trajectory on toy problems and then apply the procedure to multi-task transformers on the ~11 billion-jet ATLAS JetSet2 dataset, in both the compute- and data-constrained regimes. For the latter, we predict, to the best of our knowledge for the first time, the jointly optimal model size, training horizon, learning rate and batch size under early stopping. At compute-optimal scaling, we recover a near-equal $\sqrt{C}$ dependence of model and dataset size, and find that auxiliary objectives lower the primary jet-classification loss at equal compute budget. Expanding the inputs toward lower-level data systematically lowers the loss while leaving the scaling exponent nearly unchanged. The onset of the power-law regime is itself set by scale: below a threshold in dataset size the loss carries little information about high-compute scaling, underscoring the value of large, high-quality full-simulation datasets as a foundation for scaling studies and the development of foundation models in HEP.