AMD 在 PyTorchCon 2026 分享 MI355X 上 MXFP8 千卡级预训练方案
AMD 拿 1000 多张 MI355X 跑 MXFP8 预训练,还会讲 TorchAO 和 TorchTitan 的具体调优细节,搞训练的可以听听。
AMD 工程师 Liz 和 Shekhar 将在 10 月 20-21 日 San Jose 举办的 PyTorchCon North America 2026 上介绍 MI355X GPU 的 MXFP8 预训练工作。分享内容包括 TorchAO 中 MXFP8 关键算子的优化,以及与 TorchTitan 集成构建端到端 LLM 训练栈。议题还涵盖 Triton 和 FlyDSL 实现,以及数值验证、收敛性和性能调优经验,规模达 1000+ 张 MI355X。
At #PyTorchCon North America 2026, @lizli202503 and @indianspeedster of @AMD will discuss their work enabling scalable MXFP8 pretraining on MI355X GPUs.
Liz and Shekhar will cover optimization of key MXFP8 operators in TorchAO and how those kernels were integrated with TorchTitan to build an end-to-end upstream PyTorch training stack for large language models.
Their session, “Scaling MXFP8 Pretraining on 1K+ AMD Instinct MI355X: TorchAO Kernels and TorchTitan Training,” will also cover Triton and FlyDSL implementations and lessons from numerical validation, convergence, and performance tuning.
Join us in San Jose on October 20-21. Register today: https://t.co/jBApW8nESi