vLLM 推出弹性专家并行,可在运行中增减 GPU
vLLM 的 MoE 服务现在能在线加减 GPU 不停机,NVIDIA 的人在 PyTorch 大会上讲实现细节,做推理部署的别错过。
vLLM 新增 Elastic Expert Parallelism(弹性专家并行)功能,允许在 MoE 部署运行期间增减 GPU,且对服务中断和停机影响极小。NVIDIA 的 Itay Alroy 将在 PyTorch Conference North America 2026 上介绍 Elastic EP 的架构、关键实现细节和路线图。演讲还将讨论 EP 规模变化时的行为,以及 NIXL EP 如何在实时流量下实现扩缩容。
Elastic Expert Parallelism in @vllm_project lets you add or remove GPUs from an active Mixture-of-Experts deployment during traffic with minimal interruption to serving and minimal downtime.
At #PyTorchCon North America 2026, Itay Alroy (@nvidia) will present “Elastic Expert Parallelism in vLLM,” covering the architecture, key implementation details, open challenges, and future roadmap for Elastic EP.
Alroy will also examine what happens when EP size changes and how NIXL EP enables grow/shrink under live traffic.
Register for PyTorch Conference North America 2026: https://t.co/jBApW8ocHQ