NVIDIA 提出 NCCL M2N 集合通信原语,优化分布式张量重分片
NCCL M2N: A Layout- and Topology-Aware Collective for Distributed Tensor Resharding
做 RL 训练或大规模推理部署的可以看,DeepSeek-V3 权重同步直接从 5.78 秒砍到 2.77 秒,还是进 NCCL 的正经方案。
分布式训练与推理生成的张量布局不同,权重需要跨进程组做 M-to-N 重分片,传统 all-gather 加 broadcast 方式会产生冗余流量。论文提出 NCCL M2N,根据源和目标 mesh 与 placement 推导传输区域并生成全局通信调度,跨 NVLink 域只转发一份拷贝。在 256 张 GB200 GPU 的 NVL72 集群上,单个 FFN-MoE 层传输相比扁平直发最高提速 7.9 倍(9.8 ms 对 77.3 ms)。在 256 GPU 的 DeepSeek-V3 NeMo-RL 实验中,权重同步时间从 5.78 秒降至 2.77 秒,训练步耗时减少 12.7%。
NCCL M2N: A Layout- and Topology-Aware Collective for Distributed Tensor Resharding
Distributed training and rollout generation often use different tensor layouts, requiring model weights to be resharded across distinct process groups. This M-to-N redistribution is not directly expressed by standard collectives. Flat direct sends duplicate traffic across destination replicas, while gather-then-broadcast concentrates network injection at one root and transfers data that destinations do not need. For DeepSeek-V3 on 256 GPUs, legacy all-gather plus broadcast accounts for 29.4% of the reported reinforcement-learning step time. We present NCCL M2N, a layout- and topology-aware collective primitive for distributed tensor resharding. Given source and destination meshes and placements, it derives the required transfer regions and a global communication schedule. Its hierarchical route balances source contributions across eligible destination ranks, forwards one copy between destination NVLink domains, and completes local replication over NVLink. Network transfer and local replication overlap, avoiding replica-multiplied source egress. An aggregate data-movement model captures the limits of both stages. We evaluate NCCL M2N on up to 256 GB200 GPUs in an NVL72 cluster with NDR InfiniBand. A single FFN-MoE layer transfer achieves up to 7.9x speedup over flat direct sends (9.8 ms versus 77.3 ms). In a separate 256-GPU DeepSeek-V3 NeMo-RL experiment, NCCL M2N reduces reported weight-sync time from 5.78 s to 2.77 s, a 2.09x speedup over legacy all-gather plus broadcast, and reduces step time by 12.7%.