CTP-FL:用共同轨迹梯度预测改进联邦学习通信效率
CTP-FL: Common-Trajectory Gradient Prediction for Federated Learning
联邦学习又出新优化思路:让所有客户端沿同一条预测路径算梯度,通信量和 FedAvg-M 一样但聚合方向是无偏的,还附理论收敛证明,做分布式训练的可以看看。
论文提出 CTP-FL(Common-Trajectory Predictive Federated Learning),一种新的联邦优化方法。每轮所有客户端基于当前全局模型和上一轮聚合方向构造相同的查询点序列,沿该序列各评估 K 个随机梯度并上传平均值,服务器执行一次全局更新。其每轮计算与通信量与全参与 FedAvg-M 持平,即每客户端 K 个 mini-batch 梯度加每方向一个模型规模向量。共享查询点使聚合方向成为预测路径上平均全局梯度的无偏估计,剩余偏差由路径长度控制,无需假设客户端梯度差异或梯度有界。论文在光滑非凸目标下给出 O(√(LΔσ²/(NKR))+LΔ/R) 的平均平稳性界,并揭示了预测路径长度带来的前瞻信息与位移偏差之间的可检验权衡。
CTP-FL: Common-Trajectory Gradient Prediction for Federated Learning
Communication-efficient federated optimization commonly spends several gradient evaluations between server updates. Existing local-update methods use this computation to advance an independent model on each client. Under heterogeneous data, however, these models evaluate gradients at different locations, making the aggregated update difficult to interpret as a gradient of the global objective. We study an alternative use of the same computation budget: \emph{evaluate the global objective along a shared, predicted path}. We propose Common-Trajectory Predictive Federated Learning (\texttt{CTP-FL}). At each round, all clients construct the same sequence of query points from the current global model and the previous aggregated direction, evaluate $K$ stochastic gradients along this sequence, and upload their average. The server then performs a single global update. Thus, \texttt{CTP-FL} uses $K$ mini-batch gradients per client and one model-sized vector in each communication direction, matching the per-round computation and communication of full-participation FedAvg-M. Shared query points make the aggregated direction an unbiased estimator of the average \emph{global} gradient along the predicted path. The remaining discrepancy from the gradient at the current model is controlled by the path length, without assuming bounded client-gradient dissimilarity or bounded gradients. For smooth non-convex objectives, we establish an $\mathcal{O}\!\left( \sqrt{LΔσ^2/(NKR)}+LΔ/R \right)$ average-stationarity bound under full participation. The analysis isolates a testable trade-off: extending the prediction path provides more forward-looking gradient information but increases its displacement bias.