论文

ECO:面向注意力-FFN 分离式 LLM 服务的能耗配置优化方法

ECO: Energy-Oriented Configuration Optimization for Attention FFN Disaggregated LLM Serving

精选理由

一篇降本干货:ECO 在 Qwen、DeepSeek 的分离式部署上实测省 40.5% 能耗,做 GPU 推理优化的可以看看方法细节。

ECO 是一种针对注意力-FFN 分离(AFD)架构的 LLM 服务能耗优化方法,在有限测量预算下联合搜索部署结构与运行控制参数。它先用校准的阶段行为和流水线依赖构建结构感知的能耗先验,再用高斯过程学习残差预测误差,并以成本感知的约束贝叶斯优化挑选待测配置。在 A6000 和 A100 上用 Qwen 和 DeepSeek 测试的全部 16 个场景中,ECO 的冻结配置在满足 SLO 的前提下平均降低 40.5% 的服务能耗,输出 token 速率提升 20.7%。在 8 个 A6000 场景中,其能耗比通用约束贝叶斯优化平均低 33.1%,比遗传搜索低 25.8%。

原文 · arXiv: DeepSeek

ECO: Energy-Oriented Configuration Optimization for Attention FFN Disaggregated LLM Serving

Energy-efficient LLM serving requires minimizing serving GPU energy while meeting latency and throughput service-level objectives (SLOs). Attention--FFN disaggregation (AFD) enables separate resource allocation and operating controls for attention and expert computation, but their energy effects remain coupled through the execution pipeline. Realizing its energy-saving potential therefore requires navigating a hierarchical configuration space in which deployment structures constrain admissible controls and shape their end-to-end effects. Finding low-energy configurations that meet SLOs is challenging because physical evaluations are costly and only a small fraction of candidates can be measured. We present Energy-Oriented Configuration Optimization (ECO), which jointly searches deployment structures and their admissible operating controls under a limited measurement budget. ECO constructs a structure-aware energy prior from calibrated stage behavior and pipeline dependencies, then learns residual prediction errors with a Gaussian process. Its cost-aware constrained Bayesian optimization prioritizes measurements according to expected energy improvement while accounting for SLO feasibility, execution success, and evaluation cost, and returns the lowest-energy measured feasible configuration. Across all 16 scenarios on A6000 and A100 with Qwen and DeepSeek, ECO's frozen configurations, evaluated on disjoint requests, reduce serving energy by 40.5\% and increase output token rate by 20.7\% on average relative to baselines while meeting target SLOs. Across the 8 A6000 scenarios, its selected feasible energy averages 33.1\% below generic constrained Bayesian optimization and 25.8\% below genetic search.