EAServe 论文提出面向多模态 LLM 的 Encode 感知分离式推理服务
EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models
一篇推理系统论文,解决多模态模型 Encode 阶段拖垮 GPU 利用率的问题,goodput 比 vLLM 高 1.7 倍、比 NVIDIA Dynamo 高 4.3 倍,做 MLLM 服务的可以看看。
论文针对多模态大模型(MLLM)推理中新增的 Encode 阶段提出 EAServe 系统,将 Encode-Prefill-Decode 三阶段流水线中的 Encode 重新定位为控制点。其运行时层包含负载自适应微批处理、速率控制的 partially offload 和动态 SM 分区,配置层 HAS 通过按阶段容量画像和 TPE 贝叶斯优化搜索 GPU 分配、encode 批大小与 offload 比例。在覆盖图像、视频、音频的三种 MLLM 架构上评测,相同 SLO 约束下 EAServe 的 goodput 分别比 NVIDIA Dynamo 高 4.3 倍、比 vLLM 高 1.7 倍,并显著提升 GPU 利用均衡度。
EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models
Disaggregating the two stages, Prefill and Decode, onto separate GPU pools is now a standard optimization for (text-only) LLM serving. However, multimodal LLMs (MLLMs), which add a third phase, Encode, pose new challenges for resource allocation. Encode turns images, video, or audio into embeddings that the language model can consume, yielding a three-stage Encode-Prefill-Decode (EPD) pipeline. Existing frameworks offer only partial answers: text-only PD systems lack Encode, while EPD frameworks expose it as a separate service without regulating downstream request flow. The pipeline also carries a structural resource imbalance: every request enters through Encode before downstream work can begin, yet per-request execution leaves the encode GPU severely underutilized even at high loads, starving the downstream Prefill and Decode workers. Addressing this, we reposition Encode as the control point of the EPD pipeline, exposing three tightly coupled dimensions: when work enters downstream, where prefill executes, and how the GPU is shared. We instantiate this in EAServe across two co-designed layers. Its runtime manages load-adaptive micro-batching, rate-controlled partial offload to a co-resident prefill worker, and dynamic SM partitioning for predictable co-location. The configuration layer, Hybrid Auto Selection (HAS), navigates the joint space of GPU allocation, encode batch size, and offload ratio by pruning unbalanced allocations with per-stage capacity profiling and refining the remainder through TPE-based Bayesian optimization. Evaluated on three MLLM architectures spanning image, video, and audio, EAServe delivers up to 4.3x and 1.7x higher goodput than NVIDIA Dynamo and vLLM, respectively, under identical SLO constraints, sustains more balanced and higher GPU utilization across the EPD pipeline, and reaches near-optimal configurations faster than baseline search methods.