论文精选

MoE-CORE系统优化内存受限MoE推理

MoE-CORE: Coordinated Expert Offloading and Residency for Memory-Constrained MoE Inference

精选理由

MoE-CORE系统解决了MoE模型在内存受限设备上的推理问题,性能比vLLM Prefetch提升30倍。

MoE-CORE系统通过协调专家卸载和驻留,解决内存受限的MoE推理问题。在DeepSeek-V4-Flash-W4A8模型上,MoE-CORE的TPOT为38.0-44.8ms,远低于vLLM Prefetch的1268.9-1269.1ms。在GLM-5.2-W4A8C8模型上,MoE-CORE的TPOT为206.6-220.5ms,优于vLLM Prefetch的5941.5-5941.8ms。在84GB NPU内存限制下,最佳配置的TPOT可达21.5ms。

原文 · arXiv: DeepSeek

MoE-CORE: Coordinated Expert Offloading and Residency for Memory-Constrained MoE Inference

Sparse expert activation reduces MoE models' computation, yet expert weights can exceed limited device memory. Offloading makes inference feasible on a compact AI appliance but exposes host-to-device transfers to the inference path. We present MoE-CORE, a system that coordinates expert offloading and residency for memory-constrained MoE inference. It stages complete expert layers in alternating buffers during prefill. During decode, it combines nonuniform layer-wise cache capacity, domain-informed initialization, routing-history-aware replacement, and cross-layer prefetching. The main configuration executes router-selected experts exactly; an optional score-based substitution path handles eligible low-score misses. The main comparison uses 1K- and 128-token output caps for MoE-CORE and vLLM Prefetch, respectively. Across five workloads per model, MoE-CORE records a mean time per output token (TPOT) of 38.0-44.8 ms versus 1268.9-1269.1 ms for the evaluated vLLM Prefetch configuration on DeepSeek-V4-Flash-W4A8; the corresponding values on GLM-5.2-W4A8C8 are 206.6-220.5 and 5941.5-5941.8 ms. Under an 84-GB NPU-memory cap, the best measured DeepSeek GSM8K configuration achieves a TPOT of 21.5 ms with approximate expert substitution and multi-token prediction (MTP) at depth 2. These results support coordinated expert residency and transfer scheduling under a device-memory constraint. The code is here.