Mira:内存高效MoE推理系统
Mira: Memory-Efficient MoE Inference Using Adaptive Caching and Predictive Expert Staging
清华团队推出Mira系统,解决MoE模型在单GPU上的内存瓶颈,预测专家预取让推理速度提升5倍以上。
Mira是专为单GPU系统设计的高容量MoE推理算法-系统协同设计。该系统通过预测专家管理和定制量化格式,实现专家预取。实验显示,在内存受限GPU上,Mira平均吞吐量提升5.71倍,首token生成加速11.71倍,束搜索推理加速3.84倍。
Mira: Memory-Efficient MoE Inference Using Adaptive Caching and Predictive Expert Staging
Mixture-of-Experts (MoE) models are a compelling architecture for scaling model capacity, making them especially attractive for deployment on resource-constrained, single-GPU systems. However, this benefit is difficult to realize because expert parameters dominate memory, and token-level routing is dynamic, unpredictable, and skewed. Prior work using offloading and caching remains fundamentally reactive, as systems wait for router outputs before moving experts, leading to inefficient cache utilization and an inability to overlap transfers with compute under tight VRAM budgets. To address these challenges, we propose Mira, an algorithm-system co-design that enables high-capacity MoE inference on a single GPU. Mira shifts from a reactive to a proactive stance by coupling predictive expert management with a tailored quantization format. It introduces lightweight per-layer predictors that anticipate expert usage two layers ahead, enabling proactive prefetching. These predictions feed a two-tier HOT+STAGE GPU cache managed by token-level routing telemetry to retain frequently used experts while staging predicted ones. To minimize transfer overhead, Mira implements a custom compression for expert parameters, which reduces metadata and improves packing efficiency, while minimally degrading accuracy. Mira is implemented as a fully integrated runtime that coordinates predictors, caching policies, and quantized transfers to maximize overlap between communication and compute. Our experiments show that Mira reduces expert-induced stalls. Compared against state-of-the-art baselines, Mira achieves a 5.71x speedup in average throughput on a memory-constrained GPU. It accelerates Time-to-First-Token by 11.71x and achieves a 3.84$x average speedup in beam search inference, demonstrating its effectiveness across diverse inference scenarios.