论文

TempoKV:让 LLM KV 缓存调度更省内存的分层方案

TempoKV: Timely Staging of LLM KV Caches for Memory-Semantic Flash

精选理由

vLLM 和 LMCache 里能直接用的 KV 缓存调度优化,快速层容量从 100 GiB 砍到 25 GiB 吞吐基本不掉,还把 p95 首 token 时间降了 48%。

TempoKV 是一个面向内存语义闪存层级的 KV 缓存调度方案,通过元数据记录复用命中,待运行时估算的取回时间降至存储准备时间时才提交资源。它在 vLLM 和 LMCache 上实现,运行于 SSD 支持的 CXL 内存设备,无需改动请求调度。在 2 个模型、3 种前缀缓存比例下,相比即时预取方案每请求受保护快速层字节时间减少 63-91%。与未修改的 LMCache Device-DAX L1 配置相比,p95 TTFT 最多降低 48.0%,输出吞吐量最多提升 27.8%。

原文 · arXiv cs.AI

TempoKV: Timely Staging of LLM KV Caches for Memory-Semantic Flash

Reusable prefix key-value (KV) caches can outgrow GPU memory in large language model (LLM) serving. A memory-semantic flash hierarchy offers SSD-backed capacity with a limited fast tier, but a logical KV hit is not necessarily ready for GPU retrieval. Demand staging exposes SSD latency, whereas immediate staging can reserve fast-tier capacity long before retrieval begins. We present TempoKV, a timing-aware resource-commitment layer that separates early knowledge of reuse from the acquisition of staging resources. It records reusable-KV hits as metadata-only claims and requests commitment when the runtime-estimated time until retrieval falls to the storage-estimated time needed to make KV resident and protected against eviction. These estimates adapt to runtime progress and staging state, while commitment remains subject to available protected capacity. We implement TempoKV in vLLM and LMCache on an SSD-backed CXL memory device without changing request scheduling. Across two models and three prefix cache ratios, TempoKV reduces protected fast-tier byte-time per request by 63-91% versus immediate staging while retaining much of the serving benefit of advance staging. In a fast-tier capacity sweep, output throughput and p95 time to first token (TTFT) remain nearly unchanged as capacity decreases from 100 to 25 GiB. Compared with unmodified LMCache's Device-DAX L1 configuration, TempoKV reduces p95 TTFT by up to 48.0% and increases output throughput by up to 27.8%.