研究揭示 KV-Cache 选择性复用受请求顺序影响,文档对齐策略降低答案波动
Request Order Matters: Cache-History Sensitivity in Selective KV-Cache Reuse for Rolling Agents
跑长程智能体的开发者看看:KV-Cache 复用会导致同一提示词输出不同,论文给了文档对齐重算方案,波动从 69% 降到 26%,还保住 5.7 倍提速。
arXiv 论文研究了长程智能体滚动更新上下文时非前缀 KV-Cache 复用的稳定性问题。实验发现同一提示词可能因之前处理的请求顺序不同而产生不同答案。在 5% 重算预算下,文档对齐的重算策略把跨请求顺序的答案波动从 CacheBlend token top-k 的 69.0% 降到 26.1%。当每个提示词前面经历不同请求序列时,该策略对完整 prefill 的保真度比 token top-k 高 34.5-52.5 个百分点,两种策略均实现约 5.7 倍中位 TTFT 加速。消融实验显示连续性是稳健选择性重算的主要因素。
Request Order Matters: Cache-History Sensitivity in Selective KV-Cache Reuse for Rolling Agents
Long-running agents repeatedly call an LLM while retaining most of their document window, evicting old documents, and appending new ones. These rolling updates break exact prefix caching and motivate non-prefix KV-cache reuse with selective recomputation. We show that persistent KV-cache reuse with selective recomputation can be history-dependent: in our rolling-agent workload, an unchanged prompt can produce different answers depending on the requests processed before it. At a matched 5% recomputation budget, document-aligned recomputation reduces answer variation across request orders from 69.0% with CacheBlend's token top-$k$ policy to 26.1%. When each prompt is evaluated after a different sequence of preceding requests, document-aligned recomputation improves fidelity to full prefill by 34.5-52.5 percentage points over token top-$k$, while both policies achieve approximately 5.7$\times$ median TTFT speedup. Our ablation study shows that, in our rolling-agent workload, contiguity is the main factor associated with robust selective recomputation.