论文

PReCache:免训练的 KV Cache 共享框架,多 LoRA 智能体 TTFT 提速最高 3.1 倍

PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction

精选理由

多智能体跑长任务时每个 LoRA 都重算一遍上下文,这篇论文给了一套免训练的共享方案,提速最高 3.1 倍,做 agent 服务的值得看看实现思路。

多 LoRA 智能体系统中,每个智能体都要对不断增长的共享轨迹重复 prefill 并构建自己的 KV cache,造成冗余。论文提出 PReCache,包含 PreLRShared 和 ReBaseShared 两个设计:前者在共享上下文首次处理时预计算每个智能体的低秩 cache,后者从无 adapter 的隐状态重建基础 cache 以降低误差。实验显示 PreLRShared 相比不共享 cache 的推理最高取得 3.1 倍 TTFT 提速和 2.3 倍单请求吞吐提升,ReBaseShared 在各 cache 共享方法中准确率保持最好,相比不共享仅平均下降 1.1 分。

原文 · arXiv cs.LG

PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction

Multi-LoRA agent systems enable efficient role specialization by sharing a common backbone model. However, each agent repeatedly processes the growing shared trajectory and constructs its own KV cache, introducing substantial memory and computation redundancy in long-horizon tasks. Existing KV cache sharing methods reduce this repeated prefill, but they either require additional training or architectural constraints or retain substantial model computation. Moreover, direct cache reuse causes the current agent to rely on cache states generated by the previous agent's adapter, weakening the role-specific behavior encoded by its own LoRA. We present PReCache, a training-free KV cache sharing framework with two designs, namely PreLRShared and ReBaseShared, that share the base cache computed using the pretrained weights and precompute a compact agent-specific low-rank (LR) cache. To remove repeated prefill, PreLRShared precomputes each agent's LR cache when the shared context is first processed, allowing the current agent to use its own LR cache without reprocessing context processed by previous agents. To improve sharing accuracy, ReBaseShared reconstructs the shared base cache from adapter-free hidden states, reducing the remaining error caused by the previous agent's adapted representation. To minimize its reconstruction cost, we propose two inference schemes tailored to single-stream inference and concurrent serving, performing the same reconstruction after each agent's turn or alongside its execution, respectively. Across multiple models and agent benchmarks, PreLRShared achieves up to a 3.1x TTFT speedup and a 2.3x improvement in per-request throughput over inference without KV cache sharing. ReBaseShared best preserves accuracy overall among the evaluated cache-sharing methods, with an average drop of only 1.1 points relative to inference without cache sharing.