KV-streams 提出流式 KV 缓存方案,训练提速 2.6-5 倍
KV-streams for Efficient Compaction in Agentic Reinforcement Learning
做智能体 RL 训练的朋友可以看看,KV 缓存不清空改成流式传递,训练直接快 2.6 到 5 倍,还能当循环记忆用,即插即用。
arXiv 论文提出 KV-streams,一种可即插即用的上下文压缩加速策略。它不再在每次压缩后清空 KV 缓存,而是将其向前流式传递,兼容任意压缩策略。在三种不同的压缩策略上,KV-streams 将智能体强化学习训练的墙钟时间提速 2.6 到 5 倍。实验还发现流式传输的 KV 缓存可作为循环状态,保留早已离开上下文窗口的信息,且仅靠 RL 训练即可涌现该行为。
KV-streams for Efficient Compaction in Agentic Reinforcement Learning
Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant for a given trace. Unfortunately, most compaction strategies rely on prefilling the LLM context many times over, hindering training throughput. To alleviate this bottleneck and enable efficient trainable compaction, we propose KV-streams, a plug-and-play strategy compatible with any compaction strategy that substantially increases throughput while showing no evidence of hindering performance. KV-streams enable scalable compaction by streaming the KV cache forward rather than flushing it after each compaction. We show that KV-streams enable three different compaction strategies, achieving a 2.6 to 5x wall-clock speedup in training. Beyond efficiency, we find that the streamed KV cache can act as a recurrent state, carrying forward information that has long since disappeared from the context. Specifically, in a controlled setting we show that, contrary to prior work, RL alone is all that is needed for this behavior to emerge. Overall, we show KV-streams to be an efficient and lightweight plug-and-play addition to any post-training pipeline.