滑动窗口 KV 缓存可传递窗口外信息:Qwen、Llama、Mistral 等五模型实验
Thinking Outside the Box: Retention and Transmission of Information in Sliding-Window KV Inference
研究者测了五个开源模型,发现 Mistral 7B 和 Muse Glimmer 能从已滑出窗口的缓存里找回信息,做长上下文推理的可以看看。
一项研究考察滑动窗口 KV 推理中信息的保留与传递,即增量处理序列时只保留固定大小的 key/value 缓存。实验覆盖 Qwen、Llama、Mistral、Muse Glimmer 共五个开源权重模型。结果显示,保留此前计算的缓存状态比用原始 token 重算最后固定窗口的检索效果更好。其中 Muse Glimmer 和 Mistral 7B 的潜在信息传递最强,即使相关源 token 已离开缓存仍能恢复信息,而两者架构中本身都包含滑动窗口注意力。
Thinking Outside the Box: Retention and Transmission of Information in Sliding-Window KV Inference
Sliding-window KV inference refers to processing a sequence incrementally while retaining only a fixed-size cache of recent key and value states. It can be applied to pretrained causal transformers at inference time without additional training, while its KV-cache memory remains fixed as more tokens are processed. Because cached states are computed in the context of earlier tokens, they may carry information from beyond the current window and transmit it to later states. This study presents a series of experiments using five open-weight models spanning Qwen, Llama, Mistral, and Muse Glimmer. We investigate whether information originating outside the immediate context window can persist through a rolling KV cache and remain useful for retrieval. Initial results show that retaining previously computed states improves retrieval across the models tested compared with recomputing the final fixed window from raw tokens. We then measure how far this effect extends and find that Muse Glimmer and Mistral 7B show the strongest \emph{latent information relay}: they can recover information even after the relevant source tokens have left the cache. Both models incorporate sliding-window attention in their published architectures, an association that motivates testing whether training with sliding windows promotes more reliable information retention.