论文精选

LatentIndex:跨层共享加按层选择的稀疏注意力索引方法

LatentIndex: Cross-Layer Sharing with Layer-Specific Selection for Sparse Attention

精选理由

做长上下文推理的可以看看,这篇把 MLA 的共享思路搬进稀疏注意力索引,缓存省了 61%,速度还能快两倍多。

LatentIndex 把 Multi-head Latent Attention 的隐状态共享思路扩展到稀疏注意力的索引器层,用共享 latent cache 换掉逐层 key cache,同时保留每层独立选 token 的能力。在 DeepSeek-V3.2 上做四层共享后,索引器逻辑缓存存储减少 61.1%。训练无关的校准版本在 DeepSeek-V3.2 和 GLM-5 上,head 级注意力质量召回比 IndexCache 最高提升 3.28 个百分点,RULER 和 LongBench 成绩接近原生 DSA。分级选择(HS)变体在 8K-128K 上下文中相对 DSA 取得 2.30-2.72 倍解码索引器加速。

原文 · arXiv: DeepSeek

LatentIndex: Cross-Layer Sharing with Layer-Specific Selection for Sparse Attention

Sparse attention reduces core-attention computation, but its indexers still incur repeated selection work and per-layer key-cache storage. Reusing selected indices across layers reduces this overhead but constrains multiple layers to the same token set. We introduce LatentIndex, which extends the latent-sharing principle of Multi-head Latent Attention across indexer layers. Each layer group constructs a shared latent cache from its first layer's hidden states, while layer-specific scoring enables independent token selection. Absorbing key decoders into queries enables direct scoring of the shared cache without reconstructing historical per-layer keys. We develop training-free calibration and investigate a training-aware instantiation of this principle. To balance quality and computation, a hierarchical selection (HS) variant lets followers independently refine a shared candidate set proposed by the anchor. With four-layer sharing, LatentIndex reduces logical indexer-cache storage by 61.1% on DeepSeek-V3.2. Across DeepSeek-V3.2 and GLM-5, training-free LatentIndex improves head-wise attention-mass recall over IndexCache by up to 3.28 percentage points while maintaining RULER and LongBench performance close to native DSA. HS further achieves 2.30-2.72 times decode indexer speedups over DSA across 8K-128K contexts, retaining most of LatentIndex's recall. LatentIndex offers a new perspective on cross-layer indexing: sharing continuous representations rather than discrete selections enables efficient reuse while preserving layer-specific token selection.