论文

NoPE注意力中的局部混合编码相对位置

How Local Mixing Encodes Relative Position in Global NoPE Attention

精选理由

这篇论文解释了为什么混合模型不需要显式位置编码也能处理序列位置,对理解transformer架构很有价值。

该论文研究了混合模型如何在全局无位置编码(NoPE)层隐式编码位置信息。研究通过理论和实证证据表明,滑动窗口注意力(SWA)和门控线性注意力在残差流中引入了新近偏差,这种偏差传播到全局注意力logits并被选择。与仅含全局NoPE注意力的模型不同,混合模型中的新近偏差可在长序列中保持。

原文 · arXiv cs.LG

How Local Mixing Encodes Relative Position in Global NoPE Attention

The attention operation is naively position invariant. However, positional information is fundamental to natural language, and therefore a variety of explicit position encodings have been developed in transformer-based models, such as rotary position encoding (RoPE). Although explicit position encodings have long been assumed to be required, recent methods that interleave local mixing layers, such as sliding window attention (SWA) and gated linear attention, while not encoding position (NoPE) in global attention layers has recently been shown to be successful at scale. How and why this approach works is not well-understood. In this paper, we develop an explanation of how hybrid models of this sort can implicitly encode position at global NoPE layers. Supported by both theoretical and empirical evidence, our central argument is that SWA and gated linear attention induce a recency bias in the residual stream that propagates to, and is selected by, the global attention logits. Moreover, in contrast to the implicit position encodings found in models with only global NoPE attention, in which positional information arises solely from the causal mask, the recency bias in hybrid models can be maintained across long sequences. In addition to deepening our understanding of how hybrid models encode position, these findings may provide insights for how to encode position in a way that can extrapolate to longer sequence lengths indefinitely.