论文

MS-GLA:多时间分辨率门控线性注意力,缓解 GLA 表征瓶颈

MS-GLA: Multi-Scale Gated Linear Attention for Addressing Representational Bottlenecks via Multi-Temporal Resolution

精选理由

一篇改进 GLA 的论文,思路是把不同注意力头分到不同时间分辨率再动态融合,召回任务提升最高 18.9%,做线性注意力方向的可以看看。

MS-GLA 将 Gated Linear Attention(GLA)的注意力头分配到多个时间分辨率上:粗分辨率池化更长 token 区间,专注长程依赖;细分头组保留对局部句法结构的敏感度。一个可学习、依赖输入的融合层在每个时间步动态重组各组输出,在不增加单头状态大小的情况下扩展有效记忆容量。在匹配参数量的评测中,MS-GLA 在召回密集型任务上最高提升 18.9%,语言建模基准上平均困惑度降低 9.5%。该方法借鉴了多尺度状态空间模型(MS-SSM)的思路,将其适配到门控线性注意力框架中。

原文 · arXiv cs.AI

MS-GLA: Multi-Scale Gated Linear Attention for Addressing Representational Bottlenecks via Multi-Temporal Resolution

Gated Linear Attention (GLA) Transformers advance linear recurrent models through data-dependent gating, but face a core limitation: the fixed-capacity memory matrices across all heads operate at a single temporal resolution, where each token is processed individually, forcing them to simultaneously encode local syntactic patterns and long-range semantic structure, creating a representational bottleneck that gating alone is insufficient to resolve. We introduce Multi-Scale Gated Linear Attention (MS-GLA), which addresses this by distributing attention heads across multiple temporal resolutions. Coarser resolutions pool longer token spans naturally specializing toward long-range dependencies, while finer head groups retain sensitivity to local syntactic structure. A learnable, input-dependent fusion layer dynamically recombines head group outputs at each timestep, expanding effective memory capacity without increasing per-head state size. This multi-resolution decomposition draws on principles from Multi-Scale State-Space Models (MS-SSM), adapting them to the gated linear attention setting. We evaluate MS-GLA on language modeling, recall-intensive tasks, and long-context generalization. Across all settings, MS-GLA consistently achieves higher accuracy and lower perplexity than GLA at matched parameter counts, with up to 18.9% improvement on recall-intensive tasks and 9.5% lower average perplexity on language modeling benchmarks, validating multi-temporal resolution decomposition as a principled and effective extension of Gated Linear Attention.