Progressive Memory Transformer:多尺度记忆注意力改进时间序列学习
Progressive Memory Transformer: Memory-Aware Attention for Time-Series
一篇时间序列方向的新论文,给 transformer 加了可写记忆层,能同时捕捉局部波动和全局趋势,低标注场景下分类效果不错,做时序的朋友可以看看。
arXiv 论文提出 Progressive Memory Transformer(PMT),在 transformer 中加入可写、按窗口对齐的记忆模块,额外暴露中尺度表征。该框架在局部、中程、全局三个尺度分别施加监督目标:token 连续性、窗口级 motif、序列级一致性。在 7 个 UCR/UEA/UCI 分类基准、cue-retention 探针和预测基准上,PMT 在 1-5% 标签的低标注分类场景表现突出,预测性能在多个 horizon 上具备竞争力。
Progressive Memory Transformer: Memory-Aware Attention for Time-Series
Time-series carry structure simultaneously at multiple scales (fine-grained variation, mid-range motifs, and global properties) and downstream tasks operate at correspondingly different scales. Most existing self-supervised learning approaches supervise representations globally via instance-level contrastive losses and limited temporal neighborhood supervision, but do not explicitly exploit the structural hierarchy. We propose a learning framework that explicitly enforces a structural hierarchy across three scales independently: a local objective for token continuity, a mid-range objective for window-level motifs, and a global objective for sequence-level agreement. Realizing this framework requires the backbone to expose a representation at each scale; we introduce \textbf{Progressive Memory Transformer} (PMT), which augments a transformer with writable, window-aligned memory that exposes the mid-range scale alongside the token and sequence-level representations conventional transformers already provide. Across seven UCR/UEA/UCI classification benchmarks, a cue-retention probe, and forecasting benchmarks, PMT learns representations that probe well at the global, mid-range, and local scales---strong low-label classification (1--5\% labels), competitive forecasting performance across multiple horizons, and quantitative and qualitative evidence that memory states capture mid-range motifs.