ADPTNet:受听觉皮层启发的非线性SSM序列建模新架构
ADPTNet: Adaptive with Prescriptive Timescales Non-Linear SSM for Sequence Modelling
这篇论文用听觉皮层的固定时间尺度思路改设计 SSM,Selective Copying 上赢 Hawk,状态追踪超过 Mamba,还顺带把非线性 RNN 做成可并行的,做架构研究的可以看看。
ADPTNet 借鉴听觉皮层按固定时间尺度运作的证据,通过线性注意力与黎曼优化结合的局部拓扑共轭,实现数据自适应、长程依赖、GPU 并行和非线性递归四项特性。在 Selective Copying 基准上超过 Hawk,状态追踪优于 Mamba 等线性 SSM;在 sequential CIFAR-10 上以更少参数匹敌线性 SSM 并超过 Transformer。其脉冲神经网络版本 SpikingADPTNet 在 Spiking Speech Commands 数据集上取得 83.56%±0.15 的最高准确率。论文还提出 Conv 和 Forward 两种无 Jacobian 的 DEER 扩展,首次通过迭代卷积实现非线性 RNN 的并行化。
ADPTNet: Adaptive with Prescriptive Timescales Non-Linear SSM for Sequence Modelling
A central aim of neuromorphic computing is to provide a viable alternative to highly energy-intensive Transformer-based AI. However, efficient alternatives struggle to capture the set of qualities that have secured the Transformer's status as the de facto standard in sequence modelling. Any realistic contender must be data-adaptive, able to capture long-range dependencies, and GPU-parallelisable, but also non-linearly recurrent to enable complex reasoning. Based on evidence suggesting the auditory cortex operates on fixed timescales, this work proposes the ADaptive with Prescriptive Timescales Network (ADPTNet) as a potential solution to achieving all four properties simultaneously. ADPTNet is built around local topological conjugates, obtained by a novel combination of linear attention and Riemannian optimisation, applied to static global dynamics. This enables non-linear yet predictable long-term behaviour. Dynamical systems theory proofs provide theoretical guarantees for the parametric control of ADPTNet's timescales (its Lyapunov spectrum). ADPTNet improves performance on Selective Copying over Hawk, the existing method balancing long-range memory and adaptability, while also improving state tracking over linear SSMs like Mamba. On sequential CIFAR-10, ADPTNet matches linear SSM accuracy and outperforms existing selective models (incl. the Transformer), using fewer parameters. We also introduce a neuromorphic SpikingADPTNet, which achieves a new state-of-the-art accuracy on the Spiking Speech Commands dataset ($83.56\%\pm0.15$). Finally, ADPTNet's constant timescales enable two efficient, Jacobian-free extensions to the DEER parallel simulation algorithm (Conv and Forward DEER) that retain the same average convergence. Conv DEER adds no computational overhead beyond the network's forward pass and enables non-linear RNN parallelisation via iterated convolutions for the first time.