论文

Bearings:从一阶 Ambisonics 自监督学习声场嵌入

Bearings: Self-Supervised Soundfield Embeddings from First-Order Ambisonics

精选理由

Bearings 给音频模型补上方位感:接上后 TAU-NIGENS 定位分数从不足 4 升到 50,原模型不用重训。

Bearings 是一个自监督框架,从无标注的一阶 Ambisonics 音频中学习声场嵌入。训练采用掩码自编码器,解码器以现成单通道音频编码器输出的冻结声学嵌入为条件。学到的声场嵌入可作为可复用表示流,通过轻量可训练融合头接到冻结的声学编码器上。在声音事件定位与检测任务中,TAU-NIGENS 2021 的位置相关 F 分数从 4 以下升到 50,STARSS23 达到 39。论文称这是首个无需重训任一模型就能插入冻结声学编码器的自监督声场编码器。

原文 · arXiv cs.AI

Bearings: Self-Supervised Soundfield Embeddings from First-Order Ambisonics

Recently proposed self-supervised audio encoders learn powerful general-purpose representations of sound scenes, yet they are spatially blind. To supply the missing spatial representation of sound scenes, we introduce Bearings. Bearings is a self-supervised framework that learns soundfield embeddings from unlabeled first-order Ambisonics. We pre-train a masked auto-encoder paired with a decoder conditioned on frozen acoustic embeddings from an off-the-shelf single-channel audio encoder. Our results show that the resulting soundfield embeddings form a reusable stream that can be attached to frozen acoustic encoders with a lightweight trainable fusion head. On sound event localization and detection, concatenating our soundfield embeddings with acoustic representations provides the missing spatial information and enables joint detection and localization, raising the location-dependent F-score from below 4 to 50 on TAU-NIGENS 2021 and 39 on STARSS23. To our knowledge, Bearings is the first self-supervised soundfield encoder whose embeddings plug into frozen acoustic encoders without retraining either model.