BAT-CLIP:首个对齐脑电、音频与文本的三模态框架
BAT-CLIP: Trimodal Alignment of Brain, Audio and Text
decoded脑电信号的新思路:同时用音频和文本两个锚点对齐,比单锚点方法更稳,做脑机接口的可以看看。
arXiv 论文提出 BAT-CLIP,是首个面向 iEEG(颅内脑电图)的 CLIP 风格三模态对齐框架,将神经嵌入同时映射到预训练音频与文本锚点。现有脑-语音对齐方法只能锚定音频或文本单一模态,导致时间结构或语义可分性二选一。在自然语音 Podcast 基准上,BAT-CLIP 的表征稳健性超过双模态 CLIP 基线。论文还强调使用自监督基础模型进行 CLIP 训练的重要性。
BAT-CLIP: Trimodal Alignment of Brain, Audio and Text
Decoding and interpreting naturalistic speech from the brain increasingly relies on alignment to pretrained speech and language representation spaces. However, current CLIP-style brain-speech alignment ground neural activity to a single anchor modality-audio or text-despite the brain's inherently multimodal speech processing. This induces a trade-off: audio anchoring preserves temporal structure but weakens linguistic separability, while text anchoring captures semantics yet discards acoustic detail. We propose BAT-CLIP, the first CLIP-style trimodal alignment framework for iEEG that jointly aligns neural embeddings to both pretrained audio and text anchors in a shared, frozen audio-text manifold. On the naturalistic Podcast benchmark, BAT-CLIP yields more robust representations than bimodal CLIP baselines. We also highlight the importance of using self-supervised foundation models for CLIP training.