语义瓶颈:利用语义表示进行非侵入式语音解码
The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding
科学家们用语义中间层重建文本,比直接解码语音更准确,MEG数据转文字效果更好。
研究人员提出Brain2Semantics2Text方法,通过中间语义嵌入空间重建文本。该模型将句子层面的MEG反应映射到语义流形,然后将预测的嵌入转换为自然语言。与之前的非侵入式Brain2Text方法相比,该方法在句子级别结果上有所改进。研究利用神经科学证据,表明高级语义表示分布在皮层区域并随时间缓慢演变。
The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding
Non-invasive speech decoding remains constrained by the low signal-to-noise ratio of neural recordings, which makes fine-grained reconstruction of phonemes or individual words difficult. Motivated by neuroscientific evidence that high-level semantic representations are distributed across cortical regions and evolve over slower temporal scales, we hypothesize that semantic content may provide a more suitable target for non-invasive decoding than low-level acoustic or lexical features. We introduce Brain2Semantics2Text, a method that reconstructs text through an intermediate semantic embedding space. Our model maps sentence-level MEG responses into a semantic manifold and then inverts the predicted embeddings into natural language. This semantic bottleneck enables recovery of high-level meaning without word-level alignment. We describe the core principles of the approach, its implementation, and the strategies used to mitigate the challenges of learning a reliable neural-to-semantic mapping. Finally, we compare against prior non-invasive Brain2Text methods and show improved sentence-level results.