论文

BaLEEN:基于潜在编码实体的上下文感知ASR

BaLEEN: Biasing with Latent Encoded Entities for Context-Aware ASR

精选理由

BaLEEN让ASR在不微调模型的情况下准确识别专业术语和专有名词,错误率大幅降低。

BaLEEN是一种轻量级超网络框架,用于语音识别中的动态上下文适应。该方法通过预训练语言模型编码上下文关键词,使用Perceiver瓶颈压缩为潜在向量序列,并将上下文相关偏置向量直接注入ASR模型的中间编码器表示。在基于CTC的ASR模型测试中,BaLEEN将关键词错误率降低8.7%,同时整体词错误率改善21%,字符错误率改善28%。

原文 · arXiv cs.AI

BaLEEN: Biasing with Latent Encoded Entities for Context-Aware ASR

Transcribing domain-specific entities and rare proper nouns remains a major challenge in automatic speech recognition (ASR). In this paper, we propose BaLEEN (Biasing with Latent Encoded Entities), a lightweight, hypernetwork-based framework for dynamic contextual adaptation without fine-tuning the underlying ASR model. BaLEEN encodes variable-length contextual keywords using a pretrained language model, compresses them into a fixed sequence of latent vectors via a Perceiver bottleneck, and injects context-dependent bias vectors directly into the intermediate encoder representations of the ASR model. Because both the language model and the backbone ASR model remain entirely frozen during training, BaLEEN operates as a plug-and-play adapter that incurs zero computational overhead at inference time when context biases are precomputed. We evaluate our method on a CTC-based ASR model using a Wikipedia-derived corpus with annotated named entities and synthetic speech. Experimental results demonstrate that BaLEEN reduces keyword miss rate by 8.7% on the test set relative to the unbiased baseline while simultaneously improving overall word error rate by 21% and character error rate by 28%.