语音学驱动分词方法在德语语音识别中的跨域研究
Phonologically Informed Tokenization for German Speech Recognition: A Cross-Domain Study
研究者拿德语 ASR 做了 40 次微调实验,发现小词表时按音节切分能扛住方言和口语场景,做多语言语音的可以看看选词表的思路。
一项研究比较了三种分词器在 Omnilingual ASR wav2vec 2.0 骨干上做德语语音识别的表现:多语言字符、正字法 BPE、以及基于 Pyphen 音节划分和音素转换的语音学单元。经过 40 次微调,在域内朗读语音上三者 WER 和 CER 打平;在方言语音和自发语音的域外测试中,小词表下音节感知分词表现更好。音素混淆分析显示所有分词器犯同样的功能词错误,说明错误主要由声学编码器主导。结论是分词器选择取决于词表预算和部署时的分布偏移,而非存在普适最优解。
Phonologically Informed Tokenization for German Speech Recognition: A Cross-Domain Study
German is a morphologically rich language whose syllable structure is exceptionally well-predicted by the Knuth--Liang hyphenation algorithm. We ask whether phonologically informed tokenization can serve as a competitive target for end-to-end speech recognition. We compare three tokenizer families on the Omnilingual ASR wav2vec 2.0 backbone fine-tuned with CTC: the pretrained multilingual character inventory, a data-driven Byte-Pair Encoding (BPE) over orthography, and phonologically informed units from Pyphen syllabification and grapheme-to-phoneme conversion. Across 40 fine-tunes, we evaluate on three German test sets spanning orthogonal shifts: in-domain read speech, dialectal spontaneous speech, and standard-German spontaneous speech. In-domain, all phonologically informed tokenizers match BPE and the multilingual character baseline on both WER and CER. Under domain shift the picture splits along vocabulary size rather than the linguistic axis of variation: at small vocabularies, syllable-aware tokenization improves on dialectal speech, where phonetic surface forms vary but syllable structure is preserved, and stays ahead on spontaneous speech, where new word-forms violate vocabulary closure. A phoneme-level confusion analysis further shows that all tokenizers commit the same canonical function-word errors, indicating that the acoustic encoder, not the tokenizer, dominates the error topology. Our findings suggest that tokenizer choice may depend on the vocabulary budget as much as on the distribution shift expected at deployment rather than reducing to a single universal optimum.