论文

基于解剖标志的跨说话人声学-发音反演适配框架

Anatomy-aware cross-speaker adaptation of complete vocal-tract acoustic-to-articulatory inversion

精选理由

不用重新训练,靠脊椎和牙齿的标志点就能把声道反演结果适配到新说话人,8人测试误差降到3.19mm,做语音发音研究可以看看。

研究者提出一种几何适配框架,解决跨说话人声学-发音反演中的解剖差异问题。方法利用脊椎和牙齿结构上的解剖标志点,通过仿射变换加薄板样条(TPS)形变,把固定反演模型预测的10个声道结构轮廓映射到目标说话人几何中,无需重新训练。标志点在每个说话人选定的一帧 /u/ 图像上识别,映射可跨录音复用。模型在单人说话人 rt-MRI 数据库上训练,在另一多说话人数据库的8名说话人上评估,其中 Affine12+TPS14 配置取得最低平均点到最近点误差 3.19mm。

原文 · arXiv cs.AI

Anatomy-aware cross-speaker adaptation of complete vocal-tract acoustic-to-articulatory inversion

Cross-speaker acoustic-to-articulatory inversion requires accounting for anatomical differences between speakers. We propose a geometric adaptation framework that uses anatomical landmarks, primarily on vertebrae and dental structures,to transfer predictions from a fixed inversion model to unseen speakers. An affine transformation followed by thin-plate spline (TPS) deformation maps the predicted contours of 10 vocal-tract structures into each target speaker's geometry without retraining. Landmarks are identified in one selected /u/ frame per speaker as a common phonetic reference without assuming identical articulatory configurations across speakers, and the resulting mapping is reused across recordings. We train the model on a single-speaker rt-MRI database and evaluate adaptation on eight speakers from a separate multi-speaker rt-MRI database. We compare affine and TPS configurations using 12 or 14 landmarks. Affine12+TPS14 achieves the lowest mean point-to-closest-point error of 3.19mm. These results support the combined value of anatomical landmark information and nonrigid alignment.