Paradee:把 Kokoro-82M 蒸馏成 8M 参数的单音色 TTS 模型
Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model
一个 8.5MB 的 TTS 模型,笔记本 CPU 就能跑,速度是实时的 25 倍,音质只比原版 Kokoro 低一点,还开源了代码和音频样本。
研究人员将开源 TTS 模型 Kokoro-82M 蒸馏为 8.07M 参数的 Paradee,参数量减少 10 倍,计算量减少 15 倍。方法是用教师模型合成语料并保留其时长、音高、能量等特征,分别训练文本侧和解码器,再合并量化为 int8。量化后模型仅 8.5 MB,在单 CPU 线程上以 25 倍实时的速度运行,UTMOS 得分 4.41(教师为 4.52)。针对 2-8 kHz 浊音相位导致的杂音,用相位锁定滤波器后处理即可消除,无需额外训练。
Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model
We distill Kokoro-82M, a widely used open text-to-speech model with 54 voices, into Paradee, an 8.07M-parameter model that speaks one of them. Paradee keeps Kokoro's architecture with much narrower layers, and each of its two halves is trained separately against the frozen teacher. It has 10x fewer parameters and needs 15x less compute. We first synthesize a corpus with the teacher and keep its durations, pitch, energy and phoneme features. We then train a small text side to predict these values, and a small decoder to turn the teacher's saved values into the teacher's audio, first with spectral losses and then adversarially. Finally, we connect the two halves and quantize the weights to int8. It needs no alignment learning and no joint training, and it runs on one laptop. Stored in int8, Paradee is 8.5 MB, runs 25x faster than real time on one CPU thread, and scores 4.41 on UTMOS against the teacher's 4.52. The student initially kept a slight buzz, which we trace to the phase of voiced speech between 2 and 8 kHz. A phase-locking filter applied after synthesis removes most of it, with no training and no extra parameters. Code, model files and audio samples are at https://github.com/sahilmahendrakar/paradee