VoiceArena 发布 Monsoon:50 语言语音数据集让 Telugu 语音识别错误率从 92.7% 降到 16.1%
同一个 Whisper Medium,换个数据集微调就把 Telugu 错误率从 92.7% 干到 16.1%,还故意保留噪声做分层,思路很值得学。
VoiceArena 发布 Monsoon,一个覆盖 50 种语言、来自 23 个国家非脚本对话的语音数据集。Whisper Medium 在 IndicVoices Telugu 上的语义 WER 为 92.7%,用 Monsoon 微调同一个 769M 参数的 checkpoint 后降到 16.1%,在其自测中略胜 MAI-Transcribe-2 和 Gemini 3.1 Pro。全程没有改架构、没有换更大的模型,76.6 个百分点的提升全部来自训练数据。Monsoon 的做法是刻意保留噪声:短片段、嘈杂房间等劣质音频留在训练分布里,并按 DNSMOS 对片段分层,让劣质音频的占比成为可控的数据属性。
Low-resource ASR (automatic speech recognition) usually gets treated as a model problem.
VoiceArena just launched Monsoon, and shows it's a data collection problem.
Monsoon is a 50-language speech dataset built from unscripted conversations across 23 countries.
Off the shelf, Whisper Medium scores 92.7% semantic WER on IndicVoices Telugu. VoiceArena fine-tuned the same 769M checkpoint on Monsoon and got 16.1%, narrowly ahead of MAI-Transcribe-2 and Gemini 3.1 Pro in its own tests.
There was no architecture change and no bigger model. The 76.6-point drop comes entirely from what the model heard during fine-tuning.
Most speech data pipelines filter noise out. Monsoon keeps it in on purpose.
Short clips, noisy rooms and difficult acoustic conditions stay in the training distribution, and the mix isn't left to chance. VoiceArena stratifies segments by DNSMOS. That makes the share of degraded audio a controlled property of the data