论文

Itgan 用 LoRA 微调 Whisper 参加 NADI 2026 阿拉伯语 ASR 三项子任务

Itgan at NADI 2026 shared task: Parameter-Efficient Whisper Adaptation for Robust, Mixed-Dialect and Code-Switched Arabic ASR

精选理由

一篇很实在的参赛报告:LoRA 微调 Whisper 搞定阿拉伯语多方言和码切换识别,还公开了八个没做成的方向,做 ASR 的可以看看。

Itgan 团队为 NADI 2026 的三个 ASR 子任务开发了同一套方案:用 LoRA 在消费级 GPU 上微调 Whisper。在国家级方言识别子任务 1.1 中,系统取得 57.1% 的国家平均 WER,事后加入的线性探针可在无方言标签时恢复 44% 的 oracle 路由收益。混合方言子任务 1.2 达到 46.7% WER,结果显示基座模型的选择比适配器容量更关键。突尼斯语码切换子任务 1.3 排名第二,WER 为 14.49%,CER 5.38% 为头部提交中最低,最后 0.60 个 WER 点来自无需再训练的权重空间平均与 ROVER 投票。论文还报告了八个失败方向。

原文 · arXiv cs.AI

Itgan at NADI 2026 shared task: Parameter-Efficient Whisper Adaptation for Robust, Mixed-Dialect and Code-Switched Arabic ASR

We describe the Itgan systems for the three ASR subtasks of NADI 2026, namely robust country-level ASR (1.1), mixed-dialect ASR (1.2), and Tunisian code-switched ASR (1.3). All three share one recipe, Whisper adapted with LoRA on consumer GPUs, and each was carried by a different addition to it. On 1.1, where the dialect label is given at test time, per-dialect specialists continued from a pooled adapter gave the largest gain, and the submitted system reached 57.1% country-average WER. A post-evaluation linear probe on frozen encoder features routes utterances without the label and recovers 44% of what oracle routing gives. On 1.2 the choice of base model mattered more than adapter capacity, and system combination helped only once we added a decorrelated member, reaching 46.7% WER. On 1.3 our system placed second at 14.49% WER with the lowest CER among the leading submissions, 5.38%. Its last 0.60 WER points came without further training, mostly from an exact weight-space average of independently trained runs, with ROVER voting adding the remainder. Every comparison carries a paired-bootstrap test, and we report eight directions that did not work.