模型

BRIDGE ASR 2.0 发布:18 种印度语言加西语葡语越语的真实对话语音识别基准

精选理由

humynlabs 出了个语音识别新基准,专测真实对话和混语言场景,23 个模型已上榜,做语音方向的可以拿来自测。

humynlabs 发布 BRIDGE ASR 2.0 基准,测试 23 个语音识别模型在真实两人对话上的转写能力,每段对话 10 到 15 分钟,覆盖 18 种印度语言以及西班牙语、葡萄牙语和越南语。核心指标 code-switch F1 衡量混在印度语言句子里的英文单词是否保留英文写法,写成天城文就得零分。排行榜可按语言和指标筛选,方法与评测数据全部公开。

原文 · elvis

Recommended benchmark. I expect voice to become one of the main ways people interact with robots and physical AI. That makes speech recognition in real conversations much more important than it seems today. Real conversations are hard to transcribe. People pause, talk over each other, and switch languages mid-sentence. Robots also need to understand the languages people actually speak. More than 5.5 billion people across the Global South are non-English speakers. The community needs a good way to measure these frontier capabilities. @humynlabs built BRIDGE ASR 2.0 to measure exactly this. It tests 23 speech recognition models on real two-person conversations, each 10 to 15 minutes long, in 18 Indic languages plus Spanish, Portuguese, and Vietnamese. The metric I find most useful is code-switch F1. It checks whether English words mixed into an Indic sentence stay in English. A model that writes "data backup" in Devanagari script scores zero. You can also filter the leaderboard by language and by metric. The methodology and evaluation data are public, and the team wants researchers to test the benchmark and find where it breaks. Check out the benchmark here: humynlabs.ai/bridge/ASR/2.0 💬 2 🔄 2 ❤️ 4 👀 1216 📊 3 ⚡