论文

SCB基准测试评估语音多轮推理能力

SCB: SpeechConversationBench for Evaluating Multi-Turn Reasoning in Speech-to-Speech Models

精选理由

SCB团队用103个数学问题测试了语音系统多轮推理能力,LEGO表现优于GPT-4o Realtime。

SCB团队发布SpeechConversationBench基准,使用103个分片GSM8K数学问题评估语音系统。测试比较完整问题、拼接信息和分片披露三种条件。四个商业系统在分片条件下准确率下降5.0-25.3个百分点。LEGO语音管道在三种条件下均达到77.5%准确率,高于GPT-4o Realtime的76.6%。

原文 · arXiv cs.AI

SCB: SpeechConversationBench for Evaluating Multi-Turn Reasoning in Speech-to-Speech Models

Speech-to-speech systems must solve tasks whose requirements emerge across conversational turns. We introduce SpeechConversationBench (SCB), a focused evaluation of spoken mathematical reasoning using 103 sharded GSM8K problems. The framework compares the original problem delivered in one turn (full), its concatenated information shards delivered together (concat), and incremental spoken disclosure across turns (sharded). We report final-answer accuracy for four commercial speech systems and LEGO, a proprietary speech pipeline developed internally by the SCBX Innovation Lab team with explicit conversational context management. Relative to concat, sharded accuracy decreases by 5.0-25.3 percentage points across the four commercial systems. LEGO achieves 77.5 percent accuracy in all three conditions, compared with 76.6 percent sharded accuracy for GPT-4o Realtime. The two single-turn baselines distinguish sensitivity to problem reformulation from the additional challenges introduced by incremental spoken interaction.