papersSEP 12 04:00 UTC
Switch-Aware Evaluation of ASR and Audio Language Models on English-Yoruba Code-Switched Speech
A new arXiv preprint argues that word error rate alone hides important failures when speech recognition systems and audio language models handle code-switched speech. The authors propose an evaluation method that accounts for language switches, and apply it to English-Yoruba audio, a low-resource pair with diacritics. They find that strong monolingual benchmark scores do not carry over to this setting.