papersSEP 10 04:00 UTC
Study examines phoneme-based TTS augmentation pipeline for improving ASR training
A new arXiv paper introduces a unified pipeline that uses a single phoneme-based text-to-speech model to generate synthetic training data for automatic speech recognition. The work presents a controlled study of how text selection, reference speech, and augmentation scale affect recognition accuracy.