arXiv paper targets controllable speech generation with nonverbal vocalizations
A new arXiv preprint addresses the difficulty of synthesizing nonverbal vocalizations such as laughs, sighs, and coughs in controllable speech generation. The authors attribute the challenge to the acoustic variety of these sounds and their uneven representation in existing speech corpora, and propose modeling, scaling, and decoding methods to improve them. The work is listed under the cs.AI cross-list announcement.