papersSEP 10 04:00 UTC
Study examines whether speech-to-speech models infer gender from voice or content stereotypes
Researchers have released a study disentangling two distinct gender signals that speech-to-speech models can pick up: the acoustic characteristics of a speaker's voice and gender-related cues embedded in what is being said. This distinction matters for applications like dubbing, translation, and voice agents, where an ideal system should preserve how a speaker actually sounds rather than defaulting to stereotyped content. The work offers a framework for auditing whether these models rely on voice or on content-based assumptions when producing gendered output.