Causal Analysis and Mitigation of Spurious Speech Onsets in Full-Duplex Speech LLMs
A new arXiv paper examines why full-duplex speech models such as Moshi and its PersonaPlex derivative sometimes start talking when the user has gone quiet, including on digital-zero input. The authors trace these unwanted onsets to specific causes in the generation process and propose ways to reduce them. The work targets improving turn-taking reliability in speech-to-speech systems.