papersSEP 10 04:00 UTC
Zero-Shot Temporal Localisation of Audio Deepfakes in Multi-Speaker Conversations
A new arXiv paper addresses voice-cloning fraud in which just a sentence or two of a genuine multi-speaker conversation is swapped for synthetic audio. Rather than giving one real-or-fake verdict for an entire recording, the proposed method pinpoints the exact time spans of fake speech without needing labelled examples of the targeted fakes.