Limitations of watermarking AI-generated speech using AudioSeal

dc.contributor.authorFaziludeen, Shameeren
dc.contributor.authorSankar, Arunen
dc.contributor.authorDe Leon, Phillip L.en
dc.contributor.authorRoedig, Utzen
dc.contributor.funderScience Foundation Irelanden
dc.date.accessioned2025-12-22T10:05:12Z
dc.date.available2025-12-22T10:05:12Z
dc.date.issued2025en
dc.description.abstractAI-generated speech is currently of such high quality that it is indistinguishable from a genuine human speaker. Expert listeners or purpose-built detectors are no longer able to reliably distinguish between the two. Thus, it has been proposed that AI systems which generate speech embed a secondary signal or watermark that allows identification. AudioSeal is currently the most advanced watermarking algorithm proposed for this purpose and its resilience against common channel and coding effects has been demonstrated. In this paper, we present approaches which compromise AudioSeal, making it unusable in practical settings. First, we describe two methods that result in a shifting of the detector score distribution for watermarked speech toward the distribution for unwatermarked speech. Second, we describe a method that uses AudioSeal watermarks generated for a particular speaker’s signal on a different speaker’s signal, i.e. unmatched watermarks. These unmatched watermarks, which could be imposed on genuine human speech, are also inaudible, resilient, and result in a shift of the detector score distribution away from unwatermarked speech. Considering both approaches, we observe that AudioSeal watermarks cannot be used to reliably identify AI-generated speech from genuine human speech due to overlapping score distributions. While our results are specific to AudioSeal, it casts doubt on the approach of watermarking in general to identify AI-generated speech.en
dc.description.sponsorshipScience Foundation Ireland (13/RC/2077 P2)en
dc.description.statusPeer revieweden
dc.description.versionAccepted Versionen
dc.format.mimetypeapplication/pdfen
dc.identifier.citationFaziludeen, S., Sankar, A., De Leon, P. L. and Roedig, U. (2025) 'Limitations of watermarking AI-generated speech using AudioSeal', 7th IEEE International Conference on Trust, Privacy and Security in Intelligent Systems, and Applications (TPS 2025), Pittsburgh, PA, USA, 11-14 November 2025.en
dc.identifier.endpage10en
dc.identifier.startpage1en
dc.identifier.urihttps://hdl.handle.net/10468/18360
dc.language.isoenen
dc.publisherInstitute of Electrical and Electronics Engineers (IEEE)en
dc.relation.ispartof7th IEEE International Conference on Trust, Privacy and Security in Intelligent Systems, and Applications (TPS 2025), Pittsburgh, PA, USA, 11-14 November 2025.en
dc.relation.projectinfo:eu-repo/grantAgreement/SFI/Frontiers for the Future::Awards/19/FFP/6775/IE/Personal Voice Assistant Security and Privacy/en
dc.relation.project13/RC/2077 P2en
dc.rights© 2025, the Authors. For the purpose of Open Access, the authors have applied a CC BY public copyright licence to any Author Accepted Manuscript version arising from this submission.en
dc.rights.urihttps://creativecommons.org/licenses/by/4.0/en
dc.subjectWatermarkingen
dc.subjectAI generated speechen
dc.subjectDeep fakeen
dc.subjectAudioSealen
dc.subjectAudio watermarkingen
dc.subjectWatermark attacksen
dc.titleLimitations of watermarking AI-generated speech using AudioSealen
dc.typeConference itemen
Files
Original bundle
Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
2025384933.pdf
Size:
318.13 KB
Format:
Adobe Portable Document Format
Description:
Accepted Version
License bundle
Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
2.71 KB
Format:
Item-specific license agreed upon to submission
Description: