Limitations of watermarking AI-generated speech using AudioSeal
| dc.contributor.author | Faziludeen, Shameer | en |
| dc.contributor.author | Sankar, Arun | en |
| dc.contributor.author | De Leon, Phillip L. | en |
| dc.contributor.author | Roedig, Utz | en |
| dc.contributor.funder | Science Foundation Ireland | en |
| dc.date.accessioned | 2025-12-22T10:05:12Z | |
| dc.date.available | 2025-12-22T10:05:12Z | |
| dc.date.issued | 2025 | en |
| dc.description.abstract | AI-generated speech is currently of such high quality that it is indistinguishable from a genuine human speaker. Expert listeners or purpose-built detectors are no longer able to reliably distinguish between the two. Thus, it has been proposed that AI systems which generate speech embed a secondary signal or watermark that allows identification. AudioSeal is currently the most advanced watermarking algorithm proposed for this purpose and its resilience against common channel and coding effects has been demonstrated. In this paper, we present approaches which compromise AudioSeal, making it unusable in practical settings. First, we describe two methods that result in a shifting of the detector score distribution for watermarked speech toward the distribution for unwatermarked speech. Second, we describe a method that uses AudioSeal watermarks generated for a particular speaker’s signal on a different speaker’s signal, i.e. unmatched watermarks. These unmatched watermarks, which could be imposed on genuine human speech, are also inaudible, resilient, and result in a shift of the detector score distribution away from unwatermarked speech. Considering both approaches, we observe that AudioSeal watermarks cannot be used to reliably identify AI-generated speech from genuine human speech due to overlapping score distributions. While our results are specific to AudioSeal, it casts doubt on the approach of watermarking in general to identify AI-generated speech. | en |
| dc.description.sponsorship | Science Foundation Ireland (13/RC/2077 P2) | en |
| dc.description.status | Peer reviewed | en |
| dc.description.version | Accepted Version | en |
| dc.format.mimetype | application/pdf | en |
| dc.identifier.citation | Faziludeen, S., Sankar, A., De Leon, P. L. and Roedig, U. (2025) 'Limitations of watermarking AI-generated speech using AudioSeal', 7th IEEE International Conference on Trust, Privacy and Security in Intelligent Systems, and Applications (TPS 2025), Pittsburgh, PA, USA, 11-14 November 2025. | en |
| dc.identifier.endpage | 10 | en |
| dc.identifier.startpage | 1 | en |
| dc.identifier.uri | https://hdl.handle.net/10468/18360 | |
| dc.language.iso | en | en |
| dc.publisher | Institute of Electrical and Electronics Engineers (IEEE) | en |
| dc.relation.ispartof | 7th IEEE International Conference on Trust, Privacy and Security in Intelligent Systems, and Applications (TPS 2025), Pittsburgh, PA, USA, 11-14 November 2025. | en |
| dc.relation.project | info:eu-repo/grantAgreement/SFI/Frontiers for the Future::Awards/19/FFP/6775/IE/Personal Voice Assistant Security and Privacy/ | en |
| dc.relation.project | 13/RC/2077 P2 | en |
| dc.rights | © 2025, the Authors. For the purpose of Open Access, the authors have applied a CC BY public copyright licence to any Author Accepted Manuscript version arising from this submission. | en |
| dc.rights.uri | https://creativecommons.org/licenses/by/4.0/ | en |
| dc.subject | Watermarking | en |
| dc.subject | AI generated speech | en |
| dc.subject | Deep fake | en |
| dc.subject | AudioSeal | en |
| dc.subject | Audio watermarking | en |
| dc.subject | Watermark attacks | en |
| dc.title | Limitations of watermarking AI-generated speech using AudioSeal | en |
| dc.type | Conference item | en |
