Personal Voice Assistant security - Denial of Service

Loading...
Thumbnail Image
Files
SagiP_PhD2025.pdf(10.88 MB)
Full Text E-thesis
Date
2025-04-23
Authors
Sagi, Prathyusha
Journal Title
Journal ISSN
Volume Title
Publisher
University College Cork
Published Version
Research Projects
Organizational Units
Journal Issue
Abstract
Personal Voice Assistants (PVA) have gained significant popularity across domains ranging from smartphones and smart speakers to critical setups like Intensive Care Units (ICUs) in hospitals, educational environments, and military operations. These PVAs offer hands-free convenience, allowing users to perform tasks such as setting reminders, controlling smart home devices, and managing complex processes. Despite their utility, they present notable security risks. One key vulnerability lies in the wake word detection process. PVAs continuously listen for a specific trigger word or phrase called a Wake Word, such as ’Alexa’, ’Hey Siri’, or ’Ok Google’. This wake word detection system can be vulnerable to attacks where an attacker masks the wake word using noise, preventing the device from waking up and processing user requests, effectively causing a Denial of Service (DoS) attack. Our study demonstrates how an attacker can strategically time a carefully designed noise burst, known as a jamming signal, to interfere with wake word detection in a PVA. By superimposing the jamming signal over specific sensitive regions of the wake word, the PVA can be rendered unresponsive. This is particularly concerning in highstakes environments such as Intensive Care Units (ICUs) or military operations, where voice commands are crucial. To address this vulnerability, we explore adversarial training, where the wake word detection model is retrained on wake word samples with jamming signals superimposed on sensitive regions, making the model resilient to targeted interference. We further show that using a wider range of noise types during adversarial training improves robustness to previously unseen jamming signals, while focusing on sensitive regions reduces training time and effort. Signal processing factors such as signal energies and phonetic characteristics, along with the model’s structure, play a key role in identifying these sensitive regions. Beyond robustness, alerting users to ongoing attacks is equally essential. We extend our wake word detection model from a binary classifier to a three-class classifier, labeling audio as non-wake word, wake word, or wake word + jamming signal. To reduce false alarms, we incorporate Direction of Arrival (DOA) and Short Time Energy (STE). DOA isolates jamming signals by identifying their direction relative to the wake word, while STE detects sudden energy spikes during sensitive regions, indicating deliberate interference. Combining either with the three-class classifier reduces unnecessary alarms, making the system robust to jamming and capable of detecting actual attacks.
Description
Keywords
Denial of Service , Wake words , Jamming , Adversarial training , Wake word jamming detection
Citation
Sagi, P. 2025. Personal Voice Assistant security - Denial of Service. PhD Thesis, University College Cork.
Link to publisher’s version