Acoustic sensing for voice assistants

dc.contributor.advisorRoedig, Utz
dc.contributor.advisorDinh, Nguyen
dc.contributor.authorPham, Hoang Quoc Vieten
dc.contributor.funderScience Foundation Irelanden
dc.date.accessioned2026-05-20T13:39:46Z
dc.date.available2026-05-20T13:39:46Z
dc.date.issued2025-05-31
dc.date.submitted2025-05-31
dc.description.abstractRecent technological breakthroughs, such as AI assistants, smart devices, and IoT, are significantly enhancing the quality of human life. Over time, the commercialization of smart devices at affordable prices has made these applications more accessible, allowing users to integrate them into their daily lives more widely. Voice Assistants (VAs) in the form of smart speakers such as Amazon Alexa, Google Assistant, or Apple Siri are now commonplace. With their increasing presence in households, smart devices are becoming more versatile. Beyond serving as virtual assistants, these devices can be leveraged to create various useful applications. Transforming a smart speaker into a home alarm system allows end users to utilize their existing hardware as a valuable intrusion detection system without incurring additional costs beyond the smart speaker itself. In this work, we present the development and evaluation of a novel physical intrusion detection system GOTCHA based on human presence detection with active acoustic detection. GOTCHA can execute on simple off-the-shelf smart speaker hardware. Periodic audible chirps are employed to gather data that are then processed by a deep autoencoder trained on the acoustic profile of the empty room. As our evaluation shows, GOTCHA achieves promising results of up to 99. 2% for the F1 score. Our experiments show that a person can be detected without fail while the system barely generates false alarms. GOTCHA is a viable alternative to passive detection solutions for the detection of physical intrusion. Moreover, ensuring that GOTCHA can quickly adapt to changes in its surrounding environment is a critical challenge. We conducted experiments by retraining GOTCHA when transitioning from one background environment to another and found that with just 100 samples in the training set, the system was able to achieve an F1-score above 90% while maintaining an acceptable false alarm rate. Additionally, by applying transfer learning, the retraining time for GOTCHA was significantly reduced compared to fully retraining the entire model. This demonstrates that GOTCHA is entirely feasible for real-world deployment.en
dc.description.statusNot peer revieweden
dc.description.versionAccepted Versionen
dc.format.mimetypeapplication/pdfen
dc.identifier.citationPham, H. Q. V. 2025. Acoustic sensing for voice assistants. MRes Thesis, University College Cork.
dc.identifier.endpage64
dc.identifier.urihttps://hdl.handle.net/10468/18798
dc.language.isoenen
dc.publisherUniversity College Corken
dc.relation.projectinfo:eu-repo/grantAgreement/SFI/Frontiers for the Future::Awards/19/FFP/6775/IE/Personal Voice Assistant Security and Privacy/en
dc.relation.projectinfo:eu-repo/grantAgreement/SFI/Research Centres Programme::Phase 2/13/RC/2077_P2/IE/CONNECT_Phase 2/en
dc.rights© 2025, Hoang Quoc Viet Pham.
dc.rights.urihttps://creativecommons.org/licenses/by-sa/4.0/
dc.subjectPhysical intrusion detection
dc.subjectActive acoustic sensing
dc.subjectAnomaly detection
dc.subjectAutoencoder
dc.subjectMachine learning
dc.subjectDeep learning
dc.titleAcoustic sensing for voice assistants
dc.typeMasters thesis (Research)en
dc.type.qualificationlevelMastersen
dc.type.qualificationnameMSc - Master of Scienceen
Files
Original bundle
Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
PhamHQV_MRes2025.pdf
Size:
18.81 MB
Format:
Adobe Portable Document Format
Description:
Full Text E-thesis
License bundle
Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
5.2 KB
Format:
Item-specific license agreed upon to submission
Description: