Privacy-preserving dementia processing in speech
File(s)
Author(s)
Woszczyk, Dominika
Type
Thesis
Abstract
As voice-activated devices become integral to daily life, privacy concerns around speech data are growing. These systems can access sensitive speech attributes, such as gender, emotions, and health conditions. Notably, dementia, the seventh leading cause of death, can be detected through speech patterns well before symptoms become debilitating. Early detection, while beneficial for diagnosis, could also be exploited by insurers, or employers. With dementia's growing prevalence in ageing populations, privacy-preserving solutions are urgently needed. This thesis investigates privacy-preserving techniques for dementia detection and obfuscation in speech, focusing on Alzheimer's disease (AD) in both audio and text domains. The main research aim of this work is to protect sensitive health-related information while maintaining the functionality of voice-based systems. First, the thesis enhances adversarial robustness in dementia detection using data augmentation, demonstrating that baseline detection systems can achieve performance levels comparable to state-of-the-art (SOTA) models, in both text and audio. Next, privacy-preserving dementia detection is examined. We find that by using domain knowledge and prosody disentanglement we can effectively preserve dementia detection while anonymizing speaker embeddings that can then be used for other downstream tasks. Then, the thesis addresses dementia obfuscation in text. To this end, we explore the capacities of the SOTA large-language-model (LLM) at performing zero-shot dementia obfuscation, a task notoriously hard beforehand with the challenging low-resource setting. We find that LLMs achieve SOTA obfuscation capabilities, outperforming current work and preserving better semantics. Finally, we propose an end-to-end framework for dementia obfuscation in speech, using zero-shot synthesis to transfer speaker characteristics while minimizing dementia-related anomalies. This system demonstrates promising performance against adversaries across text, audio, and fusion while enhancing speaker characteristic transfer. Through these works, this thesis expands the fields of dementia detection, anonymization in speech, privacy-preserving text rewriting and speech synthesis.
Version
Open Access
Date Issued
2024-11-15
Date Awarded
01/03/2025
License URL
Advisor
Demetriou, Soteris
Publisher Department
Department of Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)