Computer audition for the analysis of respiratory and mental health
File(s)
Author(s)
Chang, Yi
Type
Thesis
Abstract
This thesis explores the application of computer audition methodologies to analyse respiratory and mental health, leveraging the inherent advantages of audio-based computational analysis in healthcare: non-invasive, remote, fast, efficient, and cost-effective approaches. Amidst increasing concerns over respiratory ailments like COVID-19 and rising awareness of mental health issues such as depression, this research addresses critical challenges in deploying deep neural networks (DNNs) within this domain. These challenges include data scarcity, data dependency, model inefficiency, model uninterpretability, and model vulnerability.
The thesis proposes a comprehensive set of methodologies and frameworks to enhance the effectiveness, explainability, and robustness of computer audition systems. To tackle data scarcity and enhance feature extraction, it introduces knowledge transfer techniques using pre-trained models on related large-scale datasets, and this thesis also applies data distillation to construct a synthesised, small dataset to optimise the training process. For improving model transparency, which is crucial in healthcare, a unified example-based explanation method is developed that employs adversarial attacks to select representative data and outliers, elucidating model decisions. Additionally, to fortify DNNs against adversarial attacks, a novel federated learning framework is proposed, designed to safeguard emotional speech data while enhancing privacy and resilience. This framework is complemented by an innovative black-box generator-based adversarial attacker that efficiently crafts sparse, transferable perturbations, thereby reinforcing model robustness.
Empirical evaluations across diverse datasets confirm the efficacy of the proposed approaches, suggesting their potential to considerably advance computer audition applications in supporting human well-being, both physically and mentally.
The thesis proposes a comprehensive set of methodologies and frameworks to enhance the effectiveness, explainability, and robustness of computer audition systems. To tackle data scarcity and enhance feature extraction, it introduces knowledge transfer techniques using pre-trained models on related large-scale datasets, and this thesis also applies data distillation to construct a synthesised, small dataset to optimise the training process. For improving model transparency, which is crucial in healthcare, a unified example-based explanation method is developed that employs adversarial attacks to select representative data and outliers, elucidating model decisions. Additionally, to fortify DNNs against adversarial attacks, a novel federated learning framework is proposed, designed to safeguard emotional speech data while enhancing privacy and resilience. This framework is complemented by an innovative black-box generator-based adversarial attacker that efficiently crafts sparse, transferable perturbations, thereby reinforcing model robustness.
Empirical evaluations across diverse datasets confirm the efficacy of the proposed approaches, suggesting their potential to considerably advance computer audition applications in supporting human well-being, both physically and mentally.
Version
Open Access
Date Issued
2024-08-27
Date Awarded
01/03/2025
License URL
Advisor
Schuller, Björn
Publisher Department
Department of Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
