Extracting and comparing machine learning-based representations across natural and medical modalities
File(s)
Author(s)
Cadet, Xavier
Type
Thesis
Abstract
Machine learning has led to breakthroughs across many domains, such as Computer Vision, Natural Language Processing, and Speech Recognition. Given these breakthroughs, machine learning methods are being applied to more and more fields, such as healthcare, education, and legal services. However, the data required to train such models can be sensitive and limited, making it challenging to develop models in such fields. To ensure the applicability of Artificial Intelligence in such sensitive fields, we need efficient, reliable methods that consider privacy preservation. In this thesis, I aim to address these challenges by comparing representations, evaluating their robustness to incomplete data, and preserving privacy through machine unlearning methods.
I introduce several key contributions across these domains:
I propose the Feature Impact Balance (FIB) score, a metric to assess error distribution, and the Min-Max Relative Change Quadrant (MMRCQ) plot, a visualization tool to monitor changes in extreme values. I explore the class-wise impact of data augmentation in image-based models, showing that augmentations affect different classes unequally. Data augmentations can make classes less distinguishable, and such an effect is class-dependent. I evaluate the transferability of foundation models trained on non-medical data to domains with limited labeled data, demonstrating their potential for medical applications such as dysarthria automated assessments. I show that foundation models trained on large-scale speech from healthy individuals can be used to detect dysarthria, classify words, and classify intelligibility. I also assess the robustness of self-supervised learning (SSL) models to incomplete data.
Finally, I benchmark machine unlearning methods, making models "forget" data points without losing overall performance. This benchmark shows that unlearning algorithms must be compared against stronger baselines with extensive hyper-parameter searches. This thesis advances the development of more robust, reliable, and privacy-conscious machine learning models.
I introduce several key contributions across these domains:
I propose the Feature Impact Balance (FIB) score, a metric to assess error distribution, and the Min-Max Relative Change Quadrant (MMRCQ) plot, a visualization tool to monitor changes in extreme values. I explore the class-wise impact of data augmentation in image-based models, showing that augmentations affect different classes unequally. Data augmentations can make classes less distinguishable, and such an effect is class-dependent. I evaluate the transferability of foundation models trained on non-medical data to domains with limited labeled data, demonstrating their potential for medical applications such as dysarthria automated assessments. I show that foundation models trained on large-scale speech from healthy individuals can be used to detect dysarthria, classify words, and classify intelligibility. I also assess the robustness of self-supervised learning (SSL) models to incomplete data.
Finally, I benchmark machine unlearning methods, making models "forget" data points without losing overall performance. This benchmark shows that unlearning algorithms must be compared against stronger baselines with extensive hyper-parameter searches. This thesis advances the development of more robust, reliable, and privacy-conscious machine learning models.
Version
Open Access
Date Issued
2024-10-19
Date Awarded
01/04/2025
License URL
Advisor
Haddadi, Hamed
Ahmadi-Abhari, Sara
Sponsor
UK Research and Innovation
Grant Number
EP/S023283/1
Publisher Department
Department of Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
