Preserving the privacy of sensitive attributes contained in voice signals
File(s)
Author(s)
Aloufi, Ranya
Type
Thesis
Abstract
Voice user interfaces and digital assistants are rapidly entering our lives and becoming singular touch points spanning our devices. These always-on services capture and transmit our voice data to powerful cloud services for further processing and subsequent actions. Our voices and raw audio signals collected through these devices contain a host of sensitive paralinguistic information that is transmitted to service providers regardless of deliberate or false triggers. As our emotional patterns and sensitive attributes like our identity, gender, well-being, are easily inferred using deep acoustic models, we encounter a new generation of privacy risks by using these services.
GDPR mandates the principle of data minimization, which stipulates that only data required to fulfill a specific purpose be collected. Thus, the current collection of such complex signals raises concerns regarding compliance with GDPR and other data protection regulations. Our main objective in this thesis is to lay down the foundation of designing privacy-preserving solutions for voice input at the source to satisfy the data minimization requirement.
We first present hybrid privacy-preservation approaches incorporating on-device paralinguistic information filtering with cloud-based processing. We begin with Emotionless, which filters affective attributes from raw voice data before outsourcing the analysis to cloud-based services. To do this, we use deep learning approaches, specifically CycleGAN that has been trained to learn the sensitive attribute (i.e., emotion) from the raw recording and converted neutral one before regenerating the speech signal...
GDPR mandates the principle of data minimization, which stipulates that only data required to fulfill a specific purpose be collected. Thus, the current collection of such complex signals raises concerns regarding compliance with GDPR and other data protection regulations. Our main objective in this thesis is to lay down the foundation of designing privacy-preserving solutions for voice input at the source to satisfy the data minimization requirement.
We first present hybrid privacy-preservation approaches incorporating on-device paralinguistic information filtering with cloud-based processing. We begin with Emotionless, which filters affective attributes from raw voice data before outsourcing the analysis to cloud-based services. To do this, we use deep learning approaches, specifically CycleGAN that has been trained to learn the sensitive attribute (i.e., emotion) from the raw recording and converted neutral one before regenerating the speech signal...
Version
Open Access
Date Issued
2023-01-06
Date Awarded
2023-11-01
Copyright Statement
Attribution-Non Commercial-No Derivatives 4.0 International Licence (CC BY-NC-ND)
Advisor
Boyle, David
Haddadi, Hamed
Sponsor
Saudi Arabia. Ministry of Higher Education
Publisher Department
Dyson School of Design Engineering
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
