Voice conversion and text-to-speech for privacy protection applications
File(s)
Author(s)
Nespoli, Francesco
Type
Thesis
Abstract
In recent years, the need for privacy preservation when manipulating or storing personal data has become a major issue. Particularly, speech data might contain significant Personal Identifiable Information (PII) such as full name, social security number or geographical positioning. This happens not only in the semantic but also in the acoustic domain: voice characteristics such as prosody, speaking rate, accent and intonation inherently contain a variety of PII such as personality, physical characteristics, emotional state, age and gender that can be identified and potentially used for malicious privacy attacks. This thesis explores the possibility of applying voice conversion, text-to-speech and signal processing-based approaches to remove the biometric identity from speech signals therefore protecting speaker’s privacy against identification attacks. While protecting the speaker’s identity against fraudulent access, the methods presented in this thesis aim to preserve the usefulness of speech on a variety of downstream tasks. The tools developed in this thesis explore new voice anonymization techniques with the aim to help, one one side, companies to comply with the most stringent privacy regulations, and, on the other, provide individuals the possibility to enforce their right to privacy.
Version
Open Access
Date Issued
2024-11-13
Date Awarded
01/03/2025
License URL
Advisor
Naylor, Patrick
Sponsor
European Commission
Grant Number
956369
Publisher Department
Department of Electrical and Electronic Engineering
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
