Neural speech and bio-electrical signal generation
File(s)
Author(s)
Chen, Zehua
Type
Thesis
Abstract
Recent years have witnessed the significant impact of neural networks on speech and bio-electrical signal generation. However, key problems of each task in these domains still should be addressed with innovative designs. This thesis focuses on the methods to improve the generation quality of speech and bio-electrical signals in the tasks including speech synthesis, speech enhancement, and bio-electrical signal filtering and generation.
In speech synthesis, diffusion models have achieved promising generation quality for waveform and text-to-speech synthesis. However, their slow inference speed restricts their real-time application scenarios. Hence, two fast sampling methods are proposed to accelerate the inference process while maintaining high-quality generation. In addition, considering the limited results generated by existing methods in spatial audio synthesis and text-to-audio synthesis, this research introduces cascaded diffusion models and latent diffusion models to them respectively, setting new records in these fields.
In speech enhancement, convolutional networks have shown strong results in speech denoising. However, their model architecture, such as the kernel size, has not been thoroughly investigated. Furthermore, statistical signal processing algorithms, like recursive smoothing, offer a simple way to extract temporal statistics of speech and noise. This research first studies the optimal choice of kernel size and then explores incorporating recursive smoothing into a structural convolutional kernel, both enhancing denoising results.
In bio-electrical signal generation, the focus is pulsative signal filtering and generation. As hand-crafted features are widely used for blood pressure (BP) estimation, the signal quality is crucial for accurate estimation. To evaluate the signal quality without training a network, we propose a probabilistic filtering model for Electrocardiography, Photoplethysmography, and BP recordings, improving BP estimation results. When generating the missing values, different from existing deterministic methods, we introduce diffusion models and develop augmented and personalized templates, enhancing the imputation quality in transient and extended missing cases.
In speech synthesis, diffusion models have achieved promising generation quality for waveform and text-to-speech synthesis. However, their slow inference speed restricts their real-time application scenarios. Hence, two fast sampling methods are proposed to accelerate the inference process while maintaining high-quality generation. In addition, considering the limited results generated by existing methods in spatial audio synthesis and text-to-audio synthesis, this research introduces cascaded diffusion models and latent diffusion models to them respectively, setting new records in these fields.
In speech enhancement, convolutional networks have shown strong results in speech denoising. However, their model architecture, such as the kernel size, has not been thoroughly investigated. Furthermore, statistical signal processing algorithms, like recursive smoothing, offer a simple way to extract temporal statistics of speech and noise. This research first studies the optimal choice of kernel size and then explores incorporating recursive smoothing into a structural convolutional kernel, both enhancing denoising results.
In bio-electrical signal generation, the focus is pulsative signal filtering and generation. As hand-crafted features are widely used for blood pressure (BP) estimation, the signal quality is crucial for accurate estimation. To evaluate the signal quality without training a network, we propose a probabilistic filtering model for Electrocardiography, Photoplethysmography, and BP recordings, improving BP estimation results. When generating the missing values, different from existing deterministic methods, we introduce diffusion models and develop augmented and personalized templates, enhancing the imputation quality in transient and extended missing cases.
Version
Open Access
Date Issued
2023-04-17
Date Awarded
01/03/2024
Advisor
Mandic, Danilo
Publisher Department
Electrical and Electronic Engineering
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
