Non-Intrusive POLQA estimation of speech quality using recurrent neural networks
Author(s)
Sharma, Dushyant
Hogg, Aidan
Wang, Yu
Nour-Eldin, Amr
Naylor, Patrick
Type
Conference Paper
Abstract
Estimating the quality of speech without the use of a clean reference signal is a challenging problem, in part due to the time and expense required to collect sufficient training data for modern machine learning algorithms. We present a novel, non-intrusive estimator that exploits recurrent neural network architectures to predict the intrusive POLQA score of a speech signal in a short time context. The predictor is based on a novel compressed representation of modulation domain features, used in conjunction with static MFCC features. We show that the proposed method can reliably predict POLQA with a 300 ms context, achieving a mean absolute error of 0.21 on unseen data.The proposed method is trained using English speech and is shown to generalize well across unseen languages. The neural network also jointly estimates the mean voice activity detection(VAD) with an F1 accuracy score of 0.9, removing the need for an external VAD.
Date Issued
2019-11-18
Date Acceptance
2019-06-04
Citation
2019 27th European Signal Processing Conference (EUSIPCO), 2019
Publisher
IEEE
Journal / Book Title
2019 27th European Signal Processing Conference (EUSIPCO)
Copyright Statement
© 2019 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
Source
European Signal Processing Conference (EUSIPCO)
Publication Status
Published
Start Date
2019-09-02
Finish Date
2019-09-06
Coverage Spatial
Coruna, Spain
Date Publish Online
2019-09-06