Visual-only recognition of normal, whispered and silent speech
File(s)normalwhispersilentdb.pdf (265.48 KB)
Accepted version
Author(s)
Petridis, Stavros
Shen, Jie
Cetin, Doruk
Pantic, Maja
Type
Conference Paper
Abstract
Silent speech interfaces have been recently proposed as a way to enable communication when the acoustic signal is not available. This introduces the need to build visual speech recognition systems for silent and whispered speech. However, almost all the recently proposed systems have been trained on vocalised data only. This is in contrast with evidence in the literature which suggests that lip movements change depending on the speech mode. In this work, we introduce a new audiovisual database which is publicly available and contains normal, whispered and silent speech. To the best of our knowledge, this is the first study which investigates the differences between the three speech modes using the visual modality only. We show that an absolute decrease in classification rate of up to 3.7% is observed when training and testing on normal and whispered, respectively, and vice versa. An even higher decrease of up to 8.5% is reported when the models are tested on silent speech. This reveals that there are indeed visual differences between the 3 speech modes and the common assumption that vocalized training data can be used directly to train a silent speech recognition system may not be true.
Date Issued
2018-09-13
Date Acceptance
2018-04-15
Citation
2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018, pp.6219-6223
ISSN
2379-190X
Publisher
IEEE
Start Page
6219
End Page
6223
Journal / Book Title
2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Copyright Statement
© 2018 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
Sponsor
Commission of the European Communities
Identifier
http://gateway.webofknowledge.com/gateway/Gateway.cgi?GWVersion=2&SrcApp=PARTNER_APP&SrcAuth=LinksAMR&KeyUT=WOS:000446384606076&DestLinkType=FullRecord&DestApp=ALL_WOS&UsrCustomerID=1ba7043ffcc86c417c072aa74d649202
Grant Number
645094
Source
IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Subjects
Science & Technology
Technology
Acoustics
Engineering, Electrical & Electronic
Engineering
Visual Speech Recognition
Lipreading
End-to-End Training
Whispered Speech
Silent Speech
Publication Status
Published
Start Date
2018-04-15
Finish Date
2018-04-20
Coverage Spatial
Calgary, Canada
Date Publish Online
2018-09-13