Audio-visual speech recognition with a hybrid CTC/attention architecture
File(s) 1810.00108.pdf (3.58 MB)
Accepted version
Author(s)
Petridis, Stavros
Stafylakis, Themos
Ma, Pingchuan
Tzimiropoulos, Georgios
Pantic, Maja
Type
Conference Paper
Abstract
Recent works in speech recognition rely either on connectionist temporal classification (CTC) or sequence-to-sequence models for character-level recognition. CTC assumes conditional independence of individual characters, whereas attention-based models can provide nonsequential alignments. Therefore, we could use a CTC loss in combination with an attention-based model in order to force monotonic alignments and at the same time get rid of the conditional independence assumption. In this paper, we use the recently proposed hybrid CTC/attention architecture for audio-visual recognition of speech in-the-wild. To the best of our knowledge, this is the first time that such a hybrid architecture architecture is used for audio-visual recognition of speech. We use the LRS2 database and show that the proposed audio-visual model leads to an 1.3% absolute decrease in word error rate over the audio-only model and achieves the new state-of-the-art performance on LRS2 database (7% word error rate). We also observe that the audio-visual model significantly outperforms the audio-based model (up to 32.9% absolute improvement in word error rate) for several different types of noise as the signal-to-noise ratio decreases.
Date Issued
2019-02-14
Date Acceptance
2018-12-18
Citation
2018 IEEE Spoken Language Technology Workshop (SLT), 2019, pp.513-520
ISSN
2639-5479
Publisher
IEEE
Start Page
513
End Page
520
Journal / Book Title
2018 IEEE Spoken Language Technology Workshop (SLT)
Copyright Statement
© 2018 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
Identifier
http://gateway.webofknowledge.com/gateway/Gateway.cgi?GWVersion=2&SrcApp=PARTNER_APP&SrcAuth=LinksAMR&KeyUT=WOS:000463141800072&DestLinkType=FullRecord&DestApp=ALL_WOS&UsrCustomerID=1ba7043ffcc86c417c072aa74d649202
Source
IEEE Workshop on Spoken Language Technology (SLT)
Subjects
Science & Technology
Technology
Computer Science, Artificial Intelligence
Engineering, Electrical & Electronic
Computer Science
Engineering
Audiovisual Speech Recognition
Attention Architectures
CTC
Audiovisual Fusion
Publication Status
Published
Start Date
2018-12-18
Finish Date
2018-12-21
Coverage Spatial
Athens, GREECE
Date Publish Online
2019-02-14
