Improving neural machine translation robustness via data augmentation: beyond back-translation
File(s) D19-5543.pdf (200.26 KB)
Published version
Author(s)
Li, Zhenhao
Specia, Lucia
Type
Conference Paper
Abstract
Neural Machine Translation (NMT) models have been proved strong when translating clean texts, but they are very sensitive to noise in the input. Improving NMT models robustness can be seen as a form of “domain” adaption to noise. The recently created Machine Translation on Noisy Text task corpus provides noisy-clean parallel data for a few language pairs, but this data is very limited in size and diversity. The state-of-the-art approaches are heavily dependent on large volumes of back-translated data. This paper has two main contributions: Firstly, we propose new data augmentation methods to extend limited noisy data and further improve NMT robustness to noise while keeping the models small. Secondly, we explore the effect of utilizing noise from external data in the form of speech transcripts and show that it could help robustness.
Date Issued
2019-11-01
Date Acceptance
2019-11-01
Citation
Proceedings of the 5th Workshop on Noisy User-generated Text (W-NUT 2019), 2019, pp.328-336
Publisher
Association for Computational Linguistics
Start Page
328
End Page
336
Journal / Book Title
Proceedings of the 5th Workshop on Noisy User-generated Text (W-NUT 2019)
Copyright Statement
© 2019 Association for Computational Linguistics. Materials published in or after 2016 are licensed on a Creative Commons Attribution 4.0 International License (https://creativecommons.org/licenses/by/4.0/).
Identifier
https://www.aclweb.org/anthology/D19-5543/
Source
Proceedings of the 5th Workshop on Noisy User-generated Text (W-NUT 2019)
Subjects
cs.CL
cs.CL
Publication Status
Published
Start Date
2019-11-04
Finish Date
2019-11-04
Coverage Spatial
Hong Kong, China
Date Publish Online
2019-11-01
