Grammatical error detection in transcriptions of spoken English
File(s)2020.coling-main.195.pdf (1.02 MB)
Published version
Author(s)
Caines, Andrew
Benz, Christian
Knill, Kate
Rei, Marek
Buttery, Paula
Type
Conference Paper
Abstract
We describe the collection of transcription corrections and grammatical error annotations for the CROWDED Corpus of spoken English monologues on business topics. The corpus recordings were crowdsourced from native speakers of English and learners of English with German as their first language. The new transcriptions and annotations are obtained from different crowd workers: we analyse the 1108 new crowd worker submissions and propose that they can be used for automatic transcription post-editing and grammatical error correction for speech. To further explore the data we train grammatical error detection models with various configurations including pre-trained and contextual word representations as input, additional features and auxiliary objectives, and extra training data from written error-annotated corpora. We find that a model concatenating pre-trained and contextual word representations as input performs best, and that additional information does not lead to further performance gains.
Date Issued
2020-12-08
Date Acceptance
2020-09-30
Citation
Proceedings of the 28th International Conference on Computational Linguistics, 2020, pp.2144-2162
Publisher
International Committee on Computational Linguistics
Start Page
2144
End Page
2162
Journal / Book Title
Proceedings of the 28th International Conference on Computational Linguistics
Copyright Statement
© 2020 The Author(s). This work is licensed under a Creative Commons Attribution 4.0 International Licence.Licence details: http://creativecommons.org/licenses/by/4.0/.
License URL
Source
The 28th International Conference on Computational Linguistics (COLING 2020)
Publication Status
Published
Start Date
2020-12-08
Finish Date
2020-12-13
Coverage Spatial
Virtual