MLQE-PE: A multilingual quality estimation and post-editing dataset
File(s)2010.04480v1.pdf (444.27 KB)
Working paper
Author(s)
Type
Working Paper
Abstract
We present MLQE-PE, a new dataset for Machine Translation (MT) Quality
Estimation (QE) and Automatic Post-Editing (APE). The dataset contains seven
language pairs, with human labels for 9,000 translations per language pair in
the following formats: sentence-level direct assessments and post-editing
effort, and word-level good/bad labels. It also contains the post-edited
sentences, as well as titles of the articles where the sentences were extracted
from, and the neural MT models used to translate the text.
Estimation (QE) and Automatic Post-Editing (APE). The dataset contains seven
language pairs, with human labels for 9,000 translations per language pair in
the following formats: sentence-level direct assessments and post-editing
effort, and word-level good/bad labels. It also contains the post-edited
sentences, as well as titles of the articles where the sentences were extracted
from, and the neural MT models used to translate the text.
Date Issued
2020-10-09
Citation
2020
Publisher
arXiv
Copyright Statement
© 2020 The Author(s)
Identifier
http://arxiv.org/abs/2010.04480v1
Subjects
cs.CL
cs.CL
Publication Status
Published