MSVD-Turkish: A comprehensive multimodal dataset for integrated vision
and language research in Turkish
and language research in Turkish
File(s) 2012.07098v1.pdf (3.24 MB)
Working paper
Author(s)
Type
Working Paper
Abstract
Automatic generation of video descriptions in natural language, also called
video captioning, aims to understand the visual content of the video and
produce a natural language sentence depicting the objects and actions in the
scene. This challenging integrated vision and language problem, however, has
been predominantly addressed for English. The lack of data and the linguistic
properties of other languages limit the success of existing approaches for such
languages. In this paper we target Turkish, a morphologically rich and
agglutinative language that has very different properties compared to English.
To do so, we create the first large scale video captioning dataset for this
language by carefully translating the English descriptions of the videos in the
MSVD (Microsoft Research Video Description Corpus) dataset into Turkish. In
addition to enabling research in video captioning in Turkish, the parallel
English-Turkish descriptions also enables the study of the role of video
context in (multimodal) machine translation. In our experiments, we build
models for both video captioning and multimodal machine translation and
investigate the effect of different word segmentation approaches and different
neural architectures to better address the properties of Turkish. We hope that
the MSVD-Turkish dataset and the results reported in this work will lead to
better video captioning and multimodal machine translation models for Turkish
and other morphology rich and agglutinative languages.
video captioning, aims to understand the visual content of the video and
produce a natural language sentence depicting the objects and actions in the
scene. This challenging integrated vision and language problem, however, has
been predominantly addressed for English. The lack of data and the linguistic
properties of other languages limit the success of existing approaches for such
languages. In this paper we target Turkish, a morphologically rich and
agglutinative language that has very different properties compared to English.
To do so, we create the first large scale video captioning dataset for this
language by carefully translating the English descriptions of the videos in the
MSVD (Microsoft Research Video Description Corpus) dataset into Turkish. In
addition to enabling research in video captioning in Turkish, the parallel
English-Turkish descriptions also enables the study of the role of video
context in (multimodal) machine translation. In our experiments, we build
models for both video captioning and multimodal machine translation and
investigate the effect of different word segmentation approaches and different
neural architectures to better address the properties of Turkish. We hope that
the MSVD-Turkish dataset and the results reported in this work will lead to
better video captioning and multimodal machine translation models for Turkish
and other morphology rich and agglutinative languages.
Date Issued
2020-12-13
Citation
2020
Publisher
arXiv
Copyright Statement
© 2020 The Author(s)
Sponsor
British Council (Turkey)
Commission of the European Communities
Identifier
http://arxiv.org/abs/2012.07098v1
Grant Number
352343575 - 154082
678017
Subjects
cs.CV
cs.CV
Publication Status
Published
