Object Counts! Bringing Explicit Detections Back into Image Captioning.
File(s) document(9).pdf (4.81 MB)
Published version
Author(s)
Wang, J
Madhyastha, PS
Specia, L
Type
Conference Paper
Abstract
The use of explicit object detectors as an intermediate step to image captioning - which used to constitute an essential stage in early work - is often bypassed in the currently dominant end-to-end approaches, where the language model is conditioned directly on a mid-level image embedding. We argue that explicit detections provide rich semantic information, and can thus be used as an interpretable representation to better understand why end-to-end image captioning systems work well. We provide an in-depth analysis of end-to-end image captioning by exploring a variety of cues that can be derived from such object detections. Our study reveals that end-to-end image captioning systems rely on matching image representations to generate captions, and that encoding the frequency, size and position of objects are complementary and all play a role in forming a good image representation. It also reveals that different object categories contribute in different ways towards image captioning.
Date Issued
2018
Date Acceptance
2018-06-01
Citation
2018, pp.2180-2193
ISBN
9781948087278
Publisher
Association for Computational Linguistics (ACL)
Start Page
2180
End Page
2193
Journal / Book Title
Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers)
Volume
1
Issue
N18-1
Copyright Statement
© 2018 Association for Computational Linguistics. Licensed on a Creative Commons Attribution 4.0 License (https://creativecommons.org/licenses/by/4.0/).
Identifier
https://doi.org/10.18653/v1/n18-1198
Source
Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers)
Subjects
cs.CV
cs.AI
cs.CL
Publication Status
Published
Start Date
2018-06-01
Finish Date
2018-06-06
Coverage Spatial
New Orleans, LA, USA
