Guiding visual question generation
File(s)2022.naacl-main.118.pdf (1.48 MB)
Published version
Author(s)
Vedd, Nihir
Wang, Zixu
Rei, Marek
Miao, Yishu
Specia, Lucia
Type
Conference Paper
Abstract
In traditional Visual Question Generation (VQG), most images have multiple concepts (e.g. objects and categories) for which a question could be generated, but models are trained to mimic an arbitrary choice of concept as given in their training data. This makes training difficult and also poses issues for evaluation – multiple valid questions exist for most images but only one or a few are captured by the human references. We present Guiding Visual Question Generation - a variant of VQG which conditions the question generator on categorical information based on expectations on the type of question and the objects it should explore. We propose two variant families: (i) an explicitly guided model that enables an actor (human or automated) to select which objects and categories to generate a question for; and (ii) 2 types of implicitly guided models that learn which objects and categories to condition on, based on discrete variables. The proposed models are evaluated on an answer-category augmented VQA dataset and our quantitative results show a substantial improvement over the current state of the art (over 9 BLEU-4 increase). Human evaluation validates that guidance helps the generation of questions that are grammatically coherent and relevant to the given image and objects.
Date Issued
2022-07
Date Acceptance
2023-07-10
Citation
Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp.1640-1654
ISBN
978-1-955917-71-1
Publisher
ASSOC COMPUTATIONAL LINGUISTICS-ACL
Start Page
1640
End Page
1654
Journal / Book Title
Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
Copyright Statement
©2022 Association for Computational Linguistics Materials published in or after 2016 are licensed on a Creative Commons Attribution 4.0 International License.
License URL
Identifier
https://www.webofscience.com/api/gateway?GWVersion=2&SrcApp=PARTNER_APP&SrcAuth=LinksAMR&KeyUT=WOS:000859869501053&DestLinkType=FullRecord&DestApp=ALL_WOS&UsrCustomerID=a2bf6146997ec60c407a63945d4e92bb
Source
Conference of the North-American-Chapter-of-the-Association-for-Computational-Linguistics (NAAACL) - Human Language Technologies
Subjects
Computer Science
Computer Science, Artificial Intelligence
Computer Science, Interdisciplinary Applications
Linguistics
Science & Technology
Social Sciences
Technology
Publication Status
Published
Start Date
2022-07-10
Finish Date
2022-07-15
Coverage Spatial
WA, Seattle
Date Publish Online
2022-07