Supervising model attention with human explanations for robust natural language inference
File(s)eSNLI_project___AAAI22__v2b___.pdf (1.49 MB)
Accepted version
Author(s)
Stacey, Joe
Belinkov, Yonatan
Rei, Marek
Type
Conference Paper
Abstract
Natural Language Inference (NLI) models are known to learn from biases and artefacts within their training data, impacting how well they generalise to other unseen datasets. Existing de-biasing approaches focus on preventing the models from learning these biases, which can result in restrictive models
and lower performance. We instead investigate teaching the model how a human would approach the NLI task, in order to learn features that will generalise better to previously un-seen examples. Using natural language explanations, we supervise the model’s attention weights to encourage more attention to be paid to the words present in the explanations, significantly improving model performance. Our experiments show that the in-distribution improvements of this method are also accompanied by out-of-distribution improvements, with the supervised models learning from features that gener-
alise better to other NLI datasets. Analysis of the model indicates that human explanations encourage increased attention on the important words, with more attention paid to words in the premise and less attention paid to punctuation and stop-words.
and lower performance. We instead investigate teaching the model how a human would approach the NLI task, in order to learn features that will generalise better to previously un-seen examples. Using natural language explanations, we supervise the model’s attention weights to encourage more attention to be paid to the words present in the explanations, significantly improving model performance. Our experiments show that the in-distribution improvements of this method are also accompanied by out-of-distribution improvements, with the supervised models learning from features that gener-
alise better to other NLI datasets. Analysis of the model indicates that human explanations encourage increased attention on the important words, with more attention paid to words in the premise and less attention paid to punctuation and stop-words.
Date Issued
2022-06-28
Date Acceptance
2021-12-01
Citation
AAAI 2022, 2022, 36 (10), pp.11349-11357
Publisher
Association for the Advancement of Artificial Intelligence
Start Page
11349
End Page
11357
Journal / Book Title
AAAI 2022
Volume
36
Issue
10
Copyright Statement
© 2022, Association for the Advancement of Artificial
Intelligence (www.aaai.org). All rights reserved.
Intelligence (www.aaai.org). All rights reserved.
Identifier
https://ojs.aaai.org/index.php/AAAI/article/view/21386
Source
AAAI 2022
Publication Status
Published
Start Date
2022-02-22
Finish Date
2022-03-01
Coverage Spatial
Vancouver, BC, Canada
Date Publish Online
2022-06-28