HILDIF: interactive debugging of NLI models using influence functions
File(s)2021.internlp-1.1.pdf (313.69 KB)
Published version
Author(s)
Zylberajch, Hugo
Lertvittayakumjorn, Piyawat
Toni, Francesca
Type
Conference Paper
Abstract
Biases and artifacts in training data can cause unwelcome behavior in text classifiers (such as shallow pattern matching), leading to lack of generalizability. One solution to this problem is to include users in the loop and leverage their feedback to improve models. We propose a novel explanatory debugging pipeline called HILDIF, enabling humans to improve deep text classifiers using influence functions as an explanation method. We experiment on the Natural Language Inference (NLI) task, showing that HILDIF can effectively alleviate artifact problems in fine-tuned BERT models and result in increased model generalizability.
Date Issued
2021-08-05
Date Acceptance
2021-08-01
Citation
INTERNLP 2021 - FIRST WORKSHOP ON INTERACTIVE LEARNING FOR NATURAL LANGUAGE PROCESSING, 2021, pp.1-6
Publisher
ASSOC COMPUTATIONAL LINGUISTICS-ACL
Start Page
1
End Page
6
Journal / Book Title
INTERNLP 2021 - FIRST WORKSHOP ON INTERACTIVE LEARNING FOR NATURAL LANGUAGE PROCESSING
Copyright Statement
©2021 Association for Computational Linguistics
Identifier
http://gateway.webofknowledge.com/gateway/Gateway.cgi?GWVersion=2&SrcApp=PARTNER_APP&SrcAuth=LinksAMR&KeyUT=WOS:000696714000001&DestLinkType=FullRecord&DestApp=ALL_WOS&UsrCustomerID=1ba7043ffcc86c417c072aa74d649202
Source
1st Workshop on Interactive Learning for Natural Language Processing (InterNLP)
Subjects
Science & Technology
Social Sciences
Technology
Computer Science, Artificial Intelligence
Computer Science, Theory & Methods
Linguistics
Computer Science
Publication Status
Published
Start Date
2021-08-05
Finish Date
2021-08-06
Coverage Spatial
ELECTR NETWORK
Date Publish Online
2021-08-05