Improving the robustness of natural language inference models
File(s)
Author(s)
Stacey, Joe
Type
Thesis or dissertation
Abstract
Current neural models excel on a wide range of natural language understanding tasks, often achieving human-level performance. However, unlike humans, these models generalise poorly, with substantially worse performance when tested on unseen, out-of-distribution data. In this thesis I investigate a range of strategies to improve model robustness, including model-centric methods that involve changing the model’s training procedure, and data-centric methods that can also be applied with closed-source models. In particular, I investigate methods to improve robustness while also maintaining or improving in-distribution performance, whereas most prior work involves a trade-off between the two. The methods investigated in this thesis include supervising models with human-authored natural language explanations, and generating unlabelled out-of-distribution data to be used with knowledge distillation. Additionally, I investigate data-centric methods, including upsampling challenging training examples and replacing existing training data with LLM-generated synthetic examples. The proposed methods each show significant improvements on out-of-distribution data, with the largest improvements seen when using LLM-generated synthetic data to either replace existing training examples, or using the synthetic data to improve the effectiveness of knowledge distillation. To better understand the reasons for each model prediction, I also introduce a new framework for creating inherently interpretable models. This framework decomposes inputs into different atoms, before making independent predictions for each atom. This approach provides a faithfulness guarantee for each instance, highlighting exactly which components of the input were responsible for the prediction. These interpretable methods mostly maintain model performance, with further improvements when the task is sufficiently challenging for the model. Overall, this thesis demonstrates that models can be trained to be both more robust and interpretable, without sacrificing performance.
Version
Open Access
Date Issued
2025-09-01
Date Awarded
2026-06-01
Copyright Statement
Attribution-NonCommercial 4.0 International Licence (CC BY-NC)
License URL
Advisor
Rei, Marek
Publisher Department
Department of Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
