Interval abstractions for robust counterfactual explanations
File(s) 1-s2.0-S0004370224001541-main.pdf (2.46 MB)
Published version
Author(s)
Jiang, Junqi
Leofante, Francesco
Rago, Antonio
Toni, Francesca
Type
Journal Article
Abstract
Counterfactual Explanations (CEs) have emerged as a major paradigm in explainable AI research, providing recourse recommendations for users affected by the decisions of machine learning models. However, CEs found by existing methods often become invalid when slight changes occur in the parameters of the model they were generated for. The literature lacks a way to provide exhaustive robustness guarantees for CEs under model changes, in that existing methods to improve CEs' robustness are mostly heuristic, and the robustness performances are evaluated empirically using only a limited number of retrained models. To bridge this gap, we propose a novel interval abstraction technique for parametric machine learning models, which allows us to obtain provable robustness guarantees for CEs under a possibly infinite set of plausible model changes Δ. Based on this idea, we formalise a robustness notion for CEs, which we call Δ-robustness, in both binary and multi-class classification settings. We present procedures to verify Δ-robustness based on Mixed Integer Linear Programming, using which we further propose algorithms to generate CEs that are Δ-robust. In an extensive empirical study involving neural networks and logistic regression models, we demonstrate the practical applicability of our approach. We discuss two strategies for determining the appropriate hyperparameters in our method, and we quantitatively benchmark CEs generated by eleven methods, highlighting the effectiveness of our algorithms in finding robust CEs.
Date Issued
2024-11
Date Acceptance
2024-08-28
Citation
Artificial Intelligence, 2024, 336
ISSN
0004-3702
Publisher
Elsevier
Journal / Book Title
Artificial Intelligence
Volume
336
Copyright Statement
© 2024 The Author(s). Published by Elsevier B.V. This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/).
License URL
Identifier
https://www.sciencedirect.com/science/article/pii/S0004370224001541
Publication Status
Published
Article Number
104218
Date Publish Online
2024-09-02
