Rigorous probabilistic guarantees for robust counterfactual explanations
File(s)2407.07482v1.pdf (706.88 KB)
Accepted version
Author(s)
Marzari, Luca
Leofante, Francesco
Cicalese, Ferdinando
Farinelli, Alessandro
Type
Conference Paper
Abstract
We study the problem of assessing the robustness of
counterfactual explanations for deep learning models. We focus on plausible model shifts altering model parameters and propose a novel framework to reason about the robustness property in this setting. To motivate our solution, we begin by showing for the first time that computing the robustness of counterfactuals with respect to plausible
model shifts is NP-complete. As this (practically) rules out the existence of scalable algorithms for exactly computing robustness, we propose a novel probabilistic approach which is able to provide tight estimates of robustness with strong guarantees while preserving scalability. Remarkably, and differently from existing solutions targeting
plausible model shifts, our approach does not impose requirements on the network to be analyzed, thus enabling robustness analysis on a wider range of architectures. Experiments on four binary classification datasets indicate that our method improves the state of the art in
generating robust explanations, outperforming existing methods on a range of metrics.
counterfactual explanations for deep learning models. We focus on plausible model shifts altering model parameters and propose a novel framework to reason about the robustness property in this setting. To motivate our solution, we begin by showing for the first time that computing the robustness of counterfactuals with respect to plausible
model shifts is NP-complete. As this (practically) rules out the existence of scalable algorithms for exactly computing robustness, we propose a novel probabilistic approach which is able to provide tight estimates of robustness with strong guarantees while preserving scalability. Remarkably, and differently from existing solutions targeting
plausible model shifts, our approach does not impose requirements on the network to be analyzed, thus enabling robustness analysis on a wider range of architectures. Experiments on four binary classification datasets indicate that our method improves the state of the art in
generating robust explanations, outperforming existing methods on a range of metrics.
Date Acceptance
2024-07-04
Citation
ECAI Proceedings
Publisher
IOS Press
Journal / Book Title
ECAI Proceedings
Copyright Statement
Subject to copyright. This paper is embargoed until publication. Once published the Version of Record (VoR) will be available on immediate open access.
Source
27th European Conference on Artificial Intelligence (ECAI 2024)
Start Date
2024-10-19
Finish Date
2024-10-24
Coverage Spatial
Santiago de Compostela, Spain
Rights Embargo Date
10000-01-01