Probabilistically robust counterfactual explanations under model changes
File(s) 1-s2.0-S000437022500178X-main (1).pdf (6.87 MB)
Published version
Author(s)
Marzari, Luca
Leofante, Francesco
Cicalese, Ferdinando
Farinelli, Alessandro
Type
Journal Article
Abstract
We study the problem of generating robust counterfactual explanations for deep learning models subject to model changes. We focus on plausible model changes altering model parameters and propose a novel framework to reason about the robustness property in this setting. To motivate our solution, we begin by showing for the first time that computing the robustness of counterfactuals with respect to model changes is NP-hard. As this (practically) rules out the existence of scalable algorithms for exactly computing robustness, we propose a novel probabilistic approach which is able to provide tight estimates of robustness with strong guarantees while preserving scalability. Remarkably, and differently from existing solutions targeting plausible model changes, our approach does not impose requirements on the network to be analysed, thus enabling robustness analysis on a wider range of architectures, including state-of-the-art tabular transformers. A thorough experimental analysis on four binary classification datasets reveals that our method improves the state of the art in generating robust explanations, outperforming existing methods.
Date Issued
2026-02-01
Date Acceptance
2025-11-28
Citation
Artificial Intelligence, 2026, 351
ISSN
0004-3702
Publisher
Elsevier
Journal / Book Title
Artificial Intelligence
Volume
351
Copyright Statement
© 2025 Published by Elsevier B.V
License URL
Identifier
10.1016/j.artint.2025.104459
Publication Status
Published
Article Number
104459
Date Publish Online
2025-12-20
