Enhancing the robustness of counterfactual explanations via adversarial training
File(s) SAIV_VNN_Counterfactuals.pdf (406.23 KB)
Accepted version
Author(s)
Appachi Senthilkumar, Rithik
Leofante, Francesco
Ganesh, Vijay
Type
Conference Paper
Abstract
Counterfactual explanations (CEs) provide a simple, yet powerful interpretive approach to understanding neural network behavior, envisioning hypothetical scenarios by systematically altering input features and analyzing the resulting changes in model predictions. However, CEs are useful only if they are robust, i.e., they remain consistent and meaningful even when adversarially perturbed. While prior work has explored the development of generators for robust CEs, checking the robustness of CEs produced for DNNs has not been explored within the formal verification context, nor has it been studied under adversarial training. We present a systematic study utilizing adversarial training to fortify the underlying neural network, observing its effect on the formal robustness of DNNs with respect to CEs using α, β-CROWN, a state-of-the-art NNV. Our experiments across multiple datasets, network architectures, and CE generators indicate that adversarial training has a positive impact on the number of formally verified robust CEs. We also measure the impact of adversarial training on other desirable properties of CEs, such as plausibility and proximity. While plausibility does not change, there is a trade-off with proximity when using a gradient-based CE generator.
Date Issued
2026-07-18
Date Acceptance
2026-05-15
Citation
Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 2026, 16831, pp.299-313
ISBN
978-3-032-32356-9
ISSN
0302-9743
Publisher
Springer
Start Page
299
End Page
313
Journal / Book Title
Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume
16831
Copyright Statement
© 2027 The Author(s), under exclusive license to Springer Nature Switzerland AG. This is the author’s accepted manuscript made available under a CC-BY licence in accordance with Imperial’s Research Publications Open Access policy (www.imperial.ac.uk/oa-policy)
License URL
Source
Third International Symposium on AI Verification (SAIV 2026)
Publication Status
Published
Start Date
2026-07-24
Finish Date
2026-07-25
Coverage Spatial
Lisbon, Portugal
Date Publish Online
2026-07-18
