Hidden conflicts in neural networks and their implications for explainability
File(s)3715275.3732100.pdf (1.98 MB)
Published version
Author(s)
Dejl, Adam
Zhang, Dekai
Ayoobi, Hamed
Williams, matthew
Toni, Francesca
Type
Conference Paper
Abstract
Artificial Neural Networks (ANNs) often represent conflicts between features, arising naturally during training as the network learns to integrate diverse and potentially disagreeing inputs to better predict the target variable. Despite their relevance to the “reasoning” processes of these models, the properties and implications of conflicts for understanding and explaining ANNs remain underexplored. In this paper, we develop a rigorous theory of conflicts in ANNs and demonstrate their impact on ANN explainability through two case studies. In the first case study, we use our theory of conflicts to inspire the design of a novel feature attribution method, which we call Conflict-Aware Feature-wise Explanations (CAFE). CAFE separates the positive and negative influences of features and biases, enabling more faithful explanations for models applied to tabular data. In the second case study, we take preliminary steps towards understanding the role of conflicts in out-of-distribution (OOD) scenarios. Through our experiments, we identify potentially useful connections between model conflicts and different kinds of distributional shifts in tabular and image data. Overall, our findings demonstrate the importance of accounting for conflicts in the development of more reliable explanation methods for AI systems, which are crucial for the beneficial use of these systems in the society.
Date Issued
2025-06-23
Date Acceptance
2025-04-11
Citation
FAccT '25: Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, 2025, pp.1498-1542
ISBN
9798400714825
Publisher
ACM
Start Page
1498
End Page
1542
Journal / Book Title
FAccT '25: Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency
Copyright Statement
© 2025 Copyright held by the owner/author(s). This work is licensed under a Creative Commons Attribution 4.0 International License.
License URL
Identifier
10.1145/3715275.3732100
Source
FAccT '25: The 2025 ACM Conference on Fairness, Accountability, and Transparency
Subjects
conflicts
explainable AI
feature attributions
interpretability
outof-distribution
trustworthy AI
Publication Status
Published
Start Date
2025-06-23
Finish Date
2025-06-26
Coverage Spatial
Athens, Greece