On the impact of sparsification on quantitative argumentative explanations in neural networks
File(s) paper_4.pdf (2.87 MB)
Published version
Author(s)
Peacock, Daniel
-, Mansi
Potyka, Nico
Toni, Francesca
Yin, Xiang
Type
Conference Paper
Abstract
Neural Networks (NNs) are powerful decision-making tools, but their lack of explainability limits their use in
high-stakes domains such as healthcare and criminal justice. The recent SpArX framework sparsifies NNs and
maps them to (weighted) Quantitative Bipolar Argumentation Frameworks (QBAFs) to provide an argumentative understanding of their mechanics. QBAFs can be explained by various quantitative argumentative explanation methods such as Argument Attribution Explanations (AAEs), Relation Attribution Explanations (RAEs), and Contestability Explanations (CEs) - which assign numerical scores to arguments or relations to quantify their influence on the dialectical strength of an argument to be explained. However, it remains unexplored how sparsification of NNs impacts the explanations derived from the corresponding (weighted) QBAFs. In this paper we explore two directions for impact. First, we empirically investigate how varying the sparsification levels of NNs affects the preservation of these explanations: using four datasets (Iris, Diabetes, Cancer, and COMPAS), we find that AAEs are generally well preserved, whereas RAEs are not. Then, for CEs, we find that sparsification can
improve computational efficiency in several cases. Overall, this study offers a preliminary investigation into the
potential synergy between sparsification and explanation methods, opening up new avenues for future research.
high-stakes domains such as healthcare and criminal justice. The recent SpArX framework sparsifies NNs and
maps them to (weighted) Quantitative Bipolar Argumentation Frameworks (QBAFs) to provide an argumentative understanding of their mechanics. QBAFs can be explained by various quantitative argumentative explanation methods such as Argument Attribution Explanations (AAEs), Relation Attribution Explanations (RAEs), and Contestability Explanations (CEs) - which assign numerical scores to arguments or relations to quantify their influence on the dialectical strength of an argument to be explained. However, it remains unexplored how sparsification of NNs impacts the explanations derived from the corresponding (weighted) QBAFs. In this paper we explore two directions for impact. First, we empirically investigate how varying the sparsification levels of NNs affects the preservation of these explanations: using four datasets (Iris, Diabetes, Cancer, and COMPAS), we find that AAEs are generally well preserved, whereas RAEs are not. Then, for CEs, we find that sparsification can
improve computational efficiency in several cases. Overall, this study offers a preliminary investigation into the
potential synergy between sparsification and explanation methods, opening up new avenues for future research.
Date Issued
2025-10-16
Date Acceptance
2025-09-03
Citation
CEUR Workshop Proceedings, 2025, pp.20-35
ISSN
1613-0073
Publisher
CEUR Workshop Proceedings
Start Page
20
End Page
35
Journal / Book Title
CEUR Workshop Proceedings
Copyright Statement
© 2025 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0)
License URL
Source
3rd International Workshop on Argumentation for eXplainable AI (ArgXAI@ECAI)
Subjects
Explainability
Neural Networks
Argumentative Explanations Sepal Length Activation: 0.47 Layer 1 Neuron 1 Activation: 0.83 Layer 1 Neuron 2 Activation: 0.00 Sepal Width Activation: 0.32 Petal Length Activation: 0.72 Petal Width Activation: 0.63 Iris-setosa Activation: 0.00 Iris-versicolor Activation: 0.99 Iris-virginica Activation: 1.00 Layer 2 Neuron 1 Activation: 0.02 Layer 2 Neuron 2 Activation: 0.00 Layer 2 Neuron 3 Activation: 1.00
Publication Status
Published
Start Date
2025-10-26
Coverage Spatial
Bologna, Italy
