Multi-modal information bottleneck attribution with cross-attention guidance
File(s) paper.pdf (1.65 MB)
Published version
Author(s)
Bourigault, Pauline
Bourigault, Emmanuelle
Mandic, Danilo
Type
Conference Paper
Abstract
For the progression of interpretable machine learning, particularly in the intersection of vision and language, ensuring transparency and comprehensibility in model decisions is crucial. This work introduces an enhancement to the Multi-modal Information Bottleneck attribution method by integrating cross-attention mechanisms. This targets
the core challenge of improving the interpretability of vision-language pretrained models, such as CLIP, by fostering more discerning and relevant latent representations. The
proposed method filters and retains essential information across modalities, leveraging cross-attention to dynamically focus on pertinent visual and textual features for any given
context. Through evaluations using CLIP as an example, we demonstrate improvements in attribution accuracy and interpretability over existing attribution methods, including
gradient-based, perturbation-based, attention-based, and information-theoretic methods. By providing a more nuanced understanding of model decisions, this work contributes to offer a promising avenue for deploying vision-language models in critical domains such as healthcare.
the core challenge of improving the interpretability of vision-language pretrained models, such as CLIP, by fostering more discerning and relevant latent representations. The
proposed method filters and retains essential information across modalities, leveraging cross-attention to dynamically focus on pertinent visual and textual features for any given
context. Through evaluations using CLIP as an example, we demonstrate improvements in attribution accuracy and interpretability over existing attribution methods, including
gradient-based, perturbation-based, attention-based, and information-theoretic methods. By providing a more nuanced understanding of model decisions, this work contributes to offer a promising avenue for deploying vision-language models in critical domains such as healthcare.
Date Issued
2024-11-24
Date Acceptance
2024-08-30
Citation
2024
Publisher
The British Machine Vision Association (BMVA)
Copyright Statement
© 2024. The copyright of this document resides with its authors. It may be distributed unchanged freely in print or electronic forms.
Identifier
https://bmva-archive.org.uk/bmvc/2024/papers/Paper_64/paper.pdf
Source
35th British Machine Vision Conference 2024
Publication Status
Published
Start Date
2024-11-24
Finish Date
2024-11-28
Coverage Spatial
Glasgow, United Kingdom
