Towards safer deep learning: verification, robustness, and explainability
File(s)
Author(s)
Batten, Benjamin James
Type
Thesis
Abstract
Deep neural networks have achieved remarkable success across diverse applications, from autonomous vehicles to medical diagnostics. However, their deployment in safety-critical systems is hindered by vulnerabilities to adversarial attacks, lack of robustness guarantees, and opaque decision-making. This work addresses these fundamental challenges through contributions towards verification, robustness, and explainability of neural networks in both deterministic and probabilistic settings.
We first tackle the problem of verifying neural networks against geometric perturbations, developing a novel piecewise-linear approximation method that approximates non-convex spatial transformations more precisely than existing approaches. This technique produces verification bounds up to 30% tighter than the benchmark methods, enabling formal safety guarantees for 32% more cases on vision-based baselines.
Secondly, we explore the use adversarial training in improving out-of-distribution generalisation in LiDAR-based 3D object detection. By combining a physics-inspired initialisation scheme with adversarial training, we achieve up to 6% improvement in detection accuracy under adverse weather conditions compared to simulation-based methods.
Further, we extend verification techniques to probabilistic settings by developing novel algorithms for Bayesian Neural Networks (BNNs). Our Pure Iterative Expansion and Gradient-guided Iterative Expansion methods dynamically adapt verification regions based on the parameter posterior space, yielding probabilistic robustness certificates up to 40% tighter than previous approaches, while also eliminating the need for extensive hyperparameter tuning.
Finally, we introduce the first formal framework for generating counterfactual explanations on BNNs, leveraging inherent uncertainty quantification to produce explanations with superior plausibility and robustness. Our evaluation demonstrates that BNN-based counterfactuals consistently outperform deterministic and ensemble-based alternatives across multiple metrics and datasets.
These contributions collectively advance the state of the art in neural network safety, providing both theoretical foundations and practical tools for building more dependable AI systems.
We first tackle the problem of verifying neural networks against geometric perturbations, developing a novel piecewise-linear approximation method that approximates non-convex spatial transformations more precisely than existing approaches. This technique produces verification bounds up to 30% tighter than the benchmark methods, enabling formal safety guarantees for 32% more cases on vision-based baselines.
Secondly, we explore the use adversarial training in improving out-of-distribution generalisation in LiDAR-based 3D object detection. By combining a physics-inspired initialisation scheme with adversarial training, we achieve up to 6% improvement in detection accuracy under adverse weather conditions compared to simulation-based methods.
Further, we extend verification techniques to probabilistic settings by developing novel algorithms for Bayesian Neural Networks (BNNs). Our Pure Iterative Expansion and Gradient-guided Iterative Expansion methods dynamically adapt verification regions based on the parameter posterior space, yielding probabilistic robustness certificates up to 40% tighter than previous approaches, while also eliminating the need for extensive hyperparameter tuning.
Finally, we introduce the first formal framework for generating counterfactual explanations on BNNs, leveraging inherent uncertainty quantification to produce explanations with superior plausibility and robustness. Our evaluation demonstrates that BNN-based counterfactuals consistently outperform deterministic and ensemble-based alternatives across multiple metrics and datasets.
These contributions collectively advance the state of the art in neural network safety, providing both theoretical foundations and practical tools for building more dependable AI systems.
Version
Open Access
Date Issued
2025-07-16
Date Awarded
01/02/2026
License URL
Advisor
Lomuscio, Alessio
Sponsor
UK Research and Innovation
Publisher Department
Department of Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
