Optimising data collection strategies in cyber security to address bias
File(s)
Author(s)
Highnam, Kate
Type
Thesis
Abstract
Cyber attacks on digital systems continue to increase in diversity, sophistication, and volume. Consequently, defensive measures must be adaptive, rapid, and automated. By generalising from readily available attack data, machine learning (ML) technologies have proven useful in automating the detection, identification, and characterisation of threats to reduce the adverse impact of attacks. The performance of ML depends on whether its training data contains important patterns it needs to learn or else “garbage in, garbage out.” This thesis seeks to improve the performance of ML models in cyber security by optimising data collection strategies, emphasising eliminating the potential impact of biases.
First, we craft and deploy fully instrumented cloud-based honeypots with a unique class of kernel-level sensors. We have published millions of host and network events with in-the-wild attacks. Second, to optimise honeypot deployments and any other intrusion data collection system, we provide a rigorous method to identify bias using causal graphs. We demonstrate how data collections that target bias improve ML model performance compared to only collections within the same environment. Finally, we developed an adaptive experimental design to reduce the cost of collecting additional data addressing bias. This novel method autonomously alters resource allocation based on the events seen during collection.
Our contributions provide a framework for security experts to better understand and update their datasets. The first contribution provides the data collection tool and initial dataset. The second guides experts to list actionable interventions: feasible changes to explore how an environment generates the data and remove bias that can confuse ML models as they learn. Our final contribution efficiently applies the first automated control trials, comparing the original “control” environment to a new “treated” environment where one listed intervention is applied. By adhering to our process, the updated datasets will encourage ML models to effectively defend digital systems.
First, we craft and deploy fully instrumented cloud-based honeypots with a unique class of kernel-level sensors. We have published millions of host and network events with in-the-wild attacks. Second, to optimise honeypot deployments and any other intrusion data collection system, we provide a rigorous method to identify bias using causal graphs. We demonstrate how data collections that target bias improve ML model performance compared to only collections within the same environment. Finally, we developed an adaptive experimental design to reduce the cost of collecting additional data addressing bias. This novel method autonomously alters resource allocation based on the events seen during collection.
Our contributions provide a framework for security experts to better understand and update their datasets. The first contribution provides the data collection tool and initial dataset. The second guides experts to list actionable interventions: feasible changes to explore how an environment generates the data and remove bias that can confuse ML models as they learn. Our final contribution efficiently applies the first automated control trials, comparing the original “control” environment to a new “treated” environment where one listed intervention is applied. By adhering to our process, the updated datasets will encourage ML models to effectively defend digital systems.
Version
Open Access
Date Issued
2024-05-27
Date Awarded
01/06/2025
License URL
Advisor
Jennings, Nicholas
Maffeis, Sergio
Publisher Department
Department of Electrical and Electronic Engineering
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
