Certified robustness to data poisoning in gradient-based training
File(s) 4848_Certified_Robustness_to_D.pdf (3.89 MB)
Published version
OA Location
Author(s)
Sosnin, P
Müller, MN
Baader, M
Tsay, C
Wicker, M
Type
Journal Article
Abstract
Modern machine learning pipelines leverage large amounts of public data, making it infeasible to guarantee data quality and leaving models open to poisoning and backdoor attacks. Provably bounding the behavior of learning algorithms under such attacks remains an open problem. In this work, we address this challenge by developing the first framework providing provable guarantees on the behavior of models trained with potentially manipulated data without modifying the model or learning algorithm. In particular, our framework certifies robustness against untargeted and targeted poisoning, as well as backdoor attacks, for bounded and unbounded manipulations of the training inputs and labels. Our method leverages convex relaxations to over-approximate the set of all possible parameter updates for a given poisoning threat model, allowing us to bound the set of all reachable parameters for any gradient-based learning algorithm. Given this set of parameters, we provide bounds on worst-case behavior, including model performance and backdoor success rate. We demonstrate our approach on multiple real-world datasets from applications including energy consumption, medical imaging, and autonomous driving.
Date Issued
2025-09-03
Date Acceptance
2025-08-04
Citation
Transactions on Machine Learning Research, 2025, 2025-September
Publisher
Journal of Machine Learning Research Inc.
Journal / Book Title
Transactions on Machine Learning Research
Volume
2025-September
Copyright Statement
Copyright © 2025 The Author(s). This work is licensed under a Creative Commons Attribution 4.0 International License (https://creativecommons.org/licenses/by/4.0/).
License URL
Publication Status
Published
Article Number
4848
Date Publish Online
2025-09-03
