Perceptual hashing-based client-side scanning: a flawed solution for known illegal content detection
File(s)
Author(s)
Jain, Shubham
Type
Thesis
Abstract
End-to-end encryption (E2EE) provides strong technical protection to individuals from interference. However, governments and law enforcement agencies worldwide have raised concerns that E2EE also allows illegal content to be shared undetected. Following the global pushback against key-escrow systems, tech companies, governments, and researchers have recently proposed perceptual hashing-based client-side scanning (PH-CSS) to detect illegal content in E2EE communications. PH-CSS has raised strong privacy concerns from the security and privacy community. However, the proponents of the technology have argued that the risk is limited as the technology has a single scope: detecting known illegal content. We propose two independent threat models against the PH-CSS system.
First, we propose a framework to evaluate the robustness of perceptual hashing-based client-side scanning to detect known illegal content against adversaries aiming to avoid detection by slightly modifying an image. In a large-scale evaluation, we show perceptual hashing-based client-side scanning mechanisms to be highly vulnerable to detection avoidance attacks in a black-box setting, with more than 99.9% of images successfully attacked while preserving the image's content.
Second, modern perceptual hashing algorithms are fairly flexible technology. An adversary could use this flexibility to add a secondary hidden feature to a client-side scanning system. More specifically, we show that an adversary providing the PH algorithm can ''hide" a secondary purpose of facial recognition of a target individual alongside its primary purpose of detecting known illegal content.
Our results show that perceptual hashing-based client-side scanning mechanisms are not robust in detecting known illegal content. They also carry the potential to be built with ''hidden" secondary purposes that could turn billions of user devices into tools to locate targeted individuals.
First, we propose a framework to evaluate the robustness of perceptual hashing-based client-side scanning to detect known illegal content against adversaries aiming to avoid detection by slightly modifying an image. In a large-scale evaluation, we show perceptual hashing-based client-side scanning mechanisms to be highly vulnerable to detection avoidance attacks in a black-box setting, with more than 99.9% of images successfully attacked while preserving the image's content.
Second, modern perceptual hashing algorithms are fairly flexible technology. An adversary could use this flexibility to add a secondary hidden feature to a client-side scanning system. More specifically, we show that an adversary providing the PH algorithm can ''hide" a secondary purpose of facial recognition of a target individual alongside its primary purpose of detecting known illegal content.
Our results show that perceptual hashing-based client-side scanning mechanisms are not robust in detecting known illegal content. They also carry the potential to be built with ''hidden" secondary purposes that could turn billions of user devices into tools to locate targeted individuals.
Version
Open Access
Date Issued
2023-09
Date Awarded
2024-02
Copyright Statement
Creative Commons Attribution NonCommercial Licence
License URL
Advisor
de Montjoye, Yves-Alexandre
Publisher Department
Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)