Blocked or broken? automatically detecting when privacy interventions break websites
File(s)popets-2022-0096.pdf (433.51 KB)
Published version
Author(s)
Type
Journal Article
Abstract
A core problem in the development and maintenance of crowd-sourced filter
lists is that their maintainers cannot confidently predict whether (and where)
a new filter list rule will break websites. This is a result of enormity of the
Web, which prevents filter list authors from broadly understanding the impact
of a new blocking rule before they ship it to millions of users. The inability
of filter list authors to evaluate the Web compatibility impact of a new rule
before shipping it severely reduces the benefits of filter-list-based content
blocking: filter lists are both overly-conservative (i.e. rules are tailored
narrowly to reduce the risk of breaking things) and error-prone (i.e. blocking
tools still break large numbers of sites). To scale to the size and scope of
the Web, filter list authors need an automated system to detect when a new
filter rule breaks websites, before that breakage has a chance to make it to
end users.
In this work, we design and implement the first automated system for
predicting when a filter list rule breaks a website. We build a classifier,
trained on a dataset generated by a combination of compatibility data from the
EasyList project and novel browser instrumentation, and find it is accurate to
practical levels (AUC 0.88). Our open source system requires no human
interaction when assessing the compatibility risk of a proposed privacy
intervention. We also present the 40 page behaviors that most predict breakage
in observed websites.
lists is that their maintainers cannot confidently predict whether (and where)
a new filter list rule will break websites. This is a result of enormity of the
Web, which prevents filter list authors from broadly understanding the impact
of a new blocking rule before they ship it to millions of users. The inability
of filter list authors to evaluate the Web compatibility impact of a new rule
before shipping it severely reduces the benefits of filter-list-based content
blocking: filter lists are both overly-conservative (i.e. rules are tailored
narrowly to reduce the risk of breaking things) and error-prone (i.e. blocking
tools still break large numbers of sites). To scale to the size and scope of
the Web, filter list authors need an automated system to detect when a new
filter rule breaks websites, before that breakage has a chance to make it to
end users.
In this work, we design and implement the first automated system for
predicting when a filter list rule breaks a website. We build a classifier,
trained on a dataset generated by a combination of compatibility data from the
EasyList project and novel browser instrumentation, and find it is accurate to
practical levels (AUC 0.88). Our open source system requires no human
interaction when assessing the compatibility risk of a proposed privacy
intervention. We also present the 40 page behaviors that most predict breakage
in observed websites.
Date Issued
2022-03-07
Date Acceptance
2022-06-16
Citation
Proceedings on Privacy Enhancing Technologies, 2022, 2022 (4), pp.6-23
ISSN
2299-0984
Publisher
Sciendo
Start Page
6
End Page
23
Journal / Book Title
Proceedings on Privacy Enhancing Technologies
Volume
2022
Issue
4
Copyright Statement
© 2022 The Author(s). This article is published under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 license.
Identifier
http://arxiv.org/abs/2203.03528v2
Subjects
cs.CR
cs.CR
Publication Status
Published
Date Publish Online
2022-06-16