Revisiting IoT device identification
File(s)2107.07818v1.pdf (314.82 KB)
Accepted version
Author(s)
Type
Conference Paper
Abstract
Internet-of-Things (IoT) devices are known to be the source of many security
problems, and as such, they would greatly benefit from automated management.
This requires robustly identifying devices so that appropriate network security
policies can be applied. We address this challenge by exploring how to
accurately identify IoT devices based on their network behavior, while
leveraging approaches previously proposed by other researchers.
We compare the accuracy of four different previously proposed machine
learning models (tree-based and neural network-based) for identifying IoT
devices. We use packet trace data collected over a period of six months from a
large IoT test-bed. We show that, while all models achieve high accuracy when
evaluated on the same dataset as they were trained on, their accuracy degrades
over time, when evaluated on data collected outside the training set. We show
that on average the models' accuracy degrades after a couple of weeks by up to
40 percentage points (on average between 12 and 21 percentage points). We argue
that, in order to keep the models' accuracy at a high level, these need to be
continuously updated.
problems, and as such, they would greatly benefit from automated management.
This requires robustly identifying devices so that appropriate network security
policies can be applied. We address this challenge by exploring how to
accurately identify IoT devices based on their network behavior, while
leveraging approaches previously proposed by other researchers.
We compare the accuracy of four different previously proposed machine
learning models (tree-based and neural network-based) for identifying IoT
devices. We use packet trace data collected over a period of six months from a
large IoT test-bed. We show that, while all models achieve high accuracy when
evaluated on the same dataset as they were trained on, their accuracy degrades
over time, when evaluated on data collected outside the training set. We show
that on average the models' accuracy degrades after a couple of weeks by up to
40 percentage points (on average between 12 and 21 percentage points). We argue
that, in order to keep the models' accuracy at a high level, these need to be
continuously updated.
Date Issued
2021-09-14
Date Acceptance
2021-09-01
Citation
2021, pp.1-9
Publisher
IFIP
Start Page
1
End Page
9
Copyright Statement
© 2021 Crown
Sponsor
Engineering & Physical Science Research Council (E
Engineering & Physical Science Research Council (EPSRC)
Engineering & Physical Science Research Council (EPSRC)
Engineering & Physical Science Research Council (E
Engineering & Physical Science Research Council (E
Engineering & Physical Science Research Council (E
Identifier
http://arxiv.org/abs/2107.07818v1
Grant Number
EP/R511547/1
EP/N028260/2
EP/R0222091/1
RGS128099 (EP/R03351X/1)
PO: 20232790 (Ref: 301671)
EP/V502354/1
Source
Network Traffic Measurement and Analysis Conference 2021
Subjects
cs.CR
cs.CR
cs.LG
Notes
To appear in TMA 2021 conference. 9 pages, 6 figures. arXiv admin note: text overlap with arXiv:2011.08605
Publication Status
Published
Start Date
2021-09-14
Finish Date
2021-09-15
Coverage Spatial
Virtual
Date Publish Online
2021-09-14