The pitfalls of sample selection: a case study on lung nodule
classification
classification
File(s) 2108.05386v1.pdf (413.61 KB)
Accepted version
Author(s)
Baltatzis, Vasileios
Bintsi, Kyriaki-Margarita
Folgoc, Loic Le
Manzanera, Octavio E Martinez
Ellis, Sam
Type
Conference Paper
Abstract
Using publicly available data to determine the performance of methodological
contributions is important as it facilitates reproducibility and allows
scrutiny of the published results. In lung nodule classification, for example,
many works report results on the publicly available LIDC dataset. In theory,
this should allow a direct comparison of the performance of proposed methods
and assess the impact of individual contributions. When analyzing seven recent
works, however, we find that each employs a different data selection process,
leading to largely varying total number of samples and ratios between benign
and malignant cases. As each subset will have different characteristics with
varying difficulty for classification, a direct comparison between the proposed
methods is thus not always possible, nor fair. We study the particular effect
of truthing when aggregating labels from multiple experts. We show that
specific choices can have severe impact on the data distribution where it may
be possible to achieve superior performance on one sample distribution but not
on another. While we show that we can further improve on the state-of-the-art
on one sample selection, we also find that on a more challenging sample
selection, on the same database, the more advanced models underperform with
respect to very simple baseline methods, highlighting that the selected data
distribution may play an even more important role than the model architecture.
This raises concerns about the validity of claimed methodological
contributions. We believe the community should be aware of these pitfalls and
make recommendations on how these can be avoided in future work.
contributions is important as it facilitates reproducibility and allows
scrutiny of the published results. In lung nodule classification, for example,
many works report results on the publicly available LIDC dataset. In theory,
this should allow a direct comparison of the performance of proposed methods
and assess the impact of individual contributions. When analyzing seven recent
works, however, we find that each employs a different data selection process,
leading to largely varying total number of samples and ratios between benign
and malignant cases. As each subset will have different characteristics with
varying difficulty for classification, a direct comparison between the proposed
methods is thus not always possible, nor fair. We study the particular effect
of truthing when aggregating labels from multiple experts. We show that
specific choices can have severe impact on the data distribution where it may
be possible to achieve superior performance on one sample distribution but not
on another. While we show that we can further improve on the state-of-the-art
on one sample selection, we also find that on a more challenging sample
selection, on the same database, the more advanced models underperform with
respect to very simple baseline methods, highlighting that the selected data
distribution may play an even more important role than the model architecture.
This raises concerns about the validity of claimed methodological
contributions. We believe the community should be aware of these pitfalls and
make recommendations on how these can be avoided in future work.
Date Issued
2021-09-25
Date Acceptance
2021-08-01
Citation
2021, pp.201-211
Publisher
Springer
Start Page
201
End Page
211
Copyright Statement
© 2021 Springer Nature Switzerland AG. The final publication is available at Springer via https://doi.org/10.1007/978-3-030-87602-9_19
Sponsor
Engineering & Physical Science Research Council (E
Identifier
http://arxiv.org/abs/2108.05386v1
Grant Number
RTJ13261760-1
Source
Predictive Intelligence in Medicine at MICCAI
Subjects
cs.CV
cs.CV
Notes
Accepted at PRIME, MICCAI 2021
Publication Status
Published
Start Date
2021-09-27
Finish Date
2021-10-01
Coverage Spatial
Strasbourg, France
Date Publish Online
2021-09-25
