Failure detection in medical image classification: a reality check and benchmarking testbed
File(s)186_failure_detection_in_medical_i.pdf (1.64 MB)
Published version
Author(s)
Bernhardt, Melanie
Ribeiro, Fabio De Sousa
Glocker, Ben
Type
Journal Article
Abstract
Failure detection in automated image classification is a critical safeguard for clinical deployment. Detected failure cases can be referred to human assessment, ensuring patient safety in computer-aided clinical decision making. Despite its paramount importance, there is insufficient evidence about the ability of state-of-the-art confidence scoring methods to detect test-time failures of classification models in the context of medical imaging. This paper provides a reality check, establishing the performance of in-domain misclassification detection methods, benchmarking 9 widely used confidence scores on 6 medical imaging datasets with different imaging modalities, in multiclass and binary classification settings. Our experiments show that the problem of failure detection is far from being solved. We found that none of the benchmarked advanced methods proposed in the computer vision and machine learning literature can consistently outperform a simple softmax baseline, demonstrating that improved out-of-distribution detection or model calibration do not necessarily translate to improved in-domain misclassification detection. Our developed testbed facilitates future work in this important area.
Date Issued
2022
Date Acceptance
2022-09-26
Citation
Transactions on Machine Learning Research, 2022, 2022
ISSN
2835-8856
Publisher
Transactions on Machine Learning Research
Journal / Book Title
Transactions on Machine Learning Research
Volume
2022
Copyright Statement
Creative Commons Attribution 4.0 International (CC BY 4.0)
License URL
Identifier
https://openreview.net/forum?id=VBHuLfnOMf
Publication Status
Published
Article Number
186
Date Publish Online
2022-10-23