Determining the clinical applicability of machine learning models through assessment of reporting across skin phototypes and rarer skin cancer types: a systematic review
Author(s)
Type
Journal Article
Abstract
Machine learning (ML) models for skin cancer recognition may have variable performance across different skin phototypes and skin cancer types. Overall performance metrics alone are insufficient to detect poor subgroup performance. We aimed (1) to assess whether studies of ML models reported results separately for different skin phototypes and rarer skin cancers, and (2) to graphically represent the skin cancer training datasets used by current ML models. In this systematic review, we searched PubMed, Embase and CENTRAL. We included all studies in medical journals assessing an ML technique for skin cancer diagnosis that used clinical or dermoscopic images from 1 January 2012 to 22 September 2021. No language restrictions were applied. We considered rarer skin cancers to be skin cancers other than pigmented melanoma, basal cell carcinoma and squamous cell carcinoma. We identified 114 studies for inclusion. Rarer skin cancers were included by 8/114 studies (7.0%), and results for a rarer skin cancer were reported separately in 1/114 studies (0.9%). Performance was reported across all skin phototypes in 1/114 studies (0.9%), but performance was uncertain in skin phototypes I and VI from minimal representation of the skin phototypes in the test dataset (9/3756 and 1/3756, respectively). For training datasets, although public datasets were most frequently used, with the most widely used being the International Skin Imaging Collaboration (ISIC) archive (65/114 studies, 57.0%), the largest datasets were private. Our review identified that most ML models did not report performance separately for rarer skin cancers and different skin phototypes. A degree of variability in ML model performance across subgroups is expected, but the current lack of transparency is not justifiable and risks models being used inappropriately in populations in whom accuracy is low.
Date Issued
2023-04
Date Acceptance
2022-11-09
Citation
Journal of the European Academy of Dermatology and Venereology, 2023, 37 (4), pp.657-665
ISSN
0926-9959
Publisher
Wiley
Start Page
657
End Page
665
Journal / Book Title
Journal of the European Academy of Dermatology and Venereology
Volume
37
Issue
4
Copyright Statement
© 2022 The Authors. Journal of the European Academy of Dermatology and Venereology published by John Wiley & Sons Ltd on behalf of European Academy of Dermatology and Venereology.This is an open access article under the terms of the Creative Commons Attribution-NonCommercial-NoDerivs License, which permits use and distribution in any medium,
provided the original work is properly cited, the use is non-commercial and no modifications or adaptations are made
provided the original work is properly cited, the use is non-commercial and no modifications or adaptations are made
Identifier
https://www.webofscience.com/api/gateway?GWVersion=2&SrcApp=PARTNER_APP&SrcAuth=LinksAMR&KeyUT=WOS:000905595000001&DestLinkType=FullRecord&DestApp=ALL_WOS&UsrCustomerID=a2bf6146997ec60c407a63945d4e92bb
Subjects
AI
ARTIFICIAL-INTELLIGENCE
CONVOLUTIONAL NEURAL-NETWORK
Dermatology
Life Sciences & Biomedicine
PERFORMANCE
Science & Technology
Publication Status
Published
Date Publish Online
2022-12-14