Comparison of machine learning methods for estimating case fatality ratios: an Ebola outbreak simulation study
File(s)S1 Appendix. Supplementary Information.docx (466.21 KB) journal.pone.0257005.pdf (3.24 MB)
Supporting information
Published version
Author(s)
Forna, Alpha
Dorigatti, Ilaria
Nouvellet, Pierre
Donnelly, Christl
Type
Journal Article
Abstract
Background
Machine learning (ML) algorithms are now increasingly used in infectious disease epidemiology. Epidemiologists should understand how ML algorithms behave within the context of outbreak data where missingness of data is almost ubiquitous.
Methods
Using simulated data, we use a ML algorithmic framework to evaluate data imputation performance and the resulting case fatality ratio (CFR) estimates, focusing on the scale and type of data missingness (i.e., missing completely at random—MCAR, missing at random—MAR, or missing not at random—MNAR).
Results
Across ML methods, dataset sizes and proportions of training data used, the area under the receiver operating characteristic curve decreased by 7% (median, range: 1%–16%) when missingness was increased from 10% to 40%. Overall reduction in CFR bias for MAR across methods, proportion of missingness, outbreak size and proportion of training data was 0.5% (median, range: 0%–11%).
Conclusion
ML methods could reduce bias and increase the precision in CFR estimates at low levels of missingness. However, no method is robust to high percentages of missingness. Thus, a datacentric approach is recommended in outbreak settings—patient survival outcome data should be prioritised for collection and random-sample follow-ups should be implemented to ascertain missing outcomes.
Machine learning (ML) algorithms are now increasingly used in infectious disease epidemiology. Epidemiologists should understand how ML algorithms behave within the context of outbreak data where missingness of data is almost ubiquitous.
Methods
Using simulated data, we use a ML algorithmic framework to evaluate data imputation performance and the resulting case fatality ratio (CFR) estimates, focusing on the scale and type of data missingness (i.e., missing completely at random—MCAR, missing at random—MAR, or missing not at random—MNAR).
Results
Across ML methods, dataset sizes and proportions of training data used, the area under the receiver operating characteristic curve decreased by 7% (median, range: 1%–16%) when missingness was increased from 10% to 40%. Overall reduction in CFR bias for MAR across methods, proportion of missingness, outbreak size and proportion of training data was 0.5% (median, range: 0%–11%).
Conclusion
ML methods could reduce bias and increase the precision in CFR estimates at low levels of missingness. However, no method is robust to high percentages of missingness. Thus, a datacentric approach is recommended in outbreak settings—patient survival outcome data should be prioritised for collection and random-sample follow-ups should be implemented to ascertain missing outcomes.
Date Issued
2021-09-15
Date Acceptance
2021-09-03
Citation
PLoS One, 2021, 16 (9)
ISSN
1932-6203
Publisher
Public Library of Science (PLoS)
Journal / Book Title
PLoS One
Volume
16
Issue
9
Copyright Statement
© 2021 Forna et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
License URL
Sponsor
Medical Research Council (MRC)
Wellcome Trust
Grant Number
MR/R015600/1
213494/Z/18/Z
Subjects
General Science & Technology
Publication Status
Published
Article Number
ARTN e0257005