Repository logo
  • Log In
    Log in via Symplectic to deposit your publication(s).
Repository logo
  • Communities & Collections
  • Research Outputs
  • Statistics
  • Log In
    Log in via Symplectic to deposit your publication(s).
  1. Home
  2. Faculty of Medicine
  3. School of Public Health
  4. School of Public Health
  5. Using country-level variables to classify countries according to the number of confirmed COVID-19 cases: An unsupervised machine learning approach
 
  • Details
Using country-level variables to classify countries according to the number of confirmed COVID-19 cases: An unsupervised machine learning approach
File(s)
2021-Carrillo-Larco-Clusters of people with type 2 diabetes in the general population.pdf (637.61 KB)
Published version
Author(s)
Carrillo-Larco, R
Castillo-Cara, M
Type
Journal Article
Abstract
Background: The COVID-19 pandemic has attracted the attention of researchers and clinicians whom have provided evidence about risk factors and clinical outcomes. Research on the COVID-19 pandemic benefiting from open-access data and machine learning algorithms is still scarce yet can produce relevant and pragmatic information. With country-level pre-COVID-19-pandemic variables, we aimed to cluster countries in groups with shared profiles of the COVID-19 pandemic. Methods: Unsupervised machine learning algorithms (k-means) were used to define data-driven clusters of countries; the algorithm was informed by disease prevalence estimates, metrics of air pollution, socio-economic status and health system coverage. Using the one-way ANOVA test, we compared the clusters in terms of number of confirmed COVID-19 cases, number of deaths, case fatality rate and order in which the country reported the first case. Results: The model to define the clusters was developed with 155 countries. The model with three principal component analysis parameters and five or six clusters showed the best ability to group countries in relevant sets. There was strong evidence that the model with five or six clusters could stratify countries according to the number of confirmed COVID-19 cases (p<0.001). However, the model could not stratify countries in terms of number of deaths or case fatality rate. Conclusions : A simple data-driven approach using available global information before the COVID-19 pandemic, seemed able to classify countries in terms of the number of confirmed COVID-19 cases. The model was not able to stratify countries based on COVID-19 mortality data.
Date Issued
2020-06-15
URI
http://hdl.handle.net/10044/1/86035
DOI
10.12688/wellcomeopenres.15819.3
Copyright Statement
© Author(s) (or their employer(s)) 2021. Re-use
permitted under CC BY. Published by BMJ
License URL
http://creativecommons.org/licenses/by/4.0/
Sponsor
Wellcome Trust
Grant Number
214185/Z/18/Z
Publication Status
Published
About
Spiral Depositing with Spiral Publishing with Spiral Symplectic
Contact us
Open access team Report an issue
Other Services
Scholarly Communications Library Services
logo

Imperial College London

South Kensington Campus

London SW7 2AZ, UK

tel: +44 (0)20 7589 5111

Accessibility Modern slavery statement Cookie Policy

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Cookie settings
  • Privacy policy
  • End User Agreement
  • Send Feedback