What is the suitability of clinical vignettes in benchmarking the performance of online symptom checkers? An audit study
Author(s)
Type
Working Paper
Abstract
Objective To assess the suitability of primary care vignettes in benchmarking the performance of online symptom checkers
Design Observational study using publicly available, free online symptom checkers
Participants Three symptom checkers (Healthily, Ada and Babylon) that provided consultations in English. 139 standardized patient vignettes were compiled by RCGP. Three independent GPs interpreted the vignettes to arrive at a “Gold Standard” consisting of 3 dispositions and divided into one of three categories of triage urgency: (1) emergency care required, (2) primary care required and (3) self-care.
Main outcome measures Six professional non-medical and lay inputters simulated 2774 standardized patient evaluations using 3 online symptom checkers (OSC). We recorded when OSC provided a triage recommendation and whether it correctly recommended the appropriate triage recommendation across three categories of triage urgency (emergency care, primary care or self-care). We collected data on whether the solution appeared within the first 3 dispositions in each of the standards across 2774 standardized patient evaluations.
Results When benchmarked against the Gold Standard, Healthily provided an appropriate triage recommendation 61.9% of the time compared to 45.3% and 42.4% of the time for Babylon and Ada respectively. There was poor agreement between OSC consultation outcome and Gold Standard dispositions. When compared to the Gold Standard, Healthily gave an unsafe “under-triage” recommendation 28.6% of the time overall across the three categories compared to 43.3% for Ada and 47.5% for Babylon (P<0.001).
Conclusions OSCs recommended ‘very unsafe’ triages only <4% of the time suggesting that the online consultation tools are generally working at a safe level of risk. Primary care vignettes are a helpful tool to support development of OSC, but not ideally suited to benchmark the performance of different OSC. Real-world evidence studies involving general practice are recommended to benchmark the performance of OSC in the community setting.
Strengths and limitations of this study
139 independently created primary care vignettes covering 18 subcategories of primary care were used to benchmark the performance of three online symptom checkers using 2774 unique patient simulations
A gold standard for each primary care vignette was derived using GP roundtables and single blinded testing
We investigated the extent that different inputters using the same vignette and online symptom checker received differing consultation outcomes and triage recommendations
We developed an accuracy matrix to objectively monitor online symptom checker consultation outcome and the safety of the triage recommendation
Limitations included a different number of inputters to simulate patients across the three online symptom checkers tested
Competing Interest Statement
The authors have declared no competing interest.
Funding Statement
Unconditional funding for this work was provided by Healthily (Imperial Self- Care 2020/1). The Funder did not have a role in study design or analysis. Austen El-Osta and Azeem Majeed are supported by the National Institute for Health Research (NIHR) Applied Research Collaboration (ARC) North West London. The views expressed are those of the authors and not necessarily those of the NHS, the NIHR or the Department of Health and Social Care
Design Observational study using publicly available, free online symptom checkers
Participants Three symptom checkers (Healthily, Ada and Babylon) that provided consultations in English. 139 standardized patient vignettes were compiled by RCGP. Three independent GPs interpreted the vignettes to arrive at a “Gold Standard” consisting of 3 dispositions and divided into one of three categories of triage urgency: (1) emergency care required, (2) primary care required and (3) self-care.
Main outcome measures Six professional non-medical and lay inputters simulated 2774 standardized patient evaluations using 3 online symptom checkers (OSC). We recorded when OSC provided a triage recommendation and whether it correctly recommended the appropriate triage recommendation across three categories of triage urgency (emergency care, primary care or self-care). We collected data on whether the solution appeared within the first 3 dispositions in each of the standards across 2774 standardized patient evaluations.
Results When benchmarked against the Gold Standard, Healthily provided an appropriate triage recommendation 61.9% of the time compared to 45.3% and 42.4% of the time for Babylon and Ada respectively. There was poor agreement between OSC consultation outcome and Gold Standard dispositions. When compared to the Gold Standard, Healthily gave an unsafe “under-triage” recommendation 28.6% of the time overall across the three categories compared to 43.3% for Ada and 47.5% for Babylon (P<0.001).
Conclusions OSCs recommended ‘very unsafe’ triages only <4% of the time suggesting that the online consultation tools are generally working at a safe level of risk. Primary care vignettes are a helpful tool to support development of OSC, but not ideally suited to benchmark the performance of different OSC. Real-world evidence studies involving general practice are recommended to benchmark the performance of OSC in the community setting.
Strengths and limitations of this study
139 independently created primary care vignettes covering 18 subcategories of primary care were used to benchmark the performance of three online symptom checkers using 2774 unique patient simulations
A gold standard for each primary care vignette was derived using GP roundtables and single blinded testing
We investigated the extent that different inputters using the same vignette and online symptom checker received differing consultation outcomes and triage recommendations
We developed an accuracy matrix to objectively monitor online symptom checker consultation outcome and the safety of the triage recommendation
Limitations included a different number of inputters to simulate patients across the three online symptom checkers tested
Competing Interest Statement
The authors have declared no competing interest.
Funding Statement
Unconditional funding for this work was provided by Healthily (Imperial Self- Care 2020/1). The Funder did not have a role in study design or analysis. Austen El-Osta and Azeem Majeed are supported by the National Institute for Health Research (NIHR) Applied Research Collaboration (ARC) North West London. The views expressed are those of the authors and not necessarily those of the NHS, the NIHR or the Department of Health and Social Care
Date Issued
2021-07-31
Citation
2021
ISSN
2044-6055
Publisher
Cold Spring Harbor Laboratory
Copyright Statement
© 2021 The Author(s). It is made available under a CC-BY-NC-ND 4.0 International license .
Identifier
https://www.medrxiv.org/content/10.1101/2021.07.29.21261320v1
Subjects
1103 Clinical Sciences
1117 Public Health and Health Services
1199 Other Medical and Health Sciences
Publication Status
Published