Possible sources of bias in primary care electronic health record data (re-)use
File(s)04a4489377ef7704c72e90daa8d9ffb2.pdf (535.5 KB)
Published version
Author(s)
Robert, Verheij
Vasa, Curcin
Delaney, BC
Mark, McGilchrist
Type
Journal Article
Abstract
Background - Enormous amounts of data are recorded routinely in health care as part of the care process, primarily for managing individual patient care. There are significant opportunities to use this data for other purposes, many of which would contribute to establishing a learning health system. This is particularly true for data recorded in primary care settings, as in many countries, these are the first place patients turn to for most health problems.
Objective - In this paper, we discuss whether data that is recorded routinely as part of the health care process in primary care is actually fit to use for these other purposes, how the original purpose may affect the extent to which the data is fit for another purpose and the mechanisms behind these effects. In doing so, we want to identify possible sources of bias that are relevant for the (re-)use of this type of data.
Methods –This discussion paper is based on the authors’ experience as users of electronic health records data, as a general practitioner, health informatics experts, and health services researchers. It is a product of the discussions they had during the TRANSFoRm project, which was funded by the EU and sought to develop, pilot and evaluate a core information architecture for the Learning Health System (LHS) in Europe, based on primary care electronic health records.
Results – We first describe the different stages in the processing of EHR data, as well as the different purposes for which this data is used. Given the different data processing steps and purposes, we then discuss the possible mechanisms for each individual data processing step, that can generate biased outcomes. We identified thirteen possible sources of bias. Four of them are related to the organization of a health care system, some are of a more technical nature.
Conclusions - There are a substantial number of possible sources of bias, and very little is known about the size and direction of their impact. However, any (re-)user of data that was recorded as part of the health care process (such as researchers and clinicians) should be aware of the associated data collection process and environmental influences that can affect the quality of the data. Our stepwise, actor and purpose oriented approach may help to identify these possible sources of bias. Unless data quality issues are better understood and unless adequate controls are embedded throughout the data lifecycle, data-driven healthcare will not live up to its expectations. We need a data quality research agenda to devise the appropriate instruments needed to assess the magnitude of each of the possible sources of bias, and then start measuring their impact. The possible sources of bias described in this paper serve as a starting point for this research agenda.
Objective - In this paper, we discuss whether data that is recorded routinely as part of the health care process in primary care is actually fit to use for these other purposes, how the original purpose may affect the extent to which the data is fit for another purpose and the mechanisms behind these effects. In doing so, we want to identify possible sources of bias that are relevant for the (re-)use of this type of data.
Methods –This discussion paper is based on the authors’ experience as users of electronic health records data, as a general practitioner, health informatics experts, and health services researchers. It is a product of the discussions they had during the TRANSFoRm project, which was funded by the EU and sought to develop, pilot and evaluate a core information architecture for the Learning Health System (LHS) in Europe, based on primary care electronic health records.
Results – We first describe the different stages in the processing of EHR data, as well as the different purposes for which this data is used. Given the different data processing steps and purposes, we then discuss the possible mechanisms for each individual data processing step, that can generate biased outcomes. We identified thirteen possible sources of bias. Four of them are related to the organization of a health care system, some are of a more technical nature.
Conclusions - There are a substantial number of possible sources of bias, and very little is known about the size and direction of their impact. However, any (re-)user of data that was recorded as part of the health care process (such as researchers and clinicians) should be aware of the associated data collection process and environmental influences that can affect the quality of the data. Our stepwise, actor and purpose oriented approach may help to identify these possible sources of bias. Unless data quality issues are better understood and unless adequate controls are embedded throughout the data lifecycle, data-driven healthcare will not live up to its expectations. We need a data quality research agenda to devise the appropriate instruments needed to assess the magnitude of each of the possible sources of bias, and then start measuring their impact. The possible sources of bias described in this paper serve as a starting point for this research agenda.
Date Issued
2018-05-29
Date Acceptance
2018-03-07
Citation
Journal of Medical Internet Research, 2018, 20 (5)
ISSN
1438-8871
Publisher
JMIR Publications
Journal / Book Title
Journal of Medical Internet Research
Volume
20
Issue
5
Copyright Statement
©Robert
A Verheij, Vasa Curcin,
Brendan
C Delaney
, Mark M McGilchrist.
Originally
published
in the Journal
of Medical
Internet
Research
(http://www
.jmir.org), 29.05.2018.
This is an open-access
article distributed
under the terms of the Creative
Commons
Attribution
License
(https://creativecommons.or
g/licenses/by/4.0/),
which permits
unrestricted
use, distribution,
and
reproduction
in any medium,
provided
the original
work, first published
in the Journal
of Medical
Internet
Research,
is properly
cited. The complete
bibliographic
information,
a link to the original
publication
on http://www
.jmir.org/, as well as this copyright
and license
information
must be included.
A Verheij, Vasa Curcin,
Brendan
C Delaney
, Mark M McGilchrist.
Originally
published
in the Journal
of Medical
Internet
Research
(http://www
.jmir.org), 29.05.2018.
This is an open-access
article distributed
under the terms of the Creative
Commons
Attribution
License
(https://creativecommons.or
g/licenses/by/4.0/),
which permits
unrestricted
use, distribution,
and
reproduction
in any medium,
provided
the original
work, first published
in the Journal
of Medical
Internet
Research,
is properly
cited. The complete
bibliographic
information,
a link to the original
publication
on http://www
.jmir.org/, as well as this copyright
and license
information
must be included.
Subjects
Science & Technology
Life Sciences & Biomedicine
Health Care Sciences & Services
Medical Informatics
electronic health record
data accuracy
data sharing
health information interoperability
health care systems
health information systems
medical informatics
GENERAL-PRACTICE
MEDICAL-RECORD
CLINICAL-RESEARCH
SECONDARY USE
QUALITY
SYSTEM
ASSOCIATION
COHORT
RISK
ADOLESCENTS
08 Information And Computing Sciences
11 Medical And Health Sciences
17 Psychology And Cognitive Sciences
Publication Status
Published
Article Number
e185