Journal article
Can Researchers Assess the Suitability of Datasets to Answer Their Research Questions, with Access to Metadata Only?
- Abstract:
- ObjectiveThe ability to reproduce the work of others is an essential part of the scientific disciplines. Replicating observational studies using electronic health record (EHR) data can be challenging due to complexities in data access, variations in EHR systems across institutions, and the potential for unaccounted confounding variables. Our aim is to identify the barriers to methods reproducibility for replication studies using EHR data.MethodsWe replicated a study that examined the risk of hospitalisation following a positive COVID-19 test in individuals with diabetes. Using EHR data from the NHS England’s Secure Data Environment (SDE) covering the whole of England, UK (population 57m), we sought to replicate findings from the original study, which used data from Greater Manchester (a large urban region in the UK, population 2.9m). Both analyses were conducted in Trusted Research Environments (TREs) or SDEs, containing linked primary and secondary care data, however methods reproducibility was not straightforward. Differences between the environments that contributed to the difficulties were documented, categorized into themes, and converted into a list of recommendations for TRE/SDEs.ResultsSmall differences between the environments and the data sources led to several challenges in methods reproducibility. Our recommendations of TRE/SDEs should facilitate future replication studies. The recommendations include: a need for improved machine-readable metadata for EHR data; standardization of governance processes to facilitate federated analysis; mandating of code sharing; and for environments to have a support structure for data engineers and analysts. We also propose a new theme for research, “data reproducibility”, as the ability to prepare, extract and clean data from a different database for a replication study.ConclusionEven with perfect code sharing, data reproducibility remains a challenge. Our recommendations have the potential to reduce the barriers to replication studies and therefore enhance the potential of observational studies using EHR data
- Publication status:
- Published
- Peer review status:
- Peer reviewed
Actions
Access Document
- Files:
-
-
(Version of record, html, 23.0KB, Terms of use)
-
- Publisher copy:
- 10.3233/shti220033
Authors
- Publisher:
- IOS Press
- Journal:
- Studies in Health Technology and Informatics More from this journal
- Volume:
- 290
- Pages:
- 66-70
- Publication date:
- 2022-06-06
- DOI:
- ISSN:
-
0926-9630
- Language:
-
English
- Keywords:
- Pubs id:
-
1262814
- UUID:
-
uuid_578164ce-a393-42a0-98e2-3075e2d13aa9
- Local pid:
-
pubs:1262814
- Source identifiers:
-
W4281745944
- Deposit date:
-
2026-01-30
- ARK identifier:
This ORA record was generated from metadata provided by an external service. It has not been edited by the ORA Team.
Terms of use
- Copyright date:
- 2022
If you are the owner of this record, you can report an update to it here: Report update to this record