Journal article
A standardized analytics pipeline for reliable and rapid development and validation of prediction models using observational health data
- Abstract:
- Background and objective As a response to the ongoing COVID-19 pandemic, several prediction models in the existing literature were rapidly developed, with the aim of providing evidence-based guidance. However, none of these COVID-19 prediction models have been found to be reliable. Models are commonly assessed to have a risk of bias, often due to insufficient reporting, use of non-representative data, and lack of large-scale external validation. In this paper, we present the Observational Health Data Sciences and Informatics (OHDSI) analytics pipeline for patient-level prediction modeling as a standardized approach for rapid yet reliable development and validation of prediction models. We demonstrate how our analytics pipeline and open-source software tools can be used to answer important prediction questions while limiting potential causes of bias (e.g., by validating phenotypes, specifying the target population, performing large-scale external validation, and publicly providing all analytical source code). Methods We show step-by-step how to implement the analytics pipeline for the question: ‘In patients hospitalized with COVID-19, what is the risk of death 0 to 30 days after hospitalization?’. We develop models using six different machine learning methods in a USA claims database containing over 20,000 COVID-19 hospitalizations and externally validate the models using data containing over 45,000 COVID-19 hospitalizations from South Korea, Spain, and the USA. Results Our open-source software tools enabled us to efficiently go end-to-end from problem design to reliable Model Development and evaluation. When predicting death in patients hospitalized with COVID-19, AdaBoost, random forest, gradient boosting machine, and decision tree yielded similar or lower internal and external validation discrimination performance compared to L1-regularized logistic regression, whereas the MLP neural network consistently resulted in lower discrimination. L1-regularized logistic regression models were well calibrated. Conclusion Our results show that following the OHDSI analytics pipeline for patient-level prediction modelling can enable the rapid development towards reliable prediction models. The OHDSI software tools and pipeline are open source and available to researchers from all around the world.
- Publication status:
- Published
- Peer review status:
- Peer reviewed
Actions
Access Document
- Files:
-
-
(Preview, Version of record, pdf, 3.0MB, Terms of use)
-
- Publisher copy:
- 10.1016/j.cmpb.2021.106394
Authors
- Publisher:
- Elsevier
- Journal:
- Computer Methods and Programs in Biomedicine More from this journal
- Volume:
- 211
- Article number:
- 106394
- Publication date:
- 2021-09-06
- Acceptance date:
- 2021-08-30
- DOI:
- EISSN:
-
1872-7565
- ISSN:
-
0169-2607
- Pmid:
-
34560604
- Language:
-
English
- Keywords:
- Pubs id:
-
1196044
- Local pid:
-
pubs:1196044
- Deposit date:
-
2021-11-16
- ARK identifier:
Terms of use
- Copyright holder:
- Khalid et al.
- Copyright date:
- 2021
- Rights statement:
- © 2021 The Author(s). Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
If you are the owner of this record, you can report an update to it here: Report update to this record