Journal article
Mitigating health inequities with machine learning: a nationwide cohort study developing and evaluating ethnicity-specific cardiovascular risk prediction models across 19 ethnically diverse populations in 2.5 million individuals with COVID-19
- Abstract:
-
Background
Clinical risk stratification tools based on under-representative data risk exacerbating underlying systemic health inequalities. This study leverages a nationwide ethnically and socioeconomically diverse population for risk prediction of cardiovascular events after SARS-CoV-2 infection, by developing and evaluating prediction models driven by data from ethnicity-stratified populations.
Methods
This cohort study of 2,423,777 individuals uses six linked National Health Service datasets for England, UK. For each higher-level ethnicity population (Asian, Black, White, Mixed, Other, Unknown) and their 19 sub-populations, ethnicity-specific models were developed for the prediction of the risk of a cardiovascular event within one year of a SARS-CoV-2 infection. Ethnicity-specific risk prediction models were compared with a global all-ethnicity model trained on the overall study population. Models were developed using i) features used in existing tools (e.g. QRISK3), and data-driven features selected by ii) LASSO regression, iii) Random Forest, and iv) XGBoost, with corresponding SHAP analysis. Overall, 104 ethnicity-specific models using 51 candidate risk factors were developed and internally validated using discrimination (AUROC, F1, precision, recall), calibration, and net benefit analysis using a held-out test set.
Findings
In general, ethnicity-specific models performed similarly to the global all-ethnicity model for discrimination and net benefit, with some improvement shown in calibration. For African and Arab populations, ethnicity-specific models outperformed the global all-ethnicity model by an AUROC improvement of 2.04% (IQR:1.65%–2.25%) and 5.89% (IQR:1.95%–10.11%), respectively. Ethnicity-specific models identified additional risk factors, e.g. autoimmune liver disease in the Bangladeshi population, dementia in the African population and osteoporosis in the Gypsy and Irish Traveller population. Models using on feature selection by machine learning yielded an average AUROC improvement of 3.00% (IQR:1.10%-4.35%) compared to features informed by existing tool QRISK3.
Interpretation
Ethnicity-specific models based on granular data from actual practice settings have the potential to identify unique risk factors not necessarily captured in models based on dominant or majority populations, improving overall predictive performance. Our findings suggest that population-representative and ethnically diverse data can improve clinical risk stratification tailored to ethnically diverse populations and ultimately address health inequalities through targeted intervention and resource allocation.
- Publication status:
- Accepted
- Peer review status:
- Peer reviewed
Actions
Authors
+ British Heart Foundation
More from this funder
- Funder identifier:
- https://ror.org/02wdwnk04
- Grant:
- HDRUK2023.0234
- SP/19/3/34678
+ National Institute for Health and Care Research
More from this funder
- Funder identifier:
- https://ror.org/0187kwz08
- Grant:
- NIHR303137
- NIHR203312
+ Health Data Research UK
More from this funder
- Funder identifier:
- https://ror.org/04rtjaj74
- Grant:
- Disease-HDR-23012
- HDRUK2023.0239
- 2021.0152
- Publisher:
- Elsevier
- Journal:
- Lancet Digital Health More from this journal
- Acceptance date:
- 2026-07-31
- EISSN:
-
2589-7500
- Language:
-
English
- Pubs id:
-
2452187
- Local pid:
-
pubs:2452187
- Deposit date:
-
2026-08-20
- ARK identifier:
If you are the owner of this record, you can report an update to it here: Report update to this record