Journal article icon

Journal article

Mitigating health inequities with machine learning: a nationwide cohort study developing and evaluating ethnicity-specific cardiovascular risk prediction models across 19 ethnically diverse populations in 2.5 million individuals with COVID-19

Abstract:
Background
Clinical risk stratification tools based on under-representative data risk exacerbating underlying systemic health inequalities. This study leverages a nationwide ethnically and socioeconomically diverse population for risk prediction of cardiovascular events after SARS-CoV-2 infection, by developing and evaluating prediction models driven by data from ethnicity-stratified populations.
Methods
This cohort study of 2,423,777 individuals uses six linked National Health Service datasets for England, UK. For each higher-level ethnicity population (Asian, Black, White, Mixed, Other, Unknown) and their 19 sub-populations, ethnicity-specific models were developed for the prediction of the risk of a cardiovascular event within one year of a SARS-CoV-2 infection. Ethnicity-specific risk prediction models were compared with a global all-ethnicity model trained on the overall study population. Models were developed using i) features used in existing tools (e.g. QRISK3), and data-driven features selected by ii) LASSO regression, iii) Random Forest, and iv) XGBoost, with corresponding SHAP analysis. Overall, 104 ethnicity-specific models using 51 candidate risk factors were developed and internally validated using discrimination (AUROC, F1, precision, recall), calibration, and net benefit analysis using a held-out test set.
Findings
In general, ethnicity-specific models performed similarly to the global all-ethnicity model for discrimination and net benefit, with some improvement shown in calibration. For African and Arab populations, ethnicity-specific models outperformed the global all-ethnicity model by an AUROC improvement of 2.04% (IQR:1.65%–2.25%) and 5.89% (IQR:1.95%–10.11%), respectively. Ethnicity-specific models identified additional risk factors, e.g. autoimmune liver disease in the Bangladeshi population, dementia in the African population and osteoporosis in the Gypsy and Irish Traveller population. Models using on feature selection by machine learning yielded an average AUROC improvement of 3.00% (IQR:1.10%-4.35%) compared to features informed by existing tool QRISK3.
Interpretation
Ethnicity-specific models based on granular data from actual practice settings have the potential to identify unique risk factors not necessarily captured in models based on dominant or majority populations, improving overall predictive performance. Our findings suggest that population-representative and ethnically diverse data can improve clinical risk stratification tailored to ethnically diverse populations and ultimately address health inequalities through targeted intervention and resource allocation.
Publication status:
Accepted
Peer review status:
Peer reviewed

Actions

Authors

More by this author
Institution:
University of Oxford
Division:
MSD
Department:
NDORMS
Sub department:
Centre for Statistics in Medicine
Role:
Author

Contributors


More from this funder
Funder identifier:
https://ror.org/02wdwnk04
Grant:
HDRUK2023.0234
SP/19/3/34678
More from this funder
Funder identifier:
https://ror.org/0187kwz08
Grant:
NIHR303137
NIHR203312
More from this funder
Funder identifier:
https://ror.org/04rtjaj74
Grant:
Disease-HDR-23012
HDRUK2023.0239
2021.0152


Publisher:
Elsevier
Journal:
Lancet Digital Health More from this journal
Acceptance date:
2026-07-31
EISSN:
2589-7500


Language:
English
Pubs id:
2452187
Local pid:
pubs:2452187
Deposit date:
2026-08-20
ARK identifier:


Views and Downloads






If you are the owner of this record, you can report an update to it here: Report update to this record

TO TOP