Thesis
Machine learning methods for tabular electronic health record data
- Abstract:
-
The vast majority of the world’s medical data exists in the Electronic Health Record (EHR), a document used to record demographic characteristics and discrete events throughout a patient’s life. This EHR data is often poorly suited for machine learning, consisting of manually recorded data prone to errors and missing values, multiple time-series at highly varying levels of granularity, and typically high dimensionality and low patient numbers in datasets.
Nevertheless, machine learning on EHR data has a high potential for impact, and it is necessary to evaluate the efficacy of different machine learning methods for medical data. This thesis is an accumulation of theoretical and empirical work on developing machine learning methods appropriate for EHR data. Three work components are described.
The first work component of this thesis is a systematic evaluation of recurrent neural networks and transformer models for machine learning on EHR data. This thesis also introduces novel transformer-based architectures with properties that are desirable for EHR data. Transformers are identified as models with high performance for this task, achieving accuracy scores exceeding 94% in all evaluated tasks.
The second work component of this thesis is an effort to build machine models that can predict COVID-19 vaccine adverse events with a less than 0.1% occurrence rate in training data. An Area Under Receiver Operating Curve (AUROC) of 70.1% was achieved on this task, and feature and subgroup analysis to inform clinical understanding was performed.
The third work component is an effort to build machine learning models to predict stroke deterioration events from small neurointensive care EHR datasets. We evaluate a range of machine learning architectures, and we identify architectures that perform well both on training data and external validation data. We perform further feature and subgroup analysis to help place this work in the wider context of our clinical understanding of ischaemic stroke and malignant cerebral oedema. This thesis ultimately finds that a diverse range of models are needed for EHR data in different contexts - a ’one size fits all’ approach to machine learning is insufficient to address the diverse range of tasks, requirements and data formats within this setting.
Actions
Access Document
- Files:
-
-
(Preview, Dissemination version, pdf, 13.4MB, Terms of use)
-
Authors
Contributors
- Institution:
- University of Oxford
- Division:
- MPLS
- Department:
- Engineering Science
- Role:
- Supervisor
- Institution:
- University of Oxford
- Division:
- MPLS
- Department:
- Engineering Science
- Role:
- Examiner
- Role:
- Examiner
- Grant:
- EP/Y035321/1
- Programme:
- Oxford EPSRC Centre for Doctoral Training in Healthcare Data Science
- DOI:
- Type of award:
- DPhil
- Level of award:
- Doctoral
- Awarding institution:
- University of Oxford
- Language:
-
English
- Keywords:
- Subjects:
- Deposit date:
-
2024-11-13
- ARK identifier:
Terms of use
- Copyright holder:
- O'Donoghue, O
- Copyright date:
- 2023
- Rights statement:
- © Odhran O'Donoghue 2023.
- Licence:
- CC Attribution (CC BY)
If you are the owner of this record, you can report an update to it here: Report update to this record