Thesis icon

Thesis

Machine learning methods for tabular electronic health record data

Abstract:

The vast majority of the world’s medical data exists in the Electronic Health Record (EHR), a document used to record demographic characteristics and discrete events throughout a patient’s life. This EHR data is often poorly suited for machine learning, consisting of manually recorded data prone to errors and missing values, multiple time-series at highly varying levels of granularity, and typically high dimensionality and low patient numbers in datasets.

Nevertheless, machine learning on EHR data has a high potential for impact, and it is necessary to evaluate the efficacy of different machine learning methods for medical data. This thesis is an accumulation of theoretical and empirical work on developing machine learning methods appropriate for EHR data. Three work components are described.

The first work component of this thesis is a systematic evaluation of recurrent neural networks and transformer models for machine learning on EHR data. This thesis also introduces novel transformer-based architectures with properties that are desirable for EHR data. Transformers are identified as models with high performance for this task, achieving accuracy scores exceeding 94% in all evaluated tasks.

The second work component of this thesis is an effort to build machine models that can predict COVID-19 vaccine adverse events with a less than 0.1% occurrence rate in training data. An Area Under Receiver Operating Curve (AUROC) of 70.1% was achieved on this task, and feature and subgroup analysis to inform clinical understanding was performed.

The third work component is an effort to build machine learning models to predict stroke deterioration events from small neurointensive care EHR datasets. We evaluate a range of machine learning architectures, and we identify architectures that perform well both on training data and external validation data. We perform further feature and subgroup analysis to help place this work in the wider context of our clinical understanding of ischaemic stroke and malignant cerebral oedema. This thesis ultimately finds that a diverse range of models are needed for EHR data in different contexts - a ’one size fits all’ approach to machine learning is insufficient to address the diverse range of tasks, requirements and data formats within this setting.

Actions

Access Document

Files:

Authors

More by this author
Institution:
University of Oxford
Division:
MPLS
Department:
Engineering Science
Role:
Author

Contributors

Institution:
University of Oxford
Division:
MPLS
Department:
Engineering Science
Role:
Supervisor
Institution:
University of Oxford
Division:
MPLS
Department:
Engineering Science
Role:
Examiner
Role:
Examiner


More from this funder
Grant:
EP/Y035321/1
Programme:
Oxford EPSRC Centre for Doctoral Training in Healthcare Data Science


DOI:
Type of award:
DPhil
Level of award:
Doctoral
Awarding institution:
University of Oxford


Terms of use


Views and Downloads






If you are the owner of this record, you can report an update to it here: Report update to this record

TO TOP