Thesis
On architectures and training techniques for neural networks
- Abstract:
-
The progress in deep learning over the past decade has led to the creation of models that have had a significant impact on various fields as widely ranging as biology and language. Powering this progress are methodological advances in the design of neural networks and their training procedures, alongside the availability of more data and better software and hardware.
This thesis presents four pieces of research devoted to understanding and improving the generalization of neural networks via the lenses of neural architectures and their training techniques. On the architecture design front, we introduce a new equivariant architecture based on attention, and we develop algorithms for automatically designing architectures for use in an ensemble. On the training techniques front, we propose a denoising technique for self-supervised pre-training of molecular property prediction models, and we study an optimization technique for improving generalization via periodic re-initialization of the neural network during training. Each work touches on various aspects of the overall workflow for training a neural network.
Actions
Access Document
- Files:
-
-
(Preview, Dissemination version, pdf, 8.3MB, Terms of use)
-
Authors
Contributors
- Institution:
- University of Oxford
- Division:
- MPLS
- Department:
- Statistics
- Role:
- Supervisor
- Funder identifier:
- https://ror.org/00p64v472
- Funding agency for:
- Zaidi, S
- Programme:
- Aker Scholarship
- Funder identifier:
- https://ror.org/024bc3e07
- Funding agency for:
- Zaidi, S
- Programme:
- Google PhD Fellowship
- DOI:
- Type of award:
- DPhil
- Level of award:
- Doctoral
- Awarding institution:
- University of Oxford
- Language:
-
English
- Keywords:
- Subjects:
- Deposit date:
-
2025-05-12
- ARK identifier:
Terms of use
- Copyright holder:
- Sheheryar Zaidi
- Copyright date:
- 2023
If you are the owner of this record, you can report an update to it here: Report update to this record