Thesis
Planning with learned ignorance-aware models
- Abstract:
-
One of the goals of artificial intelligence research is to create decision-makers (i.e., agents) that improve from experience (i.e., data), collected through interaction with an environment. Models of the environment (i.e., world models) are an explicit way that agents use to represent their knowledge, enabling them to make counterfactual predictions and plans without requiring additional environment interactions. Although agents that plan with a perfect model of the environment have led to impressive demonstrations, e.g., super- human performance in board games, they are limited to problems their designer can specify a perfect model. Therefore, learning models from experience holds the promise of going beyond the scope of their designers’ reach, giving rise to a self-improving vicious circle of (i) learning a model from the past experience; (ii) planning with the learned model; and (iii) interacting with the environment, collecting new experiences. Ideally, learned models should generalise to situations beyond their training regime. Nonetheless, this is ambitious and often unrealistic when finite data is used for learning the models, leading to generally imperfect models, with which naive planning could be catastrophic in novel, out-of-training distribution situations. A more pragmatic goal is to have agents that are aware of and quantify their lack of knowledge (i.e., ignorance or epistemic uncertainty).
In this thesis, we motivate and demonstrate the effectiveness of and propose novel ignorance-aware agents that plan with learned models. Naively applying powerful planning algorithms to learned models can render negative results, when the planning algorithm exploits the model imperfections in out-of-training distribution situations. This phenomenon is often termed overoptimisation and can be addressed by optimising ignorance-augmented objectives, called knowledge equivalents. We verify the validity of our ideas and methods in a number of problem settings, including learning from (i) expert demonstrations (imitation learning, §3); (ii) sub-optimal demonstrations (social learning, §4); and (iii) interacting with an environment with rewards (reinforcement learning, §5). Our empirical evidence is based on simulated autonomous driving environments, continuous control and video games from pixels and didactic small-scale grid-worlds. Throughout the thesis, we use neural networks to parameterise the (learnable) models and either use existing scalable approximate ignorance quantification deep learning methods, such as ensembles, or introduce novel planning-specific ways to quantify the agents’ ignorance.
The main chapters of this thesis are based on publications (Filos et al., 2020, 2021, 2022).
Actions
Access Document
- Files:
-
-
(Preview, Dissemination version, pdf, 22.7MB, Terms of use)
-
Authors
Contributors
- Institution:
- University of Oxford
- Division:
- MPLS
- Department:
- Computer Science
- Role:
- Supervisor
- ORCID:
- 0000-0002-2733-2078
- Role:
- Examiner
- Institution:
- University of Oxford
- Division:
- MPLS
- Department:
- Engineering Science
- Role:
- Examiner
- ORCID:
- 0000-0001-9688-2498
- Funding agency for:
- Filos, A
- Programme:
- 2021 J.P. Morgan PhD Fellowship
- DOI:
- Type of award:
- DPhil
- Level of award:
- Doctoral
- Awarding institution:
- University of Oxford
- Language:
-
English
- Keywords:
- Subjects:
- Deposit date:
-
2024-05-10
- ARK identifier:
Terms of use
- Copyright holder:
- FIlos, A
- Copyright date:
- 2022
If you are the owner of this record, you can report an update to it here: Report update to this record