Conference item icon

Conference item

Identifiability in inverse reinforcement learning

Abstract:
Inverse reinforcement learning attempts to reconstruct the reward function in a Markov decision problem, using observations of agent actions. As already observed in Russell [1998] the problem is ill-posed, and the reward function is not identifiable, even under the presence of perfect information about optimal behavior. We provide a resolution to this non-identifiability for problems with entropy regularization. For a given environment, we fully characterize the reward functions leading to a given policy and demonstrate that, given demonstrations of actions for the same reward under two distinct discount factors, or under sufficiently different environments, the unobserved reward can be recovered up to a constant. We also give general necessary and sufficient conditions for reconstruction of time-homogeneous rewards on finite horizons, and for action-independent rewards, generalizing recent results of Kim et al. [2021] and Fu et al. [2018].
Publication status:
Published
Peer review status:
Peer reviewed

Actions

Access Document

Authors

More by this author
Institution:
University of Oxford
Division:
MPLS
Department:
Mathematical Institute
Oxford college:
New College
Role:
Author
ORCID:
0000-0003-0539-6414


Publisher:
Neural Information Processing Systems Foundation
Host title:
Advances in Neural Information Processing Systems 34 (NeurIPS 2021)
Publication date:
2021-12-14
Acceptance date:
2021-09-09
Event title:
35th Conference on Neural Information Processing Systems (NeurIPS 2021)
Event location:
Virtual event
Event website:
https://nips.cc/Conferences/2021/
Event start date:
2021-12-06
Event end date:
2021-12-14


Language:
English
Keywords:
Pubs id:
1183058
Local pid:
pubs:1183058
Deposit date:
2022-01-21
ARK identifier:

Terms of use


Views and Downloads






If you are the owner of this record, you can report an update to it here: Report update to this record

TO TOP