Conference item
Distral: robust multitask reinforcement learning
- Abstract:
- Neural information processing systems foundation. All rights reserved. Most deep reinforcement learning algorithms are data inefficient in complex and rich environments, limiting their applicability to many scenarios. One direction for improving data efficiency is multitask learning with shared neural network parameters, where efficiency may be improved through transfer across related tasks. In practice, however, this is not usually observed, because gradients from different tasks can interfere negatively, making learning unstable and sometimes even less data efficient. Another issue is the different reward schemes between tasks, which can easily lead to one task dominating the learning of a shared model. We propose a new approach for joint training of multiple tasks, which we refer to as Distral (distill & transfer learning). Instead of sharing parameters between the different workers, we propose to share a "distilled" policy that captures common behaviour across tasks. Each worker is trained to solve its own task while constrained to stay close to the shared policy, while the shared policy is trained by distillation to be the centroid of all task policies. Both aspects of the learning process are derived by optimizing a joint objective function. We show that our approach supports efficient transfer on complex 3D environments, outperforming several related methods. Moreover, the proposed learning process is more robust to hyperparameter settings and more stable - attributes that are critical in deep reinforcement learning.
- Publication status:
- Published
- Peer review status:
- Reviewed (other)
Actions
Access Document
- Files:
-
-
(Preview, Version of record, pdf, 1.5MB, Terms of use)
-
- Publication website:
- https://papers.nips.cc/paper/7036-distral-robust-multitask-reinforcement-learning
Authors
- Publisher:
- Massachusetts Institute of Technology Press
- Host title:
- Advances in Neural Information Processing Systems 30 (NIPS 2017)
- Volume:
- 30
- Pages:
- 4497-4507
- Publication date:
- 2017-01-01
- Acceptance date:
- 2017-09-04
- Event title:
- 2017 Conference on Neural Information Processing Systems
- Event location:
- Long Beach, CA, USA
- Event website:
- https://nips.cc/Conferences/2017
- Event start date:
- 2017-12-04
- Event end date:
- 2017-12-09
- ISSN:
-
1049-5258
- Pubs id:
-
982776
- Local pid:
-
pubs:982776
- Deposit date:
-
2020-02-06
- ARK identifier:
Terms of use
- Copyright holder:
- Teh et al.
- Copyright date:
- 2017
- Rights statement:
- Copyright © 2017 by the authors and NIPS. This paper was presented at the 31st Annual Conference on Neural Information Processing Systems (NIPS 2017).
- Notes:
- This is the publisher's version of the item. The final version is available from MIT at: https://papers.nips.cc/paper/7036-distral-robust-multitask-reinforcement-learning.
If you are the owner of this record, you can report an update to it here: Report update to this record