Conference item icon

Conference item

Stabilising experience replay for deep multi-agent reinforcement learning

Abstract:
Many real-world problems, such as network packet routing and urban traffic control, are naturally modeled as multi-agent reinforcement learning (RL) problems. However, existing multi-agent RL methods typically scale poorly in the problem size. Therefore, a key challenge is to translate the success of deep learning on singleagent RL to the multi-agent setting. A key stumbling block is that independent Q-learning, the most popular multi-agent RL method, introduces nonstationarity that makes it incompatible with the experience replay memory on which deep RL relies. This paper proposes two methods that address this problem: 1) conditioning each agent’s value function on a footprint that disambiguates the age of the data sampled from the replay memory and 2) using a multi-agent variant of importance sampling to naturally decay obsolete data. Results on a challenging decentralised variant of StarCraft unit micromanagement confirm that these methods enable the successful combination of experience replay with multi-agent RL.
Publication status:
Published
Peer review status:
Peer reviewed

Actions

Access Document

Files:

Authors


Publisher:
PMLR
Host title:
Proceedings of the 34th International Conference on Machine Learning
Journal:
International Conference on Machine Learning More from this journal
Volume:
70
Pages:
1146-1155
Series:
Proceedings of Machine Learning Research
Publication date:
2017-07-29
Acceptance date:
2017-05-12


Pubs id:
pubs:695422
UUID:
uuid:2b650b3b-2fce-4875-b4df-70f4a4d64c8a
Local pid:
pubs:695422
Source identifiers:
695422
Deposit date:
2017-05-15
ARK identifier:

Terms of use


Views and Downloads

Views and downloads will return soon






If you are the owner of this record, you can report an update to it here: Report update to this record

TO TOP