Concrete dropout

Abstract:: Dropout is used as a practical tool to obtain uncertainty estimates in large vision models and reinforcement learning (RL) tasks. But to obtain well-calibrated uncertainty estimates, a grid-search over the dropout probabilities is necessary— a prohibitive operation with large models, and an impossible one with RL. We propose a new dropout variant which gives improved performance and better calibrated uncertainties. Relying on recent developments in Bayesian deep learning, we use a continuous relaxation of dropout’s discrete masks. Together with a principled optimisation objective, this allows for automatic tuning of the dropout probability in large models, and as a result faster experimentation cycles. In RL this allows the agent to adapt its uncertainty dynamically as more data is observed. We analyse the proposed variant extensively on a range of tasks, and give insights into common practice in the field where larger dropout probabilities are often used in deeper model layers.

Files:: Concrete dropout.pdf

(Preview, Accepted manuscript, pdf, 1.7MB, Terms of use)

Publisher:: NIPS Foundation
Host title:: Advances in Neural Information Processing Systems 31 (NIPS 2017)
Journal:: Advances in Neural Information Processing Systems 31 (NIPS 2017) More from this journal
Publication date:: 2018-07-01
Acceptance date:: 2017-09-04

Copyright holder:: Neural Information Processing Systems Foundation, Inc
Notes:: © 2018 Neural Information Processing Systems Foundation, Inc. This is the accepted manuscript version of the article. The final version is available online from Neural Information Processing Systems Foundation, Inc. at: https://papers.nips.cc/paper/6949-concrete-dropout

Licence:: Terms and Conditions of Use for Oxford University Research Archive

If you are the owner of this record, you can report an update to it here: Report update to this record

Conference item