An alternative to variance: Gini deviation for risk-averse policy gradient

Luo, Y; Liu, G; Poupart, P; Pan, Y

AI Collection

Conference item

An alternative to variance: Gini deviation for risk-averse policy gradient

Abstract:: Restricting the variance of a policy's return is a popular choice in risk-averse Reinforcement Learning (RL) due to its clear mathematical definition and easy interpretability. Traditional methods directly restrict the total return variance. Recent methods restrict the per-step reward variance as a proxy. We thoroughly examine the limitations of these variance-based methods, such as sensitivity to numerical scale and hindering of policy learning, and propose to use an alternative risk measure, Gini deviation, as a substitute. We study various properties of this new risk measure and derive a policy gradient algorithm to minimize it. Empirical evaluation in domains where risk-aversion can be clearly defined, shows that our algorithm can mitigate the limitations of variance-based risk measures and achieves high return with low risk in terms of variance and Gini deviation when others fail to learn a reasonable policy.

Publication status:: Published

Peer review status:: Peer reviewed

Actions

Email

Email this record

Send the bibliographic details of this record to your email address.

Your Email
Please enter the email address that the record information will be sent to.

-
Your message (optional)
Please add any additional information to be included within the email.
Share
Cite

Cite this record

APA Style

Luo, Y., Liu, G., Poupart, P., & Pan, Y. (2024). An alternative to variance: Gini deviation for risk-averse policy gradient. 37th Conference on Neural Information Processing Systems (NeurIPS 2023), 36, 60922–60946.

MLA Style

Luo, Y, et al. “An Alternative to Variance: Gini Deviation for Risk-Averse Policy Gradient.” 37th Conference on Neural Information Processing Systems (NeurIPS 2023), vol. 36, 2024, pp. 60922–46.

Chicago Style

Luo, Y, G Liu, P Poupart, and Y Pan. 2024. “An Alternative to Variance: Gini Deviation for Risk-Averse Policy Gradient.” In 37th Conference on Neural Information Processing Systems (NeurIPS 2023), 36:60922–46. Curran Associates.
Print

Access Document

Files:: Luo_et_al_2023_An_alternative_to.pdf

(Preview, Accepted manuscript, pdf, 6.7MB, Terms of use)

Publisher copy:: 10.52202/075280-2662

Authors

+ Luo, Y More by this author

Role:: Author

+ Liu, G More by this author

Role:: Author

+ Poupart, P More by this author

Role:: Author

+ Pan, Y More by this author

Institution:: University of Oxford
Division:: MPLS
Department:: Engineering Science
Role:: Author
ORCID:: 0009-0000-8297-9045

+ Chinese University of Hong Kong More from this funder

Funder identifier:: https://ror.org/00t33hh48
Grant:: UDF01002911

Publisher:: Curran Associates
Host title:: Advances in Neural Information Processing Systems 36
Volume:: 36
Pages:: 60922-60946
Publication date:: 2024-07-01
Event title:: 37th Conference on Neural Information Processing Systems (NeurIPS 2023)
Event location:: New Orleans, Louisiana, USA
Event website:: https://neurips.cc/Conferences/2023
Event start date:: 2023-12-10
Event end date:: 2023-12-16
DOI:: 10.52202/075280-2662
ISSN:: 1049-5258
EISBN:: 9781713899921

Language:: English
Pubs id:: 1994747
Local pid:: pubs:1994747
Deposit date:: 2026-06-16
ARK identifier:: ark:/29072/ora_8f9dddfe7c0a4b02a1136593eeb64622

Terms of use

Copyright holder:: Luo et al and NeurIPS
Notes:: This is the accepted manuscript version of the article. The final version is available online from Curran Associates at https://dx.doi.org/10.52202/075280-2662

Licence:: Terms and Conditions of Use for Oxford University Research Archive

Views and Downloads

About views and downloads

If you are the owner of this record, you can report an update to it here: Report update to this record

Conference item

An alternative to variance: Gini deviation for risk-averse policy gradient

Actions

Access Document

Authors

Terms of use

Views and Downloads

Altmetrics

Dimensions

Conference item

An alternative to variance: Gini deviation for risk-averse policy gradient

Actions

Access Document

Authors

Funding

Bibliographic Details

Item Description

Terms of use

Metrics

Views and Downloads

Altmetrics

Dimensions