Skip to main navigation Skip to search Skip to main content

Quantile-Based Policy Optimization for Reinforcement Learning

  • Peking University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

7 Scopus citations

Abstract

Classical reinforcement learning (RL) aims to optimize the expected cumulative rewards. In this work, we consider the RL setting where the goal is to optimize the quantile of the cumulative rewards. We parameterize the policy controlling actions by neural networks and propose a novel policy gradient algorithm called Quantile-Based Policy Optimization (QPO) and its variant Quantile-Based Proximal Policy Optimization (QPPO) to solve deep RL problems with quantile objectives. QPO uses two coupled iterations running at different time scales for simultaneously estimating quantiles and policy parameters. Our numerical results demonstrate that the proposed algorithms outperform the existing baseline algorithms under the quantile criterion.

Original languageEnglish
Title of host publicationProceedings of the 2022 Winter Simulation Conference, WSC 2022
EditorsB. Feng, G. Pedrielli, Y. Peng, S. Shashaani, E. Song, C.G. Corlu, L.H. Lee, E.P. Chew, T. Roeder, P. Lendermann
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages2712-2723
Number of pages12
ISBN (Electronic)9798350309713
DOIs
StatePublished - 2022
Event2022 Winter Simulation Conference, WSC 2022 - Guilin, China
Duration: Dec 11 2022Dec 14 2022

Publication series

NameProceedings - Winter Simulation Conference
Volume2022-December
ISSN (Print)0891-7736

Conference

Conference2022 Winter Simulation Conference, WSC 2022
Country/TerritoryChina
CityGuilin
Period12/11/2212/14/22

Fingerprint

Dive into the research topics of 'Quantile-Based Policy Optimization for Reinforcement Learning'. Together they form a unique fingerprint.

Cite this