Skip to main navigation Skip to search Skip to main content

Reinforcement Learning Augmented Asymptotically Optimal Index Policy for Finite-Horizon Restless Bandits

  • State University of New York Binghamton University
  • Indian Institute of Science

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

20 Scopus citations

Abstract

We study a finite-horizon restless multi-armed bandit problem with multiple actions, dubbed as R(MA)2B. The state of each arm evolves according to a controlled Markov decision process (MDP), and the reward of pulling an arm depends on both the current state and action of the corresponding MDP. Since finding the optimal policy is typically intractable, we propose a computationally appealing index policy entitled Occupancy-Measured-Reward Index Policy for the finite-horizon R(MA)2B. Our index policy is well-defined without the requirement of indexability condition and is provably asymptotically optimal. We then adopt a learning perspective where the system parameters are unknown, and propose R(MA)2B-UCB, a generative model based reinforcement learning augmented algorithm that can fully exploit the structure of Occupancy-Measured-Reward Index Policy. Compared to existing algorithms, R(MA)2B-UCB performs close to offline optimum, as well as achieves a sub-linear regret and a low computational complexity all at once. Experimental results show that R(MA)2B-UCB outperforms existing algorithms in both regret and running time.

Original languageEnglish
Title of host publicationAAAI-22 Technical Tracks 8
PublisherAssociation for the Advancement of Artificial Intelligence
Pages8726-8734
Number of pages9
ISBN (Electronic)1577358767, 9781577358763
DOIs
StatePublished - Jun 30 2022
Event36th AAAI Conference on Artificial Intelligence, AAAI 2022 - Virtual, Online
Duration: Feb 22 2022Mar 1 2022

Publication series

NameProceedings of the 36th AAAI Conference on Artificial Intelligence, AAAI 2022
Volume36

Conference

Conference36th AAAI Conference on Artificial Intelligence, AAAI 2022
CityVirtual, Online
Period02/22/2203/1/22

Fingerprint

Dive into the research topics of 'Reinforcement Learning Augmented Asymptotically Optimal Index Policy for Finite-Horizon Restless Bandits'. Together they form a unique fingerprint.

Cite this