Skip to main navigation Skip to search Skip to main content

DOPL: DIRECT ONLINE PREFERENCE LEARNING FOR RESTLESS BANDITS WITH PREFERENCE FEEDBACK

  • Guojun Xiong
  • , Ujwal Dinesha
  • , Debajoy Mukherjee
  • , Jian Li
  • , Srinivas Shakkottai
  • Harvard University
  • Texas A&M University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Restless multi-armed bandits (RMAB) has been widely used to model constrained sequential decision making problems, where the state of each restless arm evolves according to a Markov chain and each state transition generates a scalar reward. However, the success of RMAB crucially relies on the availability and quality of reward signals. Unfortunately, specifying an exact reward function in practice can be challenging and even infeasible. In this paper, we introduce PREF-RMAB, a new RMAB model in the presence of preference signals, where the decision maker only observes pairwise preference feedback rather than scalar reward from the activated arms at each decision epoch. Preference feedback, however, arguably contains less information than the scalar reward, which makes PREF-RMAB seemingly more difficult. To address this challenge, we present a direct online preference learning (DOPL) algorithm for PREF-RMAB to efficiently explore the unknown environments, adaptively collect preference data in an online manner, and directly leverage the preference feedback for decision-makings. We prove that DOPL yields a sublinear regret. To our best knowledge, this is the first algorithm to ensure Õ(T ln T) regret for RMAB with preference feedback. Experimental results further demonstrate the effectiveness of DOPL.

Original languageEnglish
Title of host publication13th International Conference on Learning Representations, ICLR 2025
PublisherInternational Conference on Learning Representations, ICLR
Pages4071-4102
Number of pages32
ISBN (Electronic)9798331320850
StatePublished - 2025
Event13th International Conference on Learning Representations, ICLR 2025 - Singapore, Singapore
Duration: Apr 24 2025Apr 28 2025

Publication series

Name13th International Conference on Learning Representations, ICLR 2025

Conference

Conference13th International Conference on Learning Representations, ICLR 2025
Country/TerritorySingapore
CitySingapore
Period04/24/2504/28/25

Fingerprint

Dive into the research topics of 'DOPL: DIRECT ONLINE PREFERENCE LEARNING FOR RESTLESS BANDITS WITH PREFERENCE FEEDBACK'. Together they form a unique fingerprint.

Cite this