Abstract
This paper studies discounted Markov Decision Processes (MDPs) with finite sets of states and actions. Value iteration is one of the major methods for finding optimal policies. For each discount factor, starting from a finite number of iterations, which is called the turnpike integer, value iteration algorithms always generate decision rules which are deterministic optimal policies for the infinite-horizon problems. This fact justifies the rolling horizon approach for computing infinite-horizon optimal policies by conducting a finite number of value iterations. This paper describes properties of turnpike integers and provides their upper bounds.
| Original language | English |
|---|---|
| Pages | 23-30 |
| Number of pages | 8 |
| DOIs | |
| State | Published - 2025 |
| Event | 2025 SIAM Conference on Control and Its Applications, CT 2025 - Montreal, Canada Duration: Jul 28 2025 → Jul 30 2025 |
Conference
| Conference | 2025 SIAM Conference on Control and Its Applications, CT 2025 |
|---|---|
| Country/Territory | Canada |
| City | Montreal |
| Period | 07/28/25 → 07/30/25 |
Fingerprint
Dive into the research topics of 'Properties of Turnpike Functions for Discounted Markov Decision Processes'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver