TY - GEN
T1 - Optimality conditions for total-cost partially observable Markov decision processes
AU - Feinberg, Eugene A.
AU - Kasyanov, Pavlo O.
AU - Zgurovsky, Michael Z.
PY - 2013
Y1 - 2013
N2 - This note describes sufficient conditions for the existence of optimal policies for Partially Observable Markov Decision Processes (POMDPs). The objective criterion is either minimization of total discounted costs or minimization of total nonnegative costs. It is well-known that a POMDP can be reduced to a Completely Observable Markov Decision Process (COMDP) with the state space being the sets of believe probabilities for the POMDP. Thus, a policy is optimal in POMDP if and only if it corresponds to an optimal policy in the COMDP. Here we provide sufficient conditions for the existence of optimal policies for COMDP and therefore for POMDP. In particular, we consider POMDPs with weakly continuous transition probabilities and bounded below K-infcompact cost functions. For a fully observable MDPs these two conditions guarantee the following three properties: (i) validity of finite-horizon and infinite-horizon optimality equations, (ii) convergence of value iterations to infinite-horizon value functions, (iii) existence of stationary optimal policies. We show that the single additional assumption, that the observation transition probability is continuous in the total variation, implies properties (i)-(iii) for the COMDP. Therefore, this condition also implies the existence of optimal policies for POMDPs. We also provide a more general and less constructive sufficient condition for the validity of (i)- (iii) for the COMDP and therefore for the existence of optimal policies for a POMDP and the possibility of finding them by transforming optimal policies for the corresponding COMDP.
AB - This note describes sufficient conditions for the existence of optimal policies for Partially Observable Markov Decision Processes (POMDPs). The objective criterion is either minimization of total discounted costs or minimization of total nonnegative costs. It is well-known that a POMDP can be reduced to a Completely Observable Markov Decision Process (COMDP) with the state space being the sets of believe probabilities for the POMDP. Thus, a policy is optimal in POMDP if and only if it corresponds to an optimal policy in the COMDP. Here we provide sufficient conditions for the existence of optimal policies for COMDP and therefore for POMDP. In particular, we consider POMDPs with weakly continuous transition probabilities and bounded below K-infcompact cost functions. For a fully observable MDPs these two conditions guarantee the following three properties: (i) validity of finite-horizon and infinite-horizon optimality equations, (ii) convergence of value iterations to infinite-horizon value functions, (iii) existence of stationary optimal policies. We show that the single additional assumption, that the observation transition probability is continuous in the total variation, implies properties (i)-(iii) for the COMDP. Therefore, this condition also implies the existence of optimal policies for POMDPs. We also provide a more general and less constructive sufficient condition for the validity of (i)- (iii) for the COMDP and therefore for the existence of optimal policies for a POMDP and the possibility of finding them by transforming optimal policies for the corresponding COMDP.
UR - https://www.scopus.com/pages/publications/84902310573
U2 - 10.1109/CDC.2013.6760790
DO - 10.1109/CDC.2013.6760790
M3 - Conference contribution
AN - SCOPUS:84902310573
SN - 9781467357173
T3 - Proceedings of the IEEE Conference on Decision and Control
SP - 5716
EP - 5721
BT - 2013 IEEE 52nd Annual Conference on Decision and Control, CDC 2013
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 52nd IEEE Conference on Decision and Control, CDC 2013
Y2 - 10 December 2013 through 13 December 2013
ER -