TY - GEN
T1 - Reinforcement control via action dependent heuristic dynamic programming
AU - Tang, K. W.
AU - Srikant, G.
PY - 1997
Y1 - 1997
N2 - Heuristic dynamic programming (HDP) is the simplest kind of adaptive critic which is a powerful form of reinforcement control. It can be used to maximize or minimize any utility function, such as total energy or trajectory error, of a system over time in a noisy environment. Unlike supervised learning, adaptive critic design does not require the desired control signals be known. Instead, feedback is obtained based on a critic network which learns the relationship between a set of control signals and the corresponding strategic utility function. It is an approximation of dynamic programming. Action-dependent heuristic dynamic programming (ADHDP) system involves two subnetworks, the action network and the critic network. Each of these networks includes a feedforward and a feedback component. A flow chart for the interaction of these components is included. To further illustrate the algorithm, we use ADHDP for the control of a simple, 2D planar robot.
AB - Heuristic dynamic programming (HDP) is the simplest kind of adaptive critic which is a powerful form of reinforcement control. It can be used to maximize or minimize any utility function, such as total energy or trajectory error, of a system over time in a noisy environment. Unlike supervised learning, adaptive critic design does not require the desired control signals be known. Instead, feedback is obtained based on a critic network which learns the relationship between a set of control signals and the corresponding strategic utility function. It is an approximation of dynamic programming. Action-dependent heuristic dynamic programming (ADHDP) system involves two subnetworks, the action network and the critic network. Each of these networks includes a feedforward and a feedback component. A flow chart for the interaction of these components is included. To further illustrate the algorithm, we use ADHDP for the control of a simple, 2D planar robot.
UR - https://www.scopus.com/pages/publications/0030682703
U2 - 10.1109/ICNN.1997.614163
DO - 10.1109/ICNN.1997.614163
M3 - Conference contribution
AN - SCOPUS:0030682703
SN - 0780341228
SN - 9780780341227
T3 - IEEE International Conference on Neural Networks - Conference Proceedings
SP - 1766
EP - 1770
BT - 1997 IEEE International Conference on Neural Networks, ICNN 1997
T2 - 1997 IEEE International Conference on Neural Networks, ICNN 1997
Y2 - 9 June 1997 through 12 June 1997
ER -