Skip to main navigation Skip to search Skip to main content

Reinforcement control via action dependent heuristic dynamic programming

  • Stony Brook University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Heuristic dynamic programming (HDP) is the simplest kind of adaptive critic which is a powerful form of reinforcement control. It can be used to maximize or minimize any utility function, such as total energy or trajectory error, of a system over time in a noisy environment. Unlike supervised learning, adaptive critic design does not require the desired control signals be known. Instead, feedback is obtained based on a critic network which learns the relationship between a set of control signals and the corresponding strategic utility function. It is an approximation of dynamic programming. Action-dependent heuristic dynamic programming (ADHDP) system involves two subnetworks, the action network and the critic network. Each of these networks includes a feedforward and a feedback component. A flow chart for the interaction of these components is included. To further illustrate the algorithm, we use ADHDP for the control of a simple, 2D planar robot.

Original languageEnglish
Title of host publication1997 IEEE International Conference on Neural Networks, ICNN 1997
Pages1766-1770
Number of pages5
DOIs
StatePublished - 1997
Event1997 IEEE International Conference on Neural Networks, ICNN 1997 - Houston, TX, United States
Duration: Jun 9 1997Jun 12 1997

Publication series

NameIEEE International Conference on Neural Networks - Conference Proceedings
Volume3
ISSN (Print)1098-7576

Conference

Conference1997 IEEE International Conference on Neural Networks, ICNN 1997
Country/TerritoryUnited States
CityHouston, TX
Period06/9/9706/12/97

Fingerprint

Dive into the research topics of 'Reinforcement control via action dependent heuristic dynamic programming'. Together they form a unique fingerprint.

Cite this