Skip to main navigation Skip to search Skip to main content

Reward hierarchical temporal memory: Model for memorizing and computing reward prediction error by neocortex

  • Hansol Choi
  • , Jun Cheol Park
  • , Jae Hyun Lim
  • , Jae Young Jun
  • , Dae Shik Kim

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In humans and animals, reward prediction error encoded by dopamine systems is thought to be important in the temporal difference learning class of reinforcement learning (RL). With RL algorithms, many brain models have described the function of dopamine and related areas, including the basal ganglia and frontal cortex. In spite of this importance, how the reward prediction error itself is computed is not understood well, including the problem of how the current states are assigned to a memorized states and how the values of the states are memorized. In this paper, we describe a neocortical model for memorizing state space and computing reward prediction error, known as reward hierarchical temporal memory (rHTM). In this model, the temporal relationships among events are hierarchically stored. Using this memory, rHTM computes reward prediction errors by associating the memorized sequences to rewards and inhibits the predicted reward. In a simulation, our model behaved similarly to dopaminergic neurons. We suggest that our model can provide a hypothetical framework of interaction between cortex and dopamine neurons.

Original languageEnglish
Title of host publication2012 International Joint Conference on Neural Networks, IJCNN 2012
DOIs
Publication statusPublished - 2012
Event2012 Annual International Joint Conference on Neural Networks, IJCNN 2012, Part of the 2012 IEEE World Congress on Computational Intelligence, WCCI 2012 - Brisbane, QLD, Australia
Duration: 10 Jun 201215 Jun 2012

Publication series

NameProceedings of the International Joint Conference on Neural Networks

Conference

Conference2012 Annual International Joint Conference on Neural Networks, IJCNN 2012, Part of the 2012 IEEE World Congress on Computational Intelligence, WCCI 2012
Country/TerritoryAustralia
CityBrisbane, QLD
Period10/06/1215/06/12

Keywords

  • HTM
  • rHTM
  • reinforcement learning
  • reward
  • reward prediction error
  • reward-HTM
  • temporal difference

Fingerprint

Dive into the research topics of 'Reward hierarchical temporal memory: Model for memorizing and computing reward prediction error by neocortex'. Together they form a unique fingerprint.

Cite this