首页 | 本学科首页   官方微博 | 高级检索  
     


Spike-based Decision Learning of Nash Equilibria in Two-Player Games
Authors:Johannes Friedrich  Walter Senn
Affiliation:Department of Physiology and Center for Cognition, Learning and Memory, University of Bern, Switzerland;Indiana University, United States of America
Abstract:Humans and animals face decision tasks in an uncertain multi-agent environment where an agent''s strategy may change in time due to the co-adaptation of others strategies. The neuronal substrate and the computational algorithms underlying such adaptive decision making, however, is largely unknown. We propose a population coding model of spiking neurons with a policy gradient procedure that successfully acquires optimal strategies for classical game-theoretical tasks. The suggested population reinforcement learning reproduces data from human behavioral experiments for the blackjack and the inspector game. It performs optimally according to a pure (deterministic) and mixed (stochastic) Nash equilibrium, respectively. In contrast, temporal-difference(TD)-learning, covariance-learning, and basic reinforcement learning fail to perform optimally for the stochastic strategy. Spike-based population reinforcement learning, shown to follow the stochastic reward gradient, is therefore a viable candidate to explain automated decision learning of a Nash equilibrium in two-player games.
Keywords:
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号