基于组合神经网络的Sarsa(λ)学习算法

Sarsa (λ) learning algorithm based on stacked neural network

下载PDF

导出

摘要标准的Sarsa(λ)算法对状态空间的要求是离散的且空间较小,而实际问题中很多系统的状态空间是连续的或尽管是离散的但空间较大,这就需要很大的内存来存储状态动作对。为此提出组合神经网络,首先用自组织映射(SOM)神经网络对状态空间进行自适应量化,然后在此基础上用BP网络拟合Q函数。该方法实现了Sarsa(λ)算法在连续和大规模状态空间的泛化。最后,实验结果表明了该方法的有效性。 The standard Sarsa （λ） algorithm requires that the state space is discrete and small. However, in real word it does not satisfy that due to the fact that it may be continuous or discrete but has large state space, so it needs large number of memory to save state-action pairs. Therefore, a stacked neural network is proposed which uses a self-organizing map （SOM） neural network to quantize adaptively the space of states, then a BP network is used to approximate a Q-function. The method realizes the generalization of Sarsa （λ） algorithm in the continuous or large-scale states. The experiment shows the validity of the proposed algorithm.

作者殷苌茗付超红薛丽华李立云

机构地区长沙理工大学计算机与通信工程学院

出处《计算机工程与设计》 CSCD 北大核心 2008年第22期5817-5819,5823,共4页 Computer Engineering and Design

关键词组合神经网络强化学习自组织映射 BP网络 Sarsa算法 stacked neural network reinforcement learning self-organizing maps back propagation network Sarsa algorithm

分类号 TP181 [自动化与计算机技术—控制理论与控制工程]

引文网络
相关文献

参考文献12

1Sutton R S,Barto A G.Reinforcement learning[M].MA:The MIT Press,1998.
2Kaelbling L P, Littman M L, Moore A W. Reinforcement leaming:A survey[J].Journal of Artificial Intelligence Research, 1996,4(2):237-285.
3Sutton R S.Learning to predit by the method of temporal differences[J].Machine Learing, 1988(3):9-44.
4Watkins CJCH,Dayan P.Q-learning[J].Machine Learning,1992 (8):279-292.
5Rummery G A,Niranjan M.On-line Q-learning using connectionist systems[R].Cambridge University Engineering Department, 1994.
6Singh S P, Sutton R S.Reinforcement learning with replacing eligibility traces[J].Machine Learning, 1996,22:123-158.
7Peng J, Williams R. Incremental multi-step Q-learning [J]. Machine Learning, 1996,22(4):283-290.
8Andrew James Smith.Applications of the self-organising map to reinforcement learning [J]. Neural Networks, 2002,15 (8-9): 1107-1124.
9林联明,王浩,王一雄.基于神经网络的Sarsa强化学习算法[J].计算机技术与发展,2006,16(1):30-32. 被引量：4
10Kohonen T.Self-organizing maps[C].Springer Series in Information Sciences.New York,USA:Springer,2001.

二级参考文献9

1阎平凡.再励学习——原理、算法及其在智能控制中的应用[J].信息与控制,1996,25(1):28-34. 被引量：30
2Astom K J. Optimal control of Markov derision processes with incomplete state estimation[J ]. Math'Anal Appl, 1998,10:174 - 205.
3Tsitsiklis J N, Roy B V. An Analysis of Temporal-Difference Learning with Function Approximation[J]. IEEE Transactions on Automatic Control, 1997,42 (5) : 674 - 690.
4Tesauro G J. TD-gammon, a self- teaching backgammon program[J]. Neural Computation, 1994, 6(2) :215 - 2192.
5Suton R S, Learning to predict by the methods of temporal diferences[J]. Machine Learning, 1988(3): 9 - 44.
6Suton R S,Barto A G. Reinforcement Learning: Introduction[M].Cambridge,MA:MIT Press,1998.
7Thrum Sebastian ,Mitcheil Tom M.Lifelong robot leaning[J].Robotics and Autonomous System.1995,15:25～46
8Ben J.A.Krose,Joris W.Mvan Dam.Adaptive state space quantisition,for reiforcement learning of collide free navigation[J].1922 IEEE/RSJ Internation Conference on Intelligent Robots and System.Rakeigh,NC.July 7～10 ,1992:1327～1332
9Watking,J.C.Hand Dayan Peter.Q-leaming[J].Machine Learning.1992,8:279～292

共引文献6

1单志超,林春生,向前.舰船水压场信号的小波能谱估计与支持向量机联合检测[J].船海工程,2008,37(3):135-138.
2李磊,孙卉,翟秋敏,郭志永.RBF神经网络在平顶山市地表水评价中的应用[J].安徽农业科学,2008,36(26):11514-11516. 被引量：2
3刘燕燕,张少白.关于DIVA模型中语速对语音生成影响的研究[J].计算机技术与发展,2011,21(12):33-35.
4尤树华,周谊成,王辉.基于神经网络的强化学习研究概述[J].电脑知识与技术,2012,8(10):6782-6786. 被引量：4
5陈晓辉,张银银,付云霞,雷帮军.自组织映射节点定位算法中邻域函数的优化方法研究[J].小型微型计算机系统,2017,38(2):213-216. 被引量：3
6刘思嘉,童向荣.基于强化学习的城市交通路径规划[J].计算机应用,2021,41(1):185-190. 被引量：8

1薛丽华,殷苌茗,李立云,胡明辉.基于多智能体的融合Sarsa(λ)学习算法[J].计算机工程与应用,2008,44(4):182-183. 被引量：2
2柴旭清,孙丽娜.基于量子粒子群和SARSA算法的蜂窝网络信道分配[J].计算机测量与控制,2015,23(10):3555-3557. 被引量：4
3陈卫东,关永贞,朱奇光,赵成龙.移动机器人模糊Sarsa(λ)学习导航研究[J].小型微型计算机系统,2013,34(11):2599-2602.
4刘云龙,吉国力.基于CMAC网络Sarsa(λ)学习的RoboCup守门员策略[J].北京工业大学学报,2012,38(9):1348-1352.
5林联明,王浩,王一雄.基于神经网络的Sarsa强化学习算法[J].计算机技术与发展,2006,16(1):30-32. 被引量：4
6李春贵,阳树洪,王萌,张增芳.基于SARSA(λ)算法的单路口交通信号学习控制[J].广西工学院学报,2008,19(2):10-14. 被引量：3
7常峰,贺元骅.基于强化学习和蚁群算法的WSN节点故障诊断[J].计算机测量与控制,2015,23(3):755-758. 被引量：1
8王志勃,毕艳茹.基于Sarsa算法和蚁群优化的监测网络路由控制设计[J].计算机测量与控制,2014,22(10):3327-3329. 被引量：2
9李学勇,欧阳柳波,李国徽.基于隐偏向信息学习的强化学习算法[J].南华大学学报（理工版）,2004,18(2):10-16. 被引量：4
10战忠丽,王强,陈显亭.强化学习的模型、算法及应用[J].电子科技,2011,24(1):47-49. 被引量：8

计算机工程与设计

2008年第22期

浏览历史

内容加载中请稍等...

基于组合神经网络的Sarsa(λ)学习算法

参考文献12

二级参考文献9

共引文献6

相关作者

相关机构

相关主题

浏览历史