A deep reinforcement learning(DRL)method based on the deep deterministic policy gradient(DDPG)algorithm is proposed to address the problems of a mismatch between the needed training samples and the actual training sam...A deep reinforcement learning(DRL)method based on the deep deterministic policy gradient(DDPG)algorithm is proposed to address the problems of a mismatch between the needed training samples and the actual training samples during the training of in-telligence,the overestimation and underestimation of the existence of Q-values,and the insufficient dynamism of the intelligence policy exploration.This method introduces the Actor-Critic Off-Policy Correction(AC-Off-POC)reinforcement learning framework and an improved double Q-value learning method,which enables the value function network in the target task to provide a more accurate evaluation of the policy network and converge to the optimal policy more quickly and stably to obtain higher value returns.The method is applied to multiple MuJoCo tasks on the Open AI Gym simulation platform.The experimental results show that it is better than the DDPG algorithm based solely on the different policy correction framework(AC-Off-POC)and the conventional DRL algorithm.The value of returns and stability of the double-Q-network off-policy correction algorithm for the deep deterministic policy gradient(DCAOP-DDPG)pro-posed by the authors are significantly higher than those of other DRL algorithms.展开更多
文摘A deep reinforcement learning(DRL)method based on the deep deterministic policy gradient(DDPG)algorithm is proposed to address the problems of a mismatch between the needed training samples and the actual training samples during the training of in-telligence,the overestimation and underestimation of the existence of Q-values,and the insufficient dynamism of the intelligence policy exploration.This method introduces the Actor-Critic Off-Policy Correction(AC-Off-POC)reinforcement learning framework and an improved double Q-value learning method,which enables the value function network in the target task to provide a more accurate evaluation of the policy network and converge to the optimal policy more quickly and stably to obtain higher value returns.The method is applied to multiple MuJoCo tasks on the Open AI Gym simulation platform.The experimental results show that it is better than the DDPG algorithm based solely on the different policy correction framework(AC-Off-POC)and the conventional DRL algorithm.The value of returns and stability of the double-Q-network off-policy correction algorithm for the deep deterministic policy gradient(DCAOP-DDPG)pro-posed by the authors are significantly higher than those of other DRL algorithms.