摘要
共同知识是多智能体系统内众所周知的知识集。如何充分利用共同知识进行策略学习,是多智能体独立学习系统中的一个挑战性问题。针对这一问题,围绕共同知识提取和独立学习网络设计,提出了一种基于观测重构的多智能体强化学习方法IPPO-CKOR。首先,对智能体的观测信息进行共同知识特征的计算与融合,得到融合共同知识特征的观测信息;其次,采用基于共同知识的智能体选择算法,选择关系密切的智能体,并使用重构特征生成机制构建它们的特征信息,其与融合共同知识特征的观测信息组成重构观测信息,用于智能体策略的学习与执行;最后,设计了一个基于观测重构的独立学习网络,使用多头自注意力机制对重构观测信息进行处理,使用一维卷积和GRU层处理观测信息序列,使得智能体能够从观测信息序列中提取出更有效的特征,有效缓解了环境非平稳与部分可观测问题带来的影响。实验结果表明,相较于现有典型的采用独立学习的多智能体强化学习方法,所提方法在性能上有显著提升。
Common knowledge is a well-known knowledge set within a multi-agent system.How to make full use of common knowledge for strategic learning is a challenging problem in multi-agent independent learning systems.In addressing this pro-blem,this paper proposes a multi-agent reinforcement learning method called IPPO-CKOR based on observation reconstruction,focusing on common knowledge extraction and independent learning network design.Firstly,the common knowledge features of agents’observation information are computed and fused to obtain fused observation information with common knowledge features.Secondly,an agent selection algorithm based on common knowledge is used to select closely related agents,and a feature generation mechanism based on reconstruction is employed to construct their feature information.The reconstructed observation information,composed of the fused observation information with common knowledge features,is utilized for learning and executing agent policies.Thirdly,a network structure based on observation reconstruction is designed,which employs multi-head self-attention mechanism to process the reconstructed observation information and uses one-dimensional convolution and GRU layers to handle observation information sequences.This enables the agents to extract more effective features from the observation information sequences,effectively alleviating the impact of non-stationary environments and partially observable problems.Experimental results demonstrate that the proposed method outperforms existing typical multi-agent reinforcement learning methods that employ independent learning in terms of performance.
作者
史殿习
胡浩萌
宋林娜
杨焕焕
欧阳倩滢
谭杰夫
陈莹
SHI Dianxi;HU Haomeng;SONG Linna;YANG Huanhuan;OUYANG Qianying;TAN Jiefu;CHEN Ying(Intelligent Game and Decision Lab(IGDL),Beijing 100091,China;College of Computer,National University of Defense Technology,Changsha 410073,China;Tianjin Artificial Intelligence Innovation Center,Tianjin 300457,China;National Innovation Institute of Defense Technology,Beijing 100071,China)
出处
《计算机科学》
CSCD
北大核心
2024年第4期280-290,共11页
Computer Science
基金
科技部科技创新2030-重大项目(2020AAA0104802)
国家自然科学基金(91948303)。
关键词
观测重构
多智能体协作策略
多智能体强化学习
独立学习
Observation reconstruction
Multi-agent cooperative strategy
Multi-agent reinforcement learning
Independent learning