摘要
提出一种具备全局供需动态感知能力、基于均值场多智能体强化学习的网约车平台订单分配算法。该算法通过将多智能体强化学习与均值场理论相结合,提升了智能体在局部空间上相互之间的协作性;通过注入全局空间上供需的动态分布信息,提升了智能体对全局供需分布的感知和优化能力。本文构建了真实历史数据驱动的模拟器,用于算法的训练和评估。实验表明,在全天时段和高峰期时段两个不同场景下,本文提出的算法在网约车司机累计收益及订单应答率两个重要指标上均显著优于现有的订单分配算法。实验结果充分验证了本文提出算法的有效性。
This paper proposes an order dispatch algorithm of online ride-hailing platform based on meanfield multi-agent reinforcement learning with the ability to globally perceive supply-demand dynamics.Our algorithm improves the collaboration between agents in the local area by integrating multi-agent reinforcement learning with mean-field theory,and enhances the ability of agents on perceiving and optimizing the global supply-demand gap across the global area by injecting the context about global supplydemand dynamics.Besides,we built a data-driven simulator for the training and evaluation of algorithms.Extensive experiments show that in two different scenarios of a whole day and rush hour,our algorithm significantly outperforms the existing order dispatch algorithms in terms of order response rate and accumulated drivers’income.The experimental results convincingly validate the effectiveness of our algorithm.
作者
宋旺
胡祥
张玉辉
卫文江
周雅诗
康傲
SONG Wang;HU Xiang;ZHANG Yuhui;WEI Wenjiang;ZHOU Yashi;KANG Ao(School of Control and Computer Engineering,North China Electric Power University,Beijing 102206,China)
出处
《数据采集与处理》
CSCD
北大核心
2023年第3期652-664,共13页
Journal of Data Acquisition and Processing
基金
国家自然科学基金(52078212)。
关键词
多智能体强化学习
均值场
全局供需动态感知
网约车平台
订单分配
multi-agent reinforcement learning
mean-field
global perceive supply-demand dynamics
online ride-hailing platform
order dispatch