[1]熊刚,蔡凯琪,陈世超,等.基于双阶段策略熵自适应Q-learning的智能配送车路径规划方法研究[J].智能系统学报,2026,21(5):1180-1193.[doi:10.11992/tis.202511006]
XIONG Gang,CAI Kaiqi,CHEN Shichao,et al.Intelligent delivery vehicle path planning based on dual-phase policy entropy adaptive Q-learning algorithm[J].CAAI transactions on intelligent systems,2026,21(5):1180-1193.[doi:10.11992/tis.202511006]
点击复制
《智能系统学报》[ISSN 1673-4785/CN 23-1538/TP] 卷:
21
期数:
2026年第5期
页码:
1180-1193
栏目:
学术论文—机器学习
出版日期:
2026-09-05
- Title:
-
Intelligent delivery vehicle path planning based on dual-phase policy entropy adaptive Q-learning algorithm
- 作者:
-
熊刚1,2, 蔡凯琪1, 陈世超2, 朱凤华2, 郭超2, 高超3, 陈德旺4
-
1. 福建理工大学 人工智能与交通工程学院, 福建 福州 350000;
2. 中国科学院自动化研究所 多模态人工智能系统全国重点实验室, 北京 100190;
3. 北京工商大学 计算机与人工智能学院, 北京 100048;
4. 福建理工大学 计算机与数据科学学院, 福建 福州 350000
- Author(s):
-
XIONG Gang1,2, CAI Kaiqi1, CHEN Shichao2, ZHU Fenghua2, GUO Chao2, GAO Chao3, CHEN Dewang4
-
1. School of Artificial Intelligence and Transportation Engineering, FuJian University of Technology, Fuzhou 350000, China;
2. The State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China;
3. School of Computer and Artificial Intelligence, Beijing Technology and Business University, Beijing 100048, China;
4. School of Computing and Data Science, FuJian University of Technology, Fuzhou 350000, China
-
- 关键词:
-
路径规划; Q-learning; 双阶段策略熵; 智能配送车; 强化学习; 动态避障; 收敛速度; 路径质量; 鲁棒性
- Keywords:
-
path planning; Q-learning; dual-phase policy entropy; intelligent delivery vehicle; reinforcement learning; dynamic obstacle avoidance; convergence speed; path quality; robustness
- 分类号:
-
TP181;U495
- DOI:
-
10.11992/tis.202511006
- 摘要:
-
针对传统Q-learning算法在智能配送车路径规划中存在收敛速度慢、路径质量不稳定以及动态环境适应性弱的问题,提出一种基于双阶段策略熵自适应Q-learning的路径规划算法(dual-phase policy entropy adaptive Q-learning,DPEA-QL)。该算法通过引入策略熵驱动的自适应学习率优化,采用softmax策略,实现在训练初期充分探索和后期高效利用,提升算法的学习效率和收敛速度;构建双阶段熵调控机制,利用sigmoid函数平滑过渡策略熵权重,有效提升策略平稳性和路径最优性;重塑基于曼哈顿距离的奖励函数,使奖励信号随配送车与目标点的距离动态变化,增强目标导向避免陷入局部收敛。在单一静态障碍、混合静态障碍和动态障碍3类仿真地图上的实验结果表明:DPEA-QL在3类地图中均能规划出一条最优路径,最短路径步数为58步,最短路径达成率较传统Q-learning算法提高24.28百分点,在收敛速度、路径质量与鲁棒性上明显优于传统Q-learning、Sarsa及SA-QL算法。
- Abstract:
-
Traditional Q-learning algorithms for intelligent delivery vehicle path planning often suffer from slow convergence, unstable path quality, and poor adaptability in dynamic environments. To address these issues, this paper proposes a dual-phase policy entropy adaptive Q-learning (DPEA-QL) algorithm. First, a policy entropy-driven adaptive learning rate mechanism combined with a softmax policy is introduced to balance sufficient early-stage exploration and efficient late-stage exploitation, thereby accelerating learning efficiency. Second, a dual-phase entropy regulation strategy is constructed, utilizing a sigmoid function to smoothly adjust the entropy weight, which effectively enhances policy stability and global optimality. Furthermore, a Manhattan distance-guided reward function is designed to dynamically reshape the reward signals based on the vehicle-to-goal distance, strengthening target orientation and preventing the algorithm from falling into local optima. Simulation experiments were conducted across three types of environments: single static, hybrid static, and dynamic obstacle maps. The experimental results demonstrate that DPEA-QL consistently plans the optimal path across three types of maps, with a shortest-path length of 58 steps. Compared with traditional Q-learning, the optimal path achievement rate is improved by 24.28 percentage points. Overall, the proposed algorithm significantly outperforms traditional Q-learning, Sarsa, and SA-QL algorithms in terms of convergence speed, path quality, and robustness.
备注/Memo
收稿日期:2025-11-4。
基金项目:福建省闽江学者讲座教授人才计划项目(GY-Z24014);国家自然科学基金项目(62461160259, U24A20277, 62306319);福建省第三批创新之星人才计划项目(003002);北京市自然科学基金项目(L241016).
作者简介:熊刚,研究员,博士生导师,兼任北京市智能化技术与系统工程技术研究中心常务副主任、中国人工智能学会高级会员、中国自动化学会会士、中国计算机学会杰出会员等。主要研究方向为人工智能、智能控制与管理。先后获得吴文俊人工智能科技进步奖二等奖、中国自动化学会科技进步奖特等奖和一等奖等10余项。通过专利合作条约(patent cooperation treaty,PCT)途径申请并获授权6项,获专利授权100余项,登记软著90余项,发表学术论文500余篇,出版专著3部。E-mail:gang.xiong@ia.ac.cn。;蔡凯琪,硕士研究生,主要研究方向为平行交通与物流系统。E-mail:397159947@qq.com。;陈德旺,博士,电气电子工程师学会高级会员,中国自动化学会高级会员,福建省“闽江学者”特聘教授。主要研究方向为人工智能算法、模糊系统和智能交通系统。获得省部级和国家一级学会科技奖励10余项。发表学术论文200余篇,出版学术专著 4部。E-mail:dwchen@fjut.edu.cn。
通讯作者:陈德旺. E-mail:dwchen@fjut.edu.cn
更新日期/Last Update:
2026-09-05