[1]熊刚,蔡凯琪,陈世超,等.基于双阶段策略熵自适应Q-learning的智能配送车路径规划方法研究[J].智能系统学报,2026,21(5):1180-1193.[doi:10.11992/tis.202511006]
 XIONG Gang,CAI Kaiqi,CHEN Shichao,et al.Intelligent delivery vehicle path planning based on dual-phase policy entropy adaptive Q-learning algorithm[J].CAAI transactions on intelligent systems,2026,21(5):1180-1193.[doi:10.11992/tis.202511006]
点击复制

基于双阶段策略熵自适应Q-learning的智能配送车路径规划方法研究

参考文献/References:
[1] 孙智威, 裴晓飞, 刘一平, 等. 无人驾驶清扫车的路径跟踪及远程控制[J]. 汽车安全与节能学报, 2022, 13(4): 729-737 SUN Zhiwei, PEI Xiaofei, LIU Yiping, et al. Path tracking and remote control of driverless sweeper[J]. Journal of automotive safety and energy, 2022, 13(4): 729-737
[2] 李玉龙, 谢辉, 宋康. 无人驾驶公交车基于循迹误差观测和目标测量误差观测的避障路径规划算法[J]. 汽车安全与节能学报, 2024, 15(4): 579-590 LI Yulong, XIE Hui, SONG Kang. An obstacle avoidance path planning algorithm for autonomous buses based on tracking error observation and target measurement error observation[J]. Journal of automotive safety and engergy, 2024, 15(4): 579-590
[3] 夏雨奇, 黄炎焱, 陈恰. 基于深度Q网络的无人车侦察路径规划[J]. 系统工程与电子技术, 2024, 46(9): 3070-3081 XIA Yuqi, HUANG Yanyan, CHEN Qia. Path planning for unmanned vehicle reconnaissance based on deep Q-network[J]. Systems engineering and electronics, 2024, 46(9): 3070-3081
[4] 成怡, 肖宏图. 融合改进A*算法和Morphin算法的移动机器人动态路径规划[J]. 智能系统学报, 2020, 15(3): 546-552 CHENG Yi, XIAO Hongtu. Mobile-robot dynamic path planning based on improved A* and Morphin algorithms[J]. CAAI transactions on intelligent systems, 2020, 15(3): 546-552
[5] QU Tianci, XIONG Gang, ALI H, et al. USV path planning under marine environment simulation using DWA and safe reinforcement learning[C]//2023 IEEE 19th International Conference on Automation Science and Engineering. Auckland: IEEE, 2023: 1-6.
[6] 李娟, 张子浩, 张宏瀚. 复杂环境下DWA与RRT算法融合的AUV局部路径规划[J]. 智能系统学报, 2024, 19(4): 961-973 LI Juan, ZHANG Zihao, ZHANG Honghan. Local path planning for AUV with fusion of DWA and RRT algorithms in a complex environment[J]. CAAI transactions on intelligent systems, 2024, 19(4): 961-973
[7] 蔡军, 钟志远. 改进蚁群算法的送餐机器人路径规划[J]. 智能系统学报, 2024, 19(2): 370-380 CAI Jun, ZHONG Zhiyuan. Path planning of a meal delivery robot based on an improved ant colony algorithm[J]. CAAI transactions on intelligent systems, 2024, 19(2): 370-380
[8] CHEN Zenghua, XIONG Gang, LIU Sheng, et al. Path planning of mobile robot based on an improved genetic algorithm[C]//2022 IEEE 2nd International Conference on Digital Twins and Parallel Intelligence. Boston: IEEE, 2022: 1-6.
[9] WANG Chunlei, YANG Xiao, LI He. Improved Q-learning applied to dynamic obstacle avoidance and path planning[J]. IEEE access, 2022, 10: 92879-92888
[10] ZHANG Kehan. Path planning of intelligent new energy vehicles based on DQN[C]//2025 5th International Conference on Electronics, Circuits and Information Engineering. Guangzhou: IEEE, 2025: 845-848.
[11] 伍锡如, 沈可扬. 基于人工势场的防疫机器人改进近端策略优化算法[J]. 智能系统学报, 2025, 20(3): 689-698 WU Xiru, SHEN Keyang. Improved proximal policy optimization algorithm for epidemic prevention robots based on artificial potential fields[J]. CAAI transactions on intelligent systems, 2025, 20(3): 689-698
[12] 张森, 代强强. 改进型深度确定性策略梯度的无人机路径规划[J]. 系统仿真学报, 2025, 37(4): 875-881 ZHANG Sen, DAI Qiangqiang. UAV path planning based on improved deep deterministic policy gradients[J]. Journal of system simulation, 2025, 37(4): 875-881
[13] 李永迪, 李彩虹, 张耀玉, 等. 基于改进SAC算法的移动机器人路径规划[J]. 计算机应用, 2023, 43(2): 654-660 LI Yongdi, LI Caihong, ZHANG Yaoyu, et al. Mobile robot path planning based on improved SAC algorithm[J]. Journal of computer applications, 2023, 43(2): 654-660
[14] CHEN Lili, LU K, RAJESWARAN A, et al. Decision transformer: reinforcement learning via sequence modeling[C]//Proceedings of the 35th International Conference on Neural Information Processing Systems. Vancouver: Curran Associates Inc. , 2021: 15084-15097.
[15] LI Zhenxin, LI Kailin, WANG Shihao, et al. Hydra-MDP: end-to-end multimodal planning with multi-target hydra-distillation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024: 15823-15833.
[16] ZHOU Zewei, CAI Tianhui, ZHAO S Z, et al. AutoVLA: a vision-language-action model for end-to-end autonomous driving with adaptive reasoning and reinforcement fine-tuning[C]//Proceedings of the 39th International Conference on Neural Information Processing Systems. San Diego: Curran Associates Inc. , 2025: 1-30.
[17] 宋丽君, 周紫瑜, 李云龙, 等. 改进Q-Learning的路径规划算法研究[J]. 小型微型计算机系统, 2024, 45(4): 823-829 SONG Lijun, ZHOU Ziyu, LI Yunlong, et al. Research on path planning algorithm based on improved Q-learning algorithm[J]. Journal of Chinese computer systems, 2024, 45(4): 823-829
[18] 田晓航, 霍鑫, 周典乐, 等. 基于蚁群信息素辅助的Q学习路径规划算法[J]. 控制与决策, 2023, 38(12): 3345-3353 TIAN Xiaohang, HUO Xin, ZHOU Dianle, et al. Ant colony pheromone aided Q-learning path planning algorithm[J]. Control and decision, 2023, 38(12): 3345-3353
[19] 王小康, 冀杰, 刘洋, 等. 基于改进Q学习算法的无人物流配送车路径规划[J]. 系统仿真学报, 2024, 36(5): 1211-1221 WANG Xiaokang, JI Jie, LIU Yang, et al. Path planning of unmanned delivery vehicle based on improved Q-learning algorithm[J]. Journal of system simulation, 2024, 36(5): 1211-1221
[20] 赵也践, 王艳红, 张俊, 等. 改进Q学习算法在作业车间调度问题中的应用[J]. 系统仿真学报, 2022, 34(6): 1247-1258 ZHAO Yejian, WANG Yanhong, ZHANG Jun, et al. Application of improved Q learning algorithm in job shop scheduling problem[J]. Journal of system simulation, 2022, 34(6): 1247-1258
[21] ZANGIROLAMI V, BORROTTI M. Dealing with uncertainty: Balancing exploration and exploitation in deep recurrent reinforcement learning[J]. Knowledge-based systems, 2024, 293: 111663
[22] MA Haozhe, LUO Zhengding, VO T V, et al. Highly efficient self-adaptive reward shaping for reinforcement learning[EB/OL]. (2025-02-28)[2025-11-04]. https://doi.org/10.48550/arXiv.2408.03029.
[23] 任伟, 朱建鸿. 改进的自校正Q-learning应用于智能机器人路径规划[J]. 机械科学与技术, 2025, 44(1): 126-132 REN Wei, ZHU Jianhong. Improved self-tuning Q-learning algorithm applied to path planning of intelligent robot[J]. Mechanical science and technology for aerospace engineering, 2025, 44(1): 126-132
[24] 温广辉, 杨涛, 周佳玲, 等. 强化学习与自适应动态规划: 从基础理论到多智能体系统中的应用进展综述[J]. 控制与决策, 2023, 38(5): 1200-1230 WEN Guanghui, YANG Tao, ZHOU Jialing, et al. Reinforcement learning and adaptive/approximate dynamic programming: a survey from theory to applications in multi-agent systems[J]. Control and decision, 2023, 38(5): 1200-1230
[25] 段建民, 陈强龙. 利用先验知识的Q-Learning路径规划算法研究[J]. 电光与控制, 2019, 26(9): 29-33 DUAN Jianmin, CHEN Qianglong. Prior knowledge based Q-learning path planning algorithm[J]. Electronics optics & control, 2019, 26(9): 29-33
[26] 蔡静雯, 马玉敏, 黎声益, 等. 基于Q学习的智能车间自适应调度方法[J]. 计算机集成制造系统, 2023, 29(11): 3727-3737 CAI Jingwen, MA Yumin, LI Shengyi, et al. Self-adaptive scheduling method for smart shop floor based on Q-learning[J]. Computer integrated manufacturing systems, 2023, 29(11): 3727-3737
[27] 杨秀霞, 高恒杰, 刘伟, 等. 基于阶段Q学习算法的机器人路径规划[J]. 兵器装备工程学报, 2022, 43(5): 197-203 YANG Xiuxia, GAO Hengjie, LIU Wei, et al. Robot path planning based on stage Q learning algorithm[J]. Journal of ordnance equipment engineering, 2022, 43(5): 197-203
[28] MAHRAN Y, GAMAL Z, EL-BADAWY A. Dynamic entropy tuning in reinforcement learning low-level quadcopter control: stochasticity vs determinism[C]//2024 34th International Conference on Computer Theory and Applications. Alexandria: IEEE, 2024: 185-192.
[29] 陈永, 康婕. 玻尔兹曼优化Q-learning的高速铁路越区切换控制算法[J]. 控制理论与应用, 2025, 42(4): 688-694 CHEN Yong, KANG Jie. Boltzmann optimized Q-learning algorithm for high-speed railway handover control[J]. Control theory & applications, 2025, 42(4): 688-694
[30] HAARNOJA T, TANG Haoran, ABBEEL P, et al. Reinforcement learning with deep energy-based policies[C]//International Conference on Machine Learning. Sydney: PMLR, 2017: 1352-1361.
[31] ALI H, XIONG Gang, WU Huaiyu, et al. Multi-robot path planning and trajectory smoothing[C]//2020 IEEE 16th International Conference on Automation Science and Engineering. Hong Kong: IEEE, 2020: 685-690.
相似文献/References:
[1]黄彦文,曹其新.RoboCup比赛环境下足球机器人路径规划研究[J].智能系统学报,2007,2(4):52.
 HUANG Yan-wen,CAO Qin-xin.Path planning for robot soccer in the RoboCup environment[J].CAAI transactions on intelligent systems,2007,2():52.
[2]秦世引,高书征.面向救援任务的地面移动机器人路径规划[J].智能系统学报,2009,4(5):414.[doi:10.3969/j.issn.1673-4785.2009.05.005]
 QIN Shi-yin,GAO Shu-zhen.Path planning for mobile rescue robots in disaster areas with complex environments[J].CAAI transactions on intelligent systems,2009,4():414.[doi:10.3969/j.issn.1673-4785.2009.05.005]
[3]曹卫华,吴净斌,吴 敏,等.无路标环境下遥操作机器人SLAM系统[J].智能系统学报,2010,5(3):240.
 CAO Wei-hua,WU Jing-bin,WU Min,et al.A system for telerobotics in environments without landmarks[J].CAAI transactions on intelligent systems,2010,5():240.
[4]薛英花,田国会,吴 皓,等.智能空间中的服务机器人路径规划[J].智能系统学报,2010,5(3):260.
 XUE Ying-hua,TIAN Guo-hui,WU Hao,et al.Path planning for service robots in an intelligent space[J].CAAI transactions on intelligent systems,2010,5():260.
[5]黄晓丹,王粉花,王志良.情感决策的智能家居虚拟人路径规划[J].智能系统学报,2010,5(4):292.
 HUANG Xiao-dan,WANG Fen-hua,WANG Zhi-liang.Using affective decisionmaking for the path planning of virtual humans in a smart home[J].CAAI transactions on intelligent systems,2010,5():292.
[6]唐小勇,于 飞,潘洪悦.改进粒子群算法的潜器导航规划[J].智能系统学报,2010,5(5):443.[doi:10.3969/j.issn.1673-4785.2010.05.011]
 TANG Xiao-yong,YU Fei,PAN Hong-yue.Submersible path-planning based on an improved PSO[J].CAAI transactions on intelligent systems,2010,5():443.[doi:10.3969/j.issn.1673-4785.2010.05.011]
[7]夏琳琳,张健沛,初妍.计算智能在移动机器人路径规划中的应用综述[J].智能系统学报,2011,6(2):160.
 XIA Linlin,ZHANG Jianpei,CHU Yan.An application survey on computational intelligence for path planning of mobile robots[J].CAAI transactions on intelligent systems,2011,6():160.
[8]蒲兴成,张军,张毅.基于神经网络的改进行为协调控制及其在智能轮椅路径规划中的应用[J].智能系统学报,2011,6(5):456.
 PU Xingcheng,ZHANG Jun,ZHANG Yi.Modified behavior coordination for intelligent wheelchair path planning based on a neural network[J].CAAI transactions on intelligent systems,2011,6():456.
[9]肖国宝,严宣辉.一种基于改进Theta *的机器人路径规划算法[J].智能系统学报,2013,8(1):58.[doi:10.3969/j.issn.1673-4785.201208032]
 XIAO Guobao,YAN Xuanhui.A path planning algorithm based on improved Theta * for mobile robot[J].CAAI transactions on intelligent systems,2013,8():58.[doi:10.3969/j.issn.1673-4785.201208032]
[10]杨茂,田彦涛.复杂环境下多机器人觅食路径规划与控制[J].智能系统学报,2013,8(2):162.[doi:10.3969/j.issn.1673-4785.201208022]
 YANG Mao,TIAN Yantao.Foraging path planning and control for multi-robot in complex environment[J].CAAI transactions on intelligent systems,2013,8():162.[doi:10.3969/j.issn.1673-4785.201208022]

备注/Memo

收稿日期:2025-11-4。
基金项目:福建省闽江学者讲座教授人才计划项目(GY-Z24014);国家自然科学基金项目(62461160259, U24A20277, 62306319);福建省第三批创新之星人才计划项目(003002);北京市自然科学基金项目(L241016).
作者简介:熊刚,研究员,博士生导师,兼任北京市智能化技术与系统工程技术研究中心常务副主任、中国人工智能学会高级会员、中国自动化学会会士、中国计算机学会杰出会员等。主要研究方向为人工智能、智能控制与管理。先后获得吴文俊人工智能科技进步奖二等奖、中国自动化学会科技进步奖特等奖和一等奖等10余项。通过专利合作条约(patent cooperation treaty,PCT)途径申请并获授权6项,获专利授权100余项,登记软著90余项,发表学术论文500余篇,出版专著3部。E-mail:gang.xiong@ia.ac.cn。;蔡凯琪,硕士研究生,主要研究方向为平行交通与物流系统。E-mail:397159947@qq.com。;陈德旺,博士,电气电子工程师学会高级会员,中国自动化学会高级会员,福建省“闽江学者”特聘教授。主要研究方向为人工智能算法、模糊系统和智能交通系统。获得省部级和国家一级学会科技奖励10余项。发表学术论文200余篇,出版学术专著 4部。E-mail:dwchen@fjut.edu.cn。
通讯作者:陈德旺. E-mail:dwchen@fjut.edu.cn

更新日期/Last Update: 2026-09-05
Copyright © 《 智能系统学报》 编辑部
地址:(150001)黑龙江省哈尔滨市南岗区南通大街145-1号楼 电话:0451- 82534001、82518134 邮箱:tis@vip.sina.com