[1]XIONG Gang,CAI Kaiqi,CHEN Shichao,et al.Intelligent delivery vehicle path planning based on dual-phase policy entropy adaptive Q-learning algorithm[J].CAAI transactions on intelligent systems,2026,21(5):1180-1193.[doi:10.11992/tis.202511006]
Copy

Intelligent delivery vehicle path planning based on dual-phase policy entropy adaptive Q-learning algorithm

References:
[1] 孙智威, 裴晓飞, 刘一平, 等. 无人驾驶清扫车的路径跟踪及远程控制[J]. 汽车安全与节能学报, 2022, 13(4): 729-737 SUN Zhiwei, PEI Xiaofei, LIU Yiping, et al. Path tracking and remote control of driverless sweeper[J]. Journal of automotive safety and energy, 2022, 13(4): 729-737
[2] 李玉龙, 谢辉, 宋康. 无人驾驶公交车基于循迹误差观测和目标测量误差观测的避障路径规划算法[J]. 汽车安全与节能学报, 2024, 15(4): 579-590 LI Yulong, XIE Hui, SONG Kang. An obstacle avoidance path planning algorithm for autonomous buses based on tracking error observation and target measurement error observation[J]. Journal of automotive safety and engergy, 2024, 15(4): 579-590
[3] 夏雨奇, 黄炎焱, 陈恰. 基于深度Q网络的无人车侦察路径规划[J]. 系统工程与电子技术, 2024, 46(9): 3070-3081 XIA Yuqi, HUANG Yanyan, CHEN Qia. Path planning for unmanned vehicle reconnaissance based on deep Q-network[J]. Systems engineering and electronics, 2024, 46(9): 3070-3081
[4] 成怡, 肖宏图. 融合改进A*算法和Morphin算法的移动机器人动态路径规划[J]. 智能系统学报, 2020, 15(3): 546-552 CHENG Yi, XIAO Hongtu. Mobile-robot dynamic path planning based on improved A* and Morphin algorithms[J]. CAAI transactions on intelligent systems, 2020, 15(3): 546-552
[5] QU Tianci, XIONG Gang, ALI H, et al. USV path planning under marine environment simulation using DWA and safe reinforcement learning[C]//2023 IEEE 19th International Conference on Automation Science and Engineering. Auckland: IEEE, 2023: 1-6.
[6] 李娟, 张子浩, 张宏瀚. 复杂环境下DWA与RRT算法融合的AUV局部路径规划[J]. 智能系统学报, 2024, 19(4): 961-973 LI Juan, ZHANG Zihao, ZHANG Honghan. Local path planning for AUV with fusion of DWA and RRT algorithms in a complex environment[J]. CAAI transactions on intelligent systems, 2024, 19(4): 961-973
[7] 蔡军, 钟志远. 改进蚁群算法的送餐机器人路径规划[J]. 智能系统学报, 2024, 19(2): 370-380 CAI Jun, ZHONG Zhiyuan. Path planning of a meal delivery robot based on an improved ant colony algorithm[J]. CAAI transactions on intelligent systems, 2024, 19(2): 370-380
[8] CHEN Zenghua, XIONG Gang, LIU Sheng, et al. Path planning of mobile robot based on an improved genetic algorithm[C]//2022 IEEE 2nd International Conference on Digital Twins and Parallel Intelligence. Boston: IEEE, 2022: 1-6.
[9] WANG Chunlei, YANG Xiao, LI He. Improved Q-learning applied to dynamic obstacle avoidance and path planning[J]. IEEE access, 2022, 10: 92879-92888
[10] ZHANG Kehan. Path planning of intelligent new energy vehicles based on DQN[C]//2025 5th International Conference on Electronics, Circuits and Information Engineering. Guangzhou: IEEE, 2025: 845-848.
[11] 伍锡如, 沈可扬. 基于人工势场的防疫机器人改进近端策略优化算法[J]. 智能系统学报, 2025, 20(3): 689-698 WU Xiru, SHEN Keyang. Improved proximal policy optimization algorithm for epidemic prevention robots based on artificial potential fields[J]. CAAI transactions on intelligent systems, 2025, 20(3): 689-698
[12] 张森, 代强强. 改进型深度确定性策略梯度的无人机路径规划[J]. 系统仿真学报, 2025, 37(4): 875-881 ZHANG Sen, DAI Qiangqiang. UAV path planning based on improved deep deterministic policy gradients[J]. Journal of system simulation, 2025, 37(4): 875-881
[13] 李永迪, 李彩虹, 张耀玉, 等. 基于改进SAC算法的移动机器人路径规划[J]. 计算机应用, 2023, 43(2): 654-660 LI Yongdi, LI Caihong, ZHANG Yaoyu, et al. Mobile robot path planning based on improved SAC algorithm[J]. Journal of computer applications, 2023, 43(2): 654-660
[14] CHEN Lili, LU K, RAJESWARAN A, et al. Decision transformer: reinforcement learning via sequence modeling[C]//Proceedings of the 35th International Conference on Neural Information Processing Systems. Vancouver: Curran Associates Inc. , 2021: 15084-15097.
[15] LI Zhenxin, LI Kailin, WANG Shihao, et al. Hydra-MDP: end-to-end multimodal planning with multi-target hydra-distillation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle: IEEE, 2024: 15823-15833.
[16] ZHOU Zewei, CAI Tianhui, ZHAO S Z, et al. AutoVLA: a vision-language-action model for end-to-end autonomous driving with adaptive reasoning and reinforcement fine-tuning[C]//Proceedings of the 39th International Conference on Neural Information Processing Systems. San Diego: Curran Associates Inc. , 2025: 1-30.
[17] 宋丽君, 周紫瑜, 李云龙, 等. 改进Q-Learning的路径规划算法研究[J]. 小型微型计算机系统, 2024, 45(4): 823-829 SONG Lijun, ZHOU Ziyu, LI Yunlong, et al. Research on path planning algorithm based on improved Q-learning algorithm[J]. Journal of Chinese computer systems, 2024, 45(4): 823-829
[18] 田晓航, 霍鑫, 周典乐, 等. 基于蚁群信息素辅助的Q学习路径规划算法[J]. 控制与决策, 2023, 38(12): 3345-3353 TIAN Xiaohang, HUO Xin, ZHOU Dianle, et al. Ant colony pheromone aided Q-learning path planning algorithm[J]. Control and decision, 2023, 38(12): 3345-3353
[19] 王小康, 冀杰, 刘洋, 等. 基于改进Q学习算法的无人物流配送车路径规划[J]. 系统仿真学报, 2024, 36(5): 1211-1221 WANG Xiaokang, JI Jie, LIU Yang, et al. Path planning of unmanned delivery vehicle based on improved Q-learning algorithm[J]. Journal of system simulation, 2024, 36(5): 1211-1221
[20] 赵也践, 王艳红, 张俊, 等. 改进Q学习算法在作业车间调度问题中的应用[J]. 系统仿真学报, 2022, 34(6): 1247-1258 ZHAO Yejian, WANG Yanhong, ZHANG Jun, et al. Application of improved Q learning algorithm in job shop scheduling problem[J]. Journal of system simulation, 2022, 34(6): 1247-1258
[21] ZANGIROLAMI V, BORROTTI M. Dealing with uncertainty: Balancing exploration and exploitation in deep recurrent reinforcement learning[J]. Knowledge-based systems, 2024, 293: 111663
[22] MA Haozhe, LUO Zhengding, VO T V, et al. Highly efficient self-adaptive reward shaping for reinforcement learning[EB/OL]. (2025-02-28)[2025-11-04]. https://doi.org/10.48550/arXiv.2408.03029.
[23] 任伟, 朱建鸿. 改进的自校正Q-learning应用于智能机器人路径规划[J]. 机械科学与技术, 2025, 44(1): 126-132 REN Wei, ZHU Jianhong. Improved self-tuning Q-learning algorithm applied to path planning of intelligent robot[J]. Mechanical science and technology for aerospace engineering, 2025, 44(1): 126-132
[24] 温广辉, 杨涛, 周佳玲, 等. 强化学习与自适应动态规划: 从基础理论到多智能体系统中的应用进展综述[J]. 控制与决策, 2023, 38(5): 1200-1230 WEN Guanghui, YANG Tao, ZHOU Jialing, et al. Reinforcement learning and adaptive/approximate dynamic programming: a survey from theory to applications in multi-agent systems[J]. Control and decision, 2023, 38(5): 1200-1230
[25] 段建民, 陈强龙. 利用先验知识的Q-Learning路径规划算法研究[J]. 电光与控制, 2019, 26(9): 29-33 DUAN Jianmin, CHEN Qianglong. Prior knowledge based Q-learning path planning algorithm[J]. Electronics optics & control, 2019, 26(9): 29-33
[26] 蔡静雯, 马玉敏, 黎声益, 等. 基于Q学习的智能车间自适应调度方法[J]. 计算机集成制造系统, 2023, 29(11): 3727-3737 CAI Jingwen, MA Yumin, LI Shengyi, et al. Self-adaptive scheduling method for smart shop floor based on Q-learning[J]. Computer integrated manufacturing systems, 2023, 29(11): 3727-3737
[27] 杨秀霞, 高恒杰, 刘伟, 等. 基于阶段Q学习算法的机器人路径规划[J]. 兵器装备工程学报, 2022, 43(5): 197-203 YANG Xiuxia, GAO Hengjie, LIU Wei, et al. Robot path planning based on stage Q learning algorithm[J]. Journal of ordnance equipment engineering, 2022, 43(5): 197-203
[28] MAHRAN Y, GAMAL Z, EL-BADAWY A. Dynamic entropy tuning in reinforcement learning low-level quadcopter control: stochasticity vs determinism[C]//2024 34th International Conference on Computer Theory and Applications. Alexandria: IEEE, 2024: 185-192.
[29] 陈永, 康婕. 玻尔兹曼优化Q-learning的高速铁路越区切换控制算法[J]. 控制理论与应用, 2025, 42(4): 688-694 CHEN Yong, KANG Jie. Boltzmann optimized Q-learning algorithm for high-speed railway handover control[J]. Control theory & applications, 2025, 42(4): 688-694
[30] HAARNOJA T, TANG Haoran, ABBEEL P, et al. Reinforcement learning with deep energy-based policies[C]//International Conference on Machine Learning. Sydney: PMLR, 2017: 1352-1361.
[31] ALI H, XIONG Gang, WU Huaiyu, et al. Multi-robot path planning and trajectory smoothing[C]//2020 IEEE 16th International Conference on Automation Science and Engineering. Hong Kong: IEEE, 2020: 685-690.
Similar References:

Memo

-

Last Update: 2026-09-05

Copyright © CAAI Transactions on Intelligent Systems