[1]XIONG Gang,CAI Kaiqi,CHEN Shichao,et al.Intelligent delivery vehicle path planning based on dual-phase policy entropy adaptive Q-learning algorithm[J].CAAI transactions on intelligent systems,2026,21(5):1180-1193.[doi:10.11992/tis.202511006]
Copy
CAAI transactions on intelligent systems[ISSN 1673-4785/CN 23-1538/TP] Volume:
21
Number of periods:
2026 5
Page number:
1180-1193
Column:
学术论文—机器学习
Public date:
2026-09-05
- Title:
-
Intelligent delivery vehicle path planning based on dual-phase policy entropy adaptive Q-learning algorithm
- Author(s):
-
XIONG Gang1; 2; CAI Kaiqi1; CHEN Shichao2; ZHU Fenghua2; GUO Chao2; GAO Chao3; CHEN Dewang4
-
1. School of Artificial Intelligence and Transportation Engineering, FuJian University of Technology, Fuzhou 350000, China;
2. The State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China;
3. School of Computer and Artificial Intelligence, Beijing Technology and Business University, Beijing 100048, China;
4. School of Computing and Data Science, FuJian University of Technology, Fuzhou 350000, China
-
- Keywords:
-
path planning; Q-learning; dual-phase policy entropy; intelligent delivery vehicle; reinforcement learning; dynamic obstacle avoidance; convergence speed; path quality; robustness
- CLC:
-
TP181;U495
- DOI:
-
10.11992/tis.202511006
- Abstract:
-
Traditional Q-learning algorithms for intelligent delivery vehicle path planning often suffer from slow convergence, unstable path quality, and poor adaptability in dynamic environments. To address these issues, this paper proposes a dual-phase policy entropy adaptive Q-learning (DPEA-QL) algorithm. First, a policy entropy-driven adaptive learning rate mechanism combined with a softmax policy is introduced to balance sufficient early-stage exploration and efficient late-stage exploitation, thereby accelerating learning efficiency. Second, a dual-phase entropy regulation strategy is constructed, utilizing a sigmoid function to smoothly adjust the entropy weight, which effectively enhances policy stability and global optimality. Furthermore, a Manhattan distance-guided reward function is designed to dynamically reshape the reward signals based on the vehicle-to-goal distance, strengthening target orientation and preventing the algorithm from falling into local optima. Simulation experiments were conducted across three types of environments: single static, hybrid static, and dynamic obstacle maps. The experimental results demonstrate that DPEA-QL consistently plans the optimal path across three types of maps, with a shortest-path length of 58 steps. Compared with traditional Q-learning, the optimal path achievement rate is improved by 24.28 percentage points. Overall, the proposed algorithm significantly outperforms traditional Q-learning, Sarsa, and SA-QL algorithms in terms of convergence speed, path quality, and robustness.