###
计算机系统应用英文版:2025,34(2):61-73
本文二维码信息
码上扫一扫!
基于SPER-TD3的无人机编队三维航迹规划
(1.南京信息工程大学 电子与信息工程学院, 南京 210044;2.国防科技大学 第六十三研究所, 南京 210007)
3D Trajectory Planning for Unmanned Aerial Vehicle Formation Based on SPER-TD3
(1.School of Electronic and Information Engineering, Nanjing University of Information Science & Technology, Nanjing 210044, China;2.Sixty-third Research Institute, National University of Defense Technology, Nanjing 210007, China)
摘要
图/表
参考文献
相似文献
本文已被:浏览 1739次   下载 1068
Received:July 10, 2024    Revised:August 01, 2024
中文摘要: 复杂地形条件下, 基于深度强化学习的无人机编队航迹规划可以完成无人机编队的轨迹寻优, 路径长度和环境适应性均优于传统启发式算法, 但仍存在训练稳定性不足、规划实时性差等问题. 面向领航者-跟随者模式的无人机集群, 本文提出了一种基于SPER-TD3算法的无人机编队实时三维航迹规划方法. 首先, 将基于SumTree的优先经验回放机制融入TD3算法, 设计了SPER-TD3算法, 确定无人机编队的轨迹; 然后, 使用基于角度队形控制方法优化跟随者的飞行轨迹, 并应用动态轨迹平滑算法优化转向角. 为了加快SPER-TD3算法的训练收敛速度和稳定性, 解决长时间依赖性问题, 设计了结合LSTM、自注意力机制以及多重感知机的网络模型结构. 在多种障碍物环境下进行了仿真实验, 结果表明, 所提方法在轨迹安全覆盖率、飞行路径平滑度、成功率、奖励大小等方面综合表现优于8种主流的深度强化学习算法, 其重要性综合评估值比当前方法提升8.5%–72.9%不等, 且训练稳定性最佳.
Abstract:In complex terrain conditions, UAV formation path planning based on deep reinforcement learning can optimize the path of UAV formation, with better path length and environmental adaptability than traditional heuristic algorithms. However, it still has problems such as insufficient training stability and poor real-time planning. For UAV clusters with a leader-follower mode, this study proposes a real-time 3D path planning method for UAV formation based on the SPER-TD3 algorithm. Firstly, the prioritized experience replay mechanism based on SumTree is integrated into the TD3 algorithm, and the SPER-TD3 algorithm is designed to determine the path of the UAV formation. Then, an angle formation control method is used to optimize the path of the followers, and a dynamic path smoothing algorithm is applied to optimize the steering angle. To accelerate the training convergence speed and stability of the SPER-TD3 algorithm, and solve the long-term dependence problem, a network model structure combining LSTM, self-attention mechanism, and multiple perceptrons is designed. Simulation experiments are conducted in environments with various obstacles. Results show that the method mentioned above is superior to eight mainstream deep reinforcement learning algorithms in terms of path safety coverage rate, flight path smoothness, success rate, and reward size. Its comprehensive evaluation value of importance is 8.5% to 72.9% higher than existing methods, and it has the best training stability.
文章编号:     中图分类号:    文献标志码:
基金项目:
引用文本:
彭博,王晓波,魏祥麟,成洁,秦华旺,范建华.基于SPER-TD3的无人机编队三维航迹规划.计算机系统应用,2025,34(2):61-73
PENG Bo,WANG Xiao-Bo,WEI Xiang-Lin,CHENG Jie,QIN Hua-Wang,FAN Jian-Hua.3D Trajectory Planning for Unmanned Aerial Vehicle Formation Based on SPER-TD3.COMPUTER SYSTEMS APPLICATIONS,2025,34(2):61-73