本文已被:浏览 698次 下载 837次
Received:May 08, 2025 Revised:May 30, 2025
Received:May 08, 2025 Revised:May 30, 2025
中文摘要: 在强化学习算法中, 经验回放(experience replay)是提高策略学习效率的关键技术之一. 然而, 如何有效地选择并重用经验, 仍然是一个具有挑战性的问题. 在经验回放中, 使用高梯度一致性的经验能更好地增强策略稳定性, 减少高时序差分(temporal difference, TD)误差带来的噪声. 基于此, 本文提出基于梯度一致性的经验回放(gradient coherence optimized experience replay, GCER)机制, 该机制使得策略网络的训练不仅基于传统的TD误差, 还基于样本的梯度一致性. GCER能够筛选出对策略提升有较高价值的经验, 同时避免因高TD误差样本的噪声而导致的策略更新不稳定. 实验结果显示, 结合梯度一致性和TD误差的经验选择机制显著提升了策略的学习效率和稳定性, 为强化学习中的样本采样策略提供了新的视角和优化方向.
Abstract:In reinforcement learning, experience replay is a key technique to improve the efficiency of policy learning. However, how to effectively select and reuse experiences remains a challenging issue. Using experiences with high gradient coherence in experience replay can better enhance policy stability and reduce the noise brought by high temporal difference (TD) errors. Based on this, gradient coherence optimized experience replay (GCER) is proposed, an algorithm that not only relies on traditional TD errors but also incorporates gradient coherence to select experiences. This method can filter out the most valuable experiences for policy improvement while avoiding unstable updates due to drastic changes in gradient direction. Experimental results show that the experience selection mechanism combining gradient coherence and TD errors significantly enhances the learning efficiency and stability of policies, providing a new perspective and direction for experience sampling strategies in reinforcement learning.
文章编号: 中图分类号: 文献标志码:
基金项目:
引用文本:
陈龙,孔燕.基于梯度一致性的强化学习经验回放机制.计算机系统应用,2025,34(12):240-248
CHEN Long,KONG Yan.Reinforcement Learning Experience Replay Mechanism Based on Gradient Coherence.COMPUTER SYSTEMS APPLICATIONS,2025,34(12):240-248
陈龙,孔燕.基于梯度一致性的强化学习经验回放机制.计算机系统应用,2025,34(12):240-248
CHEN Long,KONG Yan.Reinforcement Learning Experience Replay Mechanism Based on Gradient Coherence.COMPUTER SYSTEMS APPLICATIONS,2025,34(12):240-248

