###
计算机系统应用英文版:2025,34(12):149-155
本文二维码信息
码上扫一扫!
面向强化学习奖励变化率的动态调整机制优化与平衡
(南京信息工程大学 软件学院, 南京 210044)
Optimization and Balance of Dynamic Adjustment Mechanism for Reward Change Rate in Reinforcement Learning
(School of Software, Nanjing University of Information Science & Technology, Nanjing 210044, China)
摘要
图/表
参考文献
相似文献
本文已被:浏览 531次   下载 788
Received:May 26, 2025    Revised:June 17, 2025
中文摘要: 本文主要探讨了一种基于奖励变化的动态置信上限(upper confidence bound, UCB)系数ci(t)调整方法, 旨在改进UCB算法在动态环境中的表现. 传统的UCB算法通过固定的置信上限系数平衡探索与利用, 但在环境变化时, 固定系数会导致过度探索或过度利用, 影响算法的适应性. 为解决这一问题, 我们提出了一种基于历史奖励变化率的衰减机制, 通过动态调整每个臂的置信上限系数, 使得算法能够灵活应对环境变化, 更有效地平衡探索与利用. 实验结果表明, 基于奖励衰减机制的UCB算法能够在奖励波动较大的动态环境中表现优于标准UCB算法. 通过根据每个臂的历史表现动态调整置信上限系数, 算法减少了对不确定臂的探索, 并集中资源更快地利用表现较好的臂, 从而提高累积奖励并加速收敛. 这一改进方法在处理动态和复杂环境中具有显著优势, 能够有效提升UCB算法的长期性能.
Abstract:This study explores a method to adjust the dynamic upper confidence bound (UCB) coefficient ci(t) based on reward changes, aiming to improve the performance of the UCB algorithm in dynamic environments. The conventional UCB algorithm balances exploration and exploitation through a fixed upper confidence bound coefficient. However, in a changing environment, a fixed coefficient may lead to excessive exploration or exploitation, affecting the adaptability of the algorithm. To address this issue, a decay mechanism is proposed based on the reward change rate, which dynamically adjusts the upper confidence bound coefficient for each arm. This enables the algorithm to flexibly respond to environmental changes and more effectively balance exploration and exploitation. Experimental results show that the UCB algorithm based on the reward decay mechanism outperforms the standard UCB algorithm in dynamic environments with significant reward fluctuations. By dynamically adjusting the upper confidence bound coefficient according to the historical performance of each arm, the algorithm reduces exploration of uncertain arms and concentrates resources on exploiting better-performing arms more quickly, thereby increasing cumulative rewards and accelerating convergence. This optimization method has significant advantages in dynamic and complex environments and can effectively enhance the long-term performance of the UCB algorithm.
文章编号:     中图分类号:    文献标志码:
基金项目:
引用文本:
郑润康,孔燕.面向强化学习奖励变化率的动态调整机制优化与平衡.计算机系统应用,2025,34(12):149-155
ZHENG Run-Kang,KONG Yan.Optimization and Balance of Dynamic Adjustment Mechanism for Reward Change Rate in Reinforcement Learning.COMPUTER SYSTEMS APPLICATIONS,2025,34(12):149-155