###
计算机系统应用英文版:2026,35(2):141-153
本文二维码信息
码上扫一扫!
基于熵正则化近端策略优化的联邦客户端选择
(1.南京信息工程大学 软件学院, 南京 210044;2.浙江理工大学 信息科学与工程学院 (网络空间安全学院), 杭州 310018)
Entropy Regularization Proximal Policy Optimization for Federated Client Selection
(1.School of Software, Nanjing University of Information Science & Technology, Nanjing 210044, China;2.School of Information Science and Engineering (School of Cyber Science and Technology), Zhejiang Sci-tech University, Hangzhou 310018, China)
摘要
图/表
参考文献
相似文献
本文已被:浏览 357次   下载 525
Received:August 05, 2025    Revised:September 16, 2025
中文摘要: 近年来, 联邦学习(federated learning, FL)作为一种分布式机器学习范式, 因其能够在保护数据隐私的同时实现模型训练, 已在智能医疗、金融服务、物联网以及车联网等领域得到广泛应用. 在车联网(IoV)环境中, 由于节点高度动态和车辆资源的异构性, 并非所有客户端都适合参与联邦训练, 因此高效且鲁棒的客户端选择策略对于模型性能与系统效率至关重要. 然而, 传统FL方法大多依赖静态或启发式的客户端选择机制, 难以适应IoV场景中频繁变化的环境状态与客户端特性. 为此, 本文提出一种基于熵正则化近端策略优化(entropy regularization proximal policy optimization, ERPPO)的动态客户端选择方法, 并结合置信度加权聚合策略. 该方法通过在近端策略优化(proximal policy optimization, PPO)目标函数中引入策略熵正则项, 增强客户端选择策略的探索性, 以避免陷入局部最优. 同时, 置信度聚合机制基于客户端模型更新方差自适应调整聚合权重, 提升全局模型的收敛稳定性与鲁棒性. 实验结果表明, 所提方法在保障模型精度的前提下, 有效降低了通信开销, 并在动态环境下展现出优于传统方法的综合性能.
中文关键词: 联邦学习  车联网  客户端选择  ERPPO
Abstract:In recent years, federated learning (FL) has emerged as a distributed machine learning paradigm that enables model training while preserving data privacy. It has been widely applied in domains such as smart healthcare, financial services, the Internet of Things (IoT), and the Internet of Vehicles (IoV). However, due to the highly dynamic nature of IoV environments and the heterogeneous computing resources among vehicles, not all clients are suitable for participation in federated training. Therefore, designing an efficient and robust client selection strategy is critical for ensuring model performance and system efficiency. Traditional FL methods often rely on static or heuristic client selection mechanisms, which fail to adapt to the frequently changing states and characteristics of clients in IoV scenarios. To address this issue, this study proposes a dynamic client selection approach based on entropy regularization proximal policy optimization (ERPPO), integrated with a confidence-weighted aggregation mechanism. By incorporating a policy entropy regularization term into the PPO objective function, the proposed method enhances the exploration capability of the client selection policy, thus mitigating the risk of local optima. Furthermore, the confidence-based aggregation strategy adaptively adjusts the aggregation weights based on the variance of local model updates, which enhances the convergence stability and robustness of the global model. Experimental results demonstrate that the proposed ERPPO framework not only reduces communication overhead but also achieves superior overall performance in dynamic environments while maintaining high model precision.
文章编号:     中图分类号:    文献标志码:
基金项目:国家自然科学基金面上项目(62271264); 浙江省“尖兵领雁+X”重大科技计划 (2025C02033)
引用文本:
陈雨彤,金子龙.基于熵正则化近端策略优化的联邦客户端选择.计算机系统应用,2026,35(2):141-153
CHEN Yu-Tong,JIN Zi-Long.Entropy Regularization Proximal Policy Optimization for Federated Client Selection.COMPUTER SYSTEMS APPLICATIONS,2026,35(2):141-153