基于深度强化学习的显式拥塞通知网络优化算法
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

国家自然科学基金(61971230, 62201273)


Explicit Congestion Notification Network Optimization Algorithm Based on Deep Reinforcement Learning
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    针对数据中心网络(data center network, DCN)中突发流量频发、拥塞反馈滞后及队列震荡严重等问题, 提出一种基于深度强化学习的显式拥塞通知网络优化(deep Q-learning explicit congestion notification optimization, DQECO)算法. 首先, 该算法以PIE队列管理机制为基础, 引入深度Q网络(deep Q-learning network, DQN)实现对显式拥塞通知(explicit congestion notification, ECN)标记概率的自适应动态调节. 其次, 通过融合队列时延、队列占用率及ECN标记比例等多维度特征, 精准感知网络拥塞状态, 实现对网络拥塞的精细化动态调控. 最后, 通过设计非线性多目标奖励函数, 引导智能体在与网络环境的持续交互中最大化长期累积回报, 实现算法性能的持续迭代优化. 为验证算法性能, 本文构建基于ns3-ai的仿真环境, 在多种场景下与NewReno、QTCP算法进行对比测试; 实验结果表明, DQECO算法综合性能优于对比算法, 在高带宽场景下, 吞吐量较NewReno提升约160%、较QTCP提升约13.4%, 平均往返时延、队列占用率及丢包率分别控制在0.08 s左右、0.2以下及0.4以下, 能有效保证吞吐量稳定、抑制队列震荡, 具有更好的稳定性与自适应能力.

    Abstract:

    Aiming at the problems of frequent bursty traffic, delayed congestion feedback, and severe queue oscillation in data center networks (DCNs), this study proposes an explicit congestion notification (ECN) network optimization algorithm based on deep Q-learning (DQECO). Firstly, based on the proportional integral controller enhanced (PIE) queue management mechanism, the algorithm introduces the deep Q-learning network (DQN) to realize adaptive and dynamic adjustment of the ECN marking probability. Secondly, by fusing multi-dimensional features including the queue delay, queue occupancy rate, and ECN marking ratio, the algorithm accurately perceives the network congestion state and realizes refined dynamic control of network congestion. Finally, by designing a nonlinear multi-objective reward function, the algorithm guides the agent to maximize the long-term cumulative reward in continuous interaction with the network environment, and realizes continuous iterative optimization of algorithm performance. To verify algorithm performance, this study constructs a simulation environment based on ns3-ai and carries out comparative tests with NewReno and QTCP algorithms in multiple scenarios. Experimental results show that the comprehensive performance of the DQECO algorithm outperforms the comparison algorithms. In high-bandwidth scenarios, its throughput is improved by about 160% compared with NewReno and about 13.4% compared with QTCP. Meanwhile, its average round-trip time (RTT), queue occupancy rate, and packet loss rate are controlled at about 0.08s, below 0.2 and below 0.4, respectively. The algorithm can ensure stable throughput, suppress queue oscillation, and has better stability and adaptability.

    参考文献
    相似文献
    引证文献
引用本文

孙兵,陈辉.基于深度强化学习的显式拥塞通知网络优化算法.计算机系统应用,2026,35(8):205-220

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-12-15
  • 最后修改日期:2026-03-02
  • 录用日期:
  • 在线发布日期: 2026-06-25
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62661041 传真: Email:csa@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号