###
计算机系统应用英文版:,():1-14
本文二维码信息
码上扫一扫!
面向复杂攻防对抗场景的攻击者行为生成与策略演化算法
(1.南方电网科学研究院有限责任公司, 广州 510663;2.广东省电力系统网络安全重点实验室, 广州 510663;3.中国南方电网有限责任公司, 广州 510663)
Attacker Behavior Generation and Policy Evolution Algorithm for Complex Offensive and Defensive Adversarial Scenarios
(1.China Southern Power Grid Research Institute Co. Ltd., Guangzhou 510663, China;2.Guangdong Electric Power System Network Security Key Laboratory, Guangzhou 510663, China;3.China Southern Power Grid Co. Ltd., Guangzhou 510663, China)
摘要
图/表
参考文献
相似文献
本文已被:浏览 35次   下载 64
Received:January 30, 2026    Revised:February 14, 2026
中文摘要: 针对攻击行为建模中动态演化能力不足与战术逻辑失真导致的防御滞后问题, 本文提出一种融合多智能体强化学习(multi-agent reinforcement learning, MARL)与ATT&CK (adversarial tactics, techniques, and common knowledge)知识引导的攻击者行为生成与策略演化算法. 通过将战术-技术映射编码为约束状态空间, 确保攻击路径符合真实逻辑; 部署异构智能体协同模拟多阶段攻击, 并引入注意力机制实现意图交互; 设计分层奖励函数, 综合技术执行、战术完成与链路完整性进行策略引导; 结合动态防御扰动与经验回放机制, 支持策略在线自适应更新. 实验结果表明, 该方法在高防御强度下战术覆盖率与转移有效率分别达88.9%±2.8%和84.2%±2.2%, 任务完成率稳定在82.3%–94.2%, 收敛更快(82轮)、波动更小(标准差2.3%), 显著优于基线模型, 有效提升行为真实性、协同性与对抗鲁棒性, 为构建高保真自适应攻防系统提供新路径.
Abstract:To address delayed defense responses caused by limited dynamic evolution capabilities and distorted tactical logic in attack behavior modeling, this study proposes an attacker behavior generation and policy evolution algorithm that integrates multi-agent reinforcement learning (MARL) and ATT&CK (adversarial tactics, techniques, and common knowledge) knowledge guidance. By encoding tactical-technical mappings into a constrained state space, the proposed algorithm ensures that attack paths conform to real-world logic. Heterogeneous agents are deployed to collaboratively simulate multi-stage attacks, and an attention mechanism is introduced to achieve intent interaction. A hierarchical reward function is designed to guide policies by integrating technical execution, tactical completion, and link integrity. Combined with dynamic defense perturbation and experience replay mechanisms, online adaptive policy updates are supported. Experimental results show that the proposed method achieves tactical coverage and transfer efficiency of 88.9%±2.8% and 84.2%±2.2%, respectively, under high defense intensity, with the task completion rate remaining stable at 82.3%–94.2%. The proposed method converges faster (82 rounds) and exhibits less fluctuation (standard deviation: 2.3%), significantly outperforming the baseline model. It also effectively improves behavioral realism, coordination, and adversarial robustness, providing a new path for building high-fidelity adaptive offensive and defensive systems.
文章编号:     中图分类号:    文献标志码:
基金项目:中国南方电网有限责任公司科技项目(ZBKJXM20240163)
引用文本:
洪超,潘远,李攀登,黄开天,蒋屹新,冯国聪.面向复杂攻防对抗场景的攻击者行为生成与策略演化算法.计算机系统应用,,():1-14
HONG Chao,PAN Yuan,LI Pan-Deng,HUANG Kai-Tian,JIANG Yi-Xin,FENG Guo-Cong.Attacker Behavior Generation and Policy Evolution Algorithm for Complex Offensive and Defensive Adversarial Scenarios.COMPUTER SYSTEMS APPLICATIONS,,():1-14