面向复杂攻防对抗场景的攻击者行为生成与策略演化算法
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

中国南方电网有限责任公司科技项目(ZBKJXM20240163)


Attacker Behavior Generation and Policy Evolution Algorithm for Complex Offensive and Defensive Adversarial Scenarios
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    针对攻击行为建模中动态演化能力不足与战术逻辑失真导致的防御滞后问题, 本文提出一种融合多智能体强化学习(multi-agent reinforcement learning, MARL)与ATT&CK (adversarial tactics, techniques, and common knowledge)知识引导的攻击者行为生成与策略演化算法. 通过将战术-技术映射编码为约束状态空间, 确保攻击路径符合真实逻辑; 部署异构智能体协同模拟多阶段攻击, 并引入注意力机制实现意图交互; 设计分层奖励函数, 综合技术执行、战术完成与链路完整性进行策略引导; 结合动态防御扰动与经验回放机制, 支持策略在线自适应更新. 实验结果表明, 该方法在高防御强度下战术覆盖率与转移有效率分别达88.9%±2.8%和84.2%±2.2%, 任务完成率稳定在82.3%–94.2%, 收敛更快(82轮)、波动更小(标准差2.3%), 显著优于基线模型, 有效提升行为真实性、协同性与对抗鲁棒性, 为构建高保真自适应攻防系统提供新路径.

    Abstract:

    To address delayed defense responses caused by limited dynamic evolution capabilities and distorted tactical logic in attack behavior modeling, this study proposes an attacker behavior generation and policy evolution algorithm that integrates multi-agent reinforcement learning (MARL) and ATT&CK (adversarial tactics, techniques, and common knowledge) knowledge guidance. By encoding tactical-technical mappings into a constrained state space, the proposed algorithm ensures that attack paths conform to real-world logic. Heterogeneous agents are deployed to collaboratively simulate multi-stage attacks, and an attention mechanism is introduced to achieve intent interaction. A hierarchical reward function is designed to guide policies by integrating technical execution, tactical completion, and link integrity. Combined with dynamic defense perturbation and experience replay mechanisms, online adaptive policy updates are supported. Experimental results show that the proposed method achieves tactical coverage and transfer efficiency of 88.9%±2.8% and 84.2%±2.2%, respectively, under high defense intensity, with the task completion rate remaining stable at 82.3%–94.2%. The proposed method converges faster (82 rounds) and exhibits less fluctuation (standard deviation: 2.3%), significantly outperforming the baseline model. It also effectively improves behavioral realism, coordination, and adversarial robustness, providing a new path for building high-fidelity adaptive offensive and defensive systems.

    参考文献
    相似文献
    引证文献
引用本文

洪超,潘远,李攀登,黄开天,蒋屹新,冯国聪.面向复杂攻防对抗场景的攻击者行为生成与策略演化算法.计算机系统应用,,():1-14

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-01-30
  • 最后修改日期:2026-02-14
  • 录用日期:
  • 在线发布日期: 2026-07-10
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62661041 传真: Email:csa@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号