Abstract:To address delayed defense responses caused by limited dynamic evolution capabilities and distorted tactical logic in attack behavior modeling, this study proposes an attacker behavior generation and policy evolution algorithm that integrates multi-agent reinforcement learning (MARL) and ATT&CK (adversarial tactics, techniques, and common knowledge) knowledge guidance. By encoding tactical-technical mappings into a constrained state space, the proposed algorithm ensures that attack paths conform to real-world logic. Heterogeneous agents are deployed to collaboratively simulate multi-stage attacks, and an attention mechanism is introduced to achieve intent interaction. A hierarchical reward function is designed to guide policies by integrating technical execution, tactical completion, and link integrity. Combined with dynamic defense perturbation and experience replay mechanisms, online adaptive policy updates are supported. Experimental results show that the proposed method achieves tactical coverage and transfer efficiency of 88.9%±2.8% and 84.2%±2.2%, respectively, under high defense intensity, with the task completion rate remaining stable at 82.3%–94.2%. The proposed method converges faster (82 rounds) and exhibits less fluctuation (standard deviation: 2.3%), significantly outperforming the baseline model. It also effectively improves behavioral realism, coordination, and adversarial robustness, providing a new path for building high-fidelity adaptive offensive and defensive systems.