Abstract With the increasing complexity of network system architectures and the advancements in artificial intelligence, automated penetration testing holds crucial significance in enhancing network security. Attack path optimization is the key to automated penetration testing. However, existing methods mostly focus on maximizing the attack gain, overlooking the impact of the attack path length. To address this issue, we propose a bi-objective attack path optimization problem based on attack graphs, aiming to maximize attack gains while minimizing the attack path length. We model the problem as a Markov decision process and theoretically analyze the selection of the discount factor to ensure the consistency between the goal of the MDP and the optimization objective. To address the issues of excessive invalid actions and poor training effect of current deep reinforcement learning-based attack path optimization methods, we propose a Prior Knowledge-based Proximal Policy Optimization (PKPPO) algorithm. The algorithm first designs action filtering rules based on prior knowledge extracted from the attack graph. Based on these rules, action mask vectors are generated before eac
📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً