AI-BASED V ULNERABILITY A SSESSMENT C APABILITY AND C YBER ATTACK G RAPH A NALYSIS A P REPRINT
arXiv:2609.35414v1 [cs.CR] 28 Sep 2026
Joni Herttuainen Aalto University School of Science P.O. Box 11000, 00076 Aalto, Finland [email protected] Vesa Kuikka Aalto University School of Science P.O. Box 11000, 00076 Aalto, Finland [email protected] David Welsh Lockheed Martin [email protected]
Kirsi Hellsten Aalto University School of Science P.O. Box 11000, 00076 Aalto, Finland [email protected]
Ambrose Kam Lockheed Martin [email protected]
Arlanda Johnson Lockheed Martin [email protected]
Kimmo K. Kaski Aalto University School of Science P.O. Box 11000, 00076 Aalto, Finland [email protected]
September 29, 2026
A BSTRACT Cyber threats targeting mission-critical infrastructure are becoming more sophisticated while the barrier to launching attacks continues to fall. Traditional point solutions like antivirus and firewalls are reactive and fail to address the combinatorial complexity of modern attack surfaces. This paper presents an investigation combining two complementary methodologies: Lockheed Martin’s Vortex™/Crow™ framework, which applies multi-agent reinforcement learning (MARL) over industry-standard cyber knowledge graph to identify and prioritize attack vectors and TTPs (tactics, techniques, and procedures); and Aalto’s probabilistic attack graph model that combines network topology and its vulnerabilities to compute system-level risk metrics. The 2015 Ukraine Power Grid cyberattack serves as a well-documented validation scenario. Applied independently to the same operational technology (OT) network topology, both methodologies converge on the same attack vectors and exploit sequences as those documented in the incident record, thus providing mutual cross-validation. Attack graph analyses using node-level elimination experiments identify industrial control systems (ICS) as the most critical enablers of attack propagation, representing high-priority targets for defensive hardening. Comparison of CVSS (v2.0) and IronMiner vulnerability scoring yields in general consistent results, with IronMiner providing more actionable differentiation at network periphery nodes. The layered methodology of baseline assessment and node-level elimination proves to be scalable to large enterprise networks, thus offering defenders a structured, AI-enabled path to prioritize mitigation under realistic time and resource constraints. Keywords Attack Graphs · Multi-Agent Reinforcement Learning · Knowledge Graph · Cyber Risk Assessment · Vulnerability Assessment · Critical Infrastructure.
1
Introduction
Cybersecurity has become one of the most consequential challenges of modern societies, touching nearly every network domain, from personal devices and financial systems to industrial control systems (ICS), operational technology (OT)
AI-Based Vulnerability Assessment & Attack Graphs
A P REPRINT
and critical infrastructure such as power generation, water treatment and transportation. Unlike conventional physical threats, cyberattacks require minimal resources to initiate as malware and ready-to-use exploitation tools are freely available, and recent conflicts in Ukraine and Gaza have shown that cyber operations can be used at scale as strategic instruments, disrupting civilian infrastructure with cascading societal and economic consequences. Ransomware campaigns against utilities, hospitals, and transportation networks further confirm that cyber threats extend far beyond espionage or financial motives. Conventional defenses are poorly matched to this landscape. Traditional cybersecurity is largely a collection of reactive point solutions, e.g. antivirus, firewalls, and intrusion detection, that are narrow in scope, and even sound cyberhygiene practices can fail against sophisticated, multi-stage attacks. For an enterprise OT network that spans hundreds or thousands of nodes, each exposing many exploitable services, an exhaustive manual assessment is impractical. Attack surface analysis, Red Team exercises, and Cyber Table Top (CTT) engagements yield valuable insight but are labor-intensive, hard to scale, and inherently incomplete given the combinatorial space of possible attack paths. This motivates a shift toward proactive AI-enabled analysis that operates at the system-of-systems level, integrating vulnerability data, network topology, adversary behavior, and probabilistic risk metrics to anticipate the most impactful exploitation sequences and prioritize mitigations before an incident occurs. This paper reports a joint study by Lockheed Martin (LM) and Aalto University that pairs two independently developed but complementary methodologies toward this goal. LM’s Vortex™ platform builds a cyber knowledge graph by ingesting system configuration data together with recognized vulnerability databases (CVE [1], CWE [2], CAPEC [3]) and cyber frameworks (NIST SP 800-53B [4], MITRE ATT&CK [5], D3FEND [6], ATLAS [7], FiGHT [8]), upon which LM’s Crow™ system trains multi-agent reinforcement learning (MARL) agents to discover optimal attack sequences aligned with specified mission objectives. This is complemented by Aalto University’s probabilistic attack graph framework that models holistically network topology and exploitation of its vulnerabilities as a directed graph, enabling computation of path- and node-level risk metrics such as expected cumulative impact and inbound and outbound propagation risks. To validate and cross-compare these approaches, both were applied independently to reconstruct the 2015 Ukraine Power Grid cyberattack, as the first documented attack on critical OT infrastructure to cause a large-scale blackout, using a topology rebuilt from publicly available CIRT reporting, which provides a well-documented ground truth. The remainder of the paper is organized as follows. Section 2 describes LM’s Vortex™ attack surface assessment capability and introduces the reinforcement learning framework underlying Crow™. Section 3 formalizes the Aalto attack graph modeling methodology, including node and edge semantics and the probabilistic framework for attack propagation. Section 4 describes the use case of the 2015 Ukraine Power Grid scenario and its reconstruction. Section 5 presents the combined Vortex™/Crow™ analysis and Section 6 presents the attack graph analysis, including baseline impact estimates, vulnerability-neutralization experiments and mitigation prioritization. Section 7 summarizes the key findings of this study and presents a plan for mitigation modeling.
2
ATTACK S URFACE A SSESSMENT W ITH VORTEX™/C ROW™
In cyber incident attack vectors span every layer of a system, from embedded chipsets to network and account credentials, and extend to secondary paths through supporting systems. To address this complexity, LM developed Vortex™, a scalable platform for software attack surface analysis, threat modeling, and mitigation reporting. Vortex™ builds a customized system knowledge graph (Fig. 1) by ingesting system information, i.e. network configuration, Model-Based System Engineering (MBSE) data, Software Bills of Materials (SBOMs), and outputs from scanning tools such as ACAS, Coverity, and Fortify, and mapping them to CVE vulnerabilities, CWE weaknesses, CAPEC attack patterns, and mitigations drawn from recognized sources and frameworks such as NIST SP 800-53B and 800-160 [9]. The knowledge base is updated periodically and is extensible to new sources and formats like MITRE ATT&CK/D3FEND/ATLAS/FiGHT, and SPARTA [10]. Together CVE/CWE/CAPEC form the basis of the attack vectors that a reinforcement learning (RL) agent can apply against a given node in the network topology; drawing cyber effects from the above mentioned authoritative resources is far more efficient and accurate than constructing vulnerabilities and malware from scratch. RL agents are trained in a synthetic environment guided by a reward-bearing objective function. Whereas human operators recognize only familiar attack patterns, LM’s multi-agent reinforcement learning (MARL) framework (Fig. 2) composes attack vectors and TTPs around specified cyber objectives, such as the Deny / Degrade / Disrupt / Deceive / Destroy (5D) effects or the confidentiality-integrity-availability (CIA) triad. Trained over many thousands of episodes within the Vortex™ knowledge graph, the Crow™ agents inherit full traceability and standards compliance, and explainability is built in: each action is evaluated against the objective function during training. The approach has been validated by subject-matter experts from the Air Force Academy, Naval Academy and West Point, and integrated with commercial and government tools including STK, MATLAB, 2
AI-Based Vulnerability Assessment & Attack Graphs
A P REPRINT
EXata, and AFSIM. By automating the otherwise labor-intensive cycle of discovering vulnerabilities, mapping them to exploits and testing attack vectors, Crow™ accelerates the assessment by orders of magnitude while covering the combinatorial attack space more comprehensively than manual analysis.
Figure 1: Vortex™ Cyber Knowledge Graph.
3
Modeling Cybersecurity Defense with Attack Graphs
In cyber incidents, attackers typically compromise a network through an attack path, an ordered sequence of exploits, in which each step satisfies the prerequisites of the next, allowing lateral movement [11] from an entry point deeper into hosts and their applications. The collection of all feasible paths, constrained by reachability and system preconditions, forms an attack graph [12, 13]: a system-level representation that supports estimating an attacker’s likelihood of success and evaluating mitigation strategies under resource constraints. Existing work on attack graph analysis has mainly focused on either network-level or path-based risk assessment. Probabilistic methods derive security scores by aggregating CVSS-based vulnerability measures across the network [14, 15], while graph-based approaches such as [16] use closeness centrality metric and dependency paths to identify Desired actions result in positive rewards, undesired actions return negative rewards – user configures agent what to incentivize.
ACTION At
REWARD Rt ACTION At Environment
Agent
OBSERVATION Ot
Agent can explore the environment for new information or exploit existing information for more rewards.
REWARD Rt ACTION At Environment
Agent
OBSERVATION Ot
Agent
REWARD Rt
Environment
OBSERVATION Ot
Figure 2: Multi-Agent Reinforcement Learning (MARL) Framework. 3
AI-Based Vulnerability Assessment & Attack Graphs
A P REPRINT
worst-case attack paths. Metric suites such as [17] further enable comparisons across enterprises via categorised family scores. While these approaches have advanced network vulnerability analysis, they do not account for per-service or per-host impact, which can be critical information for risk mitigation analysis. Following the methodology of [13], where new attack graph analysis metrics expected cumulative impact (CE[I]), expected nodewise cumulative impact (CN ODE ), expected cumulative outbound impact (COU T ), and expected cumulative inbound impact (CIN ) were introduced, we model each exploit transition as an edge carrying a probabilistic exploitability weight, derived from vulnerability metrics, authentication structure, and observed attacker capabilities, which quantify the likelihood of a successful transition [14]. In this study we apply these four metrics with the following distinctions. First, the attack states that can be reached by exploiting vulnerabilities are represented as software instances present in the network devices (nodes). That is, we assume that exploiting any vulnerability in a software instance would lead to the same state during the attack. While this is a clear simplification of reality, we argue that for a high-level analysis of where to focus one’s defensive efforts, this approach is justified, particularly as the number of possible combinations becomes infeasible to compute when there are tens, if not hundreds of vulnerabilities per software instance. Since exploiting any vulnerability leads to the same state, the directed edge weight representing the likelihood of exploitation between software instances in the attack graph can be calculated as: pE (i, j) = 1 −
Y
[1 − ê(n)]
n∈Vj
ê(n) =
e(n) Emax
(1)
where Vj is the set of all vulnerabilities present in software instance j, Emax is the maximum exploitability value in the scoring system, and ê(n) ∈ [0, 1] is the normalized exploitability score of vulnerability n. Secondly, we assume full network-level connectivity between software instances that have a direct connection in the attack graph. This assumption is made because we do not have realistic data to estimate the weights of the network link (wn ). With this assumption, the probability of traversing a single attack step (originally a product of the exploitability probability pE and the network-level probability pN ) simply becomes p(i, j) = pN (i, j) ∗ pE (i, j) = pE (i, j).
wn = 1 =⇒ pN (i, j) = 1
(2)
The last distinction from the original methodology is how the impact values are derived to compute the expected impact for the cumulative metrics. As all vulnerabilities in a single software instance are aggregated, the impact values need to be aggregated as well. To explore the full range of possible impacts, we have aggregated the impact scores taking the minimum, maximum, and mean across all vulnerabilities in a software instance. The expected impact is then defined as: X E[Iik ] = I k (i) P (j, i), k ∈ {min, mean, max} (3) j∈A
where A is the set of all states immediately preceding the state i. These aggregations are used in all cumulative metrics CE[I], CN ODE , CIN , and COU T . Whereas the methodology in [13] only concentrated on acyclic graphs, in this study there is one cyclic pattern present in the attack graph between the nodes db_server_1 and engineering_ws_1. When calculating the cumulative metrics, we have taken a self-avoiding path approach. That is, only the first visit to a software instance is taken into account. Similarly, with CIN and COU T , only the first time the attack enters or exits a node is taken into account. Following the approach in [13], we utilized the exploitability and impact values of Common Vulnerability Scoring System (CVSS) version 2.0 [18] to estimate the probability and impact of the attack in its various states. Although CVSSv2 is nowadays superseded by newer versions, in 2015 it was the de facto standard, as its successor CVSSv3.0, released in June 2015, had not yet been widely adopted at the time of the attack. Moreover, a significant number of pre-2016 vulnerabilities in our scenario lack the CVSSv3.0 score. IronMiner is built upon a cyber data lake with a wealth of publicly available information about CVEs including threat intelligence reports, the number of exploits available for each CVE, indicators that a CVE is under active exploitation and more. The LM IronMiner team used machine learning and statistical techniques to build a classification model that predicts the riskiest CVEs based on USG and industry reports of active exploitation. Each CVE in the NVD is assigned an IronMiner Threat Score between 1.0-10.0 and will be updated daily. IronMiner identifies and assigns high threat scores to CVEs in CISA’s Known Exploited Vulnerability (KEV) catalog and those that are actively exploited or have other risky attributes detected with the machine learning. This methodology is found to correctly predict 87.3% high threat CVEs in the KEV catalog and other risky vulnerabilities that threat actors have already exploited. IronMiner’s Threat Score is based on researched and repeatable processes that inform defenders of CVE exploitation risks to improve the overall quality of CVE assessment. In addition, IronMiner automatically collects and processes 4
AI-Based Vulnerability Assessment & Attack Graphs
A P REPRINT
new data daily to continuously monitor changes in CVE risk profiles; as such, it provides a higher amount of due diligence, consistency, and precision than manual CVE risk scoring systems. However, as IronMiner does not estimate the impact of vulnerabilities, CVSSv2 was used not only to provide a baseline for comparison but also to supply the impact estimates required by the attack graph analysis.
4
Sample Scenario: 2015 Ukraine Power Grid
To test the combined methodology of Vortex™/Crow™ and Attack Graph against a documented ground truth, we have applied them to the 2015 Ukraine Power Grid attack, reconstructing the attack vectors and performing the analysis from open-source incident reporting, principally the E-ISAC/CIRT report ([19]). On 23 December 2015, a roughly three-hour power outage struck the Ivano-Frankivsk region of Ukraine. Subsequently, malware was found in several substations (7 × 110 kV and 23 × 35 kV), and the outage, which affected up to 225,000 customers served by three regional distribution companies, was attributed to one or more Advanced Persistent Threat (APT) groups that leveraged BlackEnergy and related tools to disrupt the targeted substations simultaneously. The incident is significant as the first confirmed power blackout caused by a cyberattack. The major sequence of events list (MSEL) drawn from the CIRT report was as follows: 1. Email phishing 2. Reconnaissance and enumeration of the network to provide an initial backdoor 3. Discovery and access Microsoft Active Directory® servers that contain corporate user accounts and credentials 4. Use of encrypted tunnel from external network to get inside control system networks 5. Discovery and access SCADA HMI due to an improperly configured firewall 6. Control override of HMI operators and breaker opening command 7. Several other actions with the intent to complicate the responses of control operators 8. KillDisk malware attempting to wipe out the control center HMI and workstations
This sequence of events is mapped onto the representative grid architecture in Fig. 3. Because the configuration of the Prykarpattya Oblenergo OT network was not disclosed in the CIRT report, we reconstructed a simplified topology under the stated hardware and software assumptions (Fig. 4). The left side of Fig. 4 reflects period-appropriate 2015 components — e.g., programmable logic controllers (PLCs) and host controller interfaces (HCIs) — while the right side reflects a more recent configuration, allowing evaluation of whether multiple Crow™ agents select attack vectors comparable to those documented in the real incident under both legacy and modern assumptions.
5
Vortex™/Crow™ Analysis
Using the assumed hardware and software inventory of the reconstructed Ukraine Power Grid topology in Fig. 4, Vortex™ performed a vulnerability assessment that listed the candidate CVEs and CAPECs for each of its components, Fig. 5. These define the action and observation spaces that the Crow™ agents explore to identify exploitable paths. CVE exploitability was scored with LM’s IronMiner, chosen for its objectivity and its incorporation of recent intelligence, though it could be substituted by any other scoring method. Because the 2015 adversary sought to disrupt power generation, the objective chosen for the reinforcement-learning reward used the specific CAPEC impacts of Bypass Protection Mechanism, Gain Privileges, and Execute Unauthorized Commands in that order for this initial pivot into the OT network. Open-source reporting attributes the incident to APT44 (Sandworm). The specific TTPs of this group (G0034) from MITRE ATT&CK were leveraged from Vortex™, and set as the prerequisite for all actions taken by the Crow™ agents. An example of a corresponding set of CVEs, CWEs, and CAPECs, as well as attack patterns for the given system configuration and 5D cyber effects, is shown in this sample set: (’G0034’, ’T1082’, ’CAPEC-313’, ’CWE-200’, ’CVE-2007-2768’) (’G0034’, ’T1539’, ’CAPEC-31’, ’CWE-20’, ’CVE-2013-2811’) (’G0034’, ’T1033’, ’CAPEC-577’, ’CWE-200’, ’CVE-2015-1602’) (’G0034’, ’T1078’, ’CAPEC-560’, ’CWE-522’, ’CVE-2021-40360’)
Trained in the resulting action and observation spaces, Crow™ MARL agents learned a reward-maximizing policy that yields not only viable attack vectors but optimal attack sequences, emulating the behavior of APT44. The two 5
AI-Based Vulnerability Assessment & Attack Graphs
A P REPRINT
Figure 3: 2015 Ukraine Power Grid Cyber Incident. highest-scoring sequences are shown in Fig. 6. Each sequence’s probability of success is the product of the per-step success probabilities of its constituent CVE–CAPEC pairs (scored with IronMiner), and the sequence with the highest overall probability is selected as the most likely course of action. From the initial starting node of this scenario, the Crow™ agents prioritized attacking the node with a software configuration similar to the actual 2015 configuration, sequence 1, emulating the behavior of APT44 in the process. An additional attack sequence 2 learned during training provides an alternative path that APT44 could have taken against a similar device with different software configurations.
6
Attack Graph Analysis
In Table 1 we show the nodewise impact values using CVSSv2 and IronMiner exploitability values. The CN ODE results are consistent across CVSS and IronMiner: in both cases, the nodes simatic_hci_1, db_server_1, and ifix_hci_1 have the highest potential impact. This is expected and explainable by the vicinity of the nodes to the entry point of the attack: the further the step is from the starting point of the attack, the lower its probability and, consequently, its expected impact. Other explaining factors are the high number of vulnerabilities present in the software of these nodes, and the high number of attack steps within node db_server_1 as can be seen in Fig. 4. The slight differences in the CN ODE values between the two scoring systems are found on the right side of the attack graph: results obtained with IronMiner indicate higher impact on nodes scalance_switch_1, siprotec_relay_1, substation_control_1, and engineering_ws_1 than those obtained with CVSS. This is clearly seen in the results represented in Fig. 7, in which the expected cumulative inbound impact per software is estimated using these two scoring systems. Also, the cumulation of impact within a node can be witnessed in Fig. 7 as, for example, the CE[I] of simatic_hci_1 increases when the attack propagates through its software Dropbear SSH, Windows XP, and en100 ethernet module. The CIN and COU T results can be seen in Fig. 8. In these box plots, the results obtained with I min and I max are represented by the low and high ends of the boxes, whereas those obtained with I mean are represented by the middle lines of the boxes. These results are consistent with the CN ODE results, and CVSS and IronMiner produce similar results. The only clear difference between these scoring systems is the slightly higher expected impacts on nodes scalance_switch_1, siprotec_relay_1, and substation_control_1, on the right side of the attack graph and the node engineering_ws_1 when calculated using IronMiner. 6
AI-Based Vulnerability Assessment & Attack Graphs
A P REPRINT
windows_8_host_1 START windows_8 8
simatic_hci_1 Windows XP Professional SP-2
Siemens Simatic PCS_7
ethernet_ module en100
moxa_converter_1
SQL Server 2014 Sinema Server
Simatic Step 7 SP1
ethernet_ module en100
Moxa nport Firmware 5100a
Sinema Remote Connect
OpenSSH 6.0
.net Framework 3.5
Windows XP Professional SP-2
Scada iFix 5.0
Windows 7 x64 SP1
Simatic Step 7 SP1
Simatic Wincc 7.0
multiprog_plc_1
powerlogic_scada _7_10 7.10
Telecontrol Basic SP2 Scalance X300eec
Scalance X-300
substation_control_1
siprotec_relay_1
ethernet_ module en100
Modbus Serial Driver 1.10
scalance_switch_1
Sinema Remote Connect
Simatic S7-1500 CPU Firmware
powerlogic_rtu_1
Windows Server 2012
engineering_ws_1
plc_controller_1
ethernet_ module en100
Moxa nport 5100a
ifix_hci_1
db_server_1
dropbear_ ssh_project
Telecontrol Basic SP2
Siemens dnp3 siemens_siprotec_ firmware_4_24 4.24
phoenixcontact_software_ multiprog_5 5.0
Siemens Sicam ak
Siemens dnp3
END
Figure 4: Assumed Power Grid and its ICS / OT Network Topology. In order to find out the significance of each node in the network, we conducted a series of analyses in which the exploitability of all vulnerabilities associated with a given node was set to zero to mimic fully securing the node. This process was repeated for each node individually, and the results of each analysis run were compared against the baseline to identify the most impactful nodes in the network. In Fig. 9 we show the resulting box plot that illustrates the decrease in overall attack impact (CE[IEN D ]) when the vulnerabilities of the individual nodes are neutralized. Those nodes that lead to a larger reduction in CE[I] can be interpreted as critical points in the attack graph, as they lie on multiple
111
windows_8_host_1 7 6
substation_control_1 siprotec_relay_1
81 64
26
simatic_hci_1 scalance_switch_1
160
9 1 3
powerlogic_rtu_1 plc_controller_1 multiprog_plc_1
116 13
131
8
39
14
moxa_converter_1
196
45 40
ifix_hci_1 engineering_ws_1
171
117 22
db_server_1 0
20
191
94 40
60
80
CVE Count
100
120
140
CAPEC Count
Figure 5: Vulnerability Assessment Results from Vortex™. 7
160
180
200
AI-Based Vulnerability Assessment & Attack Graphs
A P REPRINT
Figure 6: Crow™ results: two most probable or optimal attack sequences. Table 1: Vulnerability count and CN ODE per device. CNODE Device CVEs CVSS IronMiner I min I mean I max I min I mean I max db_server_1 34 28.4 47.2 67.9 28.0 46.9 68.3 engineering_ws_1 170 3.9 9.5 13.8 8.5 20.9 30.3 ifix_hci_1 83 15.2 23.2 32.3 16.0 25.6 35.6 moxa_converter_1 18 2.0 3.0 3.6 2.2 3.5 4.1 multiprog_plc_1 10 6.9 9.0 10.9 4.8 6.2 7.5 plc_controller_1 25 6.6 14.2 22.4 5.0 11.0 17.2 powerlogic_rtu_1 3 3.1 3.7 4.4 3.7 4.6 5.5 scalance_switch_1 13 0.8 0.9 1.0 6.2 7.3 8.0 simatic_hci_1 48 12.1 20.8 32.7 12.0 20.8 32.7 siprotec_relay_1 9 0.0 0.0 0.0 1.3 1.3 1.5 substation_control_1 10 0.0 0.0 0.0 1.9 2.0 2.2
high-impact attack paths and may serve as pivotal points for attackers. Clearly, the node simatic_hci_1 on the left side of the attack graph is the most critical node for both CVSS and IronMiner scoring systems. The next most impactful nodes are ifix_hci_1 and plc_controller_1, while securing the rest of the nodes results in moderate reductions. In particular, since db_server_1 and engineering_ws_1 are not on any of the attack paths that lead to the END node, neutralizing their vulnerabilities has no effect on the overall impact of the attack. The most notable difference between the scoring systems is that IronMiner tends to have a higher reduction on the right side of the attack graph than CVSS. Table 2 shows the reduction in CIN and COU T after neutralizing the vulnerabilities of a node. In contrast to the CE[I] results, db_server_1 stands out for both measures due to the high number of attack paths and vulnerabilities, as established in the baseline. Compared to the results in Fig. 8, neutralizing engineering_ws_1 has little effect, implying that most of the high baseline CIN and COU T have cumulated before reaching the node. Unsurprisingly, the neutralization of the vulnerabilities of simatic_hci_1 and ifix_hci_1 has the most widespread effects, as they are the pivot points for the left and right sides of the attack graph, respectively. Securing either of these nodes reduces all CIN and COU T of the downstream nodes on the attack paths towards the END node, while also partially reducing the impact on db_server_1 and engineering_ws_1. Securing the nodes downstream of these pivotal points show diminishing results the further away they are from the start of the attack due to the impact already having cumulated before reaching these nodes. Moreover, consistent with the baseline and CE[IEN D ] results, securing the peripheral nodes on the right side of the attack graph has no effect as measured by CVSS, while with IronMiner the reductions are minor but present.
7
Concluding Remarks
Here, we have reported a study using LM’s Vortex™/Crow™ integrated knowledge graph & multi-agent reinforcement learning framework and Aalto’s probabilistic Attack Graph modeling approach to identify and prioritize attack vectors and their most likely sequences among all possible ones as well as to determine system-level risk metrics of vulnerabilities / exploits in OT networks. To compare these two approaches independently, we used the 2015 Ukraine Power Grid cyberattack as a well-documented realistic validation scenario. 8
AI-Based Vulnerability Assessment & Attack Graphs
Scoring system CVSS IronMiner
120 100 80 60 40
Siemens_Simatic_PCS_7 Windows_XP_Professiona... dropbear_ssh_project ethernet_module_en100 .net_Framework_3.5 SQL_Server_2014 Sinema_Server Windows_Server_2012 OpenSSH_6.0 Scada_iFix_5.0 Sinema_Remote_Connect Windows_XP_Professiona... Moxa_nport_5100a Moxa_nport_Firmware_5100a ethernet_module_en100 Simatic_S7-1500_CPU_Fi... Simatic_Step_7_SP1 ethernet_module_en100 Simatic_Step_7_SP1 Simatic_Wincc_7.0 Sinema_Remote_Connect Windows_7__x64_SP1 Scalance_X-300 Scalance_X-300eec Telecontrol_Basic_SP2 Modbus_Serial_Driver_1.10 powerlogic_scada_7_10_... ethernet_module_en100 phoenixcontact_softwar... Siemens_dnp3 siemens_siprotec_firmw... Siemens_Sicam_ak Siemens_dnp3 Telecontrol_Basic_SP2
20 0
A P REPRINT
Systems simatic_hci_1 db_server_1 ifix_hci_1 moxa_converter_1 plc_controller_1 engineering_ws_1 scalance_switch_1 powerlogic_rtu_1 multiprog_plc_1 siprotec_relay_1 substation_control_1
Figure 7: Expected cumulative impact (CE[I]) per software. The box plots show results obtained with minimum, mean, and maximum vulnerability impact values as the low end, midline, and high end, respectively. One of the findings is that Vortex™/Crow™ and Attack Graph frameworks produced similar results; both approaches converged on the same attack vectors and exploit sequences as documented in the incident record of the Ukraine Power Grid attack in 2015, thus providing mutual cross-validation. This indicates that both holistic methodologies are useful in predicting future attacks in other scenarios or ICS / OT networks. Another observation that can be drawn from this study is that the CVSS and IronMiner scores produced similar results. IronMiner seems also to give actionable information. It should be noted that the version of CVSS used in this investigation was v2.0 which fits with the 2015 scenario timeframe, while IronMiner is based on a present-time assessment. In the future, complicated challenges still exist; given the large number of potential attack vectors and attack sequences, how do we decide which vulnerabilities we can mitigate, given the constraints of time, cost and desired efficacy or reduction of risk? As a next step, the joint LM and Aalto team aims at proposing a holistic system-based Mitigation Prioritization Model to provide a solution. This will be done by developing a framework to optimize mitigation based on the desired efficacy, time, and cost constraints.
References [1] MITRE. Common Vulnerabilities and Exposures (CVE). https://cve.mitre.org, 1999. Accessed: 2026-0601. [2] MITRE. Common Weakness Enumeration (CWE). https://cwe.mitre.org. Accessed: 2026-06-01. [3] MITRE. Common Attack Pattern Enumeration and Classification (CAPEC). https://capec.mitre.org. Accessed: 2026-06-01. [4] National Institute of Standards and Technology. Security and Privacy Controls for Information Systems and Organizations. Technical Report SP 800-53B, NIST, 2020. [5] MITRE. MITRE ATT&CK. https://attack.mitre.org. Accessed: 2026-06-01. [6] MITRE. MITRE D3FEND. https://d3fend.mitre.org. Accessed: 2026-06-01. [7] MITRE. MITRE ATLAS. https://atlas.mitre.org. Accessed: 2026-06-01. [8] MITRE. MITRE FiGHT. https://fight.mitre.org. Accessed: 2026-06-01. [9] National Institute of Standards and Technology. Engineering Trustworthy Secure Systems. Technical Report SP 800-160 Vol. 2 Rev. 1, NIST, 2021. 9
AI-Based Vulnerability Assessment & Attack Graphs
120
CVSS IronMiner
100
A P REPRINT
CVSS IronMiner
100
80
80 60 60 40
20
0
0
i_1
ser
_hc
db_
sim atic
ser
db_
sim atic
_hc
i_1 ver _1 i f i x mo xa_ _hci_1 con ver ter_ plc_ 1 con t rolle eng r_ inee ring 1 sca lanc _ws_1 e_s pow witch_ 1 erlo gic_ rtu_ mu 1 ltipr og_ p sipr lc_1 ote c_r sub elay stat _1 ion_ con trol _1
20
ver _1 i f ix_h mo c xa_ con i_1 ver ter_ plc_ 1 con t r olle eng r_ inee ring 1 sca lanc _ws_1 e_s pow witch_ 1 erlo gic_ rt mu ltipr u_1 og_ plc_ sipr 1 ote c sub _re lay_ stat ion_ 1 con trol _1
40
(a)
(b)
Figure 8: (a) CIN and (b) COU T per node. The box plots show results obtained with minimum, mean, and maximum vulnerability impact values as the low end, midline, and high end, respectively. [10] The Aerospace Corporation. Space Attack Research and Tactic Analysis (SPARTA). https://sparta. aerospace.org. Accessed: 2026-06-01. [11] Christos Smiliotopoulos, Georgios Kambourakis, and Constantinos Kolias. Detecting lateral movement: A systematic survey. Heliyon, 10(4):e26317, 2024. [12] Harjinder Singh Lallie, Kurt Debattista, and Jay Bal. A review of attack graph and attack tree visual syntax in cyber security. Computer Science Review, 35:100219, 2020. [13] Joni Herttuainen, Vesa Kuikka, and Kimmo K. Kaski. Integrating network and attack graphs for service-centric impact analysis. IEEE Access, pages 1–1, 2026. [14] Marcel Frigault, Lingyu Wang, Sushil Jajodia, and Anoop Singhal. Measuring the Overall Network Security by Combining CVSS Scores Based on Attack Graphs and Bayesian Networks, pages 1–23. Springer International Publishing, Cham, 2017. [15] John Homer, Su Zhang, Xinming Ou, David Schmidt, Yanhui Du, S Raj Rajagopalan, and Anoop Singhal. Aggregating vulnerability metrics in enterprise networks using attack graphs. Journal of Computer Security, 21(4):561–597, 2013. [16] George Stergiopoulos, Panagiotis Dedousis, and Dimitris Gritzalis. Automatic analysis of attack graphs for risk mitigation and prioritization on large-scale and complex networks in industry 4.0. International Journal of Information Security, 21(1):37–59, 2022. [17] Steven Noel and Sushil Jajodia. A Suite of Metrics for Network Attack Graph Analytics, pages 141–176. Springer International Publishing, Cham, 2017. [18] Peter Mell, Karen Scarfone, and Sasha Romanosky. A complete guide to the common vulnerability scoring system version 2.0. https://www.first.org/cvss/v2/guide, 2007. Accessed: 2026-06-01. [19] E-ISAC and SANS ICS. Analysis of the cyber attack on the Ukrainian power grid: Defense use case. Technical report, Electricity Information Sharing and Analysis Center (E-ISAC) and SANS Industrial Control Systems, March 2016.
10
AI-Based Vulnerability Assessment & Attack Graphs
60
A P REPRINT
CVSS IronMiner
50 40 30 20 10
ver _1 i f i x mo xa_ _hci_1 con ver ter_ plc_ 1 con trol eng inee ler_1 ring sca lanc _ws_1 e_s pow witch_ 1 erlo gic_ rt mu ltipr u_1 og_ plc_ sipr 1 ote c_r sub e lay_ stat ion_ 1 con trol _1
ser
db_
sim
atic
_hc
i_1
0
Figure 9: Reduction in expected cumulative impact of a successful attack (CE[I]END) after eliminating all vulnerabilities of a node. The box plots show results obtained with minimum, mean, and maximum vulnerability impact values as the low end, midline, and high end, respectively.
11
AI-Based Vulnerability Assessment & Attack Graphs
A P REPRINT
Table 2: Reduced CIN and COU T per node after securing a node Node legend: A = simatic_hci_1, B = db_server_1, C = ifix_hci_1, D = moxa_converter_1, E = plc_controller_1, F = engineering_ws_1, G = scalance_switch_1, H = powerlogic_rtu_1, I = multiprog_plc_1, J = siprotec_relay_1, K = substation_control_1 CIN COUT Affected Secured CVSS IronMiner CVSS IronMiner Node Node I min I mean I max I min I mean I max I min I mean I max I min I mean I max A 1.4 4.5 10.0 1.4 4.5 10.0 12.4 21.9 33.4 12.3 22.0 33.5 B 1.0 1.0 1.0 0.7 0.7 0.7 A D 1.0 1.0 1.0 1.0 1.0 1.0 E 1.3 3.6 5.7 1.4 3.8 5.9 A 8.8 13.7 21.0 8.5 13.4 20.7 8.8 13.7 21.0 8.6 13.6 21.0 B 17.3 25.2 34.1 15.2 22.8 32.6 23.6 37.7 56.4 21.9 37.0 56.9 B C 15.6 25.2 35.9 16.3 28.1 41.3 15.6 25.2 35.9 16.4 28.4 41.8 F 0.5 1.7 2.3 1.4 4.3 5.8 B 7.3 10.5 13.5 7.2 10.9 15.7 C C 1.4 4.8 10.0 1.4 4.8 10.0 15.9 25.5 36.1 18.8 30.6 43.8 G 0.3 0.3 0.3 2.5 2.5 2.5 A 10.1 17.3 26.7 10.2 17.5 26.9 12.6 20.5 30.5 13.4 21.5 31.8 D D 1.0 1.0 1.0 1.0 1.0 1.0 3.4 4.1 4.8 4.2 5.0 5.9 H 2.0 2.6 3.3 2.6 3.4 4.3 A 10.5 20.0 31.4 10.6 20.2 31.8 14.4 29.2 45.7 13.2 26.5 41.5 E E 1.3 3.6 5.7 1.4 3.8 5.9 5.2 12.8 19.9 4.0 10.0 15.6 I 1.2 3.3 5.2 0.9 2.3 3.6 A 8.8 13.7 21.0 8.6 13.6 21.0 8.8 13.7 21.0 8.6 13.6 21.0 B 23.6 37.7 56.4 21.9 37.0 56.9 27.2 44.4 64.1 29.2 51.3 73.4 F C 15.6 25.2 35.9 16.4 28.4 41.8 15.6 25.3 35.8 16.6 28.7 42.1 F 0.5 1.7 2.3 1.4 4.3 5.8 4.1 8.4 10.0 8.6 18.6 22.3 C 8.1 13.3 20.3 10.3 15.5 22.5 8.2 13.6 20.7 12.4 18.7 26.4 G G 0.3 0.3 0.3 2.5 2.5 2.5 0.4 0.6 0.7 4.6 5.7 6.4 K 1.0 1.0 1.0 A 12.6 20.5 30.5 13.4 21.5 31.8 13.7 21.6 31.7 14.5 22.7 33.0 H D 3.4 4.1 4.8 4.2 5.0 5.9 4.6 5.2 5.9 5.3 6.2 7.1 H 2.0 2.6 3.3 2.6 3.4 4.3 3.1 3.7 4.4 3.7 4.6 5.5 A 14.4 29.2 45.7 13.2 26.5 41.5 20.0 34.8 51.4 17.2 30.4 45.5 I E 5.2 12.8 19.9 4.0 10.0 15.6 10.9 18.4 25.6 8.0 13.9 19.6 I 1.2 3.3 5.2 0.9 2.3 3.6 6.9 9.0 10.9 4.8 6.2 7.5 C 8.2 13.6 20.7 13.7 20.2 28.3 8.3 13.6 20.7 14.5 21.0 29.1 G 0.4 0.6 0.7 5.9 7.2 8.3 0.5 0.6 0.7 6.7 8.0 9.1 J J 0.4 0.5 0.7 1.2 1.4 1.5 K 2.3 2.5 2.9 3.2 3.3 3.7 C 8.2 13.6 20.7 12.4 18.7 26.4 8.2 13.6 20.7 13.7 20.2 28.3 G 0.4 0.6 0.7 4.6 5.7 6.4 0.4 0.6 0.7 5.9 7.2 8.3 K J 0.4 0.5 0.7 K 1.0 1.0 1.0 2.3 2.5 2.9
12