Conceptio › Archive › arXiv CS
arXiv CSopen access

A Framework to Quantify the Probability of Future Cyber Loss Events

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

arXiv:2609.21717v1 [cs.CR] 18 Sep 2026

A Framework to Quantify the Probability of Future Cyber Loss Events Siem Peters

Martin Eian

Mnemonic Oslo, Norway e-mail: [email protected]

Mnemonic Oslo, Norway e-mail: [email protected]

Abstract—Cybersecurity risk quantification remains challenging due to limited operational data and difficulties in quantifying Loss Event Frequency (LEF). This paper introduces the Loss Event Frequency Security Analyser (LEFSA), a probabilistic framework that reformulates LEF estimation as machine-level Cyber Loss Event (CLE) prediction combined with hierarchical infrastructure-level aggregation. LEFSA estimates calibrated machine-level CLE probabilities from operational cybersecurity telemetry and aggregates them across infrastructure layers while accounting for machine-level dependencies. This provides a foundation for scalable, explainable, and operationally applicable cyber risk estimation at the level of machines, services, business processes, and the entire organization. The framework was evaluated using proprietary Managed Detection & Response telemetry from 23 organizations using Microsoft Defender for Endpoint. XGBoost achieved the strongest predictive performance, with a mean area under the receiver operating characteristic curve of 0.90 and consistently low calibration error across evaluation periods. The results demonstrate that operational cybersecurity telemetry contains substantial predictive information for future CLE occurrence, supporting probabilistic machine-level modeling and hierarchical aggregation as a promising foundation for quantitative, data-driven cyber risk management. Keywords-cyber risk management; cybersecurity; loss event frequency; XGBoost; probabilistic modeling.

I. I NTRODUCTION The number of cyber incidents has increased significantly in recent years, with malicious attacks nearly doubling compared to pre-pandemic levels [1]. Industry data further indicate an increasing frequency and severity of Cyber Loss Events (CLEs), particularly for large losses [2][3], while reported cyber insurance claims increased by roughly 40% year-overyear in 2024 [4]. Recent advances in AI are accelerating the cyber attack lifecycle, from vulnerability discovery to exploitation, increasing the need for proactive and data-driven cybersecurity risk management frameworks [5]. Cybersecurity risk refers to the potential for financial loss, operational disruption, or reputational damage arising from failures in an organization’s Information Technology (IT) systems caused by cyberattacks [6]. While modeling the monetary impact of CLEs has been widely studied, particularly from an insurance perspective [7][8], estimating the likelihood of a CLE affecting IT infrastructure remains largely qualitative in both academia and industry. A. State-of-the-art Established information security risk-assessment frameworks and methodologies include ISO/IEC 27005, NIST SP

800-30, OCTAVE, and CORAS, which provide structured processes for identifying, analysing, and treating cyber risks [9]– [13]. While these approaches offer well-established guidance for risk assessment, they do not prescribe a single probabilistic or data-driven method for estimating loss frequencies or financial loss magnitudes from operational data. This study therefore builds on Factor Analysis of Information Risk (FAIR) [14], one of the most established quantitative cyber-risk frameworks. FAIR is particularly suitable because it expresses cyber risk quantitatively through two components, namely the frequency and monetary impact of loss events associated with defined cyber-risk scenarios, combined representing the cyber risk of an organization. The advantages of FAIR are (1) a structured and consistent methodology, (2) quantitative (financially grounded) output, (3) transparency, and (4) compatibility with probabilistic methods. The Loss Event Frequency (LEF) component of FAIR is the estimated rate at which a specific threat scenario is expected to result in a loss event over a given time period. It represents the likelihood dimension of risk and is calculated as the combination of threat event frequency and susceptibility. In operational settings, modeling the LEF component often proves challenging due to (1) limited relevant and highquality scenario-specific data and (2) difficulties in operationalizing and estimating contact frequency and susceptibility. Moreover, while FAIR defines the relevant factors and their conceptual relationships, it does not prescribe a complete empirical procedure for acquiring organization-specific data, selecting and validating probability distributions, or learning these relationships from operational telemetry. Related work has extended FAIR using Bayesian networks to support more flexible distributions, dependencies, and model structures [15]. Although this relaxes some modelling restrictions within the FAIR framework, it still relies on scenario-level inputs that may be unavailable or difficult to estimate and validate from operational data. Many existing quantitative cyber-risk prediction approaches operate at the organization level using externally observable characteristics and public data [16][17]. Leslie et al. [18] instead used aggregated internal Managed Security Service Provider (MSSP) data to predict organization-level intrusion counts. Although these studies demonstrate that cyber incidents are predictable, organization-level models provide limited insight into which individual machines or infrastructure components generate the estimated risk. A more operationally applicable approach is to estimate fu-

ture CLE probabilities directly at the machine level. Prior work shows that endpoint behaviour, software configuration, and file-appearance patterns can accurately predict future malware infections using internal telemetry [19]–[21]. The reported predictive performance of these machine-level approaches is generally stronger than that of organization-level models, although differences in data and prediction tasks prevent direct comparison. However, they focus primarily on malware infection and do not provide a framework for predicting general CLEs, identifying their machine-level sources, and aggregating risk across all IT infrastructure layers. B. Contributions This paper proposes Loss Event Frequency Security Analyser (LEFSA), a data-driven framework to estimate the likelihood of CLEs through machine-level CLE prediction and hierarchical infrastructure-level aggregation. Rather than modeling isolated machine-level outcomes or manually constructed threat scenarios, LEFSA estimates calibrated machine-level CLE probabilities and aggregates them across infrastructure layers. This enables probabilistic cyber risk estimation at the level of machines, services, business processes, and entire organization. The main contributions of this paper are as follows: • A reformulation of LEF estimation as a measurable machine-level probabilistic prediction problem, extending existing machine-level malware prediction approaches to generalized CLE estimation. • A hierarchical cyber risk aggregation framework that propagates machine-level CLE probabilities across infrastructure layers. This would enable cyber risk estimation at all IT infrastructure levels, like machines, services, business processes, or the entire organization. • A proposed probabilistic aggregation methodology for modeling predictive distributions of aggregated CLEs. This enables estimation of not only future expected CLEs, but also tail-risk outcomes and worst-case infrastructure-level cyber risks. • A discussion of both independent and correlated aggregation settings, including dynamic dependence structures and potential copula-based approaches for modeling joint CLE distributions. • An empirical evaluation using real-world cybersecurity data from multiple organizations. C. Limitations The scope of this work is limited to the estimation of LEF and does not include modeling the loss magnitude of cyber risk. The empirical evaluation uses proprietary Managed Detection & Response (MDR) data that cannot be publicly shared; the dataset and experimental setting are described in Section III. Its empirical evaluation is limited to organizations covered by this dataset, and generalization to other environments depends on the availability and comparability of machine-level telemetry and alerts generated by intrusion detection platforms. The current aggregation procedure relies

on simplifying assumptions regarding dependencies between machines. The remainder of this paper is organized as follows: Section II describes the LEFSA framework for predicting and aggregating CLEs. Section III presents the experiments and results. Section IV discusses the results, implications, and limitations, while Section V concludes the paper and outlines future work. II. LEFSA: A G ENERALIZED L OSS E VENT F RAMEWORK A. Definitions There are T time periods, indexed by t ∈ {1, . . . , T }. Let M denote the set of all machines, where each machine has a unique identifier. At each time t, only a subset Mt ⊆ M is observed. The set of observed machines Mt may vary over time as machines are added, removed, or temporarily unobserved. M is expanded whenever a previously unseen machine appears. For each machine m ∈ Mt , let Xt,m ∈ {0, 1} denote the random variable indicating whether machine m is involved in a CLE at time t. A CLE should be defined using measurable outcomes available in the data, such as malware detections or security incidents. Ft denotes the information set, i.e., the features available up to and including period t. Organizations monitor a set of machines over time. This can be represented by the following matrix:     X1,1 X1,2 · · · X1,M 1 − 0 ···  X2,1 X2,2 · · · X2,M      X1:T =  . .. ..  = 0 1 0 · · ·  ..  .. .. .. .. . . . . .  . . . . XT,1 XT,2 · · · XT,M where rows correspond to time periods t = 1, . . . , T and columns correspond to machines m = 1, . . . , M . Missing entries indicate machines that were not observed during a given time period, reflecting that the observed machine set Mt may vary over time. For infrastructure level g, define the aggregated random P (g) variable Yt = m∈Mt,g Xt,m , where Mt,g denotes the set (g)

of machines at level g. Then Yt represents the number of machines at infrastructure level g affected by at least one CLE during period t. B. Machine-level LEF For every machine m ∈ Mt , the conditional probability that machine m will be involved in a CLE during the next time period t + 1 is denoted as pt+1,m := P(Xt+1,m = 1|Ft ). To predict the probability of a future CLE on a single machine, a model is trained to estimate pt+1,m . The conditional probability is approximated by a model fθ̂ that maps the information set Ft to the unit interval, i.e., fθ̂ : Ft → [0, 1], yielding the estimate p̂t+1,m := fθ̂ (Ft ). To aggregate machine-level predictions, LEFSA requires calibrated probability estimates satisfying: P(Xt+1,m = 1|p̂t+1,m = p) = p ∀p ∈ [0, 1],

(1)

Such estimates may be obtained either directly from probabilistic models or, when necessary, through post-hoc calibration of score-based models.

1) Probabilistic models: Inherently probabilistic models, such as logistic regression, directly estimate the mapping fθ̂ through likelihood-based inference. Under suitable conditions, these models may provide well-calibrated probability estimates. 2) Score-based models: More flexible machine learning models can offer improved predictive performance in complex or high-dimensional settings, although calibration may still need to be assessed empirically [22]. If the chosen model outputs a risk score rather than a calibrated probability satisfying (1), the probability mapping fθ̂ is implemented as a composition of two components: 1) Let hθ̂ denote the score-producing model, outputting a risk score: ŝt+1,m := hθ̂ (Ft ) 2) A calibration map v is then applied, where v : ŝ 7−→ [0, 1], yielding p̂t+1,m := v(ŝt+1,m ). Various calibration methods for v are possible, such as Platt scaling [23] or Isotonic regression [24]. Machine-level LEF estimation involves: (1) defining a CLE, (2) constructing the information set Ft , and (3) training and, if necessary, calibrating a model to estimate p̂t+1,m ∀m ∈ Mt . C. Infrastructure-level LEF The second layer of LEFSA aggregates machine-level CLE probabilities across higher infrastructure levels, yielding aggregate predictive distributions over the number of CLEs within groups of machines. Choose an infrastructure level of interest, e.g., a service or business process. Select the set of machines m that support the chosen infrastructure level g, defined as m ∈ Mt,g . The interest lies in modeling the future risk of a group of assets, P (g) i.e., Yt+1 = m∈Mt,g Xt+1,m . 1) Independent machines: Although unlikely, the machinelevel CLE variables may be conditionally independent given Ft , i.e., Cov(Xt+1,m , Xt+1,m′ | Ft ) = 0 for m ̸= m′ . In (g) that case, Yt+1 follows a Poisson-binomial distribution with parameters {p̂t+1,m }m∈Mt,g . Note that the subscript of M is t not t + 1, since machine-level predictions are made for every running machine in the current time period t. The Lindeberg–Feller Central Limit Theorem (CLT) for independent but non-identically distributed Bernoulli variables implies that (g) Yt+1 − µ̂t+1 ⇒ N (0, 1), σ̂t+1 provided no single machine dominates the total variance [25]. Therefore, a convenient P approximation is to assume normality 2 with mean µ̂ = t+1 m∈Mt,g p̂t+1,m and variance σ̂t+1 = P m∈Mt,g p̂t+1,m (1 − p̂t+1,m ). Note that since CLEs are rare (see Section III), i.e., (g) p̂t+1,m ≪ 1, the Poisson-binomial distribution of Yt+1 can alternatively P be approximated by a Poisson distribution with rate λt+1 = m∈Mt,g p̂t+1,m .

2) Correlated machines: Machine-level CLEs may exhibit conditional dependence [26], i.e., Cov(Xt+1,m , Xt+1,m′ | Ft ) ̸= 0 for m ̸= m′ . In this case, it would require modeling the joint dependence structure between machines. More generally, what must be modeled is the joint conditional distribution of the machine-level CLE vector Xt+1 = (Xt+1,m : m ∈ Mt,g ), given the available information up to time t: Ft+1 (x) = P (Xt+1 ≤ x | Ft )

(2)

It contains both the marginal risk of each machine and the dependence structure between machines. The corresponding aggregate distribution is X (g) P(Yt+1 = y | Ft ) = P (Xt+1 = x | Ft ) (3) P x: m∈M

t,g

xm =y (g)

This shows that the distribution of Yt+1 is fully determined by the joint distribution Ft+1 . In this case, accurate aggregation requires not only predicting the machine-level future marginal CLE probabilities p̂t+1,m , but also the dependence structure between machines. Additionally, the data might exhibit timevarying dependence, i.e., Xt+1 | Ft ∼ Ft+1 where Ft+1 denotes a time-varying joint distribution. Modeling dependence between machine-level CLEs is nontrivial, particularly for discrete and time-varying outcomes. One possible approach is to use copulas, which couple marginal distributions into a joint multivariate distribution [27]. Copula-based approaches, including static vine copulaGARCH formulations and vine copula models with temporal and cross-group dependence structures, have been applied to cybersecurity risk data [28][29]: Ft+1 (x1 , . . . , xd ) = Ct+1 (Ft+1,1 (x1 ), . . . , Ft+1,d (xd )) , (4) where Ct+1 captures machine-level dependence and d = |Mt,g |. Time-varying dependence may be incorporated through dynamic copula parameters [30]: Ft+1 (x1 , . . . , xd ) = Ct+1 (Ft+1,1 (x1 ), . . . , Ft+1,d (xd ); θt+1 ) (5) with θt+1 = h(θt , Zt , εt+1 ), (6) where Zt may include cyber or infrastructure information such as vulnerability disclosures, patch status, threat activity, or network changes. For short observation windows, CLEs may be modeled as binary variables Xt,m ∈ {0, 1}, while longer windows may instead use count outcomes Xt,m ∈ N0 . Although Sklar’s theorem guarantees a copula representation, uniqueness generally fails for discrete marginals, complicating both dependence interpretation and statistical inference [31]. In practice, dependence modeling with discrete marginals is often handled through approximation-based approaches,

Organization Mt,3 = Mt P (3,b) (b) Ỹt+1 = m∈Mt X̃t+1,m (3,b) B (3) {Ỹt+1 }b=1 ⇒ Yt+1 b (3) > k), VaR, ES P(Y t+1 Business process 1 Mt,1 = {M1 , M2 , M3 } P (1,b) (b) Ỹt+1 = i∈Mt,1 X̃t+1,i (1,b)

Business process 2 Mt,2 = {M4 , M5 } P (2,b) (b) Ỹt+1 = i∈Mt,2 X̃t+1,i (2,b)

(1)

{Ỹt+1 }B b=1 ⇒ Yt+1

M1 p̂t+1,1 := fθ̂ (Ft ) (b) X̃t+1,1

M2 p̂t+1,2 := fθ̂ (Ft ) (b) X̃t+1,2

(2)

{Ỹt+1 }B b=1 ⇒ Yt+1

M3 p̂t+1,3 := fθ̂ (Ft ) (b) X̃t+1,3

M4 p̂t+1,4 := fθ̂ (Ft ) (b) X̃t+1,4

M5 p̂t+1,5 := fθ̂ (Ft ) (b) X̃t+1,5

Joint simulation of machine outcomes (b) (b) (b) (b) X̃t+1 ∼ Fbt+1 , X̃t+1 = (X̃t+1,1 , . . . , X̃t+1,5 )

Figure 1. LEFSA aggregation framework. Machine-level CLE probabilities are combined through an estimated joint distribution and aggregated across infrastructure levels to obtain predictive risk distributions.

including latent continuous-variable constructions, Inference Functions for Margins (IFM), and simplified pair-copula or low-rank structures [32]–[34]. Extending these methods to dynamic settings remains an active research area and is discussed further in Section IV. Given marginal probabilities {p̂t+1,m }m∈Mt,g and an estimated joint distribution Fbt+1 , the aggregate distribution of (g) Yt+1 can be approximated with Monte Carlo simulation. For b = 1, . . . , B, generate a correlated  machine-level realization  (b) (b) (b) X̃t+1 ∼ Fbt+1 , where X̃t+1 = X̃t+1,m : m ∈ Mt,g . The corresponding aggregate realization is then (g,b)

Ỹt+1 =

X

(b)

X̃t+1,m .

m∈Mt,g

Repeating this procedure yields an empirical approximation of (g) the predictive distribution of Yt+1 . Any higher-level aggregation, such as services, business processes, or organization-level risk, is performed directly from the estimated joint machinelevel distribution Fbt+1 , thereby preserving the dependence structure across all machines and infrastructure layers. 3) Risk metrics: Having estimated the predictive distri(g) bution of Yt+1 , infrastructure-level cyber risk metrics can be computed. For example, if a business process requires four operational machines and maintains two backup servers, the probability of service disruption during period t + 1 is (g) (g) P(Yt+1 > 2 | Ft ). More generally, the distribution of Yt+1 enables estimation of Value-at-Risk (VaR), Expected Shortfall (ES), and other tail-risk measures (expressed in the number of CLEs). See Figure 1 for a complete overview of LEFSA. III. E XPERIMENTS & R ESULTS The LEFSA framework was evaluated using data from Mnemonic’s MDR service. Mnemonic AS is a Norwegian cyber security company, providing MDR and cyber risk services. Due to confidentiality agreements, the dataset cannot be

released publicly. The empirical evaluation focuses primarily on machine-level LEF, while modeling time-varying machine dependencies for infrastructure-level LEF is left for future work. 1) Data: The dataset covers 110,924 unique machines from 23 organizations, all using Microsoft Defender for Endpoint (MDE) [35]. The data span 12 weeks, although not every machine is observed in every time period. The dataset consists of MDE telemetry, alerts generated by Mnemonic’s detection systems, and incident analysis results produced by Mnemonic’s Security Operations Center (SOC). The target variable is binary and indicates whether a machine was reported in a security incident. In total, the dataset contains 952,770 machine-period instances, representing unique combinations of machines and time periods. Of these, 0.134% of the machines were reported in a security incident. 2) Features: We developed a pipeline to deploy LEFSA as a live service across 23 organizations. A preprocessing script transformed each observed machine into hourly machine profiles, represented as vectors of counts that capture machine behaviour and configuration during each hour. Following each time period, the hourly machine profiles were aggregated into a single feature vector per machine. The selected features can be grouped into four categories, ordered from highest to lowest data availability: exposure, vulnerabilities, security controls, and threats. These are denoted by Ft = (E, V, SC, T). In addition, the one-period lagged outcome Xt,m is included as a predictor to capture temporal dependence. For machines inactive or not yet observed during period t, the lagged value is set to zero. Each training and testing instance consists of machine feature vectors during week t and the corresponding target variable observed during the subsequent week t + 1. Consequently, 12 weeks of observations yield 11 consecutive rollingwindows. The objective is to predict the probability of a CLE for machine m during the subsequent time period conditional on the observed feature set, i.e., P(Xt+1,m = 1 | Ft ). 3) Model training & tuning: The first two rolling-windows were used for model selection and hyperparameter tuning. Logistic Regression, Random Forest, Histogram Gradient Boosting, and Extreme Gradient Boosting (XGBoost) were trained and evaluated to identify the best-performing model. Among these, XGBoost achieved the strongest performance. XGBoost is a scalable and regularized gradient boosting framework based on decision trees that is designed for efficient and highperformance supervised learning tasks [36]. Hyperparameter optimization was performed using Optuna [37] to identify the optimal parameter configuration for the dataset. The remaining rolling-windows were used for model evaluation. A rolling-window validation strategy was used, in which the model was iteratively trained on all available historical weeks and evaluated on the immediately following week. Specifically, for each iteration, the model was trained on weeks 1, . . . , t and evaluated on week t + 1. 4) Performance: Table I summarizes weekly predictive performance across the evaluation periods for models trained

jointly on all organizations. Joint training generally outperformed organization-specific models. Alternative variants incorporating class weight balancing and calibration were evaluated, but did not further improve performance or calibration quality. The model achieved strong discriminative performance, with a mean area under the receiver operating characteristic curve (ROC-AUC) of 0.90 and accuracy above 99.8% across all evaluation periods. Given the severe class imbalance (0.06%-0.27% positive observations per week), the area under the precision-recall curve (PR-AUC) is also reported, yielding a mean value of 0.23. Relative to the weekly positive-class baseline, this corresponds to approximately 45-500 times the performance of a random classifier. TABLE I. S UMMARY OF MODEL PERFORMANCE METRICS .

Metric

Mean Median Min

ROC-AUC 0.90 PR-AUC 0.23 Accuracy (%) 99.87 ECE (%) 0.094 Calibration slope 1.00

0.91 0.22 99.91 0.096 1.00

Max

0.82 0.95 0.07 0.37 99.75 99.95 0.023 0.187 0.72 1.30

5) Calibration: Table I shows that the trained XGBoost model exhibited good probabilistic calibration, with consistently low Expected Calibration Errors (ECE) and calibration slopes generally close to 1.0. 6) Feature importance: Feature importance was assessed using SHapley Additive exPlanations (SHAP) computed for each evaluation period. While the relative importance of individual features varied across weeks, several consistent patterns emerged. In particular, the lagged outcome variable Xt,m was the most influential predictor for all evaluation periods. 7) Machine dependencies: Although machine-level predictions demonstrated good calibration, aggregate CLE counts remained substantially overdispersed relative to the independent Bernoulli assumption. Global calibration aligned the expected and observed CLE frequencies. However, substantial overdispersion persisted (37.23, p < 0.001), indicating that the excess variability could not be explained by calibration error alone. Within the LEFSA framework, this suggests that machinelevel CLE variables might not be conditionally independent given the observed information set Ft , i.e., Xt,m ̸⊥ Xt,m′ | Ft . Although unconditional pairwise correlations between machine-level CLE variables were generally small, substantial aggregate overdispersion remained after conditioning on the predicted probabilities. This indicates residual clustering, dependence structures, or shared latent risk factors between machines that are not fully captured by the machine-level prediction model. 8) Computational requirements: LEFSA was deployed on a server equipped with a 48-core Intel Xeon Gold 5118 CPU at 2.30 GHz and 251 GB of RAM. Table II reports approximate operational resource requirements, normalized per

1,000 machines. LEFSA consists of two main computational components: (1) continuous stream processing of raw endpoint logs and alerts into hourly machine-level feature profiles, and (2) weekly feature aggregation and XGBoost model training. TABLE II. A PPROXIMATE COMPUTATIONAL REQUIREMENTS PER 1,000 MACHINES .

Resource CPU (% of one core) Peak RAM (MB) Weekly storage (MB)

Stream processing Weekly training 6.3 83 12.85

15.0 550 1.68

IV. D ISCUSSION Compared with scenario-based FAIR analyses, LEFSA operationalizes LEF estimation by predicting general machinelevel CLE probabilities directly from cybersecurity telemetry and aggregating them across infrastructure layers. It also extends existing organization-level cyber-incident forecasting and machine-level malware prediction by providing risk estimates for general CLEs at all IT infrastructure layers. The empirical evaluation demonstrated that machine-level CLE prediction using operational cybersecurity telemetry is feasible, achieving strong discrimination performance and well-calibrated probability estimates across evaluation periods. The results suggest that machine-level telemetry contains substantial predictive information for future CLE occurrence. In particular, the strong importance of the lagged outcome variable indicates temporal dependence in machine-level CLE risk. This suggests that machine-level CLE occurrence may exhibit non-memoryless temporal dynamics, limiting the suitability of simple Poisson-based formulations that assume conditionally independent events with constant intensities. The observed overdispersion and residual machine-level dependencies further suggest that infrastructure-level cyber risk aggregation may require explicit modeling of dependence structures rather than assuming conditional independence between machines. Since CLEs are rare and organizations often have many machines, the aggregation of weakly dependent machine-level predictions can amplify dependence effects at the group-level [38]. As a result, the independence assumption may underestimate aggregate variance and tail risk. Among the evaluated models, XGBoost consistently achieved the strongest predictive performance across evaluation periods. The model also demonstrated relatively strong probabilistic calibration, suggesting that machine-level CLE probabilities can be estimated with sufficient reliability to support downstream probabilistic aggregation and cyber risk estimation. The observed advantage of joint training suggests that machine-level CLE prediction benefits from shared crossorganizational telemetry patterns. This highlights the potential value of MDR-scale datasets and federated learning approaches for operational cyber risk quantification. Predictive performance and calibration were generally strong and stable across evaluation periods. However, one

evaluation period exhibited substantially reduced performance and degraded calibration. This likely reflects a temporary distribution shift in attacker behavior, organizational infrastructure, or operational security processes. This highlights the dynamic nature of cybersecurity environments and emphasizes the importance of continuous retraining and monitoring a LEF model in operational deployments. Although the empirical evaluation focused on MDE telemetry, the LEFSA framework itself is designed to be platformagnostic and extensible to other telemetry providers and cybersecurity environments. More generally, the framework provides a structured methodology for transforming operational cybersecurity data into probabilistic infrastructure-level cyber risk estimates. A. Limitations & operational considerations Several limitations should be acknowledged. The empirical evaluation was conducted using proprietary MDR telemetry collected from organizations using MDE. Consequently, the dataset cannot be publicly shared, limiting direct reproducibility of the experimental results. The observation window consisted of only 12 weeks, while CLEs remain relatively rare events. Although the machine-level dataset contained a large number of observations, the limited time horizon restricts the ability to evaluate long-term temporal dynamics and rare largescale incidents. The empirical evaluation focused primarily on machinelevel LEF estimation. While the paper proposed a generalized framework for infrastructure-level aggregation under dependence, accurate modeling of dynamic machine-level dependencies remains a challenging open problem. The empirical evaluation also has several methodological limitations. CLE labels are derived from security incident cases and may not always represent confirmed compromises, while missed detections may introduce systematic label bias. Predictions and aggregate estimates cover only machines observed through the available telemetry and detection platforms, meaning that unmonitored or newly introduced machines are therefore automatically excluded. Operational deployment at scale introduces several additional challenges. LEFSA depends on sufficiently complete and comparable endpoint telemetry, vulnerability information, and detection alerts. Thus machines without adequate monitoring are excluded, potentially causing aggregate risk to be underestimated. Changing machine populations and inconsistent identifiers further complicate tracking and predictions for newly observed assets. A production implementation must also process high-volume data streams while remaining robust to duplicate, delayed, or missing events and service outages. Finally, in the case of cross-organizational model training, it requires appropriate contractual, privacy, governance, and system-integration controls. V. C ONCLUSION & F UTURE W ORK This paper introduced LEFSA, a generalized probabilistic framework for estimating LEF through machine-level CLE

prediction and hierarchical infrastructure-level aggregation. Rather than relying primarily on manually constructed threat scenarios or externally observable organizational characteristics, LEFSA reformulates LEF estimation as a measurable machine-level prediction problem using operational cybersecurity telemetry. The empirical evaluation demonstrated that machine-level CLE prediction is feasible using real-world cybersecurity data. XGBoost achieved strong predictive and calibration performance across the evaluation periods. The results further suggest that machine-level telemetry contains substantial predictive information relevant for future CLE occurrence. Beyond empirical evaluation, the primary contribution of this work is the proposed probabilistic aggregation framework. This framework suggests aggregating calibrated machine-level CLE probabilities across infrastructure layers while preserving uncertainty and dependence structures. This provides a foundation for infrastructure-level cyber risk quantification, including estimation of aggregate CLE distributions, tail risks, and operational cyber resilience metrics. Several directions for future research remain. Most importantly, future work should investigate dependence-aware aggregation methods capable of modeling dynamic dependencies between machines. In particular, dynamic copula-based approaches in a discrete setting may provide a promising framework for combining marginal machine-level probabilities with time-varying joint dependence structures. Although a single lagged outcome variable was included, future research should investigate richer temporal dependence structures within individual machines, as well as richer telemetry sources, weak supervision, and infrastructure-aware graph representations. Additionally, federated learning could enable training across organizations while preserving data confidentiality and operational privacy constraints. Potentially, LEFSA could be adjusted to predict future operational incidents. Finally, integrating LEF estimation with probabilistic loss magnitude models would enable fully quantitative cyber risk estimation aligned with broader FAIR-style risk frameworks. Overall, this paper suggests that combining probabilistic machine-level CLE models with dependence-aware aggregation is a promising direction for operational and data-driven cyber risk management. R EFERENCES [1]

[2]

International Monetary Fund. Monetary and Capital Markets Department, “Cyber risk: A growing concern for macrofinancial stability”, in Global Financial Stability Report, April 2024: The Last Mile: Financial Vulnerabilities and Risks, International Monetary Fund, Apr. 2024, ch. 3. DOI: 10.5089/ 9798400257704.082 Global Cyber Risk Team, “Navigating the cyber claims landscape”, Chubb, Tech. Rep., Mar. 2025, https : / / www . chubb . com / content / dam / chubb - sites / chubb - com / us en / business - insurance / products / cyber / documents / chubb cyberclaimsreport-final.pdf [retrieved: August, 2026].

[3]

[4]

[5]

[6] [7]

[8]

[9]

[10]

[11]

[12]

[13] [14] [15]

[16]

[17]

Allianz Commercial, “Cyber security resilience 2024: Trends in data breach and privacy risk”, Tech. Rep., Oct. 2024, https : / / commercial . allianz . com / content / dam / onemarketing / commercial/commercial/reports/cyber- security- trends- 2024. pdf [retrieved: August, 2026]. National Association of Insurance Commissioners, “Report on the cybersecurity insurance market”, Tech. Rep., Oct. 2025, https://content.naic.org/sites/default/files/inline- files/2025_ Cybersecurity _ Insurance % 20Report . pdf [retrieved: August, 2026]. KELA Cyber Threat Intelligence, “2025 AI threat report: How cybercriminals are weaponizing AI technology”, KELA, Tech. Rep., Mar. 2025, https : / / www . kelacyber . com / resources / research/2025-ai-threat-report/ [retrieved: August, 2026]. C. Florackis, C. Louca, R. Michaely, and M. Weber, “Cybersecurity risk”, The Review of Financial Studies, vol. 36, no. 1, pp. 351–407, May 2022. DOI: 10.1093/rfs/hhac024 R. He, Z. Jin, and J. S.-H. Li, “Modeling and management of cyber risk: A cross-disciplinary review”, Annals of Actuarial Science, vol. 18, no. 2, pp. 270–309, 2024. DOI: 10 . 1017 / S1748499523000258 M. Carannante and A. Mazzoccoli, “An analytical review of cyber risk management by insurance companies: A mathematical perspective”, Risks, vol. 13, no. 8, Jul. 2025. DOI: 10.3390/risks13080144 I. D. Sánchez-García, J. Mejía, and T. San Feliu Gilabert, “Cybersecurity risk assessment: A systematic mapping review, proposal, and validation”, Applied Sciences, vol. 13, no. 1, p. 395, 2023. DOI: 10.3390/app13010395 International Organization for Standardization and International Electrotechnical Commission, ISO/IEC 27005:2022 information security, cybersecurity and privacy protection— guidance on managing information security risks, 4th ed., https://www.iso.org/standard/80585.html [retrieved: August, 2026], Oct. 2022. Joint Task Force Transformation Initiative, “Guide for conducting risk assessments”, National Institute of Standards and Technology, NIST Special Publication 800-30 Rev. 1, Sep. 2012, https://csrc.nist.gov/pubs/sp/800/30/r1/final [retrieved: August, 2026]. DOI: 10.6028/NIST.SP.800-30r1 C. J. Alberts, S. Behrens, R. D. Pethia, and W. R. Wilson, “Operationally critical threat, asset, and vulnerability evaluation (OCTAVE) framework, version 1.0”, Software Engineering Institute, Carnegie Mellon University, Tech. Rep. CMU/SEI-99TR-017, Sep. 1999, https://sei.cmu.edu/library/operationallycritical - threat - asset - and - vulnerability - evaluation - octave framework - version - 10/ [retrieved: August, 2026]. DOI: 10 . 1184/R1/6575906.v1 M. S. Lund, B. Solhaug, and K. Stølen, Model-Driven risk analysis. Springer Nature Link, Oct. 2010. DOI: 10.1007/9783-642-12323-8 J. Freund and J. Jones, Measuring and Managing Information Risk: A FAIR Approach 2nd Edition. Butterworth-Heinemann, Dec. 2025, p. 365. DOI: 10.1016/C2022-0-02439-9 J. Wang, M. Neil, and N. Fenton, “A Bayesian network approach for cybersecurity risk assessment implementing and extending the FAIR model”, Computers & Security, vol. 89, Feb. 2020. DOI: 10.1016/j.cose.2019.101659 Y. Liu et al., “Cloudy with a chance of breach: Forecasting cyber security incidents”, in 24th USENIX Security Symposium (USENIX Security 15), USENIX Association, Aug. 2015, pp. 1009–1024. A. Okutan, G. Werner, S. J. Yang, and K. McConky, “Forecasting cyberattacks with incomplete, imbalanced, and insignificant data”, Cybersecurity, vol. 1, no. 1, p. 15, 2018. DOI: 10.1186/s42400-018-0016-5

[18]

[19]

[20]

[21]

[22]

[23]

[24]

[25] [26] [27] [28]

[29]

[30]

[31] [32] [33] [34]

N. O. Leslie, R. E. Harang, L. P. Knachel, and A. Kott, “Statistical models for the number of successful cyber intrusions”, The Journal of Defense Modeling and Simulation Applications Methodology Technology, vol. 15, no. 1, pp. 49–63, Jun. 2017. DOI : 10.1177/1548512917715342 L. Bilge, Y. Han, and M. Dell’Amico, “RiskTeller: Predicting the risk of cyber incidents”, in CCS ’17: Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security, Association for Computing Machinery, Oct. 2017, pp. 1299–1311. DOI: 10.1145/3133956.3134022 M. Balduzzi, R. Reyes, J. Balaquit, R. Flores, and B. Zigh, “Forecasting future outbreaks: A behavioral and predictive approach to proactive cyber risk management”, Trend Micro, Tech. Rep., Mar. 2026, https : / / documents . trendmicro . com / assets / pdf / research - paper _ forecasting - future - outbreaks . pdf [retrieved: August, 2026], pp. 1–37. V. Zokarkar and K. Mathur, “A survey of attack prediction approaches in cyber security”, in Proceedings of the International Conference on Recent Advancements and Modernisations in Sustainable Intelligent Technologies and Applications (RAMSITA 2025), Atlantis Press, May 2025, pp. 1003–1016. DOI : 10.2991/978-94-6463-716-8_75 A. Niculescu-Mizil and R. Caruana, “Predicting good probabilities with supervised learning”, in Proceedings of the 22nd International Conference on Machine Learning, Association for Computing Machinery, 2005, pp. 625–632. DOI: 10.1145/ 1102351.1102430 J. C. Platt, “Probabilistic outputs for Support Vector Machines and comparisons to Regularized Likelihood Methods”, in Advances in Large Margin Classifiers, MIT Press, 2000, pp. 61–74. B. Zadrozny and C. Elkan, “Obtaining calibrated probability estimates from decision trees and naive Bayesian classifiers”, in Proceedings of the Eighteenth International Conference on Machine Learning, Morgan Kaufmann Publishers Inc., 2001, pp. 609–616. P. Billingsley, Probability and measure. John Wiley & Sons, Jan. 2012. Y.-Z. Chen, Z.-G. Huang, S. Xu, and Y.-C. Lai, “Spatiotemporal patterns and predictability of cyberattacks”, PLOS ONE, vol. 10, no. 5, May 2015. DOI: 10.1371/journal.pone.0124472 T. Schmidt, “Coping with copulas”, in Copulas: From Theory to Application in Finance, J. Rank, Ed., Risk Books, 2007, pp. 3–34. C. Peng, M. Xu, S. Xu, and T. Hu, “Modeling multivariate cybersecurity risks”, Journal of Applied Statistics, vol. 45, no. 15, pp. 2718–2740, Feb. 2018. DOI: 10.1080/02664763. 2018.1436701 Y. Li, X. Wang, P. Zhao, and T. Hu, “Cyber breach risk modeling for insurance: Capturing temporal and cross-group dependence”, Annals of Actuarial Science, pp. 1–25, Sep. 2025. DOI: 10.1017/s1748499525100109 A. J. Patton, “Modelling asymmetric exchange rate dependence”, International Economic Review, vol. 47, no. 2, pp. 527–556, May 2006. DOI: 10 . 1111 / j . 1468 - 2354 . 2006 . 00387.x C. Genest and J. Nešlehová, “A primer on copulas for count data”, Astin Bulletin, vol. 37, no. 2, pp. 475–515, Nov. 2007. DOI : 10.2143/ast.37.2.2024077 A. K. Nikoloulopoulos and H. Joe, “Factor copula models for item response data”, Psychometrika, vol. 80, no. 1, pp. 126– 150, Dec. 2013. DOI: 10.1007/s11336-013-9387-4 H. Joe, Dependence Modeling with Copulas (Chapman & Hall/CRC Monographs on Statistics and Applied Probability). CRC Press, Jun. 2014. DOI: 10.1201/b17116 A. Panagiotelis, C. Czado, H. Joe, and J. Stöber, “Model selection for discrete regular vine copulas”, Computational

[35]

[36]

Statistics & Data Analysis, vol. 106, pp. 138–152, Sep. 2016. DOI : 10.1016/j.csda.2016.09.007 Microsoft, Microsoft Defender for Endpoint overview, Microsoft Learn. [Online]. Available: https : / / learn . microsoft . com/en- us/defender- endpoint/microsoft- defender- endpoint, [retrieved: August, 2026], 2026. T. Chen and C. Guestrin, “XGBoost: A scalable tree boosting system”, in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Aug. 2016, pp. 785–794. DOI: 10.1145/2939672.2939785

[37]

[38]

T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama, “Optuna: A next-generation hyperparameter optimization framework”, in The 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Jul. 2019, pp. 2623– 2631. DOI: 10.1145/3292500.3330701 R. M. Cooke, C. Kousky, and H. Joe, “Micro correlations and tail dependence”, in Dependence Modeling: Vine Copulae Handbook. World Scientific Publishing, 2010, ch. 5, pp. 89– 112. DOI: 10.1142/9789814299886_0005

Record · ID 1006785 · SHA-256 1a5c0935232abe1f
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.