Information-Entropy-Driven Fault Propagation Modeling for Probabilistic Network Performance Prediction 1st Lusha Mo, 2nd Fengxiao Tang* , 3rd Xiaonan Wang, 4th Ming Zhao
arXiv:2609.08143v1 [cs.NI] 8 Sep 2026
School of Computer Science and Engineering, Central South University, Changsha, China {molusha, tangfengxiao, wangxiaonan, meanzhao}@csu.edu.cn * Corresponding Author
Abstract—Network faults can trigger cascading effects that cause abrupt and nonstationary performance degradation. Existing learning-based performance predictors mainly focus on normal operation or treat fault-induced topology and routing changes as static inputs, and typically produce deterministic point estimates. They overlook fault-propagation dynamics and uncertainty in performance evolution. The predefined-rule and purely data-driven propagation models lack a unified representation of fault definition, propagation mechanism, and impact quantification. Additionally, generic denoisers in conditional diffusion models fail to incorporate fault propagation into uncertainty modeling. To address these limitations, we propose an information-entropy-driven fault propagation paradigm (IEFP) that characterizes fault propagation via relative entropy, mutual information and transfer entropy. We then design a fault-aware graph message-passing mechanism that propagation contexts modulate network representation learning. We further develop FEMNet, which employs this mechanism as a tailored denoiser within a conditional diffusion model to enable probabilistic network performance prediction under complex fault scenarios. Compared with the strongest baselines, IEFP improves faultprediction performance, while FEMNet reduces errors in both point and probabilistic KPI prediction. Index Terms—Network fault propagation, probabilistic network performance prediction, information-entropy modeling, conditional diffusion model, graph message passing.
I. I NTRODUCTION Network performance prediction estimates how key performance indicators (KPIs) evolve under different network conditions. Faults can trigger link disruptions, route reconfiguration, and queue buildup, causing cascading effects and quality-of-service degradation [1], [2]. The deterioration is abrupt, nonstationary, and driven by structural changes in the network. Accurate and rapid assessment of fault impacts on network KPIs is therefore essential for proactive degradation detection, automated control, and intelligent maintenance, as well as for robustness evaluation and reliability enhancement during network design and deployment. Learning-based predictors [3] are prominent, which capture complex relationships among topology, routing, and traffic to predict node-, link-, and end-to-end performance. However, most assume normal operation under fixed configurations. Even when link failures are considered [4], the resulting
Fault Cascading Effect Normal link Fault link Source path Rerouting path
Nonstationary Performance Degradation
Node B
Fault Occurs
! Link failure
KPI values
KPI
Fluctuations
Node C
Node A
Node D
Node E Congestion
Node F Congestion
Existing Paradigm: Snapshot-Based Prediction Current Snapshot
Deterministic Point Prediction KPI
Topology
Routing
Time
Reality: Dynamic Fault Evolution with Uncertainty Fault Propagation
Heteroscedastic Evolution Low
High Risk
Context 1 (Mild fault) Context 2 (Severe fault)
Traffic Time
KPI
Fig. 1. Fault cascading and its impact on network performance. A local link failure can trigger rerouting and cascading congestion (top left), leading to nonstationary KPI degradation (top right). Conventional snapshot-based deterministic prediction (bottom left) cannot represent the fault propagation and context-dependent, heteroscedastic KPI distributions(bottom right).
topology, route, and other reachability variations are merely encoded as updated static inputs. Such updates omit temporal dependencies and the accumulated effects of preceding faults. Performance prediction thus remains a snapshot-based mapping from network state to deterministic KPIs, as shown in Fig. 1, without explicitly modeling how faults propagate and how the process shapes KPI evolution. This limitation becomes critical under complex faults scenarios. A central challenge is to accurately model fault propagation. Existing approaches either rely on predefined rules [5], [6] which generalize poorly to heterogeneous dependencies and stochastic triggers, or fit propagation patterns from data [7] without characterizing the persistent effect of historical faults on future fault occurrence. They lack a unified account of three fundamental questions: what constitutes a fault, how it propagates, and how its impact should be quantified. In fact, a fault implies that a component deviates from normal operation, and the abnormality may propagate along dependencies and alter the uncertainty and evolution of other components’ future states. Information theory provides a natural language for characterizing uncertainty, coupling strength, and directional
influence [8], motivating an information-entropy [9], [10] view for modeling fault propagation. Furthermore, performance prediction must also capture how fault propagation affects KPI evolution. Under fault conditions, node and link interactions depend not only on topology, routing, and traffic but also on the accumulated effects of historical faults, which can persistently raise abnormality risks in dependent components. Since fault propagation changes the importance, reliability, and effective range of information exchanged over the graph, its context should dynamically regulate message passing instead of serving as static feature. Another difficulty arises from the pronounced uncertainty and heteroscedasticity of network performance evolution under fault conditions. Fault propagation may induce complex stochastic variations through multiple joint factors, such that similar network states and fault contexts may yield markedly different degradation patterns and fluctuation magnitudes. Performance prediction should therefore move beyond deterministic point estimates to characterize conditional KPI distributions. Conditional diffusion models [11] offer a natural mechanism for this purpose. However, the generic denoisers do not model fault propagate and its effects on future states [12]. Probabilistic prediction under fault conditions requires a fault-aware denoising architecture. We address these challenges by jointly modeling fault propagation and uncertainty in KPI evolution. We first establish IEFP, an information-entropy-driven fault propagation paradigm to estimates fault risks for nodes and links. Specifically, IEFP uses relative entropy, mutual information, and transfer entropy to characterize abnormality levels, internode coupling strength, and cross-component influence from historical to future states, respectively. We then design a fault-aware graph message passing mechanism to incorporate propagation information into graph representation learning, where predictive risks gate messages, adjust attention weights, and mask weakly coupled edges. Using the mechanism as denoising network, we develop FEMNet, a probabilistic network performance prediction model under complex fault scenarios, which introduces conditional diffusion to generalize deterministic point prediction to its probabilistic counterpart. Our main contributions are as follows. 1) We propose IEFP, an information-entropy-driven fault propagation paradigm that interprets fault propagation as abnormal information transmission, accumulation and redistribution over networks, providing a new theoretical perspective for fault propagation analysis in complex network scenarios. 2) To effectively incorporate propagation information into performance prediction, we design a fault-aware graph message passing mechanism, in which predictive fault risks modulates node message and update dynamics. 3) We develop FEMNet for probabilistic network performance prediction under complex fault scenarios. By introducing conditional diffusion model, FEMNet transforms performance prediction from deterministic point regression task into conditional distribution generation task.
4) We evaluate IEFP and FEMNet under complex fault scenarios in terms of point-prediction accuracy, fault-anticipation capability, and probabilistic-prediction quality. Results demonstrate consistent improvements over the evaluated baselines. II. R ELATED W ORK A. Network Performance Prediction Among accurate, lightweight learning-based network models, RouteNet [13] based on graph neural network (GNN), models links and traffic paths to predict source-destination KPIs. xNet [14] and RouteNet-Erlang [15] incorporate queue characteristics and specialized graph message passing. HTNet [16], DGAT [17], and EAGLE [18] capture temporal behavior, attention weights, and bandwidth-aware representations. RouteNet-Fermi [4] accepts topology and routing changes after link failures as inputs but omits fault interaction and diffusion. A recent analysis [19] interprets RouteNet as highorder topological modeling that captures path and scheduling order, while RouteNet-Gauss [20] adds hardware-testbed data and configurable temporal granularity. These methods rely on configurations or state snapshots and do not model fault propagation and its persistent effects on performance evolution. B. Fault Propagation Modeling Fault propagation has been studied using Bayesian networks [21], fault trees, cascading models (e.g., threshold models, load-capacity models [22], and interdependent network [23]) and diffusion models. State explosion and predefined rules limit their ability to capture probabilistic dynamics, heterogeneous dependencies, and temporal interactions. For example, HIC model [24] combines stochastic Markovian diffusion with physics-based concepts. Probabilistic graphical models offer interpretability and causal expressiveness(e.g., Causal model [25]), but their structure learning and inference remain challenging at scale. Data-driven methods, including machine learning [7], Graph Hawkes processes [26], [27], and GNNbased models (e.g., GPIAN [28]), fit numerical variations without explicitly representing fault triggering and propagation. Existing methods lack a unified representation of how fault impacts traverse complex networks and persist over time. Information theory concerned with information, uncertainty and communication, provides a natural foundation for such theoretical framework. C. Conditional Diffusion Model Diffusion models [29], [30] have expanded from image generation to probabilistic time-series and graph prediction. TimeDiff [31] applies conditional diffusion to nonautoregressive time-series prediction, and DiffSTG [32] combines diffusion models with spatiotemporal graph modeling for probabilistic forecasting. LGD [33] formulates graph regression and classification tasks into a generic conditional generation framework. ReDiSC [34] reparameterized masked diffusion for structured node classification. SimDiff [35] improves point estimation using normalization-independent diffusion and a median-of-means estimator. These studies demonstrate
the strong potential of diffusion models for uncertainty modeling in network performance prediction. However, generic denoisers lack fault-propagation mechanisms and thus cannot fully capture fault-aware network dynamics.
The model learns to restore the original data by predicting the injected noise. Training objective is to minimize the loss between the actual and predicted noise: h i 2 min L(θ) = Ey0 ,ϵ,τ,c ∥ϵ − ϵθ (yτ , τ, c)∥2 . (7) θ
III. P ROBLEM F ORMULATION This section reviews the required information-theoretic quantities and conditional diffusion foundations, and formulates probabilistic network performance prediction under complex fault conditions. A. Preliminaries Definition 1 (Information-theoretic quantities). Let P and Q be probability distributions on Ω ⊆ Rd with density functions p(x) and q(x), respectively, where q(x) > 0 whenever p(x) > 0. Their relative entropy is Z p(x) dx, (1) DKL (P ∥Q) ≡ p(x) log q(x) Ω which is +∞ if the support condition fails. For random elements B, C, and D on standard Borel spaces with joint law PBCD , mutual information and conditional mutual information are I(B; C) ≡ DKL (PBC ∥PB ⊗ PC ) , I(B; C | D) ≡ EPD DKL PBC|D PB|D ⊗ PC|D .
(2)
The regular conditional laws in (2) are understood PD -almost (k ) (k ) surely. Let Bt B and Ct C denote the source and target history vectors, respectively. The context-conditioned transfer entropy from B to C is (k ) (k ) TEB→C|D (t) ≡ I Ct+1 ; Bt B | Ct C , Dt , (3) which measures directed predictive dependence under the specified history and context. Definition 2 (Conditional diffusion model). A conditional diffusion model is a probabilistic generative model that learns p(y0 | c), where c denotes conditional information such as labels, descriptions, attributes, or states. The forward diffusion process gradually injects Gaussian noise into data samples: q(y1:T | y0 ) =
T Y
p N yτ ; 1 − βτ yτ −1 , βτ I ,
where τ is diffusion step and βτ is a predefined noise-schedule Qτ coefficient. Let ατ = 1 − βτ and ᾱτ = s=1 αs . Then √ √ yτ = ᾱτ y0 + 1 − ᾱτ ϵ, ϵ ∼ N (0, I). (5) Starting from pure noise, the reverse denoising process gradually recovers target samples satisfying condition c: pθ (y0:T | c) = p(yT )
pθ (yτ −1 | yτ , c),
Considering a communication network, fault events may dynamically alter network topology, resource availability, routing behavior, and traffic evolution. Their effects may propagate across dependent entities and introduce substantial performance uncertainty. We aim to predict the probability distributions of network KPIs at the next timestamp under complex fault scenarios. A fault scenario is complex when performance is affected by coupled faults rather than one isolated perturbation, including at least one of: 1) concurrent faults; 2) temporally correlated faults within the observation window whose effects propagate through network dependencies; or 3) faults differing in type, duration, or severity that jointly affect multiple entities. We represent the network as a heterogeneous graph G = (V, E, A), where node sets V represent queues, links, and flows; edge sets E encode their dependencies; and A is the adjacency matrix. For node v, let zv ∈ Rdz be its static configuration and Xv (t) ∈ Rdx its time-varying feature vector, with observation xv (t). Examples include a queue’s buffer limit in zv and current occupancy in xv (t), or a link’s maximum load in zv and current bandwidth in xv (t). With x0:t v = (xv (0), . . . , xv (t)), the network attributes are Z = zv , x0:t :v∈V . (8) v The traffic states are S = {sf (t) ∈ Rdf }f ∈F , where sf (t) is flow f ’s current feature vector. Faults within observation window [0, t] form the event sequence K E0:t = εk = typek , lock , tstart , tend , (9) k k k=1 whose entries denote the fault type, affected entity, start time and end time of the k-th fault event, respectively. Given condition C = (G, Z, S, E0:t ), we estimate each node’s next-step KPI distribution: Y = {yv (t + 1) ∼ pv (· | G, Z, S, E0:t )}v∈V ,
(4)
τ =1
T Y
B. Problem Definition
(6)
τ =1
where the transition pθ (yτ −1 | yτ , c) is parameterized by a neural network and explicitly depends on c.
(10)
where yv (t + 1) and pv (·) denote node v’s target KPIs and their conditional distribution. KPIs include delay, throughput, queue length, packet loss, and related metrics. Therefore, We learn a mapping fΘ : (G, Z, S, E0:t ) → Y,
(11)
that enables the predicted distributions accurately characterize stochastic performance evolution under the joint effects of traffic dynamics, resource contention, and fault propagation. For training, applying (5) to the normalized KPI target y0 gives yτ . We optimize h i 2 min Ldiff (θ) = Ey0 ,ϵ,τ,C ∥ϵ − ϵθ (yτ , τ, C)∥2 . (12) θ
where ϵθ is the denoising network. By minimizing Ldiff , the model approximates the KPI distribution conditioned on network state and fault information. IV. I NFORMATION -E NTROPY-D RIVEN FAULT P ROPAGATION PARADIGM A fault indicates a component’s departure from normal operation and injects abnormal information into the system. This information propagates through topological dependencies, resource interactions, and protocol coupling, altering other components’ uncertainty and dynamics. IEFP model fault propagation via three information-theoretic quantities. Relative entropy measures deviation of observation distribution from normal operation and defines fault severity. Mutual information decomposes fault information into configuration vulnerability, local historical degradation, and propagated influence. Transfer entropy quantifies the directional incremental information that one component’s history provides about another’s fault state. We define Hv (t) = Xv (t − L + 1), . . . , Xv (t − 1), Xv (t) (13) (N ) as node v’s length-L history up to time t, and let pv (X) be
the reference distribution estimated from normal samples. Theorem 1 (Asymptotic Equipartition Property, AEP). If the random process {Xv (t)} is stationary and ergodic, then as n → ∞, 1 (14) − log P (Xv (1), . . . , Xv (n)) → H̄(Xv ). n Thus, average information of a long normal sequence concentrates around entropy rate H̄(Xv ), and normal observations lie in the typical set with high probability. A state unlikely under the normal distribution carries information unexplained by the nominal operating regime. Let p̂tv (X) be the observation distribution constructed from node v’s current feature. We define its anomaly information quantity as the relative entropy between the current observation and normal reference distributions: ) Sv (t) = DKL p̂tv (X) ∥ p(N (15) v (X) . Larger Sv (t) indicates greater distributional departure from normal generation mechanism and a more severe anomaly. A fault occurs when this information exceeds the system tolerance, yielding the binary state ( 1, Sv (t) ≥ λv , Fv (t) = (16) 0, Sv (t) < λv ,
with equality iff C ⊥⊥ D | B, i.e., D provides no additional information about C beyond what B already contains. Faults propagate because nodes are coupled by shared information, physical or logical constraints, routing, and resource competition. A fault at node u changes its output distribution: P (Xu (t) | Fu (t) = 1) ̸= P (Xu (t) | Fu (t) = 0).
(19)
This output enters adjacent nodes’ state dynamics and may change their posterior fault risk: P (Fv (t + 1) = 1 | Hv (t), Hu (t), zv ) ̸= P (Fv (t + 1) = 1 | Hv (t), zv ) .
(20)
Information-theoretically, faults in complex systems are conditionally stochastic. Anomalous information from neighbors’ history propagates along dependency edges, adding condition-dependent information about the target’s fault state. By Theorem 2: H(Fv (t + 1) | Hv (t), zv ) > H(Fv (t + 1) | Hv (t), Hu (t), zv ) .
(21)
We quantify this directed historical incremental information by transfer entropy: TEu→v = I(Fv (t + 1); Hu (t) | Hv (t), zv ) = H(Fv (t + 1) | Hv (t), zv )
(22)
− H(Fv (t + 1) | Hv (t), Hu (t), zv ) . A larger TEu→v indicates a stronger statistical association between neighbor u’s history and node v’s fault state, and hence stronger propagation. Theorem 3 (Chain rule of mutual information). For random variables B, C, and D, I(D; B, C) = I(D; B) + I(D; C | B).
(23)
Repeatedly applying Theorem 3 decomposes the total statistical association between node v’s fault state and the available configuration, local history, and neighborhood history: I Fv (t + 1); Hv (t), HN (v) (t), zv = I(Fv (t + 1); zv ) {z } | (conf)
Iv
+ I(Fv (t + 1); Hv (t) | zv ) | {z }
(24)
Iv(local) :local term
+ I Fv (t + 1); HN (v) (t) Hv (t), zv . | {z } Iv(prop) :propagation term
which records each fault’s timestamp and anomaly magnitude. Theorem 2 (Conditional reduction of entropy). For random variables B, C, and D, we have:
Here, HN (v) (t) collects the historical feature sequences of (conf) (local) (prop) node v’s neighbors. The terms Iv , Iv , and Iv quantify fault-risk information from static configuration, local degradation history, and neighboring histories, respectively. We convert this statistically averaged mutual information into sample-specific fault information gain using pointwise mutual information:
H(C | B, D) ≤ H(C | B),
conf ∆sF + ∆slocal (t) + ∆sprop (t). v (t) = ∆sv v v
where λv > 0 is node v’s learned tolerance threshold. The entropy-marked fault-event history is Av (t) = {(tk , Sv (tk )) : tk ≤ t, Fv (tk ) = 1},
(17)
(18)
(25)
Because the complete histories Hv (t) and HN (v) (t) are high-dimensional and costly to model, we compress them into the abnormal information carried by fault events: F ∆sF v (t; H) = ∆sv (t; A) + ζv (t),
(26)
where the compression residual is P Fv (t + 1) = 1 | Hv (t), HN (v) (t), zv . (27) ζv (t) = log P Fv (t + 1) = 1 | Av (t), AN (v) (t), zv Here, ζv (t) directly measures the information lost by replacing full feature histories with event histories. When ζv (t) ≈ 0, the fault-event sets preserve the relevant historical information: P Fv (t + 1) = 1 | Hv (t), HN (v) (t), zv (28) = P Fv (t + 1) = 1 | Av (t), AN (v) (t), zv . To obtain a graph propagation form, we decompose the joint event-history contribution into source-node contributions: X ∆su→v (t)+ηv (t), (29) ∆slocal (t)+∆sprop (t) = v v u∈N (v)∪{v}
where ηv (t) captures higher-order synergy and redundancy among sources. When them are negligible, i.e., ηv (t) ≈ 0, the gain becomes the sum of source-wise contributions: X ∆slocal (t) + ∆sprop (t) = ∆su→v (t). (30) v v
For an interpretable, estimable model, we decompose each event’s risk information as ψu→v (Su (tuk ), t − tuk ) = Wu→v φ(Su (tuk )) gu→v (t − tuk ) , (36) where Wu→v is the learned propagation weight, φ(Su (tuk )) = u 1 − e−Su (tk ) is a bounded severity function, and gu→v (·) is a temporal-decay kernel. This factorization encodes principles: stronger propagation paths and more severe faults contribute more risk information, whereas older events exert less. Let Φu→v (t) denote the effective fault-information stock transmitted from u to v, with dynamics dΦu→v (t) = −γu→v Φu→v (t)+ dt X Wu→v φ(Su (tuk )) δ(t − tuk ),
where γu→v > 0 is the learned memory-decay rate. The solution is X u Φu→v (t) = Wu→v φ(Su (tuk )) e−γu→v (t−tk ) , (38) tu k ≤t Fu (tu k )=1
corresponding to the exponential-decay kernel u
gu→v (t − tuk ) = e−γu→v (t−tk ) .
u∈N (v)∪{v}
Theorem 4 (Atomic integral property of a marked counting measure). Let δxk be the Dirac measure at event xk and ψ(x) an event-response kernel. Then Z X X ψ(x) δxk (dx) = ψ(xk ). (31) k
k
Represent source node u’s fault history by the marked counting measure X δ(tuk ,Su (tuk )) (dξ, da), Nu (dξ, da) = (32) tu k ≤t Fu (tu k )=1
which retains event time and anomaly severity as marks. The cumulative gain transmitted from source u to target v is Z ∆su→v (t) = ψu→v (a, t − ξ)Nu (dξ, da). (33) [0,t)×R+
By Theorem 4, its event-wise form is X ∆su→v (t) = ψu→v (Su (tuk ), t − tuk ) .
(34)
u∈N (v)∪{v}
tu k ≤t Fu (tu k )=1
ψu→v (Su (tuk ), t − tuk ) .
Substitution yields the history-based fault information gain conf ∆sF + v (t) = ∆sv X X Wu→v u∈N (v)∪{v}
u u 1 − e−Su (tk ) e−γu→v (t−tk ) .
tu k ≤t Fu (tu k )=1
(40) Because fault probability increases with information gain, we map the gain to a valid probability using the sigmoid function: P Fv (t + 1) = 1 Hv (t), HN (v) (t), zv = σ ∆sF v (t) X = σ µv (zv ) + Wu→v (41) u∈N (v)∪{v} ! X u u × 1 − e−Su (tk ) e−γu→v (t−tk ) .
where µv (zv ) maps static configuration zv to background abnormal information. V. FEMN ET
Consequently, the total fault information gain is X
(39)
tu k ≤t Fu (tu k )=1
tu k ≤t Fu (tu k )=1
conf ∆sF + v (t) =∆sv X
(37)
tu k ≤t Fu (tu k )=1
(35)
Building on IEFP, FEMNet couples conditional diffusion with a fault-aware denoising network to predict KPI distributions under complex fault scenarios. This section presents its overall architecture and details denoising mechanism.
Conditional Inputs
Generative Backbone: Conditional Diffusion
Fault-aware Denoising Network
Topology Structure G = ( V, E, A )
Forward Noising Process (training only)
Informantion-Entropy-Driven Fault Propagation Paragram
Gradually inject Gaussian noise
A: adjacency matrix
...
y0
Network Attributes e.g., link capacity, queue length, etc.
Relative entropy
yτ
y1
yT
...
ŷ1
...
ŷ0
P ( Fv (t +1) = 1| ℌv (t), ℌN(v) (t), zv )
ŷτ
ŷT
Fault propagate over nodes/edges
Message Passing
e.g., fault type, fault location, etc.
Faults occur and risk spread over time
Noised Samples
3. Denoising Prediction ϵө( yτ, G, Z, S, ℇ0:t , τ )
Risk Attention
Risk Gate
ρuv
πu→v
Focus on high-risk areas
Training (learn θ) 2. Forward Noising
Low
High Fault Risk
Fault-Aware Graph Massage Passing
Fault Propagation Guide Denoising
1. Sample
Iv(conf)+Iv(local)+Iv(prop)
...
Fault Probability Calculation
Historical Fault Events
Ground-Truth KPI
Fault Impact Decomposition
Gradually recover target samples
e.g., flow rates, OD demand, etc.
TEu→v
I (Fv (t+1) = 1;ℌv (t)|zv) Historical Fault Events t3 tm t t1 t2
Reverse Diffusion Process (Training & Inference)
Traffic Information
Transfer Entropy
Mutual Information
Abnormality Information Intrinsic Fault Tendency Directional Temporal Influence
...
Propagation Aggregation Mask /Update
χuv 1 1 0 1 0 0 1 0 0
Inference (prediction distribution) 4. Loss
1. Start from noise
Minimize
2. Revise Diffusion (T→0) Iterate
ϵ − ϵө
3. KPI Samples
4. Distributional KPIs
Repate S times {ŷ0}(S)~ p ( y0 | G, Z, S, ℇ0:t )
Variance Mean /Quantiles PDF
Exceedance
...
Backpropagate to update θ
Fig. 2. FEMNet architecture and end-to-end workflow. Upper blocks trace conditional information through the diffusion backbone and fault-aware denoising network, whereas lower blocks distinguish model training (left) from repeated-sampling inference for predictive KPI distributions (right).
A. Model Overview Fig. 2 shows FEMNet’s three components: conditional inputs, a generative backbone, and a fault-aware denoising network. The inputs comprise topology structure, network attributes, traffic information, and historical fault event sequences. Conditional diffusion generates probabilistic KPI distributions rather than deterministic point estimates. Faultaware graph message passing mechanism injects propagation context into denoising. The conditional diffusion backbone models the target KPI as a random variable and learns its probabilistic distribution. During training, the forward noising process gradually perturbs ground-truth KPI samples. FEMNet learns the reverse process conditioned on network and fault information. During inference, FEMNet starts from Gaussian noise and iteratively denoises it into conditional KPI samples. Repeated sampling approximates the future KPI distribution, capturing its fluctuation range, uncertainty level, and shape under fault scenarios. A key design of FEMNet lies in its fault-aware denoising network. IEFP first maps historical fault event sequences to fault risks over nodes and links. As dynamic propagation context, these risks guide graph message passing through risk gates, attention weights, and propagation masks. Consequently, the reverse process captures both the current network snapshot and the historical propagation and cumulative effects of faults.
At step τ , it predicts the noise added to KPI sample yτ : ϵ̂τ = ϵθ (yτ , τ, C). The corresponding clean-KPI estimate is √ yτ − 1 − ᾱτ ϵ̂τ √ ŷ0 = . ᾱτ
(43)
The predictor is evaluated at each reverse step, and the final sample is transformed back to the original KPI scale. Using (41), FEMNet calculates next-timestep node failure probabilities as fault risk intensity embedding rv . we initialize h0v = [xv (t), rv , εk ],
(44)
where [·] denotes concatenation. When v is fault-free, εk = 0. hiv is the learned fault-aware state after i message-passing layers. We use hif , hiq , and hil for flow, queue, and link nodes, respectively. We first introduce risk gate ρuv to control the effective amount of propagated message: ρuv = MLP([ru , rv ]).
(45)
Here, ρuv combines the endpoint risk embeddings. To capture heterogeneous influence across neighbors, we introduce risk attention. Attention bias πu→v from u to v is exp (κu→v ) , exp (κu′ →v )
(46)
κu→v = wa (hu , hv ) + wb (ru ).
(47)
πu→v = P
B. Fault-Aware Denoising Network The fault-aware denoising network serves as the reverseprocess predictor, which effectively model the fault impact.
(42)
u′ ∈N (v)
where
Here, wa and wb are learnable scoring functions. The attention mechanism emphasizes high-risk entities or edges when propagation is likely to affect the target. We further apply a dynamic propagation mask to suppress irrelevant messages under weak fault coupling. When cumulative risk or propagation relevance falls below a learnable threshold ϑ, the corresponding edge is disabled. Let χuv ∈ {0, 1} denote this mask: ( 1, ρuv ≥ ϑ χuv = . (48) 0, ρuv < ϑ Thus, ρuv controls message magnitude, while χuv removes weakly coupled propagation paths. Propagation differs by edge type. Uv (·) is implemented as GRU. For flow-node updates, queue and link messages, modulated by risk gate and propagation mask, are sequentially aggregated along the routing path. Namely, miq→f = χf q ρf q hiq ,
(49)
mil→f = χf l ρf l hil ,
(50)
hi+1 ← Uf (hif , [miq→f , mil→f ]). f
(51)
f ∈N (q)
(53)
Link’s state similarly depends on its traversing flows and connected queues, giving the message and update: X mil = πf →l ρlf hif + χlq ρlq hiq , (54) f ∈N (l)
hi+1 ← Ul (hil , mil ). l
Our ns-3 simulations use the Abilene, GEANT and Germany50 topologies and corresponding traffic matrices from SNDlib [36], a widely used data library for telecommunication network design and performance evaluation. We inject diverse faults and record propagation trajectories and KPI samples. After network failures, ns-3 simulates the redistribution of network resources and the resulting utilization changes. This process may establish a new equilibrium or trigger subsequent degradation and failure. Thus, the dataset captures both direct multi-fault impacts and cascading propagation effects. The injected faults cover five categories: node/device faults (node or interface outages and degradation), link faults (disconnection, bandwidth reduction, delay, packet loss, and bit errors), control-plane faults (adjacency loss, slow convergence, route-cost changes, and flapping-induced loops or black holes), data-plane faults (congestion, queue overflow, and RED/WRED drops), and application-level faults (flash crowds and DDoS traffic). Combining these faults with different topologies, traffic conditions, and network configurations yields scenarios with diverse propagation patterns and performance impacts for evaluating IEFP and FEMNet.
B. Experimental Settings
Queue’s state depends on its traversing flows and associated link. Risk attention is incorporated into multiflow aggregation to capture flow-specific contributions. Uq integrates queue features with risk-modulated flow and link aggregates to update the queue state: X miq = πf →q ρqf hif + χql ρql hil , (52)
hi+1 ← Uq (hiq , miq ). q
A. DataSet Generation
(55)
Collectively, these edge-type-specific updates yield faultconditioned representations for ϵθ , thereby linking the IEFPderived risk context to conditional KPI distribution generation. VI. E XPERIMENTS This section first describes the dataset generation and experimental settings, then evaluates IEFP and FEMNet in fault prediction, deterministic and probabilistic KPI prediction, component ablations, and efficiency.
We split the dataset into training, validation, and test sets at 60%:20%:20%. All methods are trained with Adam for 500 epochs, a learning rate of 0.001 and a batch size of 16, and the checkpoint with the best validation performance is selected. All diffusion models employ 100 denoising steps, and we generate 50 KPI samples per test instance. For a fair comparison, all methods use identical input features, including fault information. Experiments are implemented in PyTorch on Intel Xeon Silver 4410Y CPU and NVIDIA RTX A6000 GPU. We compare FEMNet with representative learning-based methods to access deterministic performance prediction and evaluate IEFP against GPIAN, HIC model, Causal model, and related fault-prediction methods. For probabilistic prediction, we compare FEMNet with the denoising-backbone baselines, GNN-Denoise and Rt-Denoise (RouteNet-Erlang), and diffusion models, including ReDiSC, SimDiff, and LGD. We further ablate the main components of FEMNet and IEFP, and profile memory and runtime costs. Deterministic metrics are mean squared error (MSE), mean absolute error (MAE), mean absolute percentage error (MAPE), mean percentage error (MPE) and Coefficient of Determination (R2 ). Fault prediction is assessed via accuracy, precision, recall, F1 score, and true negative rate (TNR). Probabilistic metrics are continuous ranked probability score (CRPS), normalized energy score (ES-norm), prediction interval coverage probability (PICP), mean prediction interval width (MPIW), and interval score (IS) for 90% and 95% prediction intervals (PIs). We also report inference and training time, GPU memory, and parameter count.
TABLE I D ETERMINISTIC KPI- PREDICTION PERFORMANCE OF FEMN ET AND THE BASELINES .
STGNN DGAT xNet EAGLE DCRNN Routenet-Fermi FEMNet
MSE ↓ 0.4928 0.4139 0.4281 0.3962 0.4028 0.3768 0.2282
Delay MAE ↓ MAPE ↓ 0.3278 66.53% 0.3154 63.91% 0.3263 71.83% 0.3069 63.85% 0.3003 58.48% 0.2809 61.30% 0.1678 38.29%
R2 ↑ 0.7137 0.7595 0.7512 0.7698 0.7659 0.7811 0.8674
MSE ↓ 0.6964 0.6163 0.5870 0.4854 0.5263 0.4817 0.2814
Loss-Ratio MAE ↓ MAPE ↓ 0.3848 113.01% 0.3869 123.04% 0.3833 124.48% 0.3448 113.42% 0.3524 113.37% 0.3368 101.81% 0.2147 55.17%
R2 ↑ 0.7232 0.7551 0.7667 0.8071 0.7908 0.8086 0.8881
MSE ↓ 0.6519 0.5714 0.5715 0.5686 0.5637 0.5613 0.3627
Throughput MAE ↓ MAPE ↓ 0.3844 99.47% 0.3663 91.55% 0.3319 91.04% 0.3999 80.93% 0.2989 83.27% 0.3002 83.74% 0.1868 54.02%
R2 ↑ 0.8110 0.8344 0.8343 0.8307 0.8366 0.8372 0.8949
Fig. 3. Cumulative distribution functions (CDFs) of signed MPE for test-set predictions by FEMNet, RouteNet-Fermi, and DCRNN. TABLE II P ROBABILISTIC DELAY- PREDICTION PERFORMANCE OF FEMN ET AND THE BASELINES .
GNN-Denoise Rt-Denoise ReDiSC SimDiff LGD FEMNet
CRPS ↓
ES-norm ↓
0.4077 0.2717 0.2302 0.2088 0.1577 0.1305
0.6607 0.5283 0.4496 0.4044 0.3336 0.2919
90% Prediction Interval PICP → 90% MPIW ↓ IS ↓ 0.8408 2.2957 3.8813 0.8949 1.7081 2.6046 0.8864 1.0802 2.1667 0.8989 1.0887 1.8511 0.9031 0.8576 1.4708 0.9068 0.6451 1.2640
TABLE III FAULT PREDICTION PERFORMANCE OF IEFP AND THE BASELINES .
LSTM GNN Causal model HIC model GPIAN IEFP
Accuracy 0.7149 0.7915 0.8030 0.8324 0.8813 0.9225
TNR 0.7961 0.8311 0.8369 0.8544 0.9466 0.9704
Precision 0.5653 0.6822 0.6980 0.7099 0.8673 0.9162
Recall 0.6176 0.6621 0.6949 0.8557 0.7550 0.8159
F1 score 0.5542 0.6628 0.6506 0.7319 0.8085 0.8494
Fig. 4. Precision–recall (PR) curves and Receiver operating characteristic (ROC) curves for fault prediction by IEFP and baselines. Legend entries report the corresponding areas under the curves (PR-AUCs and ROC-AUCs).
95% Prediction Interval PICP → 95% MPIW ↓ IS ↓ 0.9035 2.9657 4.7136 0.9324 2.1112 3.1674 0.9245 1.3455 2.8388 0.9354 1.3297 2.3755 0.9367 1.0683 1.8917 0.9432 0.7903 1.7043
C. Evaluation Results Deterministic performance Prediction. Table I compares FEMNet with deterministic prediction baselines on three KPIs. FEMNet consistently achieves the best performance for every KPI–metric pair. Compared with the best-performing baseline for each pair, FEMNet achieves average reductions of 38.6%, 37.9%, and 38.9% in MSE, MAE, and MAPE, respectively, and an average improvement of 9.2% in R2 across the three KPIs. Fig. 3 further shows that FEMNet generally produces more concentrated signed-MPE errors than RouteNet-Fermi and DCRNN. FEMNet’s closer concentration around zero indicates less systematic over- or underprediction. These results show that FEMNet substantially improves deterministic KPI prediction with propagation modeling, fault-aware graph reasoning, and conditional diffusion. Fault Prediction. Table III shows that IEFP achieves the highest accuracy (0.9225), TNR (0.9704), precision (0.9162), and F1 score (0.8494), with the second-highest recall (0.8159). Relative to the strongest overall baseline GPIAN, IEFP increases all indicators. The HIC model has the highest recall (0.8557), but its lower precision and F1 score indicate more false alarms. The causal model exhibits a less favorable
Fig. 5. Predictive delay distributions produced by FEMNet for three representative instances with increasing target values from left to right. TABLE IV E FFICIENCY COMPARISON OF FEMN ET AND THE REPRESENTATIVE BASELINES FOR DELAY PREDICTION . IEFP GPU Memory Usage (MB) Inference Time (ms/batch) Training Time (ms/epoch) Parameter Count
17.42 0.848 3172 13764
T=1,S=1 30.36 10.13 6657 199370
FEMNet T=100,S=1 30.36 848 6871 199370
TABLE V FEMN ET ABLATIONS . T HE W / O FAULT, M OD AND D IFF VARIANTS REMOVE FAULT INFORMATION , FAULT- AWARE MESSAGE MECHANISM , AND CONDITIONAL DIFFUSION , RESPECTIVELY
w/o Fault w/o Mod w/o Diff FEMNet
Delay MSE MAE 0.5352 0.3813 0.3481 0.2381 0.2551 0.1855 0.2282 0.1678
Loss-Ratio MSE MAE 0.7128 0.4264 0.4460 0.2763 0.3125 0.2258 0.2814 0.2147
Throughput MSE MAE 0.6738 0.3568 0.5376 0.3250 0.3941 0.1978 0.3627 0.1868
TABLE VI IEFP ABLATIONS . T HE W / O Wprop , S EVERITY AND C ONFIG VARIANTS REMOVE EDGE - SPECIFIC PROPAGATION WEIGHT, ENTROPY- DERIVED SEVERITY, AND CONFIGURATION TERM , RESPECTIVELY. S HARED D ECAY REPLACES EDGE - SPECIFIC DECAY RATES WITH SINGLE LEARNABLE RATE .
w/o Wprop w/o Severity w/o Config Shared Decay IEFP
Accuracy 0.8766 0.8896 0.9013 0.9027 0.9225
TNR 0.9276 0.9419 0.9422 0.9468 0.9704
Precision 0.8212 0.8729 0.8865 0.8947 0.9162
Recall 0.7871 0.7618 0.7844 0.7915 0.8159
F1 score 0.7973 0.8069 0.8155 0.8246 0.8494
precision-recall balance. Figs. 4 provide threshold-independent support: IEFP obtains the highest PR-AUC (0.9318) and ROCAUC (0.9576). These results show that IEFP more reliably anticipates future faults while maintaining a favorable balance between detection and false alarms. Probabilistic Performance Prediction. Table II reports that FEMNet achieves the lowest CRPS (0.1305) and ESnorm (0.2919), reducing by 17.3% and 12.5% against LGD, respectively. For the 90% PI, SimDiff has the PICP closest to nominal; FEMNet remains near nominal at 0.9068 and has lowest MPIW (0.6451) and IS (1.2640). For the 95% PI, FEMNet has the closest PICP (0.9432) and lowest MPIW (0.7903) and IS (1.7043). Although FEMNet is not closest to nominal at 90%, its narrower interval and lower score demonstrate better
T=100,S=8 30.36 5947 7003 199370
LGD (T=100,S=8) 38.72 2956 4691 295175
Routenet-Fermi 29.22 9.76 5564 105985
sharpness. Overall, FEMNet improves distributional accuracy without relying on excessively wide, conservative intervals for coverage. Fig. 5 visualizes three representative delay cases. Each target lies within both the 90% and 95% PIs, while the predictive mean and median remain consistent with the target scale. The low-delay case is relatively concentrated, whereas the higher-delay cases are broader and more dispersed, which illustrate instance-dependent, heteroscedastic uncertainty. Ablation Studies. As shown in Table V, removing fault information causes the largest errors, which demonstrates fault context is essential for understanding KPI degradation. Without fault-aware message passing, MSE and MAE increase by 52.7% and 47.4%, respectively, which verifies that fault features alone cannot replace risk-modulated graph reasoning. Conditional diffusion provides smaller but consistent gains for modeling uncertainty and complex latent distributions. Table VI shows that edge-specific propagation weights are IEFP’s most influential component: removing them lowers accuracy, TNR, precision, and F1 score by 4.59, 4.28, 9.50, and 5.21 percentage points, respectively, indicating overly broad propagation. Entropy-derived severity removal most harms recall because minor and severe fault events become less distinguishable. Configuration term removal degrades all metrics, which supports background-vulnerability modeling. Shared decay has the smallest impact, suggesting temporal heterogeneity provides a modest marginal contribution. Model Efficiency. Let N and M denote the numbers of nodes and edges, respectively. Let d, L, K, T , S, and P denote the hidden feature dimension, message-passing depth, historical fault count, reverse diffusion step count, inference sample count, and target-KPI dimension, respectively. The time complexity is O KM + ST [L(N d2 + M d) + P ] and space complexity is O(K + N d + M + P ). Table IV shows that IEFP is lightweight, using 17.42 MB and 0.848 ms/batch. At T = S = 1, FEMNet uses 30.36 MB and 10.13 ms/batch,
close to RouteNet-Fermi, indicating limited cost from fault propagation encoding and graph reasoning. Increasing T or S raises inference latency nearly linearly, consistent with multistep diffusion and repeated sampling. Memory and parameter count remain fixed because the denoiser is reused. Training time increases slightly because each instance samples one diffusion step and evaluates the denoiser once. At (T, S) = (100, 8), FEMNet trains and infers more slowly than LGD but uses less memory and fewer parameters. Thus, FEMNet trades between overhead and performance. VII. C ONCLUSION This paper studied probabilistic network performance prediction under complex faults. IEFP models fault propagation as abnormal information transmission over networks, and faultaware graph message passing injects predictive fault risk into representation learning. FEMNet converts deterministic KPI prediction into conditional distribution generation. Experiments showed improved fault prediction, deterministic and probabilistic KPI prediction; ablations confirmed each major component, while efficiency tests demonstrated lightweight IEFP and controllable FEMNet overhead. This study assumes closed-set fault types and relies on supervised historical fault events and KPI observations, whereas practical labels may be sparse, noisy, or delayed. Future work will explore openset faults, cross-domain transfer, label-efficient learning, and online adaptation. ACKNOWLEDGMENTS This work is supported by Hunan Provincial Natural Science Foundation (Grant no.2025JJ90177) and Jiangxi Provincial Natural Science Foundation (Grant no.20253BAC280098). R EFERENCES [1] J. Song, E. Cotilla-Sanchez, G. Ghanavati, and P. D. Hines, “Dynamic modeling of cascading failure in power systems,” IEEE Transactions on Power Systems, vol. 31, no. 3, pp. 2085–2095, 2015. [2] A. Varbella, K. Amara, B. Gjorgiev, M. El-Assady, and G. Sansavini, “Powergraph: A power grid benchmark dataset for graph neural networks,” Advances in Neural Information Processing Systems, vol. 37, pp. 110 784–110 804, 2024. [3] Q. Yang, X. Peng, L. Chen, L. Liu, J. Zhang, H. Xu, B. Li, and G. Zhang, “Deepqueuenet: Towards scalable and generalized network performance estimation with packet-level visibility,” in Proceedings of the ACM SIGCOMM 2022 Conference, 2022, pp. 441–457. [4] M. Ferriol-Galmés, J. Paillisse, J. Suárez-Varela, K. Rusek, S. Xiao, X. Shi, X. Cheng, P. Barlet-Ros, and A. Cabellos-Aparicio, “Routenetfermi: Network modeling with graph neural networks,” IEEE/ACM transactions on networking, vol. 31, no. 6, pp. 3080–3095, 2023. [5] I.-C. Lin, O. Yağan, and C. Joe-Wong, “Dynamic coupling strategy for interdependent network systems against cascading failures,” IEEE Transactions on Network Science and Engineering, vol. 10, no. 4, pp. 2265–2282, 2023. [6] S. Zhou, Z. Li, and J. Xiang, “Reliability analysis of dynamic fault trees with priority-and gates using conditional binary decision diagrams,” Reliability Engineering & System Safety, vol. 253, p. 110495, 2025. [7] F. Tang, L. Luo, Z. Guo, Y. Li, M. Zhao, and N. Kato, “Semidistributed network fault diagnosis based on digital twin network in highly dynamic heterogeneous networks,” IEEE transactions on mobile computing, vol. 24, no. 5, pp. 3979–3992, 2024. [8] S. Wang, Y. Li, K. Noman, D. Wang, K. Feng, Z. Liu, and Z. Deng, “Cumulative spectrum distribution entropy for rotating machinery fault diagnosis,” Mechanical Systems and Signal Processing, vol. 206, p. 110905, 2024.
[9] X. Zhang, W. Hu, F. Yang, W. Cao, and M. Wu, “A new transfer entropy approach based on information granulation and clustering for root cause analysis,” Control Engineering Practice, vol. 140, p. 105669, 2023. [10] M. Bauer, J. W. Cox, M. H. Caveness, J. J. Downs, and N. F. Thornhill, “Finding the direction of disturbance propagation in a chemical process using transfer entropy,” IEEE transactions on control systems technology, vol. 15, no. 1, pp. 12–21, 2006. [11] L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” in Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 3836–3847. [12] Z. Li, L. Xia, H. Hua, S. Zhang, S. Wang, and C. Huang, “Diffgraph: Heterogeneous graph diffusion model,” in Proceedings of the Eighteenth ACM International conference on web search and data mining, 2025, pp. 40–49. [13] K. Rusek, J. Suárez-Varela, P. Almasan, P. Barlet-Ros, and A. CabellosAparicio, “Routenet: Leveraging graph neural networks for network modeling and optimization in sdn,” IEEE Journal on Selected Areas in Communications, vol. 38, no. 10, pp. 2260–2270, 2020. [14] S. Huang, Y. Wei, L. Peng, M. Wang, L. Hui, P. Liu, Z. Du, Z. Liu, and Y. Cui, “xnet: Modeling network performance with graph neural networks,” IEEE/ACM Transactions on Networking, vol. 32, no. 2, pp. 1753–1767, 2023. [15] M. Ferriol-Galmés, K. Rusek, J. Suárez-Varela, S. Xiao, X. Shi, X. Cheng, B. Wu, P. Barlet-Ros, and A. Cabellos-Aparicio, “Routeneterlang: A graph neural network for network performance evaluation,” in IEEE INFOCOM 2022-IEEE Conference on Computer Communications. IEEE, 2022, pp. 2018–2027. [16] H. Zhou, R. Kannan, A. Swami, and V. Prasanna, “Htnet: Dynamic wlan performance prediction using heterogenous temporal gnn,” in IEEE INFOCOM 2023-IEEE Conference on Computer Communications. IEEE, 2023, pp. 1–10. [17] P. Yu, J. Zhang, H. Fang, W. Li, L. Feng, F. Zhou, P. Xiao, and S. Guo, “Digital twin driven service self-healing with graph neural networks in 6g edge networks,” IEEE Journal on Selected Areas in Communications, vol. 41, no. 11, pp. 3607–3623, 2023. [18] J. Liu, F. Tang, L. Chen, X. Li, J. Yu, Y. Zhu, Y. Yu, and Y. Yang, “Eagle: Heterogeneous gnn-based network performance analysis,” in 2023 IEEE/ACM 31st International Symposium on Quality of Service (IWQoS). IEEE, 2023, pp. 1–10. [19] G. Bernárdez, M. Ferriol-Galmés, C. Güemes-Palau, M. Papillon, P. Barlet-Ros, A. Cabellos-Aparicio, and N. Miolane, “Ordered topological deep learning: a network modeling case study,” arXiv preprint arXiv:2503.16746, 2025. [20] C. Güemes-Palau, M. Ferriol-Galmés, J. Paillisse-Vilanova, A. LópezBrescó, P. Barlet-Ros, and A. Cabellos-Aparicio, “Routenet-gauss: Hardware-enhanced network modeling with machine learning,” IEEE Transactions on Networking, 2026. [21] X. Zheng, W. Yao, Y. Xu, and N. Wang, “Algorithms for bayesian network modeling and reliability inference of complex multistate systems with common cause failure,” Reliability Engineering & System Safety, vol. 241, p. 109663, 2024. [22] T. Zhu, X. Yang, Z. Ma, H. Sun, J. Wu, Z. Gao, and J. Gao, “A small set of critical hyper-motifs governs heterogeneous flow-weighted network resilience,” Nature Communications, vol. 16, no. 1, p. 7809, 2025. [23] D. S. Roth, B. Tong, A. Bashan, and S. V. Buldyrev, “Cascading failures in networks of networks linked by directional and bidirectional hyperlinks,” Physical Review E, vol. 111, no. 1, p. 014315, 2025. [24] B. Xiang, B. Cautis, X. Xiao, O. Mula, D. Niyato, and L. V. Lakshmanan, “Predicting cascading failures with a hyperparametric diffusion model,” in Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2024, pp. 3495–3506. [25] S. S. Ghosh, A. Dwivedi, A. Tajer, K. Yeo, and W. M. Gifford, “Cascading failure prediction via causal inference,” IEEE Transactions on Power Systems, vol. 40, no. 4, pp. 3361–3373, 2024. [26] R. Cai, S. Wu, J. Qiao, Z. Hao, K. Zhang, and X. Zhang, “Thps: Topological hawkes processes for learning causal structure on event sequences,” IEEE Transactions on Neural Networks and Learning Systems, vol. 35, no. 1, pp. 479–493, 2022. [27] S. Linderman and R. Adams, “Discovering latent network structure in point process data,” in International conference on machine learning. PMLR, 2014, pp. 1413–1421. [28] K. Yang, F. Xue, T. Huang, S. Lu, L. Jiang, and X. Xu, “Lightweight model for power grid cascading failures risk evaluation based on graph
physics-informed attention network,” Expert Systems with Applications, vol. 291, p. 128468, 2025. [29] T. Hu, J. Zhang, R. Yi, Y. Du, X. Chen, L. Liu, Y. Wang, and C. Wang, “Anomalydiffusion: Few-shot anomaly image generation with diffusion model,” in Proceedings of the AAAI conference on artificial intelligence, vol. 38, no. 8, 2024, pp. 8526–8534. [30] H. Sun, Y. Cao, H. Dong, and O. Fink, “Anomaly anything: Promptable unseen visual anomaly generation,” in 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)[Forthcoming publication], 2025. [31] L. Shen and J. Kwok, “Non-autoregressive conditional diffusion models for time series prediction,” in International Conference on Machine Learning. PMLR, 2023, pp. 31 016–31 029. [32] H. Wen, Y. Lin, Y. Xia, H. Wan, Q. Wen, R. Zimmermann, and Y. Liang, “Diffstg: Probabilistic spatio-temporal graph forecasting with denoising diffusion models,” in Proceedings of the 31st ACM international conference on advances in geographic information systems, 2023, pp. 1–12. [33] C. Zhou, X. Wang, and M. Zhang, “Unifying generation and prediction on graphs with latent graph diffusion,” Advances in Neural Information Processing Systems, vol. 37, pp. 61 963–61 999, 2024. [34] Y. Li, Y. Lu, Z. Wang, Z. Wei, Y. Li, and B. Ding, “Redisc: A reparameterized masked diffusion model for scalable node classification with structured predictions,” arXiv preprint arXiv:2507.14484, 2025. [35] H. Ding, X. Wang, T. Zhou, and T. Yao, “Simdiff: Simpler yet better diffusion model for time series point forecasting,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 40, no. 25, 2026, pp. 20 781–20 789. [36] S. Orlowski, R. Wessäly, M. Pióro, and A. Tomaszewski, “Sndlib 1.0—survivable network design library,” Networks: An International Journal, vol. 55, no. 3, pp. 276–286, 2010.