ConceptioArchivearXiv CS
arXiv CSopen access

Assessing the Generalization of Graph Neural Networks for Fault Location Across Increasing Distributed Energy Resource Penetration Levels

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

Assessing the Generalization of Graph Neural Networks for Fault Location Across Increasing Distributed Energy Resource Penetration Levels Burak Karabulut†,⋆

Olayiwola Arowoloζ

arXiv:2607.29293v1 [cs.LG] 31 Jul 2026

Dept. of Information Technology, IDLab Ghent University – imec, Ghent, Belgium {burak.karabulut, chris.develder}@ugent.be

Carlo Manna⋆

Chris Develder†

Water and Energy Transition Unit VITO, Mol, Belgium {burak.karabulut, carlo.manna}@vito.be

Abstract—Accurate fault location is critical for distribution network reliability. However, increasing distributed energy resource (DER) penetration complicates fault location due to intermittent generation and bidirectional power flows that reshape fault signatures. Spatio-Temporal Graph Neural Networks (STGNNs) have shown promise by jointly modeling spatial and temporal dependencies, but their behavior under increasing DER penetration has not been studied rigorously. In this paper, we (i) systematically benchmark spatio-temporal graph attention network (STGATv2) against purely temporal (gated recurrent unit, GRU), purely spatial (GATv2) and traditional machine learning baselines, and (ii) evaluate how well models generalize across increasing DER penetration levels (10%, 25%, 50%) on a reconfigured IEEE 123-bus feeder with multiple DER injection points and moderate-to-high impedance faults. Results show that STGATv2 consistently outperforms neural baselines, achieving 92–94% macro F1 in-distribution. Notably, generalization across penetration levels is asymmetric: training at 50% penetration retains near in-distribution F1 score at lower levels, whereas training at 10% degrades considerably at 50% — with STGATv2 retaining 81–84% F1 under these drastic shifts, substantially higher than GATv2 and GRU which drop to 69–74% F1 and 73–75% F1 respectively. Under realistic measurement noise, STGATv2 maintains > 85% F1, while GRU drops as low as 33.5% F1, highlighting the critical role of topological awareness for robust fault location in active distribution networks. Index Terms—Power Systems, Fault Location, Distributed Energy Resources, Time Series, Graph Neural Networks

I. I NTRODUCTION Promptly locating faults within power distribution systems is essential for ensuring grid reliability and minimizing downtime [1]. Identifying the faulty component — typically resulting from short circuits caused by environmental factors, hardware or insulation failure — enables operators to promptly isolate the affected area and restore service [2]. Modern distribution networks are becoming larger and increasingly complex due to the increasing integration of distributed energy resources (DERs) and the electrification of demand, e.g., electric vehicle charging [3], [4]. DERs introduce inherent variability and intermittency, as weatherdriven generation can lead to significant voltage fluctuations and load imbalances, while also altering fault propagation patterns. Specifically, bidirectional power flow enables fault

Jochen L. Cremerζ

ζ

Dept. Electrical Sustainable Energy TU Delft, Delft, Netherlands {o.a.arowolo, j.l.cremer}@tudelft.nl

currents to propagate from multiple directions rather than following a radial path within the grid [5]. Consequently, fault location methods should remain accurate and robust with increasing penetration of DERs for reliable grid operation. Existing fault location methods are generally categorized into model-based and data-driven approaches [6]. Traditional model-based techniques, such as impedance, voltage sag, and traveling wavelet methods [7]–[9], rely on static assumptions regarding topology and fault characteristics. When operating conditions deviate from fixed parameters, these methods frequently suffer from increased modeling errors [10]. To overcome these limitations, data-driven methods have been extensively studied [6]. Early machine learning (ML) approaches such as support vector machines (SVMs), and random forests (RF) use hand-crafted features to locate faults [11], which require domain expertise and may not capture complex fault dynamics across changing grid conditions. More recent deep learning methods, such as Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs), have shown potential by extracting features from measurement data, with CNNs capturing spatial correlations and RNNs modeling temporal dependencies [12]. However, these methods often overlook the underlying grid topology and the non-uniform electrical connectivity between buses [13]. This lack of topological awareness may limit performance, particularly in distribution networks with varying DER penetration levels. To leverage the graph structure of distribution systems, Graph Neural Networks (GNNs) have emerged as a promising alternative for fault location [14]. By representing buses as graph nodes and grid lines as graph edges, GNNs are able to capture spatial dependencies across the feeder through the aggregation of information from neighboring buses, known as message passing [15]. For example, [13] applies Spectral Graph Convolutional Networks (GCNs), using the feeder’s Laplacian matrix in the graph Fourier domain to capture global structural information. Still, standard GNNs rely on fixed, structural normalization [16], which can limit adaptability under changing network topologies. Thus, Graph Attention Networks (GATs) [17] and GraphSAGE [18] have been explored as robust alternatives for fault location [19], [20] by

weighting the importance of neighboring nodes or using localized aggregation through neighborhood sampling, respectively. A limitation of standard GNN solutions for fault location is that they often ignore the inherent temporal dynamics of fault events. Consequently, Spatio-temporal GNNs (STGNNs) [12] have been proposed to jointly extract spatial and temporal dependencies, with recent extensions including multi-task graph-attention fault diagnosis [21] and related voltage-sag monitoring [22] under specific DER penetrations. Despite the great potential shown by GNNs in fault location, the performance of state-of-the-art GNN models has not been explored in detail for power grids with varying DER penetration, which is increasing in present-day grids, and thus a highly relevant use case. Particularly, current literature has not covered (i) a rigorous performance comparison of the various models under different DER penetration levels since existing studies often consider few injection points, nor (ii) the generalization capability of GNN-based models to increasing DER penetration levels. To address these gaps, this paper (1) Systematically and quantitatively benchmarks spatialtemporal GNN model against the spatial GNN, temporal GRU and traditional ML baselines under multiple DER injection points to assess the impact of joint spatiotemporal modeling; and (2) Evaluates the generalization capabilities of the GNNbased approaches for fault location with increasing and unseen DER penetration levels. The rest of the paper is organized as follows: Section II presents the STGNN framework and its components, Section III discusses our experimental setup, and we present the results in Section IV while Section V concludes the paper. II. M ETHODOLOGY A. Spatial Temporal Feature Extraction — STGNN framework As outlined in Section I, STGNNs have been proposed to jointly model spatial and temporal dependencies for fault location in distribution systems [12]. Specifically, this work adopts STGNN framework tailored for distribution grid fault location under increasing DER penetration levels and challenging fault conditions. The resulting pipeline (Fig. 1c) produces node-level temporal representations zi via a GRU, which are processed through an improved GAT layer to extract spatial features, yielding z′i . These embeddings are then passed to a dense classifier to produce node-level predictions. At inference time, node-level outputs are aggregated using soft voting [23], by summing the class probabilities — including a ‘no fault’ case — across nodes. The final prediction ŷ is obtained by selecting the class with the highest aggregate probability, effectively reducing the influence of outliers. The graph G = (V, E) models the distribution network, where V is the set of N nodes (buses), i.e., |V | = N , and E is the set of edges (lines) interconnecting them. This graph structure is defined by the adjacency matrix A ∈ {0, 1}N ×N , where Au,v = 1 if buses u and v are directly connected, and Au,v = 0 otherwise. However, the GNN topology does not

necessarily need to be a 1-on-1 mapping of the full feeder. Instead, the graph is constructed using only measurement locations as nodes V, following the measured only graph strategy proposed in [24], reflecting the partial observability inherent in practical distribution systems. B. Temporal Feature Extraction – Recurrent Neural Networks Fault events in distribution networks exhibit temporal behavior and form time series data. To extract these temporal features, the STGNN pipeline uses GRUs over LSTMs due to their simpler structure and lower computational cost, while maintaining sufficient capability to model the short-duration temporal dependencies typical of fault events. Specifically, a fixed-length window of per-phase root mean square (RMS) voltage measurements is input to the GRU, which produces a latent representation zi for each node. These representations are passed to the GNN to incorporate spatial dependencies. C. Spatial Feature Extraction – Graph Neural Networks In the GNN module, each node v ∈ V is associated with a feature vector hv ∈ Rd , forming the node feature matrix H ∈ RN ×d , where each row corresponds to a bus in the distribution network. Although GNNs can incorporate edge features ( e.g., line impedance or distance), in this work, only node features are considered. The objective of graph-based learning is then to learn a mapping ŷ = f (G; θ) from the graph to a fault location prediction, where θ denotes the learnable parameters. GNNs learn node representations by iteratively aggregating information from neighboring nodes u ∈ N (v) across layers: H (k+1) = f (H (k) , A; θ), (k)

(1)

N ×dk

where H ∈ R represents the node representation ma(k) trix at the output of layer k. This matrix is formed by hv , with each row representing a node v ∈ V with a dk -dimensional feature vector at layer k. Note that the neighborhood N (v) includes the node itself via self-loops (v ∈ N (v)), ensuring that each node preserves its own features during aggregation. GNN architectures vary in how message passing is performed in eq. (1), particularly in how information from neighboring nodes is aggregated. In this work, we use an improved Graph Attention Network (GATv2) [25] to adaptively weight neighbors, as the relative importance of nodes shifts dynamically with the spatial distribution of DERs. Unlike fixedweight GNNs [16], this approach captures the evolving fault signatures inherent in active grids. The representation of a node v at layer k + 1 is computed as:   X , h(k+1) = ϕ αvu W h(k) (2) v u u∈N (v)

Here, W is a learnable projection matrix and the attention coefficient αvu ∈ R captures the relative importance of neighbor u for updating node v, and is calculated as:   (k) (k) exp aT · LeakyReLU(W1 hv + W2 hu )  , αvu = P (k) (k) T j∈N (v) exp a · LeakyReLU(W1 hv + W2 hj ) (3)

Fig. 1. Model architectures for fault location: (a) Shared GRU for temporal feature extraction per node; (b) Shared GATv2, where measurement sequences are treated as features (Fin = F × S) to capture spatial dependencies; and (c) STGNN pipeline, where GRU-based temporal embeddings are refined via GNN message passing. All models conclude with a classification head and soft voting to aggregate node-level probabilities. N : number of nodes, F : number of features, S: sequence length of measurement windows, Z: GRU hidden state dimension (and output for GRU(a) and GNN(b)), Z ′ : STGNN output dimension.

where a is a learnable attention vector and LeakyReLU denotes a nonlinear activation. III. E XPERIMENTAL S ETUP A. Simulation Setup and Data Collection Due to the scarcity of real-world fault data and the limited observability of distribution networks, we generate synthetic data1 via dynamic time series simulations in OpenDSS [26] using PyDSS [27]. We use the IEEE 123-bus feeder, a standard benchmark in fault diagnosis studies [12], [13], which operates at a frequency of 60 Hz and a nominal voltage of 4.16 kV. For this study, the feeder is reconfigured by opening the tie switch at (60, 160) and closing the one at (54, 94). Rerouting power through lateral branches with loads connected yields more subtle voltage variations across the measured nodes, thereby providing fault signatures that are more challenging to detect. The IEEE 123-bus feeder has a total active load of 3.49 MW. To evaluate models’ generalization capability under increasing penetration levels, DERs, including PV, wind, and battery energy storage systems (BESS), are integrated into the reconfigured feeder at nominal penetrations of 10%, 25%, and 50% of total load, representing current to near-future highpenetration deployment regimes [28]. Higher generation is achieved by progressively expanding the hosting node set, such that lower level configurations are subsets of higher ones. Consistent with typical deployment practices, DERs are first placed at single-phase lateral leaf nodes as small single-phase PV units and gradually extended toward central and upstream nodes with higher-capacity 3-phase units (3-phase PV, BESS, and wind) as aggregate generation increases. To assess the impact of spatial density, two systematic DER placement configurations are considered: (i) Localized: a configuration with more spatially biased DER placement, mostly focusing on specific feeder sections at each penetration level and, (ii) Dispersed: a configuration with a more uniform DER distribution across the network at each penetration level. Notably, in the local configuration, DER injections are primarily observed by a few proximate measurement nodes in the graph, whereas in 1 Generating synthetic data is standard practice in distribution grid fault

studies as real-world datasets are scarce and typically restricted due to security, and proprietary concerns [12], [13].

the dispersed configuration, their impact is distributed across multiple monitoring points. (see Table I for DER locations).2 All 11 short-circuit fault types [2] are simulated across 25 locations. Fault duration is set to 20 ms, corresponding to the lower bound of typical primary protection clearing times. System responses are recorded at N = 25 idealized measurement nodes with 1 ms resolution. We capture threephase RMS voltage magnitudes, as they are more accessible via existing infrastructure compared to current and provide stable fault signatures across the feeder [12]. This results in sequences of 20 samples per fault event. For each of the 25 fault scenarios, 100 simulations are run to capture a wide range of operating conditions. Specifically, bus loads are independently scaled by L ∼ U(0.5, 1.3), representing off-peak to peak demand. Similarly, DER outputs are scaled using a shared factor M ∈ {0, 0.25, 0.5, 0.75, 1}, where M = 1 corresponds to peak generation conditions, such as sunny weather with strong wind. BESS follow a probabilistic operating behavior conditioned on M . For M ≥ 0.5, the probabilities of charging, discharging, and idle operation respectively are 0.6 / 0.3 / 0.1, while for M < 0.5, they are 0.3 / 0.6 / 0.1. This models higher charging activity under high generation and increased discharging under low generation. Fault resistances are sampled from Rf ∈ {10, 20, 40, 80, 100} Ω. Coupled with the switch reconfiguration, these moderate-to-high impedance faults yield subtle voltage drops (1–30 V), making fault location significantly more challenging than standard bolted faults. From the resulting time series, 59 ms windows are extracted, producing 40 sliding windows of S = 20 timesteps (20 ms). The dataset contains 2.5 million samples (50% no fault, 2% per fault location), split at the simulation level into 70% training, 15% validation, and 15% test sets, while preserving temporal alignment across measurement nodes. B. Model Training and Evaluation Fault location is formulated as a 26-class problem (25 locations and a no-fault class) and optimized using cross2 In the Localized configuration, initially a larger set of hosting nodes (9 vs. 6) is used to increase spatial density while maintaining same aggregate penetration as the Dispersed case. These partially distinct node sets provide varied injection patterns and prevent location bias.

TABLE I B US INDICES FOR FAULT, MEASUREMENT DEVICE AND DER PLACEMENT. DER S ADDED TO THE SUBSETS ARE TYPESET IN GREEN . Elements

Bus Locations

Fault Nodes

7, 13, 18, 21, 25, 29, 35, 42, 47, 51, 53, 55, 57, 62, 65, 72, 80, 83, 86, 89, 93, 97, 99, 101, 108

Number 25

Measured Nodes

1, 6, 13, 18, 25, 26, 33, 35, 36, 45, 47, 54, 60, 69, 75, 76, 81, 84, 91, 97, 103, 105, 109, 113, 300

25

10% Penetration

4, 11, 20, 24, 37, 43, 71, 88, 104

9

25% Penetration

4, 11, 16, 17, 20, 24, 37, 43, 48, 71, 76, 88, 104

13

50% Penetration

4, 11, 16, 17, 20, 24, 37, 43, 48, 49, 65, 66, 71, 76, 88, 90, 92, 96, 104, 300

20

Localized DER Configuration

Dispersed DER Configuration 10% Penetration

12, 24, 43, 66, 71, 88

6

25% Penetration

12, 16, 24, 39, 43, 48, 66, 71, 76, 88, 114

11

50% Penetration

11, 12, 16, 24, 32, 39, 43, 48, 59, 65, 66, 71, 76, 85, 88, 92, 96, 114, 300

19

Fig. 2. IEEE 123-node feeder with fault, DER, and measurement locations, shown for the 50% penetration under dispersed configuration.

entropy loss. To evaluate the contribution of spatial and temporal modeling, we consider various baselines: (i) traditional ML models, namely RF and SVM with Principal Component Analysis (PCA), to establish a performance reference on clean data; (ii) a purely temporal model based on a GRU, where each measurement node is processed independently, (iii) a purely spatial model based on GATv2 with residual connections, where each time window is flattened into a node feature vector (Fin = F × S, see Fig. 1). All neural network models, including the STGNN, produce node-level predictions that are aggregated into graph-level outputs using a soft voting scheme, as described in Section II-A. Hyperparameters are selected empirically based on architectural characteristics. The STGNN and GRU baseline use a recurrent hidden size of 128 for temporal feature extraction. In the STGNN, this is followed by a GNN layer with a hidden dimension of 64. The GATv2 baseline has a hidden dimension of 128, which is then projected to 64 dimensions. Neural network models are implemented in PyTorch, using PyTorch Geometric [29] for GNN architectures, and trained on an Intel i5-9500 CPU with 64 GB RAM. Input features (V1 , V2 , V3 ) are Z-score normalized. Node embeddings are mapped to 26-dimensional logits via a fully connected layer, with ReLU activations. During training, PyTorch’s crossentropy loss internally applies log-softmax for numerical stability. We apply dropout of 0.35 and batch normalization; GATv2 models use attention dropout of 0.3 with four heads. All models are trained until convergence, using AdamW with learning rate 0.0005 and weight decay 1 × 10−4 . Performance is evaluated using the macro F1-score on the test set. IV. R ESULTS AND D ISCUSSION A. Comparing Spatio-Temporal GNN with baseline models Table II reports fault location performance for all neural models across DER penetration levels for both configurations.

While traditional ML baselines achieve high in-distribution performance (87–93% F1 for RF; 83–87% F1 for PCASVM), their inherent sensitivity to distribution shifts leads to severe degradation under unseen DER penetration levels3 , precluding them from detailed discussion in subsequent subsections. Across both configurations, STGATv2 maintains indistribution F1 scores (diagonals in Table II) roughly 5-10 points higher than GATv2 and 8-19 points higher than GRU baselines. This performance gap suggests that jointly modeling spatial and temporal dependencies is more informative than either alone: while the GRU captures the temporal dynamics of fault-induced voltage drop, it lacks the topological awareness to correlate these drops across the feeder topology. Conversely, GATv2 leverages graph topology to capture fault signatures but lacks the temporal memory to distinguish them from load variations. By extracting temporal features before propagating them via attention-based aggregation, STGATv2 jointly captures spatio-temporal characteristics of active networks. Looking at in-distribution performance more closely, GATv2 consistently outperforms GRU across both configurations (e.g., 86.13% vs. 74.43% F1 at 25% localized, with the gap most pronounced under higher penetration in this setting), suggesting that attention-based message passing over flattened temporal features may provide a stronger inductive bias than sequential temporal processing alone. In particular, STGATv2 and GRU achieve a higher F1 score in the dispersed configuration, particularly at higher training penetration levels, as its uniform DER distribution ensures that fault-induced voltage drops are observed more consistently across measurement nodes. Furthermore, having fewer DER hosting nodes requires higher per-node injections that act as significant negative loads. This leads to more pronounced voltage shifts during faults, making signatures easier to capture. However, GATv2 shows no consistent direction between 3 Under increasing DER levels (e.g., train on 10% and test on 50%), F1 scores drop to as low as 47% (RF) and 60% (PCA-SVM).

TABLE II FAULT L OCATION M ACRO F1 (%) ± S TANDARD D EVIATION ACROSS 3 RANDOM SEEDS , BY T RAIN /T EST DER P ENETRATION L EVELS

Model

Localized DER Configuration

Dispersed DER Configuration

Test (%)

Test (%)

Training (%) 10

25

50

10

25

50

STGATv2

10 25 50

94.07 ± 1.82 93.77 ± 0.37 90.70 ± 0.71

92.97 ± 1.53 93.63 ± 0.42 89.60 ± 0.85

81.47 ± 0.97 85.77 ± 0.72 92.37 ± 0.55

94.03 ± 0.24 93.73 ± 0.46 91.67 ± 2.44

93.93 ± 0.20 94.03 ± 0.10 91.03 ± 2.90

83.53 ± 0.89 85.90 ± 2.05 93.03 ± 1.07

GATv2

10 25 50

85.83 ± 2.05 86.70 ± 1.95 83.30 ± 3.05

84.53 ± 2.40 86.13 ± 1.50 82.73 ± 3.38

73.93 ± 2.41 75.37 ± 2.13 86.70 ± 2.70

85.50 ± 1.66 84.17 ± 0.97 86.23 ± 2.40

83.93 ± 1.81 83.83 ± 0.77 85.33 ± 2.18

69.30 ± 2.19 77.63 ± 3.21 87.93 ± 1.90

GRU

10 25 50

84.53 ± 2.85 77.57 ± 3.46 76.83 ± 6.31

83.73 ± 2.94 74.43 ± 4.16 76.53 ± 6.76

73.50 ± 2.36 70.70 ± 0.57 76.53 ± 6.30

82.47 ± 2.47 83.07 ± 2.82 85.27 ± 1.59

81.40 ± 1.91 82.73 ± 2.98 84.03 ± 1.80

74.73 ± 2.06 76.10 ± 3.06 85.13 ± 1.44

configurations, with slightly better performance in the 25% localized in-distribution case suggesting that spatial locality can provide a concentrated feature for attention mechanisms to leverage, while the uniform distribution of the dispersed configuration can dilute this spatial focus. B. Asymmetric Generalization to DER Penetration Level Beyond superior in-distribution performance, STGATv2 generalizes better across varying penetration levels compared to baseline models. However, this generalization is inherently asymmetric: models trained at 50% penetration retain their performance at lower levels, dropping at most 4 points. Conversely, training at 10% for 50% testing leads to substantial degradation; while STGATv2 in the localized configuration drops to 81.47% F1, GATv2 and GRU drop more significantly to 73.93% and 73.50% F1, respectively. Interestingly, this drop is not gradual, as models trained at 10% lose at most 1.6 points at 25% whereas 16.2 at 50%, for our configurations. This suggests high penetration, with complex bidirectional flows, includes fault patterns seen at lower levels. Thus, training at high penetration provides a diverse feature set that generalizes downward, while models trained at low penetration fail to adapt to complex signatures under severe DER shifts. Notably, the models generalize more effectively in the dispersed configuration compared to the localized setup. Consistent with the discussion in Section IV-A, the uniform DER distribution provides relatively more stable measurement patterns across measurement nodes. This allows STGATv2 to retain 83.53% F1 in the dispersed configuration versus 81.47% in localized under the 10% to 50% shift, whereas GRU experiences a narrower gap between configurations (74.73% F1 vs. 73.50% F1). In contrast, GATv2 shows a larger gap between configurations under this severe shift, landing at 69.30% F1 in the dispersed configuration versus 73.93% in localized, as its attention mechanism may become diluted across distributed DER signatures. Ultimately, integrating spatial reasoning with temporal dynamics is essential to maintain fault location performance as DER penetration increases.

C. Impact of Measurement Noise on Model Performance Following the assessment of generalization capabilities, we investigate model robustness to measurement noise while maintaining baseline DER penetration levels, isolating noise as the sole source of distribution shift. This is particularly critical as moderate to high fault impedance combined with the reconfigured feeder leads to weaker fault signatures, yielding notably subtle voltage drops (1–30 V). Measurement noise is introduced via zero-mean Gaussian perturbations with signalto-noise ratios (SNRs) of 50 dB and 45 dB. The standard deviation of the noise is calculated as σnoise = 10−SNR/20 . Noise is injected for each sample in the batch independently, resulting in unseen test perturbations. At 50 dB, model performance decreases across the board. STGATv2 remains the most consistent with an F1 drop of up to 6 points, maintaining performance between 88% and 90% F1 across both configurations. The purely spatial GATv2 model drops up to 7 points, landing at around 81% F1. On the other hand, the sequential GRU shows higher sensitivity to noise; its performance falls up to 15 points to 70%– 72% F1 on localized and up to 23 points to 62.5% F1 on dispersed. This sharper decline occurs because GRU lacks spatial awareness to isolate subtle fault-induced voltage drops masked by noise. When noise is increased to 45 dB, the importance of topological awareness becomes more evident. STGATv2 retains a score between 85% and 87% F1, as the joint spatio-temporal modeling allows it to effectively filter out the measurement noise. GATv2 experiences a higher drop compared to STGATv2 as it lacks the temporal memory needed to distinguish the fault dynamics, landing at around 74% F1. Meanwhile, GRU performance drops significantly, falling to 62% F1 (localized) and 33.5% F1 (dispersed). Surprisingly, GRU suffers a steeper F1 drop in the dispersed configuration under measurement noise. This may be attributed to uniform DER distribution, combined with the noise masking already subtle fault signatures across the network, whereas, in the localized configuration, nodes away from the DER injection points remain more distinguishable.

In summary, we observe that topological (spatial) awareness appears to be fundamental for model robustness in distribution grid fault location, especially under subtle fault signal conditions. While modeling both spatial and temporal correlations provides the highest robustness to noise, the temporal GRU model struggles to identify fault patterns once the signal-tonoise ratio degrades local measurement integrity. V. C ONCLUSION AND F UTURE W ORK This work systematically (i) benchmarks STGATv2 against purely spatial, purely temporal, and traditional ML baselines under multiple DER injection points to assess the impact of joint spatio-temporal modeling, and (ii) assesses model generalization to increasing and unseen DER penetration levels for distribution network fault location. Our results confirm that jointly modeling spatial and temporal dependencies outperforms baselines (92–94% F1). Notably, generalization across DER penetration is asymmetric: models trained at high penetration (50%) retain near in-distribution performance at lower levels, whereas models trained at low penetration (10%) degrade severely under high penetration shifts. Despite this degradation, STGATv2 retains 81–84% F1, while baselines drop sharply. Furthermore, under realistic measurement noise, STGATv2 maintains robust performance (> 85% F1) compared to severe baseline degradation, suggesting that topological awareness is critical for robustness. Future work includes extending this framework to larger, diverse networks to evaluate model generalization and sensitivity to DER placement configurations. Additionally, evaluating extreme DER penetration regimes (up to 100%) will provide insights into robustness under severe voltage volatility and operational limits, while incorporating detailed grid and inverter-based dynamics will support validation toward real-world deployment. ACKNOWLEDGMENT Research reported in this publication was supported by VITO grant number VITO UGENT PhD 2301 and partially funded by the Flemish Government (under the ’Onderzoeksprogramma Artificiële Intelligentie (AI) Vlaanderen’ programme). This research was also supported by the Research Foundation - Flanders (FWO) under grant number V425326N. R EFERENCES [1] IEEE, Piscataway, NJ, USA, IEEE Guide for Electric Power Distribution Reliability Indices, 2022. IEEE Std 1366-2022 (Revision of IEEE Std 1366-2012). [2] J. J. Grainger and W. D. Stevenson, Power System Analysis. New York: McGraw-Hill, 1994. [3] International Energy Agency, “Electricity grids and secure energy transitions,” tech. rep., IEA, Paris, France, 2023. [4] J. Liu, H. Shi, Y. Chen, C. Yang, M. Ma, and Y. Li, “Graphgan-based fault detection and location for complex power grids,” in Proc. Int. Conf. Energy Power Electr. Technol. (CEPET 2025), (Wuhan, China), pp. 465– 469, IEEE, 2025. [5] B. B. Adetokun, C. M. Muriithi, J. O. Ojo, and O. Oghorada, “Impact assessment of increasing renewable energy penetration on voltage instability tendencies of power system buses using a QV-based index,” Sci. Rep., vol. 13, p. 9782, 2023. [6] H. Rezapour, S. Jamali, and A. Bahmanyar, “Review on artificial intelligence-based fault location methods in power distribution networks,” Energies, vol. 16, no. 12, p. 4636, 2023.

[7] S. Chandran, R. Gokaraju, and K. Narendra, “An extended impedancebased fault location algorithm in power distribution system with distributed generation using synchrophasors,” IET Gener. Transm. Distrib., vol. 18, no. 3, pp. 479–490, 2024. [8] R. F. Buzo, H. M. Barradas, and F. B. Leão, “A new method for fault location in distribution networks based on voltage sag measurements,” IEEE Trans. Power Deliv., vol. 36, no. 2, pp. 651–662, 2021. [9] F. Liu, L. Xie, K. Yu, Y. Wang, X. Zeng, L. Bi, and X. Tang, “A novel fault location method based on traveling wave for multi-branch distribution network,” Electr. Power Syst. Res., vol. 224, p. 109753, 2023. [10] M. MansourLakouraj, R. Hossain, H. Livani, and M. Ben-Idris, “Application of graph neural network for fault location in PV penetrated distribution grids,” in Proceedings of the North American Power Symposium (NAPS 2021), (College Station, TX, USA), pp. 1–6, IEEE, 2021. [11] K. Chen, C. Huang, and J. He, “Fault detection, classification and location for transmission lines and distribution systems: a review on the methods,” High Volt., vol. 1, no. 1, pp. 25–33, 2016. [12] B. L. H. Nguyen, T. V. Vu, T.-T. Nguyen, M. Panwar, and R. Hovsapian, “Spatial-temporal recurrent graph neural networks for fault diagnostics in power distribution systems,” IEEE Access, vol. 11, pp. 46039–46050, 2023. [13] K. Chen, J. Hu, Y. Zhang, Z. Yu, and J. He, “Fault location in power distribution systems via deep graph convolutional networks,” IEEE J. Sel. Areas Commun., vol. 38, no. 1, pp. 119–131, 2020. [14] F. Scarselli, M. Gori, A. Tsoi, M. Hagenbuchner, and G. Monfardini, “The graph neural network model,” IEEE Trans. Neural Netw., vol. 20, no. 1, pp. 61–80, 2009. [15] W. Liao, B. Bak-Jensen, J. R. Pillai, and Y. Wang, “A review of graph neural networks and their applications in power systems,” J. Mod. Power Syst. Clean Energy, vol. 10, no. 2, pp. 345–360, 2022. [16] T. N. Kipf and M. Welling, “Semi-supervised classification with graph convolutional networks,” in Proc. Int. Conf. Learn. Represent. (ICLR 2017), (Toulon, France), pp. 1–14, OpenReview.net, 2017. [17] P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Liò, and Y. Bengio, “Graph attention networks,” in Proc. Int. Conf. Learn. Represent. (ICLR 2018), (Vancouver, BC, Canada), pp. 1–12, OpenReview.net, 2018. [18] W. L. Hamilton, R. Ying, and J. Leskovec, “Inductive representation learning on large graphs,” in Proc. 31st Conf. Neural Inf. Process. Syst. (NIPS 2017), (Long Beach, CA, USA), pp. 3345–3355, Neural Information Processing Systems Foundation, 2017. [19] Z. Wang, B. Huang, B. Zhou, J. Chen, and Y. Wang, “An enhanced fault localization technique for distribution networks utilizing cost-sensitive graph neural networks,” Processes, vol. 12, no. 11, p. 2312, 2024. [20] M. Fan, J. Xia, H. Zhang, and X. Zhang, “Fault location method of distribution network based on VGAE-GraphSAGE,” Processes, vol. 12, no. 10, p. 2179, 2024. [21] W. Huang, P. Chen, Y. Huang, and S. Chen, “Fault diagnosis in active distribution networks with renewable energy using multi-task learning and graph attention networks,” Electr. Power Syst. Res., vol. 263, p. 113513, 2027. [22] S. Pan and S. Xue, “A GCN-GRU-based framework for voltage sag detection and early warning in distribution networks,” Electr. Power Syst. Res., vol. 260, p. 113353, 2026. [23] T. K. Ho, J. J. Hull, and S. N. Srihari, “Decision combination in multiple classifier systems,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 16, no. 1, pp. 66–75, 1994. [24] B. Karabulut, C. Manna, and C. Develder, “Robustness of spatiotemporal graph neural networks for fault location in partially observable distribution grids,” arXiv preprint arXiv:2401.12345, 2024. [25] S. Brody, U. Alon, and E. Yahav, “How attentive are graph attention networks?,” in Proc. Int. Conf. Learn. Represent. (ICLR 2022), (Virtual Event), pp. 1–26, OpenReview.net, 2022. [26] The Electric Power Research Institute (EPRI), “OpenDSS.” [Online]. Available: https://www.epri.com/pages/sa/opendss, 2024. [27] National Renewable Energy Laboratory (NREL), “PyDSS interface.” [Online]. Available: https://www.nrel.gov/grid/pydss.html, 2024. [28] European Environment Agency, “Share of energy consumption from renewable sources in europe,” tech. rep., EEA, 2025. [29] M. Fey and J. E. Lenssen, “Fast graph representation learning with PyTorch Geometric.” arXiv preprint arXiv:1903.02428, 2019.

Record · ID 422272 · SHA-256 932f1e725af48215
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.