Conceptio › Archive › arXiv CS
arXiv CSopen access

Offline Reinforcement Learning for Distribution-Grid Protection

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Offline Reinforcement Learning for Distribution-Grid Protection Julian Oelhaf1†* , Alexander Luce1† , Christian Bergler2 , Andreas Maier1 , Siming Bayer1

arXiv:2609.24703v1 [eess.SY] 21 Sep 2026

1

Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen–Nürnberg, Erlangen, Germany 2 Department of Electrical Engineering, Media and Computer Science, Ostbayerische Technische Hochschule Amberg-Weiden, Amberg, Germany * Corresponding author: [email protected]

Abstract—Data-driven protection may complement conventional relays in distribution grids whose operating conditions vary with distributed generation, switching events, and changing short-circuit levels. We study line-selective tripping from static trajectories of a realistically simulated CIGRE medium-voltage network using offline reinforcement learning. A convolutional Q-network receives causal voltage-current phasor and apparentimpedance features, optionally together with raw waveforms, and is trained with conservative Q-learning (CQL). A controlled sensitivity study evaluates two observation windows, reward variants, and three CQL weights under a common split and training protocol; one exploratory post-hoc run additionally increases the discount factor from γ = 0.95 to 0.99. On 225 held-out episodes, the best per-timestep result is obtained with combined input and CQL weight α = 0.9, reaching precision 0.9993, recall 0.9496, and F1-score 0.9738. Because dense pertimestep scores do not encode the terminal semantics of relay operation, we also evaluate the first non-wait action in each episode. The default combined-input agent selects the correct line-trip action first in 98.13 % of 214 fault episodes, but trips in 72.73 % of the 11 non-fault episodes. In the post-hoc run, the corresponding rates are 98.60 % and 54.55 %, respectively. The results show that dense predictive performance and terminal protection behavior can lead to different model rankings. Offline CQL therefore demonstrates strong faulted-line selection on the simulated fault episodes, while the static trajectories, small nonfault set, and single-seed post-hoc design preclude conclusions about practical relay security or deployment readiness. Index Terms—conservative Q-learning, distribution-grid protection, offline reinforcement learning, protective relaying

(CQL) addresses this setting by suppressing unsupported action values [5]; related work has applied physics-guided CQL to generator tripping [6]. A. Related Work and Positioning

Data-driven protection has predominantly been studied through supervised fault detection, classification, and localization [7]. These approaches map voltage or current measurements to fault labels or protection decisions using raw waveforms or phasor-domain quantities. Yoon and Yoon employ segmented voltage waveforms in a Transformer-based model for fault and disturbance classification [1], while Shanmugapriya and Baskaran use synchronized phasor measurements from phasor measurement units (PMUs) with a deep residual network [8]. Related controlled work has evaluated machinelearning methods for fault detection and line identification under a common experimental protocol [9]. Such objectives can yield strong predictive performance, but they generally treat observations as classification samples and do not directly encode the asymmetric consequences of waiting, tripping a healthy line, or selecting the wrong line. RL has been investigated for protection coordination and breaker control. Wu et al. formulate protective relaying as an interactive multi-agent RL problem [10]. Kordowich et al. similarly employ closed-loop PowerFactory simulations, but combine centralized teacher agents with decentralized backup I. I NTRODUCTION agents [3]. In both cases, breaker actions alter the subsequent Protection devices must isolate faults rapidly and selectively trajectory, unlike the fixed trajectories considered here. while remaining secure during switching and other benign Offline RL addresses learning from such fixed datasets but is transients. Conventional overcurrent schemes can be challenged vulnerable to overestimating weakly represented actions. CQL by bidirectional power flows, changing short-circuit levels, and mitigates this effect through conservative value estimation [5]. operating states introduced by distributed energy resources. Gao et al. apply this principle to physics-guided generatorThese developments have motivated both data-driven fault tripping control [6], whereas the present work considers lineanalysis [1] and reinforcement learning (RL)-based protection selective distribution-grid protection from static fault and schemes, ranging from relay coordination [2] to direct breaker switching-event trajectories. control in closed-loop grid simulations [3]. This leaves a methodological gap between supervised The data are precomputed EvEMTBench transient simuprotection models, which are typically evaluated as repeated lations [4], rather than interactions with a live grid. The classification decisions, and interactive RL formulations, in task is therefore offline RL: the agent must estimate relaywhich actions affect future states. Our setting lies between action values from a fixed dataset and cannot improve unsafe these cases: the policy is learned from a fixed archive, but decisions through further exploration. Conservative Q-learning its operational output is a terminal line-trip decision. The † These authors contributed equally to this work. evaluation of such offline protection policies when dense

per-timestep predictions and the first irreversible trip lead to different conclusions has received limited attention. We address that gap with a controlled study of offline CQL on static CIGRE medium-voltage trajectories. We compare causal phasor and impedance features with combined rawwaveform inputs, and vary observation length, reward design, and CQL regularization under a common split and training protocol. Beyond per-timestep precision, recall, F1-score, and false-positive rate, we evaluate the first non-wait action in each episode to distinguish correct operation, wrong-line trips, missed trips, nuisance trips, and operating delay.

TABLE I S UMMARY OF THE DATASET SUBSET USED FOR TRAINING AND EVALUATION . Property

Value

Network buses / line segments Line cubicles / raw channels Episodes: development / held-out Held-out: fault / non-fault Episode duration / sampling rate Event onset / grid frequency Trip actions / wait actions

14 / 15 29 / 174 4,282 / 225 214 / 11 0.5 s / 9.6 kHz 0.1 s / 50 Hz 15 / 1

II. S YSTEM , DATA , AND M ETHOD

rows sampled before onset, densely around onset, and sparsely thereafter. Excluding incomplete phasor windows, 1,656,784 The CIGRE medium-voltage benchmark network in its rows remain for window size W = 48 and 1,626,810 European configuration [11], shown in Fig. 1, is simulated with for W = 96. The development episodes are divided with DIgSILENT PowerFactory [12]. Trajectories are taken from seed zero into 3,853 optimization episodes and 429 internal the EvEMTBench dataset [4]. The network has 14 buses and monitoring episodes; no episode crosses a partition boundary. 15 line segments. Measurements are recorded at both terminals A comparison with the current EvEMTBench metadata showed of each line except one cubicle at MainBus 8 on line 8–14, that the project dataset contains 4,507 simulations with IDs yielding 29 line cubicles. Three-phase currents and voltages 0–4506, whereas the cleaned online metadata contains 4,509 therefore provide 174 raw channels per timestep. We use 4,507 simulations. The two additional records, IDs 4507 and 4508, are episodes of 0.5 s sampled at 9.6 kHz. This subset includes non-fault switch_ibr_trip events and were not included nine fault families, including single-phase-to-ground (1ph-G), in training or evaluation. For the 4,507 shared simulations, event 2ph, 2ph-G, and 3ph faults, and eleven non-fault switching types and targets match exactly. The experiments therefore use or disturbance types. Processed episodes align events at 0.1 s. 4,353 fault and 154 non-fault episodes, whereas the complete The agent receives neither a time index nor an onset indicator; cleaned metadata contains 156 non-fault episodes. Differences in event-time origin, arc-time-constant units, and directory alignment is used only for labels and latency evaluation. prefixes are consistent with differing metadata and file-location conventions; they do not change the event-class assignments used here. A. Simulated Distribution-Grid Episodes

B. Causal Signal Representation Raw inputs contain the 174 line-cubicle waveform channels. Under normal operation, these voltage and current waveforms oscillate at the 50-Hz fundamental. Phasors provide a more interpretable fundamental-frequency representation, while apparent impedance may supply a physically motivated inductive bias for line-selective protection. The phasor representation applies a full-cycle sliding discrete Fourier transform using N = 9600/50 = 192 timesteps. For signal p, N −1

P [n] =

2 X p[n − k]e−j2πk/N . N k=0

Voltage and current phasors give the apparent impedance Zc,ϕ [n] =

Uc,ϕ [n] = Rc,ϕ [n] + jXc,ϕ [n], Ic,ϕ [n]

for cubicle c and phase ϕ [13]. Division by very low current magnitudes is suppressed using the causal reference Fig. 1. CIGRE benchmark medium-voltage grid in the European configuration.

mc [n] = max max |Ic,ϕ [τ ]|. The held-out evaluation set contains 225 episodes (214 fault and 11 non-fault); the other 4,282 episodes form the development pool. The transition archive contains 1,686,758

τ ≤n ϕ∈{a,b,c}

All experiments use this causal maximum; a full-trajectory maximum would leak future information. The impedance is

retained when |Ic,ϕ [n]| ≥ 0.005mc [n] and replaced by zero otherwise. For each phase, the model receives |U |, |I|, R, and X, totaling 348 phasor features. The combined representation uses separate convolutional branches for the 174 raw and 348 phasor channels. At decision timestep n, the state sn is a matrix containing the most recent W timesteps across the C input channels, i.e., sn ∈ RW ×C . Thus, features at n use only measurements through n. This finite window approximates the Markov property because longer dependencies cannot be excluded. We evaluate observation lengths W ∈ {48, 96} feature timesteps, corresponding to one-quarter and one-half of a 50-Hz cycle at the feature-sequence level. Because each phasor uses a trailing 192-sample cycle, the causal raw-signal support spans 239 samples for W = 48 and 287 samples for W = 96. The action space has 16 choices: trip both breakers associated with one of the 15 line segments, or wait. C. Offline Conservative Q-Learning

LTD = E(s,a,r,s′ )∼D



2  Qθ̄ (s , a ) . Qθ (s, a) − r − γ max ′ ′

′

a

CQL adds the conservative regularization term " LCQL = Es∼D log

# X

e

Qθ (s,a)

− Qθ (s, aD ) ,

a

where aD denotes the action contained in the offline dataset. The total objective is L = LTD + αLCQL . The counterfactual examples provide direct supervision for selected unsafe actions, but they cover only a subset of the possible state–action combinations. The CQL term complements this augmentation by suppressing high Q-values for actions that are weakly supported by the offline data distribution. At inference, the policy selects

The sequential protection task is formulated as a finiteπθ (s) = arg max Qθ (s, a). horizon Markov decision process M = (S, A, P, R, γ) [14]. a∈A Here, S contains the finite-window measurement states defined All prespecified models are trained for 30 epochs with above, A comprises the 15 line-trip actions and the wait Adam at 10−3 , batch size 1024, γ = 0.95, and τ = 0.005 action, R assigns the protection-dependent rewards in (1), on one A100 MIG GPU with 16 CPU cores; the 853,604and γ discounts future rewards. In an interactive protection parameter combined model trains in about 63 min. The postenvironment, the transition model P (s′ |s, a) would describe hoc model differs only in γ = 0.99. The 429-episode internal how a trip or wait action changes the subsequent electrical monitoring partition is used to track training losses but not state. The available archive is static, however: s′ is the next to select checkpoints or compute the reported performance recorded window in the PowerFactory trajectory and is not metrics; those metrics use the separate 225-episode evaluation regenerated in response to the selected action. The problem is set. Random-number generators and data loading use seed zero; therefore an offline, action-independent approximation of the strict deterministic CUDA kernels are not enforced. interactive decision process [15]. A dilated one-dimensional convolutional network maps each TABLE II observation to 16 Q-values. Each input has shape W × C, M ODEL AND TRAINING SETTINGS . is batch-normalized channel-wise, and is rearranged to the Setting Value (B, C, W ) layout required by the one-dimensional convolutions. Phasor channels 348 For combined input, the raw-waveform and phasor branches are Combined channels 174 raw + 348 phasor pooled separately, concatenated, and passed to a fully connected Convolution kernel / dilations 7 / 1, 3, 9, 27 action head. Pooled features per branch 128 Optimizer / learning rate Adam / 10−3 With the default settings, the reward is  5,    5, r(s, a) = 0,    −100,

correct fault trip after onset, wait after onset in a non-fault episode, wait before onset or during a fault, pre-event, non-fault, or wrong-line trip. (1) The heuristic reward values reflect the asymmetric cost of disconnecting a healthy line or selecting the wrong line. Zero reward for waiting before onset or during a fault allows the policy to defer its decision without an immediate penalty, while the positive reward reinforces secure post-event waiting and correct selective tripping. Counterfactual wrong-line actions augment the fixed transition set. The temporal-difference loss uses the target network Qθ̄ :

Epochs / batch size Discount γ (default / post hoc) Target update τ

30 / 1024 0.95 / 0.99 0.005

III. E XPERIMENTAL P ROTOCOL The controlled sensitivity study has three blocks. Block A compares phasor and combined representations at W = 48 and 96 with α = 0.5. Block B uses combined input and W = 48: two variants change the false-positive penalty to −10 and −200, while a third retains the default penalties and increases the shared correct reward from 5 to 50 for both correct fault trips and post-event non-fault waits. Block C changes α to 0.1 and 0.9 around the default 0.5. Each run uses the same partitions, inputs, seed, 30-epoch stopping point, and evaluation episodes. Following review of the initial results,

(b) Non-fault: nuisance trip

(a) Fault: correct first trip

Action

Other line

TRIP

Correct line WAIT

WAIT 80

90

100

110

120

80

1ph-G short circuit, MainLn3–8 (sim. 655); +0.10 ms Agent Required

90

100

110

120

Time (ms)

Time (ms) Onset

First trip

Line energization, MainBus8 (sim. 4449); +0.42 ms Post-trip interval

Fig. 2. Representative held-out action sequences for the default combined-input W = 48 policy. (a) A correct first trip fixes the terminal fault outcome despite a later wrong-line prediction. (b) A single nuisance trip fixes a false non-fault outcome despite predominantly correct waiting. Shading marks predictions ignored by the first-trip evaluation but retained in the dense metrics.

we conducted one exploratory post-hoc run using combined Wilson 95 % confidence intervals are reported for the highinput with W = 48 and γ = 0.99. It otherwise retains the lighted non-fault false-trip rates. Correct-trip latency is summainputs, split, seed, architecture, reward, CQL weight, and 30- rized by its median and 95th percentile. The terminal evaluation epoch stopping point. The run is reported separately from the definition was fixed before computing the results. prespecified three-block matrix. IV. R ESULTS Every run saves its configuration, environment, training Table III reports the nine prespecified runs and the exhistory, checkpoints, predictions, metrics, and file hashes. ploratory post-hoc discount-factor run. Window length does Checkpoints retain Python, NumPy, and CPU random-numbernot have a representation-independent effect. Phasor-only F1 generator state so that interrupted seeded training resumes decreases from 0.9632 at W = 48 to 0.9454 at W = 96, consistently. Evaluation uses the final epoch for every configuwhereas combined-input F1 increases from 0.9637 to 0.9686. ration; neither dense scores nor first-trip outcomes guide early Combining raw and phasor data is therefore nearly neutral at stopping or checkpoint selection. W = 48 but improves F1 by 0.0231 at W = 96. The longer combined model also reduces FPR from 2.05 % to 0.32 %. A. Per-Timestep Evaluation Changing only the false-positive penalty has no measured After the fault onset, the correct line-trip action is a true effect: the −10, default −100, and −200 settings yield identical positive (TP); waiting or selecting another line-trip action predictions and metrics. The sampled transition archive contains is a false negative (FN). Before fault onset and throughout no pre-event or non-fault trip rows, so this parameter changes non-fault episodes, waiting is a true negative (TN) and any none of the sampled training rewards. Increasing the shared trip is a false positive (FP). We report precision, recall, F1- correct reward from 5 to 50 affects both correct fault trips score, and false-positive rate (FPR). Recall quantifies how and post-event non-fault waits; it lowers FPR to 0.13 % but reliably the policy selects the faulted line after fault onset, also lowers recall and F1. No monotonic reward conclusion is whereas FPR quantifies unnecessary trips before event onset supported. or during non-fault episodes. These metrics characterize dense Performance is sensitive to the tested CQL weight. Reducing action predictions, not a deployed terminal relay. To exclude α from 0.5 to 0.1 lowers recall to 0.8236 and F1 to 0.9023. incomplete phasor windows, evaluation begins only once the Increasing it to 0.9 gives the highest recall (0.9496) and F1 full causal support is available: after 239 raw samples for (0.9738), with an FPR of 0.28 %. Because the study does not W = 48 and 287 raw samples for W = 96. The post-fault include α = 0, these results characterize sensitivity only within interval is identical, but the longer-window configurations the tested CQL range. Compared with the default combinedcontain 48 fewer pre-event decisions; their FPR and first-trip W = 48 model, the exploratory γ = 0.99 run improves security therefore use a slightly shorter negative-event horizon. precision from 0.9946 to 0.9997, recall from 0.9347 to 0.9422, and F1-score from 0.9637 to 0.9701, while reducing FPR from 2.05 % to 0.10 %. It remains below the α = 0.9 model’s B. Terminal First-Trip Evaluation F1-score; one post-hoc setting cannot establish a general trend. For each saved sequence of evaluation predictions, the first action other than wait is the first trip; all later actions are First-Trip Behavior ignored. In a fault episode, it is classified as a premature trip, The terminal results in Table IV differ from the per-timestep correct-line trip, wrong-line trip, or right-censored no-trip. In a ranking. Figure 2 illustrates the asymmetry of this evaluation: non-fault episode, it is classified as a false trip or no-trip. For later prediction errors cannot undo an earlier correct trip, correct post-fault trips, latency is whereas a single nuisance trip determines the terminal outcome of a non-fault episode. All configurations have zero premature ntrip − 960 tlat = s. (2) trips. The default combined W = 48 agent selects the correct 9600

TABLE III P ER - TIMESTEP EVALUATION PERFORMANCE . P, R, AND F1 ARE PROPORTIONS ; FPR IS REPORTED IN PERCENT. Block

Configuration

W

α

P

R

F1

FPR (%)

A A A A

Phasor Combined Phasor Combined

48 48 96 96

0.5 0.5 0.5 0.5

0.9922 0.9946 0.9992 0.9992

0.9359 0.9347 0.8972 0.9397

0.9632 0.9637 0.9454 0.9686

2.97 2.05 0.31 0.32

B B B

Combined, FP penalty −10 Combined, FP penalty −200 Combined, correct reward 50

48 48 48

0.5 0.5 0.5

0.9946 0.9946 0.9996

0.9347 0.9347 0.9252

0.9637 0.9637 0.9610

2.05 2.05 0.13

C C Post hoc

Combined Combined Combined, γ = 0.99

48 48 48

0.1 0.9 0.5

0.9977 0.9993 0.9997

0.8236 0.9496 0.9422

0.9023 0.9738 0.9701

0.76 0.28 0.10

line-trip action first in 210 of 214 fault episodes (98.13 %), episodes and introduces none, while correct first trips on fault with one wrong-line trip and three no-trip episodes. Its median episodes increase from 210 to 211. Nevertheless, the 11 noncorrect-trip latency is one timestep (0.104 ms), and its 95th fault episodes, one training seed, and post-hoc design preclude percentile is 1.667 ms. The post-hoc γ = 0.99 model is correct a general conclusion that a larger γ improves security. first in 211 fault episodes (98.60 %), with one wrong-line trip and two no-trip episodes; its median latency remains 0.104 ms, Fault-Family Analysis while its 95th percentile increases to 1.927 ms. By contrast, Table V groups the first-trip outcomes into conventional the per-timestep F1 leader at α = 0.9 is correct first in 197 short circuits, high-impedance fault (HIF), and incipient faults. fault episodes (92.06 %), selects a wrong-line trip in 15, and The default combined W = 48 agent is correct first on all has two no-trip episodes. This distinction follows directly from 193 conventional short-circuit episodes, but only on 11 of 13 terminal semantics: a later wrong action cannot undo a correct HIF and 6 of 8 incipient-fault episodes. The post-hoc γ = first trip, while an early wrong-line trip cannot be repaired by 0.99 model is correct first in 192 of 193 conventional shortsubsequent predictions. circuit episodes, all 13 HIF episodes, and 6 of 8 incipient-fault episodes. Its net gain of one correct fault episode therefore TABLE IV comprises two additional correct HIF operations offset by one T ERMINAL FIRST- TRIP EVALUATION PERFORMANCE . FAULT OUTCOMES ARE fewer correct conventional short-circuit operation. The small PERCENTAGES OF 214 EPISODES . NF TRIP IS THE FALSE - TRIP PERCENTAGE HIF and incipient-fault counts preclude strong conclusions. AMONG 11 NON - FAULT EPISODES . p95 IS CORRECT- TRIP LATENCY IN MS . They nevertheless contradict a blanket conclusion that incipient Configuration Correct Wrong line No trip NF trip p95 faults are always missed: several policies select the correct line-trip action first in most or all of these eight episodes. Phasor W = 48 87.38 12.15 0.47 72.73 1.219 Combined W = 48 γ = 0.99 (post hoc) Phasor W = 96 Combined W = 96 FP penalty −10 FP penalty −200 Correct reward 50 α = 0.1 α = 0.9

98.13 98.60 94.86 89.72 98.13 98.13 97.66 93.46 92.06

0.47 0.47 5.14 9.81 0.47 0.47 0.93 0.47 7.01

1.40 0.93 0.00 0.47 1.40 1.40 1.40 6.07 0.93

72.73 1.667 54.55 1.927 45.45 1.354 72.73 1.042 72.73 1.667 72.73 1.667 54.55 2.125 36.36 6.896 54.55 2.750

TABLE V C ORRECT- FIRST- TRIP RATE (%) BY FAULT FAMILY. PARENTHESES GIVE THE EPISODE COUNT SHARED BY EVERY ROW. Configuration Phasor W = 48 Combined W = 48 Phasor W = 96 Combined W = 96 Combined α = 0.9 γ = 0.99 (post hoc)

Short circuit (193) HIF (13) Incipient (8) 89.64 100.00 95.34 93.78 93.26 99.48

61.54 84.62 84.62 38.46 92.31 100.00

75.00 75.00 100.00 75.00 62.50 75.00

Non-fault security remains the main limitation. The default combined W = 48 agent trips in 8 of 11 non-fault episodes (72.73 %; Wilson 95 % CI 43.44–90.25 %), despite its pertimestep FPR of 2.05 %. The lowest observed nuisance-trip rate occurs for α = 0.1, with 4 of 11 episodes (36.36 %; V. D ISCUSSION AND L IMITATIONS CI 15.17–64.62 %). This apparent security improvement is accompanied by only 200 correct first trips among 214 fault Per-timestep and terminal metrics support different concluepisodes, 13 no-trip fault episodes, and the largest p95 latency sions. Phasor features were sufficient for strong performance, of 6.896 ms; it is therefore consistent with greater reluctance while raw waveforms helped only with the longer observato trip rather than uniformly better protection. The α = 0.9 tion window. The tested CQL weight affected performance, and post-hoc γ = 0.99 models each trip in 6 of 11 non-fault yet the best dense F1 did not produce the best terminal episodes (54.55 %; CI 28.01–78.73 %). In the paired discount- behavior. Reward conclusions were also limited by actions factor comparison, γ = 0.99 removes the nuisance trip in two absent from the fixed archive. This gap between nominal and

operational rankings complements prior findings that cleandata performance may not predict behavior under degraded measurements [16]. The post-hoc discount-factor result has a plausible but limited explanation. Wait actions receive bootstrapped future value, whereas trip actions terminate a sampled transition, so increasing γ can favor waiting. However, the recorded next state is action-independent and the archive contains no preevent or non-fault trip rows. The nuisance-trip reduction is therefore an indirect policy effect, not evidence that a larger discount factor generally solves the security problem. These experiments establish feasibility, not deployment readiness. The static trajectories do not respond to trip actions, and development and evaluation cover one simulated network and operating distribution. Only 11 non-fault episodes are available, first-trip evaluation omits breaker dynamics, coordination time, communication delay, hardware uncertainty, and distribution shifts, and each configuration uses one training seed. The reported intervals therefore do not quantify seed variability or the low false-trip probabilities required in practice. Future work should test frozen policies in interactive or controller-hardware-in-the-loop settings with terminal breaker actions, broader non-fault coverage, topology and contingency shifts, and explicit security targets. Conventional-relay baselines and calibrated supervisory logic are needed. Dense metrics should be retained for diagnostic tracking, but model selection should optimize terminal protection costs because the first trip represents the irreversible relay command. VI. C ONCLUSION We studied offline CQL for line-selective tripping on simulated CIGRE medium-voltage trajectories using causal phasor and impedance features. Across the controlled sensitivity study, combined input with W = 96 achieved an F1-score of 0.9686, while increasing the CQL weight to 0.9 at W = 48 gave the best per-timestep F1-score of 0.9738. Window and representation effects were coupled, while the false-positivepenalty variants were identical because the sampled archive contains no transitions on which that penalty is applied. Terminal evaluation changed the interpretation: the default combined W = 48 model selected the correct line-trip action first in 98.13 % of fault episodes, yet false-tripped in 72.73 % of the small non-fault set. The exploratory γ = 0.99 followup reaches an F1-score of 0.9701 and the lowest dense FPR, 0.10 %, while changing the terminal outcomes from 210 to 211 correct fault trips and from eight to six non-fault trips. These changes are encouraging, but the small non-fault set and singleseed post-hoc design do not support a general discount-factor conclusion. Overall, offline CQL yields strong faulted-line selection on these simulations, but substantially broader security testing and an interactive environment are required before the approach can be considered for practical grid protection.

DATA AND C ODE AVAILABILITY The EvEMTBench data used in this study are publicly available via FAUDataCloud. Code for data preparation, training, evaluation, and reproduction is available at github.com/julianoelhaf/offline-cql-protection. ACKNOWLEDGMENT This project was funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) - 535389056. R EFERENCES [1] D.-H. Yoon and J. Yoon, “Development of a real-time fault detection method for electric power system via transformer-based deep learning model,” International Journal of Electrical Power & Energy Systems, vol. 159, p. 110069, Aug. 2024. [2] H. C. Kilickiran, B. Kekezoglu, and N. G. Paterakis, “Reinforcement Learning for Optimal Protection Coordination,” in 2018 International Conference on Smart Energy Systems and Technologies (SEST). IEEE, Sep. 2018, pp. 1–6. [3] G. Kordowich, M. Jaworski, T. Lorz, C. Scheibe, and J. Jaeger, “A hybrid Protection Scheme based on Deep Reinforcement Learning,” in 2022 IEEE PES Innovative Smart Grid Technologies Conference Europe (ISGT-Europe). IEEE, Oct. 2022, pp. 1–6. [4] G. Kordowich, J. Loebel, J. Oelhaf, A. Maier, S. Bayer, C. Bergler, and J. Jaeger, “A simulation based dataset of faults and events for machine learning in power systems,” 2026. [5] A. Kumar, A. Zhou, G. Tucker, and S. Levine, “Conservative Q-Learning for Offline Reinforcement Learning,” in Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, Eds., vol. 33. Curran Associates, Inc., 2020, pp. 1179–1191. [6] J. Gao, S. Chen, S. Fan, J. Zhang, K. Jiang, J. Hao, and D. W. Gao, “Conservative Q-learning with physics guidance for generator tripping control,” International Journal of Electrical Power & Energy Systems, vol. 172, p. 111181, Nov. 2025. [7] J. Oelhaf, G. Kordowich, M. Pashaei, C. Bergler, A. Maier, J. Jäger, and S. Bayer, “A Scoping Review of Machine Learning Applications in Power System Protection and Disturbance Management,” International Journal of Electrical Power & Energy Systems, vol. 172, p. 111257, Nov. 2025. [8] J. Shanmugapriya and K. Baskaran, “Rapid Fault Analysis by Deep Learning-Based PMU for Smart Grid System,” Intelligent Automation & Soft Computing, vol. 35, no. 2, pp. 1581–1594, 2023. [9] J. Oelhaf, G. Kordowich, P. A. Pérez-Toro, T. Arias-Vergara, A. Maier, J. Jäger, and S. Bayer, “A Systematic Evaluation of Machine Learning Methods for Fault Detection and Line Identification in Electrical Power Grids,” in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, Apr. 2025, pp. 1–5. [10] D. Wu, X. Zheng, D. Kalathil, and L. Xie, “Nested Reinforcement Learning Based Control for Protective Relays in Power Distribution Systems,” in 2019 IEEE 58th Conference on Decision and Control (CDC). IEEE, Dec. 2019, pp. 1925–1930. [11] CIGRE Task Force C6.04.02, “Benchmark Systems for Network Integration of Renewable and Distributed Energy Resources,” CIGRE, Tech. Rep. 575, Apr. 2014. [12] DIgSILENT GmbH, PowerFactory 2024 User Manual. DIgSILENT GmbH, 2024. [13] S. H. Horowitz, A. G. Phadke, and C. F. Henville, Power System Relaying, 5th ed. John Wiley & Sons, Ltd, 2023. [14] R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed., ser. Adaptive Computation and Machine Learning. The MIT Press, Nov. 2018. [15] S. Levine, A. Kumar, G. Tucker, and J. Fu, “Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems,” 2020. [16] J. Oelhaf, M. Pashaei, G. Kordowich, C. Bergler, A. Maier, J. Jäger, and S. Bayer, “Robustness evaluation of machine learning models for fault classification and localization in power system protection,” IET Conference Proceedings, vol. 2026, no. 3, pp. 188–193, Jun. 2026.

Record · ID 1028702 · SHA-256 e024df88af0716c5
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.