Conceptio › Archive › arXiv CS
arXiv CSopen access

Detecting False Data Injection and Unstable Operation in Smart Grid via System-Aware Graph Boundary Learning

Emad Efatinasab et al. · arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

arXiv:2609.35506v1 [cs.CR] 28 Sep 2026

Detecting False Data Injection and Unstable Operation in Smart Grid via System-Aware Graph Boundary Learning Emad Efatinasab

Denis Donadel

University of Padua Padova, Italy [email protected]

Fondazione Bruno Kessler (FBK) Trento , Italy [email protected]

Mirco Rampazzo

Chuadhry Mujeeb Ahmed

University of Padua Padova, Italy [email protected]

Newcastle University Newcastle upon Tyne, UK [email protected]

Abstract

1

Cyber-physical power systems increasingly rely on data-driven tools to detect instability and support reliable grid operation. However, reliable stability prediction is difficult when unstable operating configurations are rare, sensitive, or unavailable during model development, since collecting such data safely and at scale is often impractical. At the same time, False Data Injection (FDI) attacks can manipulate reported system parameters to trigger false instability alarms or conceal genuinely unsafe operation, and such threats in Decentral Smart Grid Control (DSGC) systems remain largely unexplored. These two challenges are rarely addressed jointly, leaving a gap between stability prediction and attack detection that this work aims to close. In this paper, we introduce StarGNN, a graph learning framework that learns the stable operating region exclusively from clean stable configurations and uses a single abnormality score to flag reported configurations that should not be trusted as evidence of safe operation, covering both genuine instability and unseen FDI manipulations. Each configuration is represented as a producer– consumer star graph and processed by a role aware graph neural network, with physics constrained pseudo-negatives generated by perturbing reaction time and price response parameters standing in for the unavailable unstable and attack data. A single threshold, calibrated only on held-out stable data, is used without task or attack specific adjustment. Evaluated on nine unseen FDI scenarios, StarGNN detects between 0.780 and 0.973 of attacks on stable configurations and retains a post attack instability recall between 0.972 and 0.999 on unstable ones, with 0.899 recall against a stronger adaptive attacker, showing that stable-only boundary learning can support both stability prediction and attack detection without access to genuine unstable labels or attack samples during training.

Cyber-physical power systems are now widely adopted across the electricity sector, where they play a central role in supporting reliable, efficient, and secure grid operation. The operation of cyberphysical power systems relies on interdependencies among the physical power grid, physical communication network, and logical communication network, which support the transmission of operational data and control commands between grid components and control centres [25]. The increasing reliance on cyber-physical technologies has accelerated the evolution of conventional electrical networks toward smart grids. Maintaining a reliable match between electricity generation and consumption is a key objective of smart grid operation. To achieve this, predictive tools and proactive control measures are employed to anticipate imbalances and limit the likelihood of system disturbances [13]. A commonly used approach for aligning electricity consumption with prevailing grid conditions is demand response [18, 31]. Within this context, Decentral Smart Grid Control (DSGC) [33] provides a local mechanism for dynamic demand response. It uses grid frequency deviations to indicate the prevailing power imbalance: a decrease signals insufficient generation, whereas an increase indicates excess supply. Consequently, electricity prices are determined from local frequency measurements. However, introducing an electricity price dependent on the frequency causes the economic response of grid participants to interact directly with the physical grid dynamics. Whether the resulting operating point remains locally stable depends on the intensity of this price response, delays in participant adaptation, and the time window used to average frequency measurements [9, 33]. Since the response coefficients and adaptation delays need not be identical across grid participants, identifying whether perturbations around equilibrium will diminish or increase is a challenging task [9]. Unstable grid operation may compromise both power system reliability and economic activity, potentially triggering severe voltage disturbances and, in extreme cases, cascading blackouts across large areas [15]. Previous research has demonstrated that machine learning can effectively support smart grid stability prediction, with consistently high prediction accuracy reported across different model architectures [3, 10, 35]. Despite this strong performance, such supervised approaches share a critical limitation: their practical deployment is constrained by the limited availability of representative instability

CCS Concepts • Security and privacy → Intrusion/anomaly detection and malware mitigation; • Computer systems organization → Embedded and cyber-physical systems; • Computing methodologies → Machine learning.

Introduction

Emad Efatinasab, Denis Donadel, Mirco Rampazzo, and Chuadhry Mujeeb Ahmed

data. Real operational datasets containing both stable and unstable grid conditions are difficult to obtain because instability events are uncommon, undesirable, and potentially unsafe to reproduce deliberately [14]; their collection may also require substantial expertise, resources, and time [19]. Consequently, methods that assume access to unstable examples may be difficult to apply under realistic data constraints, which is the case for most models in the literature. Beyond this data scarcity, DSGC systems face a second and largely separate challenge: their exposure to cyber threats. Integrating computational and communication capabilities with physical processes can improve system efficiency; however, the use of open, networked communication channels also increases cyber-physical systems’ exposure to malicious external interference [23]. Attacks on the cyber infrastructure of electrical grids may interrupt energy services, undermine critical assets, and generate substantial risks for human well-being and environmental security [5]. Among these threats, False Data Injection (FDI) attacks are widely regarded as one of the most critical cyber threats to the integrity of modern power grids [7], in which adversaries manipulate system data to influence monitoring or control decisions. Although supervised FDI attack detectors can achieve high accuracy when the attack patterns encountered during testing are represented in their training data, detection performance can deteriorate substantially when the characteristics of a new FDI attack differ from those learned during training [37]. Despite this plausible attack surface, FDI threats in DSGC systems remain largely unexplored, motivating a detector that addresses natural instability and adversarial manipulation jointly rather than as separate problems. Contributions. In this paper, we bridge these gaps by proposing StarGNN, the first system model-guided graph neural network for unified detection of natural instability and FDI attacks in DSGC systems, trained on stable configuration samples only. StarGNN represents each operating configuration as a producer-consumer star graph with role-aware node processing and assigns a single risk score to how far it departs from stable operation. The model learns the boundary of the stable region without ever observing unstable or attacked configurations during training. This is a key advantage over supervised detectors in the literature, which require instability or attack data that are scarce, unsafe to collect, or, for FDI attacks, largely unexplored in DSGC systems. Our contributions can be summarized as follows: • We propose StarGNN, a DSGC model-guided graph neural network that encodes each operating configuration as a producer–consumer star graph with role-aware node processing, assigning a single risk score to departures from stable operation. • We introduce a learning setting in which neither unstable configurations nor attack samples are available during training, along with a generation strategy that produces physically plausible synthetic samples by perturbing reaction time and price-response parameters. • We formulate a single detector and threshold, calibrated only on held-out stable data, for both natural instability and FDI manipulation, without task or attack-specific adjustment. Moreover, we provide a security evaluation of

StarGNN covering nine unseen FDI scenarios and a constrained adaptive attacker with binary alarm feedback. • We show that StarGNN reaches 0.989 accuracy and a macro F1-score of 0.97 on stability prediction. Under attack, it maintains a mean attack recall of 0.988 against generic FDI scenarios, while still reaching an attack recall of 0.899 against an adaptive attacker, consistently outperforming established anomaly detection baselines. Organization. The remainder of this paper is organized as follows. Section 2 reviews related work. Section 3 and Section 4 introduce the considered system and threat model, respectively. Section 5 introduces the attacks we employed in this study. StarGNN is illustrated in Section 6 and evaluated in Section 7. Finally, Section 8 offers some discussion insights, while Section 9 concludes this paper with some final remarks.

2

Related Work

AI-based Stability Prediction. Recent advances in artificial intelligence have significantly expanded the capabilities of smart-grid stability assessment and control [35]. For instance , Wang et al. [41] compared six machine learning models such as Random Forest (RF), Extreme Gradient Boosting (XGBoost), Support Vector Machine (SVM), K-Nearest Neighbors (KNN), Logistic Regression (LR), and Categorical Boosting (CatBoost) for smart grid stability classification and prediction. Another study introduced a pruned onedimensional time-aware CNN for grid stability analysis and energy cost optimization [1]. A Multidirectional LSTM model is proposed by [3] for stability prediction. Numerous studies have investigated smart grid stability using a range of machine learning and artificial intelligence techniques [6, 16, 30]. However, most existing approaches rely on supervised learning and assume access to representative samples from both stable and unstable classes, which may limit their practical deployment when instability data are scarce or unavailable in real world situations. Furthermore, to the best of our knowledge, prior work has not represented DSGC operating configurations as graph structured data to explicitly capture the dependencies between the producer and participating consumers. False Data Injection Attacks. FDI attacks compromise measurement integrity and pose a serious threat to supervisory control and data acquisition systems [20]. These concerns have motivated growing interest in developing effective detection mechanisms for FDI attacks. For instance, Zhang et al. [43] proposed a data-driven method for detecting FDI attacks in distribution system using Autoencoders for dimensionality reduction and feature extraction, and a GAN-based framework that identified attacks by learning discrepancies between secure and manipulated measurements. Takiddin et al. [36] proposed a graph autoencoder-based detector that learns spatio-temporal power system features from multiple topology realizations, improving generalization to unseen FDI attacks. Despite the growing literature on FDI detection in power systems, the corresponding threat has received limited attention in the context of DSGC systems. Previous studies have examined adversarial attacks against AI-based stability predictors [14, 15]. Such attacks represent a related form of data integrity manipulation, but they are typically designed to exploit the decision boundary

Detecting False Data Injection and Unstable Operation in Smart Grid via System-Aware Graph Boundary Learning

A

genuine physical instability or manipulated system information through the same learned risk score.

v2

DSGC

Decision

4

B v1

v4

v3

Figure 1: Overview of the system and threat model. An attacker can compromise one of the participants’ instrumentation (A) or compromise the communication through substation weaknesses or MitM attacks (B). of a particular prediction model. In contrast, the non-adaptive FDI scenarios considered in this work are motivated by power system attack mechanisms and the structure of the DSGC setting. Their construction reflects participant roles, producer–consumer interactions, admissible operating ranges, power balance relationships, and realistic forms of measurement or control manipulation. The resulting attacks are therefore system-driven rather than detectordriven and can be generated without knowledge of the model. To the best of our knowledge, this broader, physics-aware FDI threat has not been systematically studied for DSGC systems.

3

System model

We consider a power system based on the DSGC framework [9, 33]. DSGC combines the physical behavior of the grid with a price-based demand-response mechanism. Grid frequency serves as a local indicator of the balance between generation and demand. A frequency drop indicates insufficient generation, whereas a frequency increase reflects excess supply. The electricity price is adjusted according to this signal so that producers and consumers can modify their behavior without relying on centralized price computation [33]. Stability, therefore, depends on the combined effect of the participant reaction times, nominal power values, and price-response coefficients. For each parameter configuration, local stability is assessed around the synchronous operating point through the characteristic equation obtained from the linearized DSGC model [9]. The system used in this study represents a four participant DSGC system. It consists of one producer, denoted by 𝑣 1 , and three consumers, denoted by 𝑣 2 , 𝑣 3 , and 𝑣 4 . The producer is connected to each consumer, resulting in the star shaped interaction structure, as shown in Figure 1. No direct interaction between the consumer nodes is included in this system representation. Each DSGC configuration is represented as x = [𝝉 ⊤, p⊤, g⊤ ] ⊤ ∈ R12 , where 𝝉, p, g ∈ R4 denote the reaction-time, power, and price-response variables, respectively. The power values satisfy 𝑝 1 = |𝑝 2 + 𝑝 3 + 𝑝 4 | ,

(1)

Importantly, characteristic roots are not part of the dataset and are not supplied to the proposed model during training or inference. The stability prediction model receives only the operational pa4 . The detector is designed to identify either rameters {𝜏𝑖 , 𝑝𝑖 , 𝑔𝑖 }𝑖=1

Threat Model

We consider FDI attacks against a deployed DSGC stability monitoring pipeline, as shown in Figure 1. The monitored system receives a reported parameter vector e x, defined in Section 3, and uses it to decide whether the current DSGC configuration should be treated as normal or abnormal. The security property of interest is the integrity of this reported input. If the input is manipulated before it reaches the detector, the stability decision may no longer reflect the actual operating condition. Attackers capabilities. The attacker is assumed to compromise part of the measurement or reporting path. Such a compromise may occur through the firmware of microprocessor-based remote terminal units [21], weaknesses in IEC 61850-based substation communication [26], or interception of data exchanged between field devices and the control system through a man-in-the-middle attack [34]. The attacker does not need to physically compromise the grid. Instead, the attacker tampers with the values that describe the DSGC configuration before those values are processed by the monitoring system. We consider an integrity attacker with partial control over the reported parameters. Depending on the scenario, the adversary may alter a single parameter, several parameters belonging to one participant, or parameters reported by multiple participants. The attacker may modify reaction-time parameters, reported power values, price-response coefficients, or combinations of these features. This captures both localized compromises, such as a single reporting unit, and coordinated compromises affecting several reported attributes. The attacks are performed at inference time. The adversary is not assumed to modify the training data, the learned detector parameters, the classification threshold, or the stability labels. The focus is instead on data integrity attacks that provide a syntactically valid but manipulated parameter vector. We consider four attacker capability profiles, as summarized in Table 1. Many of the Generic non adaptive FDI attacks are blackbox with respect to the detector, as they require no access to its architecture, parameters, gradients, decision threshold, or anomaly score. Balance-preserving FDI attacks follow a grey-box setting in which the attacker has limited knowledge of the DSGC physical model, including the power-balance relation 𝑝 1 = |𝑝 2 + 𝑝 3 + 𝑝 4 | but has no internal knowledge of the detector. Replay-style FDI attacks represent a second grey-box capability in which the attacker can observe and reuse previously valid participant reports. We also consider a constrained adaptive grey-box attacker (single node adaptive concealment attacker) that targets instability concealment. The attacker compromises a single participant and can modify only its reported reaction time and price response parameters, 𝜏𝑖 and 𝑔𝑖 , while all power measurements remain unchanged. It has no access to the model architecture, parameters, gradients, continuous anomaly score, or decision threshold, but it can submit a limited number of candidates and observe the corresponding binary alarm response.

Emad Efatinasab, Denis Donadel, Mirco Rampazzo, and Chuadhry Mujeeb Ahmed

Training Phase

Table 1: Attacker capability profiles. Attacker

Access

Capability

Generic non-adaptive Balance-preserving Replay style Adaptive single node

Black-box Grey-box Grey-box Grey-box

Modify 𝜏, 𝑝, 𝑔 reports Knows power-balance relation Reuses valid reports Modifies 𝜏𝑖 , 𝑔𝑖 ; binary feedback

Testing Phase

SmartGNN Pseudo-negative samples generator

Attacker Objective. The main attacker objective is the concealment of an unreliable or manipulated configuration to be accepted as normal. In the most safety-critical case, an actually unstable configuration is reported in a way that makes it appear stable, which may delay or suppress corrective actions [13, 15]. In severe cases, this may allow disturbances to propagate and contribute to cascading failures or wider outages, with consequences for dependent services such as telecommunications, transportation, and emergency response [42]. The second objective is to create false alarms by making stable configurations appear abnormal. Such false alarms can reduce operator confidence in the monitoring system and may trigger unnecessary interventions. The detector uses a single abnormal class. Both genuinely unstable DSGC configurations and attacked parameter reports are mapped to this class, since manipulated reports cannot be trusted as evidence of stable operation. This conservative design provides an operational warning that corrective action is needed since either case invalidates the assumption that the reported configuration represents safe operation, particularly when an attacker attempts to conceal an unstable condition. It also remains effective for attacks that create false instability alarms, as both cases indicate that the reported grid state is unreliable and needs further investigation. None of the attack samples is used to train the detector or to construct the physics-constrained pseudo-negative samples. The evaluation, therefore, tests whether the detector can identify previously unseen manipulations using only stable training data and the physical structure encoded in the proposed method.

5

False Data Injection Attacks

Based on the capabilities of different attackers, we defined and evaluated nine non-adaptive FDI attacks that cover parameter bias, power manipulation, participant compromise, coordinated manipulation, load redistribution, and replay style substitution. Table 2 summarizes the attacker scope and the system constraints preserved by each scenario. The attacks are generated only at inference time and are independent of the StarGNN detector. All manipulated values are restricted to admissible feature ranges derived from the clean training data. Detailed perturbation magnitudes and generation procedures are reported in Appendix C. The first six attacks provide generic DSGC-specific manipulations over reaction time, price response, and power reports. The remaining three adapt established attack patterns from the power system security literature: load redistribution, localized load redistribution, and replay. These scenarios range from isolated parameter falsification to coordinated and constraint-preserving manipulation. The attack IDs in Table 2 map to the capability profiles in Table 1 as follows: A1–A3 and A5 use the generic non-adaptive black-box

Stable -------Unstable -------Samples with Attacks

Stable samples

Decision

Figure 2: Overview of training and testing of StarGNN.

setting; A4, A7, and A8 use the balance-preserving grey-box setting; A9 follows the replay style grey-box setting; and A10 uses he adaptive single-node grey-box setting. A6 can follow either the generic black-box or balance-preserving grey-box profile depending on whether its power component is based on A3 or A4. We additionally consider a constrained adaptive attacker whose objective is to conceal an unstable operating point. The adversary compromises one participant and modifies only its reaction time and price response reports, 𝜏𝑖 and 𝑔𝑖 , while all power reports remain unchanged. Each mutable feature may deviate by at most 50% of its training-derived range and remains within its admissible bounds. The attacker has no access to the StarGNN architecture, parameters, continuous risk score, or decision threshold. It can only submit candidate reports and observe the corresponding binary alarm. Starting from an initial random perturbation, it performs a bounded randomized search using 25 additional queries, for a maximum of 26 detector evaluations per sample. The search combines local exploration around the current candidate with global sampling within the admissible perturbation region. Among alarm free candidates, the attacker retains the one with the largest normalized displacement from the original configuration. Further implementation details are provided in Appendix C.

6

Methodology

The proposed training and testing pipeline of StarGNN is represented in Figure 2. Stable configuration samples are used to generate physics-constrained pseudo-negatives, after which both are represented as producer–consumer star graphs and processed by StarGNN. The model learns a scalar risk score through separation, ranking, compactness, and variance-based objectives. A threshold calibrated solely from stable validation scores is then used for both instability prediction and detection of unseen FDI attacks. A representation of the different layers of the proposed step is shown in the Appendix, Figure 4.

6.1

Stable-Only Data Preparation

The dataset is first divided according to the binary stability label. Let D𝑠 and D𝑢 denote the sets of stable and unstable configurations, respectively. Only samples in D𝑠 are used during model development. The stable configurations are randomly divided into three non-overlapping partitions, where 60% of the stable samples are assigned to training, 20% to validation, and the remaining 20% to testing. All unstable configurations are excluded from the training

Detecting False Data Injection and Unstable Operation in Smart Grid via System-Aware Graph Boundary Learning

Table 2: Summary of the evaluated FDI attacks. Detailed perturbation rules and parameter ranges are provided in Appendix C. ID

Attack

Modified reports

Compromise scope

Constraint / characteristic

Motivation

A1

Reaction-time bias

𝜏

1–4 parameters

Bounded additive bias

A2

Price-response bias

𝑔

1–4 parameters

Bounded additive bias

A3 A4

Power imbalance Balance-preserving power

1–3 consumers 1–3 consumers

A5

Node takeover

Consumer 𝑝 Consumer and producer 𝑝 𝜏, 𝑝, 𝑔

Producer power remains unchanged Producer power recomputed to preserve power balance Complete local report is modified

A6

Coordinated FDI

𝜏, 𝑔, 𝑝

Multiple feature groups

A7

Load redistribution

Consumer 𝑝

Three consumers

Time-delay-inspired manipulation [40] Falsification of reported priceresponse behavior Load-altering attack [28] Stealthy FDI through constraintpreserving manipulation Compromise of one participant’s reporting unit Coordinated falsification across feature groups Load-redistribution attack [39]

A8

Local load redistribution

Consumer 𝑝

1–2 consumers

A9

Replay style node substitution Single node adaptive concealment

𝜏, 𝑝, 𝑔

One participant

𝜏, 𝑔

One participant

A10

One participant

and validation stages. Instead, the complete unstable subset is retained for final evaluation. The final stability prediction test set is formed by combining the held-out stable samples with all available unstable samples. The stable test partition is also kept separately because unseen attacks are generated only from held-out stable configurations. Feature standardization is carried out using statistics computed only from the stable training partition. The same statistics are then applied to the stable validation set, the final test set, and all attacked samples. No information from unstable or attacked configurations is used to fit the scaler. The minimum and maximum values of each feature are also estimated from the stable training partition. These empirical ranges are used only to define bounded perturbation regions for the later generation of pseudo-negatives and attack samples. For feature 𝑘, the expanded interval is written as h i tr tr 𝑥𝑘,min − 𝛿𝑅𝑘 , 𝑥𝑘,max + 𝛿𝑅𝑘 , (2)

Reaction-time, price-response, and power reports are jointly modified Aggregate consumer demand preserved before clipping; producer power recomputed Total consumer demand approximately preserved; producer power recomputed Participant report replaced using values from another stable configuration Power unchanged; 50% range bound; 25 additional binary-alarm queries

𝑔

𝑔e𝑖 = 𝑔𝑖 + 𝑑𝑖 𝛼𝑖 𝐷𝑖 ,

6.2

Physics-Constrained Pseudo-Negative Generation

sponding parameter limit: ( 𝑔max − 𝑔𝑖 , 𝑑𝑖 = +1, 𝑔 𝐷𝑖 = 𝑔𝑖 − 𝑔min, 𝑑𝑖 = −1.

x = [𝜏1, . . . , 𝜏4, 𝑝 1, . . . , 𝑝 4, 𝑔1, . . . , 𝑔4 ] ⊤ ∈ D𝑠tr

(4)

denote a stable training configuration. A pseudo-negative e x is obtained by perturbing selected reaction time parameters, price response coefficients, or both. The nominal power values are not intentionally perturbed during this stage. Three perturbation types

(5)

(6)

The admissible interval is 𝑔𝑖 ∈ [0.05, 1.0]. Reaction time perturbations are generated in the same way. For a selected 𝜏𝑖 , 𝛽𝑖 ∼ U (0.10, 0.35).

(7)

and 𝐷𝑖𝜏 =

Training only on stable configurations provides no direct indication of how the risk score should behave away from the observed stable data. To introduce such contrast without using genuine unstable samples, we generate a set of pseudo-negative configurations from each stable training sample. These samples are used to shape the learned decision boundary, but they are not treated as physically verified unstable operating points. Let

𝛼𝑖 ∼ U (0.20, 0.65).

𝑔 and 𝐷𝑖 is the available distance from the current value to the corre-

𝜏e𝑖 = 𝜏𝑖 + 𝑑𝑖 𝛽𝑖 𝐷𝑖𝜏 , (3)

Static adaptation of replay attacks [27] Adaptive instability concealment

are considered. A price response perturbation is selected with probability 2/5, a reaction time perturbation with probability 1/5, and a combined perturbation with probability 2/5. For each selected parameter group, between one and four participant variables are chosen uniformly without replacement. For a selected price response coefficient 𝑔𝑖 , a direction 𝑑𝑖 ∈ {−1, +1} is sampled with equal probability. The modified value is

where 𝛿 = 0.25 and   tr tr 𝑅𝑘 = max 𝑥𝑘,max − 𝑥𝑘,min , 10−6

Local load-redistribution attack [24]

( 𝜏max − 𝜏𝑖 , 𝑑𝑖 = +1, 𝜏𝑖 − 𝜏min, 𝑑𝑖 = −1.

(8)

The reaction times are restricted to 𝜏𝑖 ∈ [0.5, 10.0]. For a combined perturbation, the price response and reaction time transformations are applied sequentially. The perturbations are non-directional. In particular, the generator does not assume that increasing a reaction time or price response coefficient always moves the system toward instability. Both upward and downward changes are considered so that the model is not tied to a predetermined monotonic relation between a parameter and the stability condition. After each transformation, the variables are clipped to their admissible domains: 𝜏𝑖 ∈ [0.5, 10.0], 𝑔𝑖 ∈ [0.05, 1.0], 𝑝𝑖 ∈ [−2.0, −0.5] for 𝑖 ∈ {2, 3, 4}, and 𝑝 1 ∈ [1.5, 6.0]. The producer power is then projected onto the producer–consumer balance constraint 𝑝e1 = 𝑝e2 + 𝑝e3 + 𝑝e4 . This step ensures that the

Emad Efatinasab, Denis Donadel, Mirco Rampazzo, and Chuadhry Mujeeb Ahmed

pseudo-negatives do not rely on an obvious power-balance violation and remain consistent with the basic structure of the DSGC data. For every stable training configuration, 𝐾 = 16 independently generated variants are stored: P (x𝑛 ) = {e x𝑛(1) , . . . , e x𝑛(𝐾 ) }. During each training iteration, one of the 𝐾 variants is selected at random for every stable sample in the batch. This allows the same stable configuration to be paired with different perturbations over the course of training without increasing the batch size. The generated samples receive a pseudo-negative target only for the purpose of contrastive boundary learning. They should therefore not be interpreted as synthetic ground-truth unstable samples. Their role is to expose the model to admissible departures from observed stable configurations while preserving the main parameter limits and the producer–consumer power relation.

6.3

Topology-Aware Graph Representation

The original dataset stores each DSGC configuration as a twelve dimensional tabular vector. Although this form is convenient for data storage, it does not explicitly preserve the producer–consumer structure of the system. We therefore reorganize every configuration as a graph whose nodes correspond to the four DSGC participants. Before graph construction, each feature is standardized using the mean and standard deviation obtained from the stable training partition, as described in Section 6.1. The three parameters belonging to participant 𝑖 are grouped into a node feature vector: ⊤  𝑖 ∈ {1, 2, 3, 4}. (9) h𝑖(0) = 𝜏 𝑖 , 𝑝 𝑖 , 𝑔𝑖 , Here, 𝜏 𝑖 represents the standardized reaction time, 𝑝 𝑖 the standardized nominal power, and 𝑔𝑖 the standardized price response coefficient of participant 𝑖. The four node vectors are stacked to form the initial node feature matrix  (h (0) ) ⊤   1   (0) ⊤   (h2 )  (0) H =  (0) ⊤  ∈ R4×3 . (10)  (h3 )   (0)   (h ) ⊤   4  The first row corresponds to the producer, while the remaining three rows correspond to the consumers. The node feature matrix is associated with the star-shaped graph introduced in Section 3. G = (V, E),

V = 𝑣 1, 𝑣 2, 𝑣 3, 𝑣 4 .

(11)

E = {(𝑣 1, 𝑣 2 ), (𝑣 1, 𝑣 3 ), (𝑣 1, 𝑣 4 )} .

(12)

and Node 𝑣 1 represents the producer, and nodes 𝑣 2 , 𝑣 3 , and 𝑣 4 represent the consumers. The graph, therefore, preserves the producer– consumer interaction pattern contained in the DSGC model rather than treating the twelve input variables as unrelated features. The same transformation is applied to stable samples, pseudo-negative configurations, genuine unstable samples, and attacked observations.

6.4

StarGNN Risk Model

The proposed model follows the general message passing neural network framework, in which a node updates its representation using its own features and information received from neighboring

nodes [17]. We adapt this idea to the fixed producer–consumer structure of the DSGC system. The resulting architecture, referred to as StarGNN, uses role-aware message passing and graph aggregation. The producer receives information from all three consumers, while each consumer receives information only from the producer. The producer representation is also retained separately when the node embeddings are combined at the graph level. Standard operations such as linear transformations, activation functions, normalization, dropout, and pooling are used throughout the network. The DSGCspecific contribution lies in how these operations are arranged to reflect the different roles of the producer and consumers. The model consists of a shared node encoder, two star structured message-passing layers, a graph level aggregation stage, and two output heads. One head produces a graph embedding used by the stable-only training objective, while the other produces the scalar risk score used for detection. Let h i⊤ H (0) = h1(0) , h2(0) , h3(0) , h4(0) ∈ R4×3 (13) denote the node feature matrix of one DSGC configuration. Each row contains the standardized reaction time, nominal power, and price response coefficient of one participant. A shared multilayer encoder maps each three-dimensional node vector to the hidden space:     h𝑖(0,𝑒 ) = ReLU W𝑒,2 ReLU W𝑒,1 h𝑖(0) + b𝑒,1 + b𝑒,2 . (14) The same encoder parameters are used for the producer and the three consumers. The hidden dimension is set to 𝑑ℎ = 128. The encoded node representations are then processed by 𝐿 = 2 messagepassing layers. At layer ℓ, the transformed message of node 𝑖 is (ℓ ) (ℓ ) (ℓ ) u𝑖(ℓ ) = W𝑚 h𝑖 + b𝑚 .

(15)

The producer receives the mean of the three consumer messages: 4

m1(ℓ ) =

1 ∑︁ (ℓ ) u . 3 𝑗=2 𝑗

(16)

Each consumer receives the message generated by the producer: m𝑖(ℓ ) = u1(ℓ ) ,

𝑖 ∈ {2, 3, 4}.

(17)

The node update combines the incoming message with a separate transformation of the current node representation:  h  i h𝑖(ℓ+1) = Dropout LN ReLU W𝑠(ℓ ) h𝑖(ℓ ) + b𝑠(ℓ ) + m𝑖(ℓ ) , (18) (ℓ ) where W𝑠(ℓ ) and W𝑚 are trainable self and message transformations, respectively. The operator LN(·) denotes layer normalization. A dropout rate of 0.10 is applied after each message-passing layer. This communication rule follows the DSGC star structure directly. The producer gathers information from all consumers, while the consumers do not exchange messages with one another. So we avoid introducing edges that are not present in the adopted system model. After the final message-passing layer, the producer embedding is retained as

h𝑝 = h1(𝐿) .

(19)

Detecting False Data Injection and Unstable Operation in Smart Grid via System-Aware Graph Boundary Learning

The three consumer embeddings are summarized using mean and element-wise maximum pooling:

configuration and the variant generated from it. We therefore add a ranking loss:

4

h𝑐mean =

1 ∑︁ (𝐿) h , 3 𝑖=2 𝑖

𝐵

h𝑐max = max h𝑖(𝐿) . 𝑖 ∈ {2,3,4}

(20)

The graph level representation is then formed as q = h𝑝 ∥ h𝑐mean ∥ h𝑐max ∈ R3𝑑ℎ ,

(21)

where ∥ denotes concatenation. This role-aware aggregation differs from a conventional global mean or maximum pooling operation, which treats all nodes in the same way. Keeping the producer embedding separate preserves its central role in the DSGC system. The mean consumer embedding represents their shared behavior, while maximum pooling retains the strongest response observed among the three consumers. A graph head maps q to a lower-dimensional embedding:   z𝜙 (x) = W𝑧,2 Dropout ReLU W𝑧,1 q + b𝑧,1 + b𝑧,2, (22) where z𝜙 (x) ∈ R𝑑𝑧 . and 𝑑𝑧 = 64. This embedding is used by the compactness and variance terms introduced in Section 6.5. The risk maps the graph embedding to a scalar logit:   ⊤ 𝑟𝜙 (x) = w𝑟,2 Dropout ReLU W𝑟,1 z𝜙 (x) + b𝑟,1 + 𝑏𝑟,2 . (23) The logit is used directly as the risk score and is not converted into a probability at inference time. A larger value indicates that the input provides a stronger instability cue relative to the stable configurations observed during training. The model, therefore, returns  𝑓𝜙 (x, G) = z𝜙 (x), 𝑟𝜙 (x) . (24)

Lrank =

1 ∑︁ 𝑝 max 0, 𝑚 + 𝑟𝑛𝑠 − 𝑟𝑛 , 𝐵 𝑛=1

where 𝑚 = 1 is the ranking margin. This term encourages the pseudo-negative associated with each stable configuration to receive a risk score at least 𝑚 units higher than the score of its clean counterpart. In addition to score separation, the stable embeddings are encouraged to form a compact region. Following the one class representation learning principle used in Deep SVDD [32], a center c ∈ R𝑑𝑧 is calculated from the embeddings of all stable training samples: 𝑁

c=

𝑠 1 ∑︁ z𝜙 (x𝑛 ) , 𝑁𝑠 𝑛=1

6.5

𝐵

LBCE = −

 1 ∑︁ h log 1 − sigmoid(𝑟𝑛𝑠 ) 2𝐵 𝑛=1 i 𝑝  + log sigmoid(𝑟𝑛 ) .

(26)

The pseudo-negative target is used only to provide contrast during training. The binary term separates the two groups at the batch level, but it does not explicitly preserve the pairing between a stable

(28)

𝐵

LSVDD =

1 ∑︁ 𝑠 2 z −c 2. 𝐵 𝑛=1 𝑛

(29)

Only stable embeddings contribute to this term. Pseudo-negative embeddings are not forced toward the stable center. A compactness objective alone can reduce the diversity of the learned representation. To avoid all stable embeddings collapsing to nearly the same point, we impose a minimum standard deviation on every embedding dimension. This follows the variance-floor idea used to prevent representation collapse in variance-regularized learning [11]. Let √︂  

Stable-Only Training Objective

The first pair denotes the embedding and risk score of the stable sample x𝑛 , whereas the second corresponds to the selected pseudonegative e x𝑛 . The training objective contains four terms: a binary separation loss, a pairwise ranking loss, a stable-embedding compactness term, and a variance regularizer. The binary separation term assigns target 0 to stable samples and target 1 to pseudo-negatives:

x𝑛 ∈ D𝑠tr .

The center is treated as a fixed, non-trainable quantity while the network parameters are updated. The compactness loss is

𝑠𝑘 =

The StarGNN is trained using stable configurations and the pseudonegatives derived from them. For each stable sample in a minibatch, one of its precomputed pseudo-negative variants is selected at random. The clean configuration and the selected variant are converted to graphs and passed through the same network. For a mini-batch of size 𝐵, let  𝑝 𝑝 z𝑠𝑛 , 𝑟𝑛𝑠 = 𝑓𝜙 (x𝑛 , G) , z𝑛 , 𝑟𝑛 = 𝑓𝜙 (e x𝑛 , G) . (25)

(27)

Var 𝑧𝑠1𝑘 , . . . , 𝑧𝑠𝐵𝑘 + 𝜖

(30)

denote the batch standard deviation of embedding dimension 𝑘. The variance penalty is 𝑑

Lvar =

𝑧 1 ∑︁ max (0, 𝑠 0 − 𝑠𝑘 ) , 𝑑𝑧

(31)

𝑘=1

where 𝑠 0 = 0.5 and 𝜖 = 10−6 . Dimensions whose standard deviation is already above the target value, receive no penalty. The complete objective is L = 𝜆BCE LBCE + 𝜆rank Lrank + 𝜆SVDD LSVDD + 𝜆var Lvar,

(32)

The loss weights were set to 𝜆BCE = 1.0, 𝜆rank = 0.5, 𝜆SVDD = 0.05, and 𝜆var = 0.10. The stable center is initialized from the embeddings produced by the initial network and is recomputed every ten training epochs. This periodic update allows the center to follow the developing stable representation without treating it as a trainable model parameter. The model is trained for 150 epochs using AdamW with a learning rate of 10−3 , weight decay of 10−5 , and a mini-batch size of 128. Gradient norms are clipped at 5 during optimization. The detector uses only the raw StarGNN risk logit 𝑟𝜙 (x), while the SVDD and variance terms act as regularizer during training. After training, the model produces a scalar risk logit 𝑟𝜙 (x) for each input configuration. This raw logit is used directly as the detection score. No sigmoid transformation or SVDD distance is

Emad Efatinasab, Denis Donadel, Mirco Rampazzo, and Chuadhry Mujeeb Ahmed

added during inference. The decision threshold is calibrated using only the held-out stable validation set:    R𝑠val = 𝑟𝜙 (x) : x ∈ D𝑠val , 𝜂 = 𝑄𝑞 R𝑠val . (33) Here, 𝑄𝑞 (·) denotes the empirical quantile operator, and 𝑞 = 0.90 in the implemented configuration. Because the threshold is selected from stable validation scores alone, neither genuine unstable configurations nor attacked samples influence its value. The final test set is used only after the model and threshold have been fixed. The same risk score and decision threshold are used for both evaluation tasks. In the stability prediction experiment, samples above the threshold are classified as unstable. In the attack detection part, a manipulated report above the same threshold is flagged as containing an instability cue. The use of the 0.90 quantile corresponds to a relatively permissive operating point. In the absence of tied scores, approximately 10% of the stable validation configurations lie above the threshold.

7

Evaluation

In this section, we evaluate StarGNN. First, we present the dataset used in Section 7.1, which contains both stable and unstable data but no attacks. Then, Section 7.2 investigates the impact of the attacks introduced in Section 5 on the dataset. Section 7.3 showcases the capabilities of StarGNN in detecting unstable configurations, while Section 7.4 discusses the capabilities in identifying FDI attacks. The reported metrics for evaluation of our approach are accuracy, macro-averaged precision, macro-averaged recall, macro-averaged F1-score, and the area under the receiver operating characteristic curve. Let TP, TN, FP, and FN denote the numbers of true positives, true negatives, false positives, and false negatives, respectively. Accuracy measures the proportion of correctly classified samples: TP + TN . (34) TP + TN + FP + FN For each class 𝑐 ∈ {0, 1}, precision and recall are calculated as Accuracy =

Precision𝑐 =

TP𝑐 , TP𝑐 + FP𝑐

Recall𝑐 =

TP𝑐 . TP𝑐 + FN𝑐

(35)

The class-specific F1-score is the harmonic mean of precision and recall: Precision𝑐 Recall𝑐 . (36) F1𝑐 = 2 Precision𝑐 + Recall𝑐 Because the two classes may contain different numbers of samples, precision, recall, and F1-score are reported using macro averaging, so that each class contributes equally. For a metric 𝑀 ∈ {Precision, Recall, F1}, the macro-averaged value is computed as 1

Macro-𝑀 =

1 ∑︁ 𝑀𝑐 , 2 𝑐=0

(37)

where 𝑀𝑐 denotes the corresponding class-wise metric for class 𝑐. The receiver operating characteristic curve is obtained by varying the decision threshold and plotting the true-positive rate against the false-positive rate. The ROC AUC summarizes this curve as ∫ 1 ROC AUC = TPR(FPR) 𝑑 (FPR). (38) 0

7.1

Dataset

The experiments use an augmented version of the Electrical Grid Stability Simulated Dataset from the UCI Machine Learning Repository [8]. The data set has been used in previous studies on power system stability and adversarial manipulation [2, 13, 14, 29] and is, to our knowledge, the only publicly available data set derived from the DSGC system. It contains 60,000 simulated configurations of a four-node DSGC system composed of one producer and three consumers, including 21,720 stable and 38,280 unstable configurations. Each configuration is described by the twelve input variables discussed in Section 3. The data set contains no attack samples; all FDI scenarios used in our experiments are generated separately and are excluded from model training.

7.2

Attack Impact Evaluations

To examine whether the considered FDI attacks can meaningfully alter a conventional AI-based stability decision, we train an auxiliary supervised multilayer perceptron (MLP) to classify clean DSGC configurations as stable or unstable. MLP-based models have been widely used for data-driven power system stability prediction [4, 38]. We therefore use the MLP as a representative neural classifier for quantifying attack impact. It is used only for this auxiliary evaluation and is not part of the proposed StarGNN framework. Specifically, we apply each FDI attack to configurations that the MLP initially classifies correctly and measure how often the manipulated input changes the prediction from stable to unstable (S→U) which represents a false alarm, in which a stable operating condition is incorrectly classified as unstable or from unstable to stable (U→S) represents concealment, in which an actually unstable configuration is incorrectly reported as stable. The model contains three hidden layers with 128, 64, and 32 neurons and uses GELU activation, layer normalization, and dropout. The attacks were applied separately to held-out stable and unstable configurations. The MLP also outputs a predicted probability for the unstable class. A value close to zero indicates that the model considers the configuration stable, whereas a value close to one indicates a strong unstable prediction. Before the attacks were applied, the mean predicted instability probability was approximately 0.0297 for stable samples and 0.9717 for unstable samples. On the exact StarGNN attack-source set, the clean model correctly classified 98.96% of the unstable configurations. Table 3 summarizes the effect of the attacks on the auxiliary supervised stability predictor. False-alarm and concealment effects. Table 3 shows that the effectiveness of an attack depends strongly on its objective and the manipulated feature group. Coordinated FDI (A6) produces the strongest non-adaptive concealment effect, changing 48.43% of initially correct unstable predictions to stable and reducing the mean instability probability from approximately 0.9717 to 0.4991. Reaction-time bias (A1) produces a similar effect, with a 45.10% unstable-to-stable flip rate despite modifying only reaction-time variables. Node takeover (A5) and replay style substitution (A9) also produce substantial concealment, with flip rates of 39.57% and 24.91%, respectively. The behavior is different for the false alarm objective. Price-response bias (A2) produces the largest stable-tounstable flip rate at 27.65%, while its concealment rate is lower at

Detecting False Data Injection and Unstable Operation in Smart Grid via System-Aware Graph Boundary Learning

Table 3: Impact of the evaluation attacks on the auxiliary supervised stability predictor. S→U denotes the percentage of initially correct stable predictions changed to unstable, while U→S denotes the percentage of initially correct unstable predictions changed to stable. 𝑀𝑆 and 𝑀𝑈 denote the mean predicted probability of the unstable class after attacks on stable and unstable configurations, respectively.

Stable-only GNN score distribution

3000

Stable test Unstable test Stable-only threshold

2500

Count

2000 1500 1000 500

Attack

S→U (%)

U→S (%)

𝑀𝑆

𝑀𝑈

A1 A2 A3 A4 A5 A6 A7 A8 A9 A10

16.73 27.65 1.73 0.69 20.66 22.29 1.04 0.66 17.74 N/A

45.10 18.40 1.70 0.23 39.57 48.43 0.56 0.28 24.91 39.73

0.1648 0.2684 0.0344 0.0326 0.2025 0.2182 0.0327 0.0324 0.1753 N/A

0.5305 0.8007 0.9518 0.9701 0.5875 0.4991 0.9676 0.9696 0.7280 0.5871

18.40%. Reaction-time bias (A1) shows the opposite tendency, being considerably more effective at concealing unstable configurations than at creating false alarms. This indicates that different DSGC feature groups affect the stability decision in different directions. Limited impact of power-based manipulations. Attacks that primarily modify the power profile have a much smaller effect on the auxiliary classifier. Power imbalance (A3) produces flip rates of only 1.73% and 1.70% in the two directions, while balance-preserving power (A4), load redistribution (A7), and local load redistribution (A8) remain close to or below 1%. Thus, modifying more reported values does not necessarily lead to a larger change in the predicted stability state. Under the considered setting, the choice of manipulated feature appears more important than the number of modified entries. Adaptive single-node concealment. The adaptive attack (A10) is evaluated only in the unstable-to-stable direction because it is designed specifically to conceal unstable operating points. It changes 39.73% of initially correct unstable MLP predictions to stable and reduces the mean instability probability to 0.5871. This is a substantial effect despite the attacker controlling only a single participant. The result provides an additional stress test showing that the manipulated reports found through adaptive search can also meaningfully affect a conventional supervised stability classifier. The auxiliary MLP confirms that several of the evaluation attacks are capable of producing meaningful changes in a supervised stability decision. The strongest attack depends on the adversarial objective as price-response bias is most effective at inducing false alarms, whereas coordinated FDI and reaction-time bias produce the strongest non-adaptive concealment effects. These results provide an attack-impact reference for the StarGNN evaluation in Section 7.4, where the main question is whether the proposed stable region approach continues to identify such manipulated configurations as abnormal.

0

0

10

20

30

One-class GNN score

40

50

Figure 3: Distribution of the raw StarGNN risk scores for held-out stable and unstable DSGC configurations.

7.3

StarGNN Stability Prediction Evaluation

First, we investigated StarGNN’s ability to detect instability in our dataset. Our system is trained, and the threshold is calibrated without using unstable configurations. The model is trained only on stable samples and their physics-constrained pseudo-negative counterparts. The stability prediction results are reported using the threshold obtained from the 0.90 quantile of the stable validation scores (threshold value is discussed in the ablation studies in Section 8.1). StarGNN achieves an accuracy of 0.989, with macro-averaged precision, recall, and F1-score of 0.991, 0.952, and 0.970, respectively. The ROC AUC reaches 0.999. The macro-averaged scores are important because the test set contains substantially more unstable than stable configurations. The macro F1-score of 0.970 indicates that the detector performs well across both classes. The ROC AUC of 0.999 shows that the continuous risk score provides a very strong ordering of stable and unstable configurations across different possible thresholds. In addition to the full ROC AUC, we report the standardized partial ROC AUC over the low-false-alarm region FPR ≤ 0.05 which was 0.993. This measure focuses the comparison on the operating range most relevant to practical monitoring, where high detection sensitivity is required without causing excessive false alarms. Moreover, StarGNN’s performance on the unseen unstable set suggests that the pseudo-negatives used during training provided a useful contrastive signal to shape the boundary around the stable operating region. Figure 3 provides a clearer view of the separation produced by the learned risk score. The stable test samples are concentrated within a relatively narrow low score region, whereas the unstable configurations are distributed mainly at considerably higher scores. Most samples, therefore, lie on the expected side of the stable-only threshold. Furthermore, to examine whether the generated pseudonegatives provide a consistent training direction, we generated fresh perturbations from held-out stable test configurations and compared each pseudo-negative score with the score of its clean source sample. As shown in Appendix Fig. 5, 96.55% of the clean– pseudo-negative pairs produced a positive score difference, while 86.54% satisfied the full ranking margin of 1.0. The median score increase was 7.19. A one-sided Wilcoxon signed-rank test also confirmed that the pseudo-negative scores were significantly higher

Emad Efatinasab, Denis Donadel, Mirco Rampazzo, and Chuadhry Mujeeb Ahmed

Table 4: Stable-origin unseen-FDI detection results at 𝑞 = 0.90. The attacks are applied to clean stable configurations and represent the false-alarm objective. Attack

Acc.

M Prec.

M Rec.

M F1

Attack Rec.

ROC AUC

A1 A2 A3 A4 A5 A6 A7 A8 A9

0.913 0.939 0.842 0.886 0.915 0.938 0.920 0.849 0.904

0.913 0.941 0.848 0.886 0.915 0.940 0.920 0.853 0.904

0.913 0.939 0.842 0.886 0.915 0.938 0.920 0.849 0.904

0.913 0.939 0.842 0.886 0.915 0.938 0.920 0.848 0.904

0.921 0.973 0.780 0.867 0.926 0.971 0.935 0.793 0.904

0.966 0.985 0.906 0.944 0.969 0.985 0.969 0.912 0.960

Mean

0.900

0.902

0.900

0.900

0.896

0.955

M Prec., M Rec., and M F1 denote macro precision, macro recall, and macro F1-score. Attack Rec. is the proportion of attacked stable configurations classified as abnormal.

Table 5: Post-attack instability detection at 𝑞 = 0.90 for attacks applied to unstable configurations. This evaluation represents the instability-concealment objective. Attack

Acc.

M Prec.

M Rec.

M F1

U Rec.

ΔRec.

ROC AUC

A1 A2 A3 A4 A5 A6 A7 A8 A9 A10

0.938 0.950 0.952 0.951 0.945 0.943 0.952 0.952 0.946 0.902

0.940 0.954 0.956 0.955 0.948 0.945 0.956 0.956 0.949 0.902

0.938 0.950 0.952 0.951 0.945 0.943 0.952 0.952 0.946 0.902

0.938 0.950 0.951 0.951 0.945 0.942 0.952 0.951 0.946 0.902

0.972 0.996 0.998 0.998 0.985 0.980 0.999 0.998 0.987 0.899

0.027 0.003 0.000 0.001 0.013 0.018 0.000 0.000 0.011 0.100

0.987 0.997 0.998 0.999 0.992 0.990 0.999 0.999 0.994 0.975

Mean

0.943

0.946

0.943

0.942

0.981

0.017

0.993

U Rec. denotes the recall of attacked unstable configurations. ΔRec. denotes the reduction relative to the clean unstable recall of 0.9995. M Prec., M Rec., and M F1 denote macro precision, macro recall, and macro F1-score. For the adaptive attack A10, we used 25 additional binary-alarm queries.

than their paired clean scores (𝑝 < 10−16 ). These results show that the generator produces controlled deviations that the trained model consistently places outside the low risk region associated with stable operation. This confirms that they provide informative directional stress samples for learning the boundary around the observed stable region.

7.4

Unseen FDI Attack Detection Results

The attack evaluation uses the same StarGNN risk score and decision threshold as for stability prediction. The threshold is calibrated once at the 0.90 quantile of the stable validation scores and is not adjusted for any attack family. We evaluate the attacks under the two objectives defined in the threat model. For the false-alarm objective, attacks are applied to clean stable configurations, and the detector must distinguish the manipulated reports from normal operation. For the concealment objective, attacks are applied to unstable configurations, and the relevant question is whether the manipulated input can be moved into the learned stable region. Detection under the false alarm induction objective. Table 4 shows that StarGNN detects most manipulated stable configurations despite never observing these attacks during training. Price response bias and coordinated FDI are the most readily detected cases, with

attack recalls of 0.973 and 0.971, respectively. Reaction time bias, node takeover, load redistribution, and replay node substitution are also detected reliably. The replay result (A9) is particularly relevant because the substituted values originate from another valid stable configuration, yet the resulting participant level combination is still frequently recognized as inconsistent with the learned stable operating region. The more difficult cases are those that make relatively limited changes to the power profile. Power imbalance (A3) and local load redistribution (A8) reach attack recalls of 0.780 and 0.793, respectively, while the balance-preserving power attack (A4) reaches 0.867. These are also among the attacks with the smallest effect on the auxiliary supervised MLP used in Section 7.2 to quantify attack impact. The attacks that remain closest to the learned stable region are therefore also the least effective at changing a conventional supervised stability decision. This suggests that the lower detection rates are associated with comparatively mild manipulations rather than with attacks that strongly alter the predicted system state while remaining undetected. Resistance to instability-concealment attacks. Table 5 considers the more critical setting in which the source configuration is genuinely unstable. Before manipulation, StarGNN identifies 99.9% of the selected unstable configurations. After the nine non-adaptive attacks (A1-A9), unstable recall remains above 97% in all cases, indicating that these manipulations generally fail to move an unsafe configuration into the stable region learned by StarGNN. This result is particularly relevant when considered together with the attack impact reported in Table 3. Coordinated FDI (A6), reaction time bias (A1), node takeover (A5), and replay style node substitution (A9) cause unstable to stable flips in the MLP in 48.43%, 45.10%, 39.57%, and 24.91% of the corresponding cases, respectively. Within the same attack families, StarGNN retains unstable recall values of 0.980, 0.972, 0.985, and 0.987. Thus, attacks that can substantially alter a conventional supervised classification still leave most manipulated unstable configurations outside the stable operating region learned by StarGNN. Reaction time bias (A1) produces the largest reduction in unstable recall among the non-adaptive attacks, from 0.999 on clean unstable configurations to 0.972 after manipulation. Coordinated FDI (A6) follows at 0.980, while node takeover (A5) and replay style substitution (A9) retain recalls of 0.985 and 0.987, respectively. The remaining attacks have only a small effect on instability detection. Price response bias (A2) retains an unstable recall of 0.996, while power imbalance (A3), balancepreserving power (A4), load redistribution (A7), and local load redistribution (A8) all remain above 0.998. The contrast between the two attacker objectives is also notable. Manipulations of otherwise stable configurations are harder to detect than attempts to conceal already unstable configurations. In the latter case, the source configuration already lies outside the learned stable region, and most non-adaptive manipulations are insufficient to move it back into that region. This is reflected in the consistently high post-attack unstable recall, compared with the wider variation observed for stable-origin attack detection. Impact of adaptive binary feedback. The single node adaptive concealment attack (A10) provides a more demanding test. With 25 additional binary alarm queries, unstable recall decreases to 0.899,

Detecting False Data Injection and Unstable Operation in Smart Grid via System-Aware Graph Boundary Learning

corresponding to a reduction of 0.100 from the clean value. The ROC AUC remains 0.975, indicating that the continuous StarGNN risk score still provides substantial separation between clean, stable, and adaptively manipulated, unstable configurations. The auxiliary MLP provides a complementary measure of attack impact. Under the same adaptive manipulation, 39.73% of its initially correct unstable predictions flip to the stable class, confirming that the attack can substantially affect a conventional supervised stability classifier. Against StarGNN, the detector evasion rate is 10.06%. The MLP result establishes that the manipulation can change a supervised stability decision, whereas the StarGNN result assesses whether the proposed stable region approach continues to recognize the manipulated configuration as abnormal. Even with adaptive binary feedback, most attacked unstable configurations remain outside the stable region learned by StarGNN. The adaptive setting should be interpreted as a stress test rather than as an attacker’s capability that is always available in practice. It assumes that the adversary can repeatedly submit modified reports, observe the resulting binary alarm, and associate each response with the corresponding candidate. Such feedback may be delayed or unavailable in deployments where alarms are visible only to system operators. Repeated probing may also create a detectable sequence of related submissions. Stateful defenses for query-based black-box attacks have shown that query histories can be used to identify iterative or highly similar inputs [12, 22], while rate limiting would further restrict the number of opportunities available for adaptive search. None of the evaluated attacks is used during training, and all attack samples are generated only after the StarGNN model and its decision threshold have been fixed. The physics-constrained pseudo-negative generator modifies the reaction time and price response variables, whereas the evaluation also includes power manipulation, node compromise, load redistribution, replay-style substitution, coordinated FDI, and adaptive binary feedback search. The observed detection performance, therefore, cannot be explained simply by reproducing the perturbation patterns used during training.

8

Discussion

In this section, we start by ablation studies presented in Section 8.1. Section 8.2 compares StarGNN with baselines and Section 8.3 discusses computational complexity.

8.1

Ablation Study

We first examine the effect of the threshold quantile and then study the main components of the proposed framework. Threshold Sensitivity. The decision threshold is obtained from the empirical distribution of the stable validation scores. A higher quantile produces a more permissive threshold, reducing false alarms on stable configurations at the cost of lower sensitivity to unstable or manipulated samples near the learned boundary. The complete sensitivity analysis for clean stability prediction, the six generic attacks, the three literature-inspired attacks, and the adaptive attack is reported in Appendix Fig. 6. The fixed unseen attacks remain detectable across most of the examined threshold range, whereas the adaptive attack is more

Table 6: Component ablation at 𝑞 = 0.90.

Configuration

Acc.

M Prec.

M Rec.

M F1

ROC AUC

Full framework 0.989 0.991 0.952 0.970 0.999 Without ranking 0.987 0.982 0.948 0.964 0.998 MLP, no topology 0.969 0.903 0.940 0.920 0.989 Without pseudo-negatives 0.101 0.051 0.500 0.092 0.494

sensitive to increasingly permissive thresholds. At the selected operating point of 𝛼 = 0.90, clean stability prediction achieves a macro F1-score of 0.970. The generic attacks obtain a mean attack recall of 0.988, with a worst-case recall of 0.972, while the literatureinspired attacks achieve a mean recall of 0.995 and a worst-case recall of 0.987. The threshold-specific adaptive attack yields a macro F1-score of 0.902 and an attack recall of 0.899. These results indicate that 𝛼 = 0.90 provides a reasonable balance between reducing false alarms on normal configurations and retaining sensitivity to both fixed unseen attacks and the stronger binary feedback adaptive attack. Contribution of the Main Components. We next remove or replace individual components while keeping the data partitions, optimization settings, and threshold quantile unchanged. Table 6 reports the main results on the genuine stability prediction task. The non-topological baseline is a parameter-matched MLP that receives the same 4 × 3 node tensor after flattening it into a twelve dimensional vector. It therefore has access to the same input values, but it does not perform producer–consumer message passing. In the clean-only variant, pseudo-negatives are removed, and the model is trained using only the SVDD compactness and variance terms. Since its risk head receives no supervision, the distance from the stable center is used as its test score. Removing the pseudo-negatives causes the largest drop. The macro F1-score falls from 0.970 to 0.092, and the ROC AUC decreases from 0.999 to 0.494. In this setting, training only on stable samples does not provide enough information to establish which departures should receive higher risk scores. The model consequently fails to distinguish the unseen unstable configurations from the stable region. This result shows that the pseudo-negatives are not a minor data augmentation step; they provide the contrast required to learn a useful boundary without using genuine unstable labels. Replacing StarGNN with the parameter-matched MLP also reduces performance. The macro F1-score falls to 0.920, despite both models receiving the same twelve standardized features and using the same pseudo-negative samples and loss terms. The difference supports the use of the producer–consumer structure rather than treating the participant variables as an unstructured vector. Removing the ranking term produces a smaller, but consistent, decline. The macro F1-score decreases to 0.964, while ROC AUC falls to 0.998. Binary separation already distinguishes stable samples from pseudo-negatives, but the pairwise ranking term further encourages every selected pseudo-negative to receive a higher score than the stable configuration from which it was generated. The ablation results identify the pseudo-negative generator as the main source of contrastive boundary information and the topology-aware architecture as an important part of the final performance.

Emad Efatinasab, Denis Donadel, Mirco Rampazzo, and Chuadhry Mujeeb Ahmed

Table 7: Threshold-independent comparison of the proposed method and clean stable baselines. Attack ROC AUC values are averaged across the nine non-adaptive evaluation attacks.

Method PCA reconstruction MLP autoencoder Deep SVDD Isolation Forest One-Class SVM StarGNN (proposed)

Stability AUC

False-alarm AUC

Concealment AUC

0.786 0.755 0.725 0.659 0.541 0.999

0.657 0.798 0.797 0.721 0.811 0.955

0.817 0.886 0.783 0.811 0.825 0.993

Stability AUC compares clean stable and unstable configurations. False-alarm AUC compares clean stable configurations with attacked stable configurations. Concealment AUC compares clean stable configurations with attacked unstable configurations.

8.2

Comparison with Stable-Only Baselines

The proposed model was compared with established anomaly detection baselines such as One-Class SVM, Isolation Forest , PCA reconstruction, an MLP autoencoder, and Deep SVDD. All methods were trained using the same clean stable training partition, and the feature scaler was fitted only on this partition. Genuine unstable configurations and attacked observations were excluded from baseline training and threshold calibration. The conventional baselines were retained in their standard form and were not provided with the physics-constrained pseudo-negatives used to train StarGNN. The comparison is based on ROC AUC because it evaluates the complete ranking of the anomaly scores independently of a particular operating threshold. This is important because the methods produce scores on different numerical scales and because the complete stability test set is strongly imbalanced. Three evaluations are reported in Table 7. StarGNN achieves a ROC AUC of 0.999 for genuine instability, 0.955 for stable-origin attack detection, and 0.993 for attacked unstable configurations. It therefore maintains a strong score ordering across all three settings. The difference between the two attack columns is also interesting as most conventional baselines achieve higher ROC AUC in the unstable-origin setting than in the stable-origin setting. For example, the MLP autoencoder increases from 0.798 to 0.886, while PCA reconstruction increases from 0.657 to 0.817. This does not necessarily indicate stronger identification of the injected manipulation. The underlying inputs are already unstable, and their physical abnormality may remain visible even after the attack. One-Class SVM shows a different limitation. It reaches a stable-origin attack ROC AUC of 0.811 despite obtaining only 0.541 for stability prediction task. Similarly, reconstruction-based methods retain some anomaly ranking ability, but their results remain below those of StarGNN, particularly for genuine instability and subtle stable-origin manipulations.

8.3

Computational Complexity

Let 𝑛 denote the number of stable training configurations, 𝐾 the number of precomputed stress variants per configuration, 𝑉 the number of graph nodes, 𝐻 the hidden dimension, and 𝐿 the number of message-passing layers. Constructing the bounded random stress

bank requires O (𝑛𝐾𝐷) operations and memory, where 𝐷 is the number of input variables. In the present implementation, 𝐾 = 16 and 𝐷 = 12 are fixed; therefore, stress-bank generation and storage scale linearly with the number of training configurations, i.e., O (𝑛). Although all 𝐾 variants are retained, only one stress configuration is sampled for each stable observation in a training iteration, so the training cost does not increase by a factor of 𝐾. For one configuration, the node encoder and linear transformations in the message-passing layers require approximately O (𝐿𝑉 𝐻 2 ) operations. Message aggregation over the star topology is linear in the number of nodes, O (𝐿𝑉 𝐻 ), since the producer receives an aggregate of the consumer messages and each consumer receives the producer message; no all pairs node interaction is computed. Because training evaluates both the clean and selected stressed configurations, one epoch has complexity O (𝑛𝐿𝑉 𝐻 2 ), up to a constant factor of two. Recomputing the SVDD center every 𝑅 epochs introduces an additional O ((𝐸/𝑅)𝑛𝐿𝑉 𝐻 2 ) cost over 𝐸 training epochs. Hence, the  complete training procedure has complexity O 𝐸𝑛𝐿𝑉 𝐻 2 + 𝑅𝐸 𝑛𝐿𝑉 𝐻 2 which is linear in 𝑛 when 𝐸, 𝑅, 𝐿, 𝑉 , and 𝐻 are fixed. In the evaluated model, 𝐸 = 150, 𝑅 = 10, 𝐿 = 2, 𝑉 = 4, and 𝐻 = 64. During inference, each observation is converted into a four-node graph and processed through a single forward pass. Scoring 𝑛 test configurations therefore requires O (𝑛 test 𝐿𝑉 𝐻 2 ) operations. Furthermore, the proposed model required an average of 1.0864 seconds per epoch, with a total training time of 163.1861 seconds. All experiments were conducted on the free Kaggle cloud platform using an Intel Xeon CPU at 2.20 GHz, 32 GB of RAM, and an NVIDIA Tesla T4 GPU under Ubuntu Linux. Scalability. For a producer–consumer star graph with 𝑉 nodes, the number of edges is 𝑉 − 1. The message-passing cost therefore grows linearly with the number of participants rather than quadratically, since StarGNN does not evaluate all-pairs node interactions. More precisely, the per-configuration forward cost can be written as O (𝐿𝑉 𝐻 2 + 𝐿𝑉 𝐻 ), where the first term accounts for the shared linear transformations and the second for message aggregation. For fixed 𝐿 and 𝐻 , this gives linear computational growth with 𝑉 . The model size itself does not grow with the number of participants. The same node encoder and message-passing transformations are reused across nodes, and the consumer representations are combined through permutation invariant aggregation. Thus, increasing the number of consumers increases the amount of computation required for a forward pass, but does not require a separate set of parameters for each additional participant.

9

Conclusion

This paper presented a stable-only framework that uses one model and one threshold for both DSGC stability prediction and unseen attack detection. The model is trained on stable configurations and physics-constrained pseudo-negatives, which provide controlled departures from the stable region without being treated as genuinely unstable samples. StarGNN further preserves the producer– consumer structure instead of processing the twelve inputs as an unstructured vector. The same risk score performs well on stability prediction and provides acceptable detection across the nine unseen attacks considered in the threat model. The ablation results show that pseudo-negative generation is essential for learning a useful

Detecting False Data Injection and Unstable Operation in Smart Grid via System-Aware Graph Boundary Learning

boundary, while the lower performance of the parameter-matched MLP confirms the benefit of topology-aware graph learning. Limitation and Future Work. The main limitation concerns the size of the evaluated DSGC system. The benchmark used in this work contains one producer and three consumers and is the only publicly available data set derived from the DSGC system. The reported detection results are therefore necessarily validated in this four-node setting. The StarGNN architecture does not rely on all-pairs message passing, and its shared transformations and consumer aggregation lead to linear computational growth with the number of participants for fixed model dimensions. This provides a favorable computational scaling property, but it does not establish that the same detection performance will be retained in larger DSGC systems. Larger systems may exhibit operating patterns and attack interactions that are not represented in the available benchmark. Validation on larger DSGC configurations will therefore require a suitable larger-scale simulation environment, and is foreseen as future work.

References [1] Love Allen Chijioke Ahakonye, Cosmas Ifeanyi Nwakanma, Jae-Min Lee, and Dong-Seong Kim. 2024. Low computational cost convolutional neural network for smart grid frequency stability prediction. Internet of Things 25 (2024), 101086. doi:10.1016/j.iot.2024.101086 [2] Alaa Alaerjan and Randa Jabeur. 2025. An online learning method for assessing smart grid stability under dynamic perturbations. Scientific Reports 15, 1 (2025), 10559. [3] Mamoun Alazab, Suleman Khan, Somayaji Siva Rama Krishnan, Quoc-Viet Pham, M. Praveen Kumar Reddy, and Thippa Reddy Gadekallu. 2020. A Multidirectional LSTM Model for Predicting the Stability of a Smart Grid. IEEE Access 8 (2020), 85454–85463. doi:10.1109/ACCESS.2020.2991067 [4] Zaid Allal, Hassan N. Noura, Ola Salman, and Khaled Chahine. 2024. Leveraging the power of machine learning and data balancing techniques to evaluate stability in smart grids. Engineering Applications of Artificial Intelligence 133 (2024), 108304. doi:10.1016/j.engappai.2024.108304 [5] Manuel S. Alvarez-Alvarado, Christhian Apolo-Tinoco, Maria J. Ramirez-Prado, Francisco E. Alban-Chacón, Nabih Pico, Jonathan Aviles-Cedeno, Angel A. Recalde, Felix Moncayo-Rea, Washington Velasquez, and Johnny Rengifo. 2024. Cyber-physical power systems: A comprehensive review about technologies drivers, standards, and future perspectives. Computers and Electrical Engineering 116 (2024), 109149. doi:10.1016/j.compeleceng.2024.109149 [6] Gulfaraz Anis, Naila Samar Naz, Taher M. Ghazal, Muhammad Sajid Farooq, Muhammad Saleem, Chan Yeob Yeun, Munir Ahmad, and Khan Muhammad Adnan. 2026. Smart and transparent grid stability prediction for efficient energy management using explainable AI. Energy Strategy Reviews 64 (2026), 102083. doi:10.1016/j.esr.2026.102083 [7] Souhila Aoufi, Abdelouahid Derhab, and Mohamed Guerroumi. 2020. Survey of false data injection in smart power grid: Attacks, countermeasures and challenges. Journal of Information Security and Applications 54 (2020), 102518. doi:10.1016/j. jisa.2020.102518 [8] Vadim Arzamasov. 2018. Electrical Grid Stability Simulated Data . UCI Machine Learning Repository. DOI: https://doi.org/10.24432/C5PG66. [9] Vadim Arzamasov, Klemens Böhm, and Patrick Jochem. 2018. Towards Concise Models of Grid Stability. In 2018 IEEE International Conference on Communications, Control, and Computing Technologies for Smart Grids (SmartGridComm). 1–6. doi:10.1109/SmartGridComm.2018.8587498 [10] Kemal Aygul, Mostafa Mohammadpourfard, Mert Kesici, Fatih Kucuktezcan, and Istemihan Genc. 2024. Benchmark of machine learning algorithms on transient stability prediction in renewable rich power grids under cyber-attacks. Internet of Things 25 (2024), 101012. doi:10.1016/j.iot.2023.101012 [11] Adrien Bardes, Jean Ponce, and Yann LeCun. 2021. Vicreg: Varianceinvariance-covariance regularization for self-supervised learning. arXiv preprint arXiv:2105.04906 (2021). [12] Steven Chen, Nicholas Carlini, and David Wagner. 2020. Stateful Detection of Black-Box Adversarial Attacks. In Proceedings of the 1st ACM Workshop on Security and Privacy on Artificial Intelligence (Taipei, Taiwan) (SPAI ’20). Association for Computing Machinery, New York, NY, USA, 30–39. doi:10.1145/3385003. 3410925

[13] Emad Efatinasab, Nahal Azadi, Gian Antonio Susto, Chuadhry Mujeeb Ahmed, and Mirco Rampazzo. 2025. Fortifying smart grid stability: Defending against adversarial attacks and measurement anomalies. Sustainable Energy, Grids and Networks 43 (2025), 101799. doi:10.1016/j.segan.2025.101799 [14] Emad Efatinasab, Alessandro Brighente, Denis Donadel, Mauro Conti, and Mirco Rampazzo. 2025. Towards robust stability prediction in smart grids: GAN-based approach under data constraints and adversarial challenges. Internet of Things 33 (2025), 101662. doi:10.1016/j.iot.2025.101662 [15] Emad Efatinasab, Alessandro Brighente, Mirco Rampazzo, Nahal Azadi, and Mauro Conti. 2024. GAN-GRID: A Novel Generative Attack on Smart Grid Stability Prediction. In Computer Security – ESORICS 2024, Joaquin Garcia-Alfaro, Rafał Kozik, Michał Choraś, and Sokratis Katsikas (Eds.). Springer Nature Switzerland, Cham, 374–393. [16] Arman Fathollahi. 2025. Machine Learning and Artificial Intelligence Techniques in Smart Grids Stability Analysis: A Review. Energies 18, 13 (2025). doi:10.3390/ en18133431 [17] Justin Gilmer, Samuel S Schoenholz, Patrick F Riley, Oriol Vinyals, and George E Dahl. 2017. Neural message passing for quantum chemistry. In International conference on machine learning. Pmlr, 1263–1272. [18] Marian B. Gorzałczany, Jakub Piekoszewski, and Filip Rudziński. 2020. A Modern Data-Mining Approach Based on Genetically Optimized Fuzzy Systems for Interpretable and Accurate Smart-Grid Stability Prediction. Energies 13, 10 (2020). doi:10.3390/en13102559 [19] Wei Guo, Xiang Zha, Kun Qian, and Tao Chen. 2019. Can Active Learning Benefit the Smart Grid? A Perspective on Overcoming the Data Scarcity. In 2019 IEEE 2nd International Conference on Electronics and Communication Engineering (ICECE). 346–350. doi:10.1109/ICECE48499.2019.9058539 [20] Youbiao He, Gihan J. Mendis, and Jin Wei. 2017. Real-Time Detection of False Data Injection Attacks in Smart Grid: A Deep Learning-Based Intelligent Mechanism. IEEE Transactions on Smart Grid 8, 5 (2017), 2505–2516. doi:10.1109/TSG.2017. 2703842 [21] Charalambos Konstantinou and Michail Maniatakos. 2016. A Case Study on Implementing False Data Injection Attacks Against Nonlinear State Estimation. In Proceedings of the 2nd ACM Workshop on Cyber-Physical Systems Security and Privacy (Vienna, Austria) (CPS-SPC ’16). Association for Computing Machinery, New York, NY, USA, 81–92. doi:10.1145/2994487.2994491 [22] Huiying Li, Shawn Shan, Emily Wenger, Jiayun Zhang, Haitao Zheng, and Ben Y Zhao. 2022. Blacklight: Scalable defense for neural networks against { QueryBased } { Black-Box } attacks. In 31st USENIX Security Symposium (USENIX Security 22). 2117–2134. [23] Zirui Liao, Jian Shi, Shaoping Wang, Yuwei Zhang, Rui Mu, and Zhiyong Sun. 2025. A Survey of Resilient Coordination for Cyber–Physical Systems Against Malicious Attacks. IEEE Internet of Things Journal 12, 22 (2025), 46500–46525. doi:10.1109/JIOT.2025.3612764 [24] Xuan Liu and Zuyi Li. 2014. Local Load Redistribution Attacks in Power Systems With Incomplete Network Information. IEEE Transactions on Smart Grid 5, 4 (2014), 1665–1676. doi:10.1109/TSG.2013.2291661 [25] Yigu Liu, Alexandru Ştefanov, and Peter Palensky. 2024. Generating Large-Scale Synthetic Communication Topologies for Cyber–Physical Power Systems. IEEE Transactions on Industrial Informatics 20, 11 (2024), 13463–13472. doi:10.1109/TII. 2024.3438232 [26] Richard Macwan, Christopher Drew, Prosper Panumpabi, Alfonso Valdes, Nitin Vaidya, Pete Sauer, and Dmitry Ishchenko. 2016. Collaborative defense against data injection attack in IEC61850 based smart substations. In 2016 IEEE Power and Energy Society General Meeting (PESGM). 1–5. doi:10.1109/PESGM.2016.7741376 [27] Yilin Mo and Bruno Sinopoli. 2009. Secure control against replay attacks. In 2009 47th Annual Allerton Conference on Communication, Control, and Computing (Allerton). 911–918. doi:10.1109/ALLERTON.2009.5394956 [28] Amir-Hamed Mohsenian-Rad and Alberto Leon-Garcia. 2011. Distributed Internet-Based Load Altering Attacks Against Smart Power Grids. IEEE Transactions on Smart Grid 2, 4 (2011), 667–674. doi:10.1109/TSG.2011.2160297 [29] John Mulo, Pu Tian, Adamu Hussaini, Hengshuo Liang, and Wei Yu. 2023. Towards an adversarial machine learning framework in cyber-physical systems. In 2023 IEEE/ACIS 21st International Conference on Software Engineering Research, Management and Applications (SERA). IEEE, 138–143. [30] Mithat Önder, Muhsin Ugur Dogan, and Kemal Polat. 2023. Classification of smart grid stability prediction using cascade machine learning methods and the internet of things in smart grid. Neural Computing and Applications (2023), 1–19. [31] Peter Palensky and Dietmar Dietrich. 2011. Demand Side Management: Demand Response, Intelligent Energy Systems, and Smart Loads. IEEE Transactions on Industrial Informatics 7, 3 (2011), 381–388. doi:10.1109/TII.2011.2158841 [32] Lukas Ruff, Robert Vandermeulen, Nico Goernitz, Lucas Deecke, Shoaib Ahmed Siddiqui, Alexander Binder, Emmanuel Müller, and Marius Kloft. 2018. Deep one-class classification. In International conference on machine learning. PMLR, 4393–4402. [33] Benjamin Schäfer, Moritz Matthiae, Marc Timme, and Dirk Witthaut. 2015. Decentral smart grid control. New journal of physics 17, 1 (2015), 015002.

Emad Efatinasab, Denis Donadel, Mirco Rampazzo, and Chuadhry Mujeeb Ahmed

[34] Omer Sen, Dennis van der Velde, Philipp Linnartz, Immanuel Hacker, Martin Henze, Michael Andres, and Andreas Ulbig. 2021. Investigating Man-in-theMiddle-based False Data Injection in a Smart Grid Laboratory Environment. In 2021 IEEE PES Innovative Smart Grid Technologies Europe (ISGT Europe). 01–06. doi:10.1109/ISGTEurope52324.2021.9640002 [35] Zhongtuo Shi, Wei Yao, Zhouping Li, Lingkang Zeng, Yifan Zhao, Runfeng Zhang, Yong Tang, and Jinyu Wen. 2020. Artificial intelligence techniques for stability analysis and control in smart grids: Methodologies, applications, challenges and future directions. Applied Energy 278 (2020), 115733. doi:10.1016/j.apenergy. 2020.115733 [36] Abdulrahman Takiddin, Rachad Atat, Muhammad Ismail, Osman Boyaci, Katherine R. Davis, and Erchin Serpedin. 2023. Generalized Graph Neural NetworkBased Detection of False Data Injection Attacks in Smart Grids. IEEE Transactions on Emerging Topics in Computational Intelligence 7, 3 (2023), 618–630. doi:10.1109/TETCI.2022.3232821 [37] Abdulrahman Takiddin, Muhammad Ismail, Rachad Atat, Katherine R. Davis, and Erchin Serpedin. 2024. Graph Autoencoder-Based Detection of Unseen False Data Injection Attacks in Smart Grids. In Intelligent Systems and Applications, Kohei Arai (Ed.). Springer Nature Switzerland, Cham, 234–244. [38] Ferhat Ucar. 2023. A Comprehensive Analysis of Smart Grid Stability Prediction along with Explainable Artificial Intelligence. Symmetry 15, 2 (2023). doi:10. 3390/sym15020289 [39] Praveen Verma and Chandan Chakraborty. 2024. Load Redistribution Attacks Against Smart Grids–Models, Impacts, and Defense: A Review. IEEE Transactions on Industrial Informatics 20, 8 (2024), 10192–10208. doi:10.1109/TII.2024.3393005 [40] J. K. Wang and Chunyi Peng. 2017. Analysis of Time Delay Attacks against Power Grid Stability. In Proceedings of the 2nd Workshop on Cyber-Physical Security and Resilience in Smart Grids (Pittsburgh, PA, USA) (CPSR-SG’17). Association for Computing Machinery, New York, NY, USA, 67–72. doi:10.1145/3055386.3055392 [41] Xuan Wang, XiaoFeng Zhang, Feng Zhou, and Xiang Xu. 2025. Smart grid stability prediction using artificial intelligence: A study based on the UCI smart grid stability dataset. Sustainable Computing: Informatics and Systems 47 (2025), 101175. doi:10.1016/j.suscom.2025.101175 [42] Grace Wickerson, Autumn Burton, and John A. Laitner. 2024. A Call for Immediate Public Health and Emergency Response Planning for Widespread Grid Failure Under Extreme Heat. Federation of American Scientists. https://fas.org/publication/gridfailure-extreme-heat/ [43] Ying Zhang, Jianhui Wang, and Bo Chen. 2021. Detecting False Data Injection Attacks in Smart Grids: A Semi-Supervised Deep Learning Approach. IEEE Transactions on Smart Grid 12, 1 (2021), 623–634. doi:10.1109/TSG.2020.3010510

A

Open Science

Figure 4: Overview of the proposed framework.

to their predefined admissible bounds to avoid invalid reported values.

C.1

Non-Adaptive FDI Attacks

Reaction-time bias attack. The attacker selects between one and We make the code and experimental data publicly available at: https: four reaction-time variables and modifies each selected 𝜏𝑖 according //anonymous.4open.science/r/System-Aware-Graph-Boundary-Learning-to 8E0E Δ𝜏 = 𝑠 𝜌 𝑅 , 𝜌 ∼ U (0.15, 0.50), (39) 𝑖

B

StarGNN framework details

More details regarding the configuration of StarGNN are illustrated in Figure 4. Stable configurations and pseudo-negative samples, all representing four-node star graphs, are used to train the StarGNN encoder. It employs separation and ranking loss, SVDD compactness, and regularization to learn the shape of the stable data, as discussed in Section 6. Then, a validation set from the same stable and pseudo-negative sample dataset is used to determine a threshold for the testing phase, where the model is asked to flag anomalous behaviors.

C

Detailed FDI Attack Construction

This appendix provides the exact construction of the FDI attacks summarized in Table 2. All attacks are generated only after the StarGNN model and its decision threshold have been fixed. No attack samples are used during model training or threshold calibration. Let 𝑅𝑘 denote the range of feature 𝑘 computed from the clean training data. These ranges are used to scale the perturbation magnitudes. After each manipulation, the affected features are clipped

𝑖 𝑖 𝜏𝑖

𝑖

where 𝑠𝑖 ∈ {−1, +1} determines the direction of the perturbation. The manipulated value is therefore 𝜏e𝑖 = 𝜏𝑖 + Δ𝜏𝑖 .

(40)

This attack represents falsification of the reported response-time behavior of DSGC participants. It adapts the idea of time-delay attacks studied in power-grid control systems [40]. Since the DSGC dataset does not contain packet-level transmission delays, the attack is implemented through the reported reaction-time parameter. Price-response bias attack. The attacker selects between one and four price-response coefficients and modifies each selected 𝑔𝑖 as Δ𝑔𝑖 = 𝑠𝑖 𝜌𝑖 𝑅𝑔𝑖 ,

𝜌𝑖 ∼ U (0.15, 0.55),

(41)

with 𝑠𝑖 ∈ {−1, +1}. The resulting value is 𝑔e𝑖 = 𝑔𝑖 + Δ𝑔𝑖 .

(42)

The attack represents falsification of the reported participant response to frequency-dependent prices and changes how the reported economic-control behavior is presented to the monitoring system.

Detecting False Data Injection and Unstable Operation in Smart Grid via System-Aware Graph Boundary Learning

Power-imbalance attack. The attacker modifies between one and three consumer-power values. For each selected consumer, Δ𝑝𝑖 = 𝑠𝑖 𝜌𝑖 𝑅𝑝𝑖 ,

𝜌𝑖 ∼ U (0.15, 0.50),

(43)

where 𝑠𝑖 ∈ {−1, +1}. The producer power report 𝑝 1 is left unchanged. The attack is related to load-altering attacks in which compromised demand measurements are modified to affect grid operation [28]. In this case, the adversary controls consumer side reports but does not coordinate the changes to preserve producer– consumer power balance. Balance-preserving power attack. The attacker again modifies between one and three consumer-power values, but uses a smaller perturbation interval, Δ𝑝𝑖 = 𝑠𝑖 𝜌𝑖 𝑅𝑝𝑖 ,

𝜌𝑖 ∼ U (0.10, 0.40).

(44)

After modifying the selected consumers, the producer power is recomputed as 𝑝e1 = 𝑝e2 + 𝑝e3 + 𝑝e4 . (45) The manipulated report therefore remains consistent with the basic producer–consumer power balance relation. This attack follows the general idea of constraint-aware or stealthy FDI, where several measurements are jointly changed so that simple system consistency checks remain satisfied. Node-takeover attack. A single participant 𝑖 is selected. All three of its reported attributes, 𝜏𝑖 , 𝑝𝑖 , and 𝑔𝑖 , are modified simultaneously. Each feature receives an independently selected perturbation direction, and the magnitude is sampled between 20% and 60% of the corresponding training-derived feature range. All modified values are clipped to their admissible bounds. This attack models compromise of a participant’s reporting unit, allowing the attacker to falsify the complete local report rather than only one measurement type. Coordinated FDI attack. The attacker combines reaction time and price response manipulation with either the power-imbalance attack or the balance-preserving power attack. The attack therefore alters multiple feature groups within the same reported DSGC configuration. This scenario represents a stronger adversary capable of coordinating falsified reports across control-related and power-related variables rather than manipulating a single measurement category.

C.2

Literature-Inspired FDI Scenarios

The attack generates unequal random redistribution weights 𝑤 2, 𝑤 3, 𝑤 4 satisfying 𝑤 2 + 𝑤 3 + 𝑤 4 = 1, (47) and constructs the modified consumer reports as 𝑖 ∈ {2, 3, 4}.

𝑝e1 = 𝑝e2 + 𝑝e3 + 𝑝e4 .

(48)

(49)

Before boundary clipping, the transformation therefore preserves the aggregate consumer demand while changing its distribution across participants. The scenario models an attacker that attempts to hide the manipulation at the aggregate demand level while altering local reports. Local load-redistribution attack. Local load-redistribution attacks assume that the adversary can manipulate only a limited part of the system [24]. In our implementation, one or two consumer-power reports are selected and modified. Another consumer report is then adjusted in the opposite direction so that the total consumer demand remains approximately unchanged. After the local redistribution, the producer power is recomputed as 𝑝e1 = 𝑝e2 + 𝑝e3 + 𝑝e4 . (50) This attack models a spatially limited compromise in which the attacker controls only a subset of consumer measurements rather than all reported loads. Replay style node substitution. A conventional replay attacker records valid measurements and later presents them again to the monitoring system [27]. The DSGC dataset contains static operating configurations rather than temporal measurement sequences, so a direct temporal replay cannot be implemented. We therefore use a static replay style adaptation. One participant 𝑖 is selected, and its complete reported attribute vector (𝜏𝑖 , 𝑝𝑖 , 𝑔𝑖 )

(51)

is replaced with the corresponding participant attributes taken from another clean stable configuration. The remaining participants retain their original values. The resulting input is therefore assembled from values that are individually valid and were observed in clean operation, although the substituted participant may no longer be consistent with the rest of the current configuration.

C.3

Adaptive Single Node Concealment Attack

The adaptive attack targets an unstable operating point x and seeks a manipulated report that does not trigger the StarGNN abnormality alarm. For each attacked sample, one participant 𝑖 is selected and only its reaction time and price-response reports, I𝑖 = {𝜏𝑖 , 𝑔𝑖 },

Load-redistribution attack. Following the load-redistribution principle in [39], the attacker changes the distribution of consumer demand while attempting to preserve the aggregate reported demand. For a clean configuration, the original aggregate consumer demand is 𝑃 load = 𝑝 2 + 𝑝 3 + 𝑝 4 . (46)

𝑝e𝑖 = 𝑤𝑖 𝑃load,

The producer power is then recomputed according to

(52)

are mutable. All power reports and all attributes belonging to the other participants remain unchanged. For each mutable feature 𝑘 ∈ I𝑖 , the maximum allowed change is defined as |𝑥𝑘′ − 𝑥𝑘 | ≤ 𝜖𝑅𝑘 , (53) where 𝑅𝑘 is the corresponding training-derived feature range. In the reported experiments, 𝜖 = 0.5,

(54)

so each mutable feature may change by at most 50% of its trainingderived range. Candidate values are additionally clipped to the admissible feature bounds.

Emad Efatinasab, Denis Donadel, Mirco Rampazzo, and Chuadhry Mujeeb Ahmed

The attacker has no access to the StarGNN architecture, learned parameters, continuous risk score, gradients, or decision threshold. Its only feedback is the binary alarm returned for each submitted candidate. Starting from an initial random perturbation, the attacker performs a limited budget randomized search. Candidate generation combines two forms of exploration: • Local exploration, which samples new perturbations around the currently retained candidate; and • Global exploration, which samples candidates over the complete admissible perturbation region. Each candidate is submitted to StarGNN and the attacker observes only whether the abnormality alarm is triggered. The attacker is allowed 25 additional adaptive queries after evaluating

the initial random candidate, resulting in at most 26 detector evaluations per attacked sample. Among candidates that do not trigger the StarGNN alarm, the attacker retains the candidate with the largest normalized displacement from the original configuration. The displacement is measured as   ′ 𝑥𝑘 − 𝑥𝑘 . (55) 𝑆 (x′ ) = 𝑅𝑘 𝑘 ∈ I𝑖 2

The search, therefore, favors successful alarm free candidates that differ as much as possible from the original unstable operating point within the allowed single-participant perturbation region.

Detecting False Data Injection and Unstable Operation in Smart Grid via System-Aware Graph Boundary Learning

0.08

Paired score gaps No score increase Median score gap (7.19) Training ranking margin (1.00)

0.07

Density

0.06 0.05 0.04 0.03 0.02 0.01 0.00

20

0

20

40

60

Paired risk-score increase s(x) s(x)

80

Figure 5: Distribution of the paired risk score increase 𝑠 (e 𝑥 ) − 𝑠 (𝑥) for pseudo-negatives. 1.0

0.8

Metric value

Metric value

0.8

0.6

0.4

0.6 0.4 Mean Macro F1 Mean attack recall Worst-case attack recall Selected threshold quantile (0.90)

0.2

0.2

0.0

Threshold sensitivity: generic attacks

1.0

Macro F1 Balanced accuracy Selected quantile (0.90) 0.80

0.82

0.85

0.90

0.95

Stable-validation threshold quantile

0.97

0.0

0.99

0.80

0.82

(a) Clean stability prediction.

0.97

0.99

0.97

0.99

0.8

Metric value

Metric value

0.95

Threshold sensitivity: adaptive attack

1.0

0.8 0.6 0.4 Mean Macro F1 Mean attack recall Worst-case attack recall Selected threshold quantile (0.90)

0.2 0.0

0.90

(b) Generic unseen attacks.

Threshold sensitivity: literature-inspired attacks

1.0

0.85

Stable-validation threshold quantile

0.80

0.82

0.85

0.6 0.4 0.2

0.90

0.95

Stable-validation threshold quantile

(c) Literature inspired unseen attacks.

0.97

0.99

0.0

Macro F1 Adaptive attack recall Selected threshold quantile (0.90)

0.80

0.82

0.85

0.90

0.95

Stable-validation threshold quantile

(d) Single node adaptive concealment attack.

Figure 6: Threshold sensitivity results for clean stability prediction, generic unseen attacks, literature inspired unseen attacks, and the single node adaptive concealment attack. The vertical dotted line denotes the selected operating point, 𝛼 = 0.90. The adaptive attack is regenerated independently at each threshold using the corresponding binary detector feedback.

Record · ID 1108603 · SHA-256 59f2fa316079a0fd
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.