NOT ALL RELATIONS ARE EQUAL: RELATION-BALANCED AND CALIBRATED GRAPH LEARNING FOR PROVENANCE-BASED INTRUSION DETECTION Lijie Zheng1
Alessandro Brighente2
Yulong Shen1
Mauro Conti2,3
School of Computer Science and Technology, Xidian University, Xi’an, China 2 Department of Mathematics, University of Padova, Padova, Italy 3 School of Science and Technology, Örebro University, Örebro, Sweden ABSTRACT
Provenance-Based Intrusion Detection Systems (PIDSs) detect Advanced Persistent Threats (APTs) by analyzing system interactions. However, existing methods largely treat relations uniformly, overlooking statistical heterogeneity; in CADETS, relation frequencies differ by approximately 140,000×. This may cause PIDSs to focus more on frequent relations and overlook differences in normal error levels across relations, increasing the risk of false alarms and missed detections. We present RECAL, an unsupervised framework using relation-balanced masked graph learning to better capture rare interaction patterns. It further calibrates reconstruction errors against each relation’s benign error distribution to produce comparable anomaly evidence, helping distinguish attacks from benign behavior and reduce false alarms. On three DARPA E3 datasets, RECAL achieves F1 scores of 99.99%, 99.93%, and 99.99%, outperforming the best baseline on each dataset by 0.88, 0.82, and 0.42 percentage points, respectively. Compared with the baseline reporting the lowest FPR, RECAL reduces mean FPR by approximately 105×, 4×, and 41×.
single global threshold may then over-alert on high-score relations and overlook anomalous behavior on low-score relations [7, 10, 11]. To address these problems, we present RECAL, an unsupervised, relation-calibrated PIDS framework that models both frequent and rare interactions through relation-balanced masked graph learning and converts within-relation behavioral deviations into comparable anomaly evidence across relations, thereby reducing false positive rates. The main contributions of this work are summarized as follows. • We identify capacity misallocation and cross-relation score miscalibration as two challenges arising from relation heterogeneity in existing PIDSs. • We present RECAL, integrating relation-balanced masked graph learning with relation-calibrated detection. Relationstratified masking, independent decoder heads, and balanced reconstruction mitigate frequent-relation dominance. Empirical quantile calibration and Fisher fusion align and aggregate cross-relation evidence, improving detection and reducing false alarms.
Index Terms— Provenance graph, intrusion detection, masked graph autoencoder, relation heterogeneity, anomaly detection
• We systematically evaluate RECAL on three DARPA E3 datasets. The results confirm its effectiveness, with the highest F1 and lower false positive rates among compared methods. Our implementation is available at h t t p s : //github.com/Jiex2001/RECAL to support further research.
1. INTRODUCTION
≈140,000-fold 106 104
CL O EA O SE TE PE _O N M O RE BJ D_ PR MMAD O AP EX CE EC SS U FO TE R EX K W IT RI CO LS TE N EE SE NECK CH G A ND T _P C T R C O RE IN EP CV CIP T M A O D_ UFRO L FI N M LE LI N _ RE AT K NA TR M L E TR FC INK UN N RE C TL A SE CVM TE ND S G M FL SIG SG O N W A S_ L T B O O IND TH ER
102
CR
An Advanced Persistent Threat (APT) blends stealthy malicious actions into normal system activity over extended periods, posing challenges for signature-based or rule-based detection [1]. Kernel audit logs record interactions between system entities, and provenance graphs constructed from these logs are widely used to detect such attacks [2, 3]. Using these provenance graphs, some ProvenanceBased Intrusion Detection Systems (PIDSs) apply graph learning to model benign behavior and identify deviations to detect potential attacks [4, 5, 6, 7, 8]. However, these PIDSs are largely insensitive to the heterogeneity of interaction relations. Across relation types (system-call semantics such as read, write, and execute), audit data exhibits drastically different statistics. As shown in Fig. 1, event counts across interaction relations in the CADETS benign training graphs differ by approximately 140,000×. Existing methods nevertheless treat all relations uniformly in both learning and judgment, even when some encode relation types as input features [6, 7]. This insensitivity raises two problems. (1) Capacity misallocation. The learning signal is dominated by high-frequency relations, so rare but security-critical interactions, such as payload execution and privilege escalation, may be insufficiently modeled [4, 5, 6]. (2) Cross-relation score miscalibration. When benign score distributions differ across relations, raw scores do not indicate comparable degrees of deviation [9]. A
Edge count
arXiv:2609.16462v1 [cs.CR] 15 Sep 2026
1
Ji He1
Fig. 1. CADETS benign training graphs. The x-axis shows audit-log relations; the y-axis shows event counts (log scale).
2. METHODOLOGY The RECAL pipeline comprises four stages, as shown in Fig. 2. First, graph construction (§2.1) converts kernel audit logs into a typed provenance graph. Second, node featurization (§2.2) encodes
§ 2.1 Graph Construction
§ 2.3. Relation-Balanced Masked Graph Learning
§ 2.2 Node Featurization Masked Graph
𝒙𝒗 = 𝒕𝒗 ⊕ 𝒔𝒗 ⊕ 𝒑𝒗
System Audit Logs Parse
f
𝒕𝒗 : type (one-hot) process file IP
a
⊕
Reduce
c
Provenance Graph
𝒆𝟏(recv)
exec
exec v
Node embeddings
Loss Aggregation
…
rR
xˆ
Reconstruction errors
Stage 2. Per-Relation Calibration & Fusion
𝒆𝟑 (exec)
1
Empirical CDF
Normal
read
…
time
flatten
1
𝐷𝑒𝑥𝑒𝑐
…
Stage 1. KNN Candidate Screening
𝒆𝟐 (write)
1
xˆvrecv
§ 2.4. Relation-Calibrated Anomaly Detection
recv write exec recv write
𝐷𝑟𝑒𝑐𝑣
Masked node
⊕ IP
Encoder
…
𝒑𝒗 : transition profile File
GAT
xˆvwrite
……
Word2Vec
Process
xˆ
𝐷𝑤𝑟𝑖𝑡𝑒
a b c d e f
Backpropagation
read v
𝐷𝑟𝑒𝑎𝑑
a b c d e f
b d
𝒔𝒗 : semantic embedding /bin/bash 128.55.12.110 ……
Initial Embedding xv
e
Reconstruction Embedding
Decoder Output Embedding
Fisher Fusion
sv
write
1
Benign
Alarm
Candidate Anomaly
Fig. 2. Overview of RECAL. Graph construction and node featurization transform audit logs into attributed provenance graphs for relationbalanced masked graph learning. Node embeddings support KNN candidate screening, while relation-specific reconstruction errors support candidate calibration and evidence fusion for final detection. Orange dashed arrows indicate data flow between modules. entity types, attribute semantics, and temporal relation transitions as node features. Third, relation-balanced masked graph learning (§2.3) improves rare-relation modeling and produces per-relation reconstruction errors. Finally, relation-calibrated anomaly detection (§2.4) combines K-Nearest Neighbor (KNN) candidate screening with relation calibration and evidence fusion for two-stage node-level detection. 2.1. Graph Construction We convert kernel audit logs into a directed heterogeneous provenance graph G = (V, E, R), where V denotes the set of system entities, including processes, files, and network flows; E denotes the set of interaction edges; and R denotes the relation vocabulary defined by the audit logs, including read, write, execute, and connect. During construction, we apply standard provenance graph reduction techniques [12] to control graph size while preserving causal structure. 2.2. Node Featurization For each node v, the initial feature is xv = tv ⊕ sv ⊕ pv , where T is the set of entity types, tv ∈ {0, 1}|T | is the one-hot encoding of the entity type, and sv ∈ Rd is a semantic embedding, obtained by tokenizing the attribute strings of the node (process names, file paths, IP addresses) and averaging their Word2Vec vectors [13]. The third 2 part, pv ∈ R|R| , is a temporal relation-transition profile. Let Bv = (e1 , . . . , eL ) be the L most recent events involving v in ascending temporal order, and let r(ei ) and θ(ei ) denote the relation type and the timestamp of event ei . The profile entry for a transition (r, r′ ) is pv [r, r′ ] =
L X I r(ei−1 ) = r ∧ r(ei ) = r′ e−λ (θv −θ(ei )) , (1) i=2
where I[·] is the indicator function, θv is the timestamp of the most recent event involving v, and λ controls the temporal decay. The profile gives each interaction a short-range behavioral context. When the transitions receive→write and write→execute are active together, a download-then-execute pattern becomes visible at the feature level. Maintaining the profile costs an O(L) sliding buffer per node and requires no training.
2.3. Relation-Balanced Masked Graph Learning This module learns normal interaction patterns through masked graph autoencoding [14], using node features from benign training graphs as reconstruction targets without attack labels. To mitigate the dominance of frequent relations, we organize reconstruction into relationspecific learning tasks and balance their contributions through masking and reconstruction optimization. Masking selects the nodes whose features will be reconstructed. To increase reconstruction opportunities for rare relations, we assign each relation observed in the training data a masking rate based on its event frequency, pr = clip p0 (f¯/fr )γ , pmin , pmax , (2) where fr is the training event count of relation r, f¯ is the median nonzero frequency, p0 is the base masking rate, and pmin and pmax are its bounds. The exponent γ controls the emphasis on rare relations, with γ = 0 giving equal per-relation rates. For each relation r, we independently sample its participating nodes with probability pr ; a node is masked if selected under any relation. Its overall masking probability therefore depends on both relation frequencies and the number of participating relations. We replace the input features of selected nodes with a learnable mask vector. On the resulting masked graph, the encoder uses edge-typeconditioned attention [15, 16] to produce node embeddings. We then average the embeddings of node v’s neighbors connected by relation r, including both incoming and outgoing neighbors, to obtain mrv . Each relation has an independent lightweight linear decoder Dr that reconstructs the masked node features as x̂rv = Dr (mrv ). This allows different interaction patterns to be modeled separately using shared node representations. Because the number of reconstruction samples still varies across relations, we average the loss within each relation and sum these averages with equal weight, X 1 X L= ℓ x̂rv , xv , (3) |Mr | v∈M r∈R,|Mr |>0
r
where Mr is the set of masked nodes participating in relation r and ℓ is the scaled cosine error [14]. The objective weights each nonempty relation’s mean reconstruction loss equally, reducing the bias due to unequal sample counts. Per-relation reconstruction also retains errors associated with different interactions, allowing subsequent detection to interpret these errors against relation-specific benign baselines.
Table 1. Statistics of the DARPA E3 scenarios. Dataset Nodes Edges Malicious Mal. Size CADETS 1,452,123 7,164,818 12,857 0.89% 19.3 GB THEIA 1,558,101 3,574,557 25,338 1.63% 18.8 GB TRACE 3,197,278 4,661,252 68,135 2.13% 16.2 GB Total 6,207,502 15,400,627 106,330 1.71% 54.3 GB 2.4. Relation-Calibrated Anomaly Detection To address cross-relation score miscalibration, we interpret reconstruction errors against each relation’s benign error distribution. We reserve the tail of the benign period for calibration and exclude it from masked graph model training. After training, we obtain reconstruction errors on the calibration graph through batched masking and estimate a relation-specific empirical Cumulative Distribution Function (CDF) F̂r [17]. During the detection stage, a KNN detector [6, 18] measures the average nearest-neighbor distance between node embeddings and a benign reference set. A deliberately relaxed threshold yields a highrecall candidate set. For each candidate node v, we map its reconstruction error erv to a within-relation benign quantile qvr = F̂r (erv ) using the corresponding reference distribution. This transformation places errors from different relations on a common quantile scale, allowing their deviations from normal behavior to be compared. We aggregate anomaly evidence across the relations in which each candidate node participates. Let R(v) denote the set of relation types involving v and kv = |R(v)| their number. We interpret prv = 1 − qvr as an empirical upper-tail probability estimate and use Fisher’s method [19] to fuse the evidence, P sv = Fχ2 − 2 r∈R(v) ln prv , (4) 2kv
where Fχ2 is the CDF of the χ2 distribution with 2kv degrees 2kv of freedom. The score accounts for the number of participating relations and accumulates anomaly evidence across them. Since relation-specific errors share node embeddings and may be dependent, the χ2 distribution serves as an approximate reference. An alarm is raised when sv exceeds the decision threshold τ . Reconstruction errors for test nodes are obtained through batched masked inference. Subsequent calibration and fusion require only quantile-table lookups and χ2 CDF evaluations, adding little computational overhead. 3. EXPERIMENTS This section evaluates RECAL’s detection performance, component contributions, parameter sensitivity, and computational efficiency.
Fraction ofof nodes Fraction nodes
gn Benign
100 100
before calibration
Maliciou Malicious
after calibration
10−2 10−2 10−4 10−4 −2
10−2 10
−1
10−1 10 Raw error
Raw error
(a) Before calibration
0 100 0.0 10 0.0
0.5
0.5 score s Calibrated v Calibrated score sv
1.0 1.0
(b) After calibration
Fig. 3. Score distributions of benign and malicious nodes before and after relation-conditional calibration on CADETS. Baselines. We select representative baselines and State-of-the-Art (SOTA) methods, including Log2vec [21], THREATRACE [4], Unicorn [10], FLASH [5], MAGIC [6], STGAN [22], and AEGIS [23], to evaluate RECAL’s detection performance and false positive control. Experiments run on a single NVIDIA RTX 5090 GPU. Our results in tables 2 and 3 are means over five seeds (0–4); other experiments use seed 0. 3.2. Main Results Table 2 shows that RECAL achieves the highest F1 among the compared methods on CADETS, THEIA, and TRACE, reaching 99.99%, 99.93%, and 99.99%, respectively. Since recall is nearly saturated for several strong baselines, performance differences mainly lie in false positive control and the resulting precision gains. RECAL models different interaction patterns through per-relation reconstruction, then calibrates and aggregates anomaly evidence against each relation’s benign error distribution. This helps reduce the influence of differing error scales across relations on detection decisions and limit false alarms caused by normal behavior. RECAL achieves mean FPRs of 0.0004%, 0.0041%, and 0.0019% on CADETS, THEIA, and TRACE, respectively. Compared with STGAN, the baseline with the lowest reported FPRs in the table, RECAL achieves approximately 105×, 4×, and 41× reductions in mean FPR on CADETS, THEIA, and TRACE, respectively. Across all five seeds, 18 nodes on THEIA remain undetected, including 17 associated with /home/admin/profile. Further inspection reveals substantial overlap between their local interaction patterns and benign behavior. Although these nodes enter the candidate set, their calibrated and fused anomaly evidence remains below the final detection threshold. This highlights a limitation of deviationbased detection when the local behavior of attack-associated entities closely resembles normal activity. 3.3. Ablation Study
3.1. Experimental Setup Datasets. We evaluate on three scenarios of DARPA TC Engagement 3 (CADETS, THEIA, and TRACE) [20], whose scale and maliciousnode ratios are summarized in Table 1. Purely benign time spans are used for training, the tail of the benign span is held out as the calibration set (Sec. 2.4), and time spans containing attacks are used for testing. We adopt the node-level labels of THREATRACE [4] and follow its evaluation protocol, reporting Precision (Prec.), Recall (Rec.), F1 score (F1), and False Positive Rate (FPR). Inspired by MAGIC’s evaluation procedure [6], we perform a linear search over candidate thresholds for each dataset and select τ at the operating point with the highest test F1.
Table 3 compares five-seed mean performance across learning and detection configurations under the same training budget. The variant using uniform masking and a shared decoder no longer produces per-relation reconstruction errors, so detection falls back to KNN. This joint variant yields lower mean F1 on CADETS and TRACE and slightly higher mean F1 on THEIA, with higher mean FPRs than the full model across all three scenarios. With Relation-Balanced Learning (BL) retained, removing Relation-Calibrated Detection (CD) reduces mean F1 by 3.94, 0.50, and 0.48 percentage points on CADETS, THEIA, and TRACE, respectively, supporting this module’s role in both anomaly detection and false positive control. Fig. 3 provides a qualitative illustration on CADETS, where benign and malicious
Table 2. Performance Comparison. CADETS THEIA TRACE System Prec. ↑ Rec. ↑ F1 ↑ FPR ↓ Prec. ↑ Rec. ↑ F1 ↑ FPR ↓ Prec. ↑ Rec. ↑ F1 ↑ FPR ↓ Log2vec [21] 49.20% 84.59% 62.21% 1.59% 62.49% 66.05% 64.22% 0.29% 54.39% 78.27% 64.18% 1.83% THREATRACE [4] 90.42% 99.97% 94.96% 0.19% 87.04% 99.74% 92.96% 0.11% 71.56% 99.99% 83.42% 1.11% Unicorn [10] 31.00% 100.00% 47.00% – 67.00% 67.00% 67.00% – 28.00% 100.00% 34.00% – FLASH [5] 94.69% 99.99% 97.27% 0.10% 93.10% 99.83% 96.35% 0.05% 94.65% 99.99% 97.25% 0.16% MAGIC [6] 94.40% 99.77% 97.01% 0.22% 98.23% 99.99% 99.11% 0.14% 99.17% 99.98% 99.57% 0.09% STGAN [22] 98.40% 99.83% 99.11% 0.0426% 94.49% 99.83% 97.09% 0.0156% 99.50% 98.83% 99.16% 0.0777% AEGIS [23] 100% 99% 99% – 96% 97% 97% – 99% 99% 99% – RECAL (ours) 99.99% 99.99% 99.99% 0.0004% 99.95% 99.92% 99.93% 0.0041% 99.98% 100.00% 99.99% 0.0019% Arrows indicate whether higher (↑) or lower (↓) values are preferred; “–” denotes unreported values.
CADETS F1 (left)
10−2
0.98
10−4 1
10
Calib. ratio (%)
10−2
0.98
10−3 0.96
10
1.00
−1
10−2
0.98
10−4 2
5
R
10
20
(b) Inference masking rounds R
(a) Calibration-set ratio
10−1
1.00
10−2
0.98
10−3 0.96
10−4 0
0.5
γ
1
1.5
(c) Masking exponent γ
10−3 0.96
10−4 10−4
101
17 s
50 s 13 s
100 10−1
CADETS THEIA
TRACE
103
98.9%
102 101 100 10−1
95.8%
96.2%
0.6% 0.5% 1.7% 2.4% 1.8% 2.1%
CADETS THEIA
TRACE
(b) Inference time by stage
Fig. 5. Training and inference costs of RECAL (three-run average). Percentages are each stage’s share of total inference time.
10−1
1
102
284 s 217 s
769 s
stage-1 KNN
stage-2 calibration
FPR = 0
1.00
100
103
(a) Training and inference time
10−3 0.96
F1
TRACE
FPR (right) 10−1
1.00
F1
THEIA
TRACE F1 ↑ FPR ↓ 99.44 0.0513 99.51 0.0128 99.86 0.0311 99.98 0.0039 99.57 0.0714 99.99 0.0019
Wall clock (s)
THEIA F1 ↑ FPR ↓ 99.94 0.0043 99.44 0.0838 99.50 0.0741 99.48 0.0801 95.49 0.7652 99.93 0.0041
FPR (%)
CADETS F1 ↑ FPR ↓ 99.26 0.0569 96.05 0.3581 97.09 0.2261 99.08 0.0699 98.67 0.0958 99.99 0.0004
FPR (%)
Variant w/o BL w/o CD z-score calib. max fusion mean fusion RECAL
masked forward
inference
Inference (s)
train
Table 3. Results of Ablation Study (%).
10−3
Learning rate
10−2
(d) Learning rate
Fig. 4. Parameter sensitivity of RECAL. Solid lines show F1 (left axis) and dashed lines show FPR (right axis, log scale). node scores are more clearly separated after relation-conditional calibration. Replacing empirical quantile calibration with z-scores also reduces F1 across all three scenarios and raises the mean FPR on CADETS to 0.2261%, indicating that standardization based only on the mean and variance does not provide the same false positive control. Maximum and mean aggregation both lower F1 and increase FPR. The former retains only the strongest single-relation evidence, while the latter may dilute localized anomalies; Fisher’s method accumulates evidence across relations. 3.4. Parameter Sensitivity Fig. 4 shows F1 above 99.5% across all three scenarios with only 1% of the reserved calibration nodes. Using more benign calibration samples generally helps lower the FPR. Increasing masking rounds
preserves more visible neighborhood context, and performance improves as R rises from 1 to 5. Between 5 and 20 rounds, F1 varies by less than 0.01 percentage points per scenario, while forward-pass cost increases. For the masking exponent, γ = 0.5 yields the highest F1 and lowest FPR on CADETS and THEIA; larger values increase their FPR. F1 on TRACE varies by less than 0.02 percentage points. Across tested learning rates from 10−4 to 10−2 , F1 exceeds 99.7% in all scenarios, although FPR varies. The default 10−3 achieves the highest F1 and lowest FPR on CADETS and THEIA, while F1 on TRACE changes little. 3.5. Efficiency Fig. 5 reports the runtime of RECAL. On TRACE, the largest scenario by node count, training for 200 epochs takes about 217 seconds. Inference over approximately 3.3 million nodes across its five test graphs takes about 284 seconds. KNN reference index construction and queries account for 98.9% of this inference time and are the main computational cost. In contrast, the score calibration and Fisher fusion stage accounts for only approximately 0.5% to 2.1% of inference time across all three scenarios, indicating its low computational overhead. The reported inference times exclude preprocessing, data loading, and one-time calibration reference construction. 4. CONCLUSION We present RECAL for unsupervised provenance-based intrusion detection through relation-balanced masked graph learning and relationcalibrated anomaly detection. Across three DARPA E3 scenarios, RECAL achieves the highest F1 among the compared methods while maintaining low false positive rates. Future work will explore adaptive relation calibration to address concept drift in long-running deployments.
5. COMPLIANCE WITH ETHICAL STANDARDS This study uses publicly available DARPA E3 system-audit datasets and involves no human or animal subjects. No ethical approval was required. 6. REFERENCES [1] Michael Zipperle, Florian Gottwalt, Elizabeth Chang, and Tharam Dillon, “Provenance-based intrusion detection systems: A survey,” ACM Computing Surveys, vol. 55, no. 7, pp. 1–36, 2023. [2] Md Nahid Hossain, Sadegh M. Milajerdi, Junao Wang, Birhanu Eshete, Rigel Gjomemo, R. Sekar, Scott Stoller, and V. N. Venkatakrishnan, “SLEUTH: Real-time attack scenario reconstruction from COTS audit data,” in Proceedings of the USENIX Security Symposium (USENIX Security), 2017, pp. 487–504. [3] Sadegh M. Milajerdi, Rigel Gjomemo, Birhanu Eshete, R. Sekar, and V. N. Venkatakrishnan, “HOLMES: Real-time APT detection through correlation of suspicious information flows,” in Proceedings of the IEEE Symposium on Security and Privacy (S&P), 2019, pp. 1137–1152. [4] Su Wang, Zhiliang Wang, Tao Zhou, Hongbin Sun, Xia Yin, Dongqi Han, Han Zhang, Xingang Shi, and Jiahai Yang, “THREATRACE: Detecting and tracing host-based threats in node level through provenance graph learning,” IEEE Transactions on Information Forensics and Security, vol. 17, pp. 3972–3987, 2022. [5] Mati Ur Rehman, Hadi Ahmadi, and Wajih Ul Hassan, “Flash: A comprehensive approach to intrusion detection via provenance graph representation learning,” in Proceedings of the IEEE Symposium on Security and Privacy (S&P), 2024, pp. 3552–3570. [6] Zian Jia, Yun Xiong, Yuhong Nan, Yao Zhang, Jinjing Zhao, and Mi Wen, “MAGIC: Detecting advanced persistent threats via masked graph representation learning,” in Proceedings of the USENIX Security Symposium (USENIX Security), 2024, pp. 5197–5214. [7] Zijun Cheng, Qiujian Lv, Jinyuan Liang, Yan Wang, Degang Sun, Thomas Pasquier, and Xueyuan Han, “Kairos: Practical intrusion detection and investigation using whole-system provenance,” in Proceedings of the IEEE Symposium on Security and Privacy (S&P), 2024, pp. 3533–3551. [8] Fan Yang, Binyan Xu, Di Tang, and Kehuan Zhang, “Beyond nodes vs. edges: A multi-view fusion framework for provenance-based intrusion detection,” in Proceedings of the IEEE Symposium on Security and Privacy (S&P), 2026, pp. 3739–3758. [9] Hans-Peter Kriegel, Peer Kröger, Erich Schubert, and Arthur Zimek, “Interpreting and unifying outlier scores,” in Proceedings of the SIAM International Conference on Data Mining (SDM), 2011, pp. 13–24. [10] Xueyuan Han, Thomas Pasquier, Adam Bates, James Mickens, and Margo Seltzer, “UNICORN: Runtime provenance-based detector for advanced persistent threats,” in Proceedings of the Network and Distributed System Security Symposium (NDSS), 2020.
[11] Wajih Ul Hassan, Shengjian Guo, Ding Li, Zhengzhang Chen, Kangkook Jee, Zhichun Li, and Adam Bates, “NoDoze: Combatting threat alert fatigue with automated provenance triage,” in Proceedings of the Network and Distributed System Security Symposium (NDSS), 2019. [12] Zhang Xu, Zhenyu Wu, Zhichun Li, Kangkook Jee, Junghwan Rhee, Xusheng Xiao, Fengyuan Xu, Haining Wang, and Guofei Jiang, “High fidelity data reduction for big data security dependency analyses,” in Proceedings of the ACM SIGSAC Conference on Computer and Communications Security (CCS), 2016, pp. 504–516. [13] Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean, “Efficient estimation of word representations in vector space,” arXiv preprint arXiv:1301.3781, 2013. [14] Zhenyu Hou, Xiao Liu, Yukuo Cen, Yuxiao Dong, Hongxia Yang, Chunjie Wang, and Jie Tang, “GraphMAE: Selfsupervised masked graph autoencoders,” in Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), 2022, pp. 594–604. [15] Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio, “Graph attention networks,” in Proceedings of the International Conference on Learning Representations (ICLR), 2018. [16] Dan Busbridge, Dane Sherburn, Pietro Cavallo, and Nils Y. Hammerla, “Relational graph attention networks,” arXiv preprint arXiv:1904.05811, 2019. [17] Zheng Li, Yue Zhao, Xiyang Hu, Nicola Botta, Cezar Ionescu, and George H. Chen, “ECOD: Unsupervised outlier detection using empirical cumulative distribution functions,” IEEE Transactions on Knowledge and Data Engineering, vol. 35, no. 12, pp. 12181–12193, 2023. [18] Fabrizio Angiulli and Clara Pizzuti, “Fast outlier detection in high dimensional spaces,” in Proceedings of the European Conference on Principles of Data Mining and Knowledge Discovery (PKDD), 2002, pp. 15–27. [19] Ronald A. Fisher, Statistical Methods for Research Workers, Oliver and Boyd, 4th edition, 1932. [20] Angelos D. Keromytis, “Transparent computing engagement 3 data release,” https://github.com/darpa-i2o/Transparent-Compu ting/blob/master/README-E3.md, 2018. [21] Fucheng Liu, Yu Wen, Dongxue Zhang, Xihe Jiang, Xinyu Xing, and Dan Meng, “Log2vec: A heterogeneous graph embedding based approach for detecting cyber threats within enterprise,” in Proceedings of the ACM SIGSAC Conference on Computer and Communications Security (CCS), 2019, pp. 1777–1794. [22] Anyuan Sang, Xuezheng Fan, Li Yang, Yuchen Wang, Lu Zhou, Junbo Jia, and Huipeng Yang, “STGAN: Detecting host threats via fusion of spatial-temporal features in host provenance graphs,” in Proceedings of the ACM Web Conference (WWW), 2025, pp. 1046–1057. [23] Baihang Liu, Fangjiao Zhang, Yun Feng, Canhua Chen, Xiaoyu Wang, Yaqin Cao, and Qixu Liu, “AEGIS: Enhancing provenance-based intrusion detection system with LLMpowered deep semantic representation,” in Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026, pp. 14162–14166.