Conceptio › Archive › arXiv CS
arXiv CSopen access

Beyond Nodes vs. Edges: A Multi-View Fusion Framework for Provenance-Based Intrusion Detection

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

Beyond Nodes vs. Edges: A Multi-View Fusion Framework for Provenance-Based Intrusion Detection Fan Yang† , Binyan Xu† , Di Tang‡* , Kehuan Zhang†* † ‡ The Chinese University of Hong Kong Sun Yat-Sen University

arXiv:2604.14685v1 [cs.CR] 16 Apr 2026

{yf020, binyxu, khzhang}@ie.cuhk.edu.hk

Abstract—Provenance-based intrusion detection has emerged as a promising approach for analyzing complex attack behaviors through system-level provenance graphs. However, existing defense methods face an inherent granularity limitation. Node-centric detectors, which evaluate anomalies using entities’ attributes and local structural patterns, may misclassify benign behavioral changes or configuration modifications as suspicious. In contrast, edge-centric detectors, which focus more on interactions, may lack sufficient contextual awareness of the involved entities, leading to missed detections when compromised entities perform seemingly ordinary operations. These analytical biases highlight a persistent gap between node-centric and edge-centric analyses. To mitigate this gap, we present P ROV F USION, a multiview detection framework that integrates anomaly signals from three distinct views (i.e., attribute, structure, and causality). The framework fuses heterogeneous anomaly signals through lightweight fusion schemes and determines the final anomaly decisions through a voting-based integration process, providing a more consistent and context-aware assessment of system behavior. This design enables P ROV F USION to capture both entity-level deviations and interaction-level anomalies within a consistent analytic pipeline. Experiments on nine widely used benchmark datasets demonstrate that P ROV F USION achieves higher detection accuracy and lower false-positive rates than single node- and edge-centric baselines, maintaining stable performance across scenarios. Overall, the results suggest that our multi-view anomaly fusion together with voting-based decision aggregation offers a practical and effective direction for advancing provenance-based intrusion detection.

1. Introduction Intrusion Detection Systems (IDSs) [1], [2], [3], [4], [5] have long served as a cornerstone of system defense, particularly against sophisticated threats such as Advanced Persistent Threats (APTs) [6]. Among them, provenancebased intrusion detection systems (PIDSs) [7], [8] have received growing attention in recent years [9], [10], [11], [12]. These systems construct provenance graphs from system audit logs, where nodes represent entities such as processes * Corresponding authors

[email protected]

and files, and directed edges represent causal or informationflow interactions (e.g., read, write). By preserving the causal chains of system behavior, provenance graphs provide a structured and interpretable foundation for attack detection, forensic analysis, and damage assessment, thereby playing an increasingly indispensable role in modern IDSs. With the development of Graph Neural Networks (GNNs) [13], provenance-based detection has evolved from handcrafted or statistical rules toward learning-based analysis [14], [15], [16]. GNNs integrate node/edge attributes with local topology via message passing [17], [18], producing contextualized representations that support fine-grained local reasoning over dependencies within a L-hop neighborhood, where L is the number of GNN layers. These advances have improved PIDSs’ capacity to model local dependencies and to score entities/edges individually, supporting a move from coarse-grained, subgraph-level alerts to finer-grained attack localization [19], [20]. Despite these advances, existing fine-grained GNNbased PIDSs adopt a single detection view—either nodecentric [17], [18], [21], [22], [23] or edge-centric [24], [25], [26]. Node-centric methods assess whether an entity is anomalous based on attribute semantics and the structural context of its L-hop neighborhood; benign dependency changes (e.g., software or configuration updates) can shift representations away from the training distribution and inflate false positives. Edge-centric methods evaluate individual interactions based on aggregated representations of the involved entities; high-frequency edges may be regarded as normal even when the involved entities are compromised, leading to missed detections (e.g., a compromised process performing ordinary-looking actions that are classified as normal). Thus, each detection view is informative but inherently limited, providing a partial view of system behavior. The inherent limitations of node- and edge-centric detection views reflect a fundamental challenge: the behavioral information encoded in provenance graphs is only partially utilized. Therefore, an effective intrusion detection system should integrate these views to cross-reference entity- and interaction-level anomalies, enabling a more context-aware assessment of system execution. Achieving such integration, however, is non-trivial. Different detection views naturally produce heterogeneous anomaly scores, each emphasizing distinct behavioral fac-

© 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. Accepted for publication in the 2026 IEEE Symposium on Security and Privacy (IEEE S&P 2026).

tors and exhibiting different statistical scales. Naı̈ve aggregation (e.g., summation or averaging) can miscalibrate each view’s contribution, and a single averaging rule cannot adapt to different operating contexts, leading to unstable or biased detection outcomes across datasets and environments. To address the difficulty of integration, we develop P ROV F USION, a unified provenance-based intrusion detection framework that embodies multi-view fusion with adaptive, voting-based detection. It operates at both the score and decision stages. At the score stage, heterogeneous anomaly scores from different detection views are integrated through a multi-view fusion paradigm that evaluates each system entity across seven dimensions of anomaly, each intended to capture a distinct behavioral characteristic relevant to potential anomalies. At the decision stage, to ensure practicality and stability, we introduce a voting-based adaptive detection mechanism that makes the final anomaly decisions based on the vote result of multiple scoring dimensions. This two-stage design enhances the utilization of behavioral information while stably producing effective results. We evaluate P ROV F USION on six datasets from the DARPA Transparent Computing program (E3 [27] and E5 [28]), and three host subsets of the DARPA OpTC dataset [29], which are widely used in prior PIDS research [17], [18], [24], [25], [26]. On the TC datasets, P ROV F USION achieves higher detection quality—more true positives (TPs) with fewer false positives (FPs)—than SOTA node- and edge-centric baselines (on average, 24.8 TPs and 6.2 FPs vs. the second-best method’s 7.2 TPs and 136.5 FPs at the node level), and higher attack-level recall (detecting 13/14 attacks vs. 12/14 for the second best). On OpTC, P ROV F USION achieves the best detection–FP trade-off (avg 7 TPs / 13 FPs), while node-centric baselines incur 104 –105 FPs and edge-centric baselines collapse to < 3 TPs. In addition, an in-depth ablation study shows that P ROV F USION outperforms its single-view variants, further supporting the benefit of combining distinct detection views. Our main contributions are as follows: (1) Characterizing the granularity biases. We analyze the biases of node- and edge-centric detection view in GNN-based PIDSs, and provide an empirical and conceptual understanding of how these views capture distinct but incomplete aspects of provenance behavior. (2) A multi-view detection framework. We present P ROVF USION, a multi-view framework that unifies node- and edge-centric views within a single pipeline, combining three distinct detection views—attribute, structure, and causality—to capture diverse facets of system behavior. (3) A multi-dimensional fusion and voting-based detection mechanism. We design a two-level fusion-anddetection process that first integrates heterogeneous anomaly scores through monotone fusion schemes, and then achieves stable detection through voting-based integration. (4) Empirical evaluation and reproducibility. We implement P ROV F USION and evaluate it on nine DARPA TC and OpTC datasets. The results demonstrate that P ROVF USION achieves consistently high accuracy, lower false

positive rates, and comparable efficiency across multiple datasets, supporting its practical applicability. To foster reproducibility, we will release the source code: (https: //github.com/Joney-Yf/ProvFusion) along with a manually refined ground-truth label set based on official DARPA reports to support future research and benchmarking.

2. Motivation: Understanding How Different Detection Views Focus Differently Provenance-based intrusion detection systems (PIDSs) typically analyze system behavior from two complementary granularities: the state or structure of system entities (nodes) [17], [18], [23] and the interactions (edges) between them [24], [25]. Both detection views have demonstrated effectiveness, yet each exhibits different focus that can lead to false alarms or missed detections. These differences reflect the intrinsic emphasis of each detection view rather than particular implementations. To ensure representativeness, we reproduced several state-of-the-art systems under the standardized Velox benchmark [24]. Specifically, we included MAGIC [17], NodLink [23], and FLASH [18] as three node-centric methods, together with Velox [24], Kairos [26] and Orthrus [25] as three edge-centric baselines. Detailed experimental setting is in Section §5.1.2 and all our analyses are based on results in Table 2.

2.1. Node-Centric Perspective and Challenges Node-centric PIDSs evaluate whether a system entity (e.g., a process or file) deviates from normal patterns, based on its attributes and neighborhood structure, usually by leveraging Graph Neural Networks (GNNs) [13], [30], [31]. Our reproductions suggest two recurring difficulties that limit their robustness. 2.1.1. Sensitivity to Benign Novelty. MAGIC and NodLink frequently produce false positives when encountering benign but previously unseen behaviors (e.g., MAGIC and NodLink have over 90k false positives in THEIA-E3 dataset). Here, “novelty” refers to normal variations such as software updates, configuration changes, or new user activities. In our reproductions, entities introducing unseen attributes or structural patterns—such as a file with a new path—often received high anomaly scores despite being benign. This observation suggests that these node-centric models primarily focus on how unusual an entity and its structural patterns appear, which makes them sensitive to benign but previously unseen behaviors that are not necessarily malicious. 2.1.2. Objective–Security Misalignment in Type Prediction. A similar tendency appears in FLASH, which predicts the type of each node. While this task captures categorical semantics, it does not necessarily align with security relevance. For instance, a compromised process may still exhibit process-like characteristics and thus be assigned a low anomaly score. In our reproductions, several malicious processes exhibited anomalous interactions that received high anomalous scores under edge-centric methods such as Orthrus. Yet their node-level anomaly scores remained low,

Node Score (MAGIC)

1.0 0.8

Node anomalies

0.6 0.4

Edge anomalies

0.2 0.0 0.0

Attack A1 A2 A3

0.2

0.4 0.6 0.8 Edge Score (Orthrus)

1.0

Figure 1: Visualization of node- and edge-centric detection. No malicious nodes show high scores in both methods. suggesting that node-type prediction alone tends to evaluate local patterns rather than interaction abnormality, providing limited evidence of compromise. Summary: Overall, node-centric analysis primarily focuses on an entity’s local patterns or structural consistency. This focus allows it to capture statistical irregularities but limits its awareness of whether such irregularities correspond to truly suspicious interactions in context.

2.2. Edge-Centric Perspective and Challenges Edge-centric methods shift the analytical focus from entities to their interactions. The detection of anomalous interactions is often formulated as edge-type prediction [24], [25], [26]. Given the larger sample of edges and the more uniform distribution of edge types relative to node’s structural patterns, this view tends to yield more precise results. However, since it models interactions, it has limited awareness of the broader behavioral context in which the involved entities operate. On one hand, edge-centric detectors may yield false positives from rare-but-benign events, where legitimate yet infrequent operations are assigned high anomaly scores due to their rarity. On the other hand, they may incur false negatives from common-but-malicious events, as attackers often exploit high-frequency system calls to blend into normal activity. In our reproductions, some attack interactions are structurally anomalous and detected by node-centric methods (e.g., MAGIC). Yet under the edge-centric view, they received low anomaly scores because their event types are common in benign traces. These results suggest that assessing interaction-level plausibility alone may not fully capture the anomalous state of the participating entities. Summary: Edge-centric analysis focuses on evaluating the plausibility of interactions between entities. While this view enables stable learning of interaction patterns, it provides limited awareness of the broader behavioral context in which those entities participate.

2.3. Visualization and Design Implications We align the anomaly scores of different detection views on the same set of ground-truth entities to visualize how their focuses differ. We take MAGIC (node-centric) and Orthrus (edge-centric) as representative examples. As shown in Figure 1, most detected malicious points fall into exactly one high-score region. Some lie above the

red line but left of the blue line (flagged only by MAGIC); others lie to the right of the blue line but below the red line (flagged only by Orthrus). The upper-right quadrant—where both views would agree—is visibly empty. This pattern indicates that the two views capture different facets of abnormality and that either view alone provides incomplete attack coverage. Together, these findings motivate the multi-view framework introduced in Section 4, which integrates both views to form a more balanced and context-aware foundation for intrusion detection.

3. Threat Model Being consistent with prior work [17], [18], [24], [25], [26], we consider adversaries whose objective is to compromise a host system and maintain unauthorized access or control over time. Our analysis focuses on activities captured by standard kernel-level monitoring frameworks [32], [33], [34], and therefore excludes threats that operate outside this scope (e.g., hardware-level side channels). We assume that data collection and model training occur in a trusted environment, as in prior studies [24], [25], [26], thereby excluding data and model poisoning from our threat model. The Trusted Computing Base (TCB) includes the P ROV F USION software, the provenance capture mechanism, and the underlying operating system. These components are assumed to be protected from direct compromise through system-hardening techniques described in prior research [35], [36]. Finally, we assume that the integrity of the provenance record is ensured by tamper-evident or append-only logging mechanisms [37], [38], which make any modification attempts detectable. 4. P ROV F USION’s Framework This section introduces P ROV F USION, an anomalybased intrusion detection system. As established in Section §2, methods that rely on a single view—either focusing on node states (node-centric) or interactions (edgecentric)—exhibit distinct strengths and weaknesses. Our design goal is to leverage these distinct views to improve detection performance and reduce avoidable errors. Based on this goal, we design P ROV F USION: a multiview anomaly fusion and voting-based detection framework. To avoid conflating signals, we treat attributes and structure as separate views (details in §4.2). Specifically, P ROV F USION characterizes each system entity from three distinct views: (i) an attribute view that captures deviations in an entity’s intrinsic features, (ii) a structural view that identifies abnormalities in the entity’s role and structure patterns, and (iii) a causal view that evaluates the plausibility of interactions (i.e., edges) the entity participates in. We quantify each view independently, fuse them across seven anomaly dimensions, and finalize decisions via a voting-based mechanism, yielding stable performance across heterogeneous score scales. Figure 2 illustrates the overall architecture. The framework has three stages: (1) Provenance Graph Preparation (§4.1), (2) Multi-view Anomaly Characterization (§4.2), and (3) a Multi-Dimensional Anomaly Fusion and Voting-based Detection Framework (§4.3).

Stage 1. Provenance Graph Preparation

1

LOG

2

3

4

5

6

7

Features

b.dot

5

cmd.exe

6

2 7

4

x.exe c2.com:80 y.dll

3

process file netflow

one-hot/ word2vec 5

1

4

5

2 3 7 Type Graph

6

6

K-NN Density Estimator

𝑖𝑖=1

𝑆𝑆𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎 (𝑣𝑣)

B. Node-centric Structure Anomaly: “find abnormal entities in topology”

word.exe a.doc

1

Stage 3. Multi-Dimensional Anomaly Fusion & Detection

Stage 2. Multi-View Anomaly Characterization A. Node-centric Attribute Anomaly: “find abnormal paths” 2 𝑑𝑑𝑖𝑖 𝐾𝐾 4 1 new 1 � 𝑑𝑑𝑖𝑖 3 𝐾𝐾 5 Attribute Graph

1

5

1

4

2 3 7 Masked Type Graph

Predict

6

grad

GMAE

4 6 7

K-NN

Predicted Node Types

K-NN (like A)

𝑆𝑆struc (𝑣𝑣)

C. Edge-centric Causal Anomaly: “find unmatched entity-event pairs”

4

5

2 3 7 Attribute Graph

6

1

4

2 3 7 Attribute Graph

MLP

Predict grad

Trained GMAE (like B)

1 2 1 1

2 3 5 6

𝑆𝑆𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐 (𝑣𝑣)

Prediction Error

Percentile Ranked Score 𝑆𝑆𝑆𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎 𝑣𝑣 𝑆𝑆′struc 𝑣𝑣 𝑆𝑆′𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐 (𝑣𝑣)

LOG

Benign Data

Worst Case Threshold 𝜏𝜏𝑖𝑖

Multi-Dimensional Analysis Specialist: V1, V2, V3 ′ 𝑆𝑆𝑣𝑣𝑣𝑣𝑣𝑣𝑣𝑣 ≥ 𝜏𝜏𝑖𝑖 Pairwise Fusion: V4, V5 𝑓𝑓(Top-2 Scores) ≥ 𝜏𝜏𝑖𝑖 Holistic Fusion: V6, V7 𝑓𝑓(All-3 Scores) ≥ 𝜏𝜏𝑖𝑖 False

Benign

True

Malicious

Figure 2: The overall architecture of P ROV F USION. It takes raw system logs, prepares a provenance graph, characterizes node anomaly from three complementary views (Attribute, Structural, Causal), and finally fuses these scores to detect anomalies.

4.1. Provenance Graph Preparation We transform raw system logs into an attributed provenance graph that supports subsequent analysis. 4.1.1. Graph Construction. We parse audit logs to construct a directed multi-graph G = (V, E), where nodes represent system entities and edges represent causal or information-flow events among them (e.g., read). Each node is associated with a categorical type, a set of semantic attributes (e.g., command line), and a timestamp, while each edge is annotated only with its type and timestamp. Following standard practice [24], [25], we focus on three major types of entities and their corresponding interactions: process, file, and netflow. A complete list of the considered entity and event types is provided in Appendix A. 4.1.2. Graph Featurization. We adopt distinct featurization for nodes and edges to encode type and semantic properties. ①Node Features. For each node v , we construct a feature vector xv by concatenating a one-hot vector encoding its type (e.g., process) with a semantic attribute embedding derived from textual information such as file paths or command lines. Following prior work [18], [25], [39], the attribute embedding is obtained using Word2Vec  [40], [41]:  xv,attr = W ord2V ec(xtext ), xv = one-hot(type) ∥ xv,attr ②Edge Features. For each edge (u, v), the feature vector euv is represented as a multi-hot indicator over event types (e.g., read). This representation captures the fact that multiple event types may occur between the same pair of entities, preserving the diversity of interactions while substantially reducing edge count; we empirically validate that this aggregation does not degrade detection in Appendix O.

4.2. Multi-view Anomaly Characterization Rather than collapsing all information into a single representation and then obtaining a final anomaly score, we compute three per-view anomaly scores for each node v ∈ V : Sattr (v), Sstruc (v), and Scausal (v), then combine

the three scores through a simple voting-based fusion approach (Section §4.3). Each view is modeled and scored independently for two main reasons. (1) Disentanglement and contamination control. In our preliminary experiments, we observed that message passing [42] in GNNs can allow attribute-level novelty to spread through the graph, making normal structural patterns appear abnormal. This observation originally motivated separating the node-centric perspective into attribute and structural views (see ablation study in §5.3.4). Per-view scoring keeps each view calibrated to its own feature distribution and helps mitigate cross-signal interference. (2) Task heterogeneity and learning stability. The three views differ primarily in their analytical objectives rather than in the input features they use. While all views rely on the same provenance graph, each defines a distinct prediction task—attribute outlier detection, structural pattern assessment, and edge-type plausibility estimation. A unified optimization objective would force these heterogeneous tasks to share a single loss landscape, where the gradients associated with one objective could interfere with those of another. Such coupling risks obscuring the semantics of each task and destabilizing the learning process, ultimately producing a score whose meaning is difficult to interpret. Modeling and scoring each view separately avoids these conflicts and allows each view to focus on its respective objective; exploring more integrated optimization strategies is left for future work. 4.2.1. View 1: Node-centric Attribute Anomaly. Goal. The attribute view focuses on detecting anomalies based on an entity’s semantic attributes. Method. This view focuses purely on each entity’s semantic attributes, independent of its structural context. Given the attribute vector xv,attr obtained in Section §4.1, we measure how far a node’s semantics deviate from those observed in benign data. Specifically, we employ a density-based KNearest Neighbor (KNN) detector [43], which estimates the local density of a sample by computing its average distance

to the k nearest neighbors in the semantic feature space. A node is assigned a higher anomaly score Sattr (v) when its representation lies in a sparse region—indicating that it is semantically dissimilar to most other entities. Rationale. This design avoids assumptions about the underlying data distribution and can be efficiently accelerated on GPUs [44], [45], providing both efficiency and scalability for large provenance graphs. 4.2.2. View 2: Node-centric Structural Anomaly. Goal. The structural view focuses on detecting anomalies from the perspective of an entity’s role and structure pattern, independent of its semantic attributes. Method. To capture purely structural behavior, we construct a simplified type-only provenance graph Gtype , where each node is represented as a one-hot vector and each edge is represented as a multi-hot event-type vector defined in Section §4.1. This abstraction removes semantic noise while retaining the topological information needed to learn normal behavior patterns. We train a self-supervised Graph Masked Autoencoder (GMAE) [46] on Gtype to model typical structures patterns. Its encoder employs an edge-aware graph-attention mechanism [47] that integrates event-type information into message passing: cuv = attn-MLP(hu ∥ hv ∥ euv ) , (1) exp(cuv ) (l) , (2) αuv =P k∈Nv ∪{v} exp(cvk )   X (l) (l) . h(l+1) = σ αuv Wval h(l) (3) v u u∈Nv ∪{v}

where Nv represents neighbors of node v , and hlv represents the representation of node v at layer l. Including edge features allows the encoder to adjust attention weights according to event semantics, improving context awareness without additional supervision. During training, a subset of nodes is randomly masked, and the decoder reconstructs their original type vectors from the encoder outputs. The reconstruction objective is optimized using a cosine-similarity loss, which encourages the reconstructed vectors to align with the ground-truth representations. This self-supervised objective promotes learning of common structural relations among entity types. After training, the encoder output hstruc (v) serves as the structural embedding for node v . We then estimate its anomaly score Sstruc (v) using a KNN density estimator [43] in the embedding space. Nodes that are structurally dissimilar to most others—i.e., having large distances from their nearest neighbors—are assigned higher anomaly scores. Rationale. By focusing solely on structural patterns, which are captured through the connectivity between entity types and their associated event types, this view reduces the influence of attribute semantics on structural embeddings and helps prevent benign yet novel attribute variations from causing normal structures to be misidentified as abnormal. In this way, it provides a clear basis for detecting deviations from typical structural patterns.

4.2.3. View 3: Edge-centric Causal Anomaly. Goal. The causal view focuses on detecting anomalies in the plausibility of interactions between entities—that is, whether the edge linking two nodes aligns with normal causality. Method. To evaluate interaction plausibility, we formulate an edge-type prediction task on the full, attribute-rich provenance graph Gfull . We reuse the Graph Masked Autoencoder (GMAE) [46] framework from the structural view but train it on richer semantic features (xv ) with attributes to obtain semantically informed embeddings hsem (v). After self-supervised pretraining, a lightweight 2-layer MLP [48] decoder predicts the event types for each edge (u, v) using the concatenated embeddings of its involved nodes: ŷuv = MLP([hsem (u) ∥ hsem (v)]) . (4) Since an edge may correspond to multiple event types, the ground truth yuv is represented as a multi-hot vector, and the model is trained on benign data using a weighted binary cross-entropy (BCE) loss LBCE (ŷuv , yuv ). The weight for each eventptype i is computed from the benign training data as wi = (N − ni )/ni , where N is the total number of edges and ni is the count of class i, to mitigate class imbalance during training. During inference, the prediction loss of each edge serves as its anomaly score, and a node’s causal anomaly score Scausal (v) is defined as the maximum loss among its adjacent edges, highlighting entities involved in implausible interactions. Edges (and the associated nodes) that yield higher prediction losses are therefore considered more anomalous. Rationale. This view models how entities interact through events and evaluates whether such interactions are behaviorally plausible. By learning interaction semantics through an event-type prediction objective on benign data, it captures behavioral regularities distinct from those characterized in the attribute and structural views. This separation thus allows the framework to identify abnormal causal flows—such as unexpected information transfers or execution chains—by detecting deviations from these learned patterns.

4.3. A Multi-Dimensional Anomaly Fusion and Voting-based Detection Framework Goal. Given three view-specific anomaly scores per node—attribute (Sattr ), structure (Sstruc ), and causality (Scausal )—our objective is to integrate them into a unified, transparent decision framework. The design should (i) normalize heterogeneous score scales, (ii) combine anomaly scores from multiple views under a principled rule space, and (iii) provide transparency for post-hoc diagnosis. All thresholds are derived solely from benign validation data without reliance on test data [49], [50]. Method. Our approach consists of three steps: ① percentile normalization, ② fusion through seven anomaly dimensions, and ③ ensemble voting. ①Step 1: Percentile Normalization. For each view ∈ {attr, struc, causal}, we map the raw score of node v to its empirical quantile rank within the benign validation distribution:   benign ′ Sview (v) = Quantile Sview (v) | Sview ∈ [0, 1].

This converts heterogeneous scores into a comparable, scale′ free space, where a higher Sview (v) consistently reflects a stronger relative anomaly signal. ②Step 2: Seven Anomaly Dimensions. We define seven binary detectors {Di }7i=1 , each corresponding to a specific linear combination pattern of the three normalized scores. A detector flags node v when its combined score is not smaller than a threshold τi that is set as the maximum benign score for that dimension. The detectors are organized by the number of contributing views. Group 1 - Single-view Specialists (D1–D3): capture strong ′ anomalies confined to one view, where D1 = [Sattr (v) ≥ τ1 ], ′ ′ D2 = [Sstruc (v) ≥ τ2 ], and D3 = [Scausal (v) ≥ τ3 ]. Group 2 – Pairwise Corroboration (D4–D5): capture cases where two views jointly indicate abnormality. Let ′ ′ Smax 1 (v) and Smax 2 (v) denote the two largest values ′ ′ ′ among {Sattr , Sstruc , Scausal }. ′ ′ D4 : [Smax 1 (v) + Smax 2 (v) ≥ τ4 ].

For a weighted variant, we compute a dynamic fusion with a weighting coefficient α, which is set to 5 by default. A hyperparameter study of α is reported in Section §5.5. ′ X eαSi (v) , Sfused = wi Si′ (v), wi = P αSj′ (v) e j∈Top-2 i∈Top-2 D5 : [Sfused ≥ τ5 ].

Group 3 – All-View Aggregation (D6–D7): represent joint evidence from all three views: ′ ′ ′ D6 : [Sattr (v) + Sstruc (v) + Scausal (v) ≥ τ6 ], and its weighted counterpart: ′ eαSi (v) wi = P , αSj′ (v) j∈{attr,struc,causal} e

Sfused =

X

wi Si′ (v),

i

D7 : [Sfused ≥ τ7 ].

③Step 3: Ensemble Voting. Each detector provides a binary decision, and their collective votes determine the final outcome: 7 X V (v) = [Di (v) flags], i=1

is malicious(v) = True ⇐⇒ V (v) ≥ Tv .

We use Tv = 4 by default, corresponding to a simple majority among the seven detectors. A sensitivity analysis on Tv is reported in the evaluation (see Section §5.3.3). Analyst-facing output. For each alerted node, we report both ′ ′ ′ the normalized triplet (Sattr , Sstruc , Scausal ) and the detector outcomes [D1 , . . . , D7 ] with V (v), so that analysts can trace which evidence sources and thresholds jointly contributed to the decision. Rationale. To effectively integrate the three heterogeneous anomaly scores, we first normalize them. We adopt percentile normalization as it provides a scale-free and distribution-agnostic basis for integration. Section §5.3.1 confirms that this choice is critical for a stable, balanced fusion.

Building on this foundation, we adopt a linear-fusion paradigm because it naturally preserves the monotonic relationship among anomaly scores—where higher viewspecific scores consistently indicate stronger anomaly evidence—and helps keep the contribution of each view directly transparent. In contrast, density- or probability-based integration emphasizes distributional deviations, which may violate this monotonic property and assign high anomaly likelihoods to samples with moderate scores simply because they lie in sparse regions. This theoretical rationale guided our design choice, and subsequent evaluation (Section §5.3.2) was consistent with our analysis, showing that linear fusion remained stable and effective across datasets, whereas density- or probability-based alternatives tended to exhibit the expected sensitivity to distributional variations. Building on this principle, we adopt linear fusion to combine the scores from the three views. Conceptually, any linear-fusion strategy can be grouped into three categories according to the number of participating views—using one, two, or all three—which correspond to Groups 1–3 in our design. Single-view detectors (Group 1) apply a direct threshold on individual scores, whereas the multiview detectors in Groups 2 and 3 are implemented in both unweighted and adaptive weighted forms. Accordingly, we design seven detectors that form a small family of representative linear-combination patterns among the three normalized scores, capturing single-view, pairwise, and allview anomaly evidence. All detectors follow a monotonic reasoning principle: higher view-specific scores indicate stronger anomaly evidence. For instance, Group 2 fuses the two highest-scoring views, and in its adaptive variants larger weights are assigned to stronger signals, reflecting this monotonic rule. All detectors remain lightweight; the weighted forms use only a single weighting coefficient α. Together, these outputs clarify how multi-view evidence is combined and calibrated, offering decision-level transparency that helps analysts trace the origin of each alert and perform post-hoc diagnosis or auditing when needed.

5. Evaluation In this section, we evaluate P ROV F USION in comparison with recent state-of-the-art (SOTA) provenance-based intrusion detection systems to assess its detection performance and efficiency. The evaluation provides a systematic and transparent analysis of how well P ROV F USION detects attacks and operates under realistic system settings. We first introduce the research questions that guide our study and then present the experimental setup, including the datasets, baselines, and evaluation metrics. We investigate the following four questions: RQ1: Effectiveness. How effective is P ROV F USION in detecting a broad range of attacks while maintaining comparatively low false positives relative to recent SOTA systems? RQ2: Component analysis. How do the key design components of P ROV F USION—particularly the multi-view analysis, and the multi-dimensional anomaly fusion and detection framework—shape its overall detection behavior? RQ3: Efficiency. Is P ROV F USION computationally efficient and scalable?

RQ4: Hyperparameter study. How do key hyperparameters—such as the weighting coefficient α, learning rate—affect the overall performance of P ROV F USION? Implementation. P ROV F USION is implemented in Python. We use Word2Vec [40] for attribute embedding and adopt a Graph Masked Autoencoder (GMAE) [46] with edge-aware graph attention [47] as the graph learning backbone. Source code and configurations will be released upon publication to facilitate reproducibility. All experiments were conducted with fixed random seeds and consistent data splits following prior work [24], [25]. Key hyperparameters were selected based on validation performance for each dataset, following the same tuning protocol for all compared methods. A detailed hyperparameter study is presented in Section §5.5. All experiments were performed on a server running Ubuntu 22.04, equipped with a 128-core AMD EPYC 7543 processor, 512 GB of RAM, and an NVIDIA A100 GPU.

5.1. Experimental Setup 5.1.1. Datasets. Following recent SOTA works [24], [25], we use large-scale datasets from the DARPA Transparent Computing (TC) program [27], which serve as standard benchmarks for PIDSs. Other public datasets are either substantially smaller in scale [11], [51], [52] or not publicly accessible at the time of writing [10], [53]. The DARPA TC datasets originate from red-team engagements that simulated realistic, multi-stage attacks (e.g., intrusions on SSH, email, and web servers) amid a high volume of benign user activities. Data were collected across three operating systems—FreeBSD [54], Linux [55], and Android [56]—using distinct capture agents: CADETS, THEIA, and CLEARSCOPE (abbreviated as CLRSCP in tables when space is limited). To assess generalization beyond the TC family, we evaluate on the DARPA Operationally Transparent Cyber (OpTC) [29] dataset—an independent program collected via Windows ETW at enterprise scale (500 hosts, 10 days). Following prior work [24], we use the three targeted hosts (H201, H501, H051) as evaluation subsets. A salient characteristic of these datasets is the severe class imbalance, where malicious events are vastly outnumbered by benign ones. Table 9 in Appendix B provides detailed statistics. Data Splitting. We follow the splitting protocols adopted in prior work [24], [25] to enable comparable evaluation, using consistent train/validation/test partitions as specified by the corresponding benchmarks [24]. The details of our data splitting are provided in Appendix C. 5.1.2. Compared Baselines. We compare P ROV F USION with six representative and publicly reproducible PIDSs: Kairos [26], NodLink [23], MAGIC [17], FLASH [18], Orthrus [25], and Velox [24]. We exclude SLOT [57] and R-CAID [21] because their official implementations were unavailable during this study; moreover, the reported performance of SLOT on the DARPA datasets is comparable to MAGIC and FLASH, so its exclusion is unlikely to affect conclusions. For each baseline, we started from the configurations recommended in the corresponding papers or repositories and further performed validation-based tuning under

the same protocol as P ROV F USION. Other SOTA PIDSs are subgraph-level [58] or rule-based [59], and are therefore excluded. We report the best-performing results observed across both sources to ensure fairness and reproducibility. We additionally include a score-level ensemble baseline, Magic+Velox (M+V in tables), which linearly fuses the pernode anomaly scores of Magic and Velox to test whether P ROV F USION’s gains can be obtained by trivially combining existing detectors. 5.1.3. Evaluation Metrics. Given the extreme data imbalance and the operational constraints of a practical IDS, standard metrics like AUC [60] and accuracy can be misleading (e.g., a system with 99.99% accuracy may still produce an impractically high number of false alarms). We therefore report the Matthews Correlation Coefficient (MCC) as a more informative thresholded metric for imbalanced binary classification. Because any thresholded metric ultimately depends on the chosen anomaly threshold, we further complement MCC with a threshold-independent, ranking-oriented measure—Area under the Detection–Precision curve (ADP)—to assess how well a method prioritizes malicious entities above benign ones under varying precision levels. Matthews Correlation Coefficient (MCC) [61]. MCC summarizes binary classification quality and remains robust under severe imbalance by accounting for all four entries of the confusion matrix: TP × TN − FP × FN MCC = p (T P + F P )(T P + F N )(T N + F P )(T N + F N )

An MCC of +1 indicates perfect prediction, 0 random prediction, and −1 inverse prediction. Area under the Detection–Precision Curve (ADP). Proposed by Velox [24], ADP evaluates how effectively a system detects attacks across a spectrum of precision levels, mitigating reliance on any single threshold. Let D(p) denote the fraction of detected attacks at precision p. The Detection–Precision curve plots D(p) against p, and Z 1 |{Ai | Ai ∩ R(p) ̸= ∅}| ADP = D(p) dp, D(p) = , k 0 where k is the number of attack campaigns, Ai is the set of malicious nodes for attack i, and R(p) is the set of nodes flagged as anomalous when the decision threshold is swept to achieve precision p. Ranking for P ROV F USION. Since P ROV F USION produces seven detector decisions from three normalized viewspecific scores, we derive a unified ranking by P first ordering 7 nodes according to their vote count V (v) = i=1 Di (v), with ties resolved using the maximum score among the three views. This process yields a deterministic total order for ADP computation.

5.2. Overall Detection Performance (RQ1) To answer our first research question, we evaluate P ROVF USION’s detection performance under the refined ground truth introduced below, focusing on two aspects: (1) its ability to detect every attack campaign (attack coverage), and (2) its precision in identifying individual malicious entities (node-level performance).

5.2.1. Refined Ground Truth Construction. While Orthrus [25] provides the most comprehensive ground truth available for the DARPA TC datasets, our cross-check with the official DARPA engagement reports [62] and the REAPr dataset [63] revealed that several documented campaigns (e.g., Phishing Email w/ Executable Attachment in THEIAE3) were not represented in the released labels. This observation indicates that even strong benchmarks still leave room for refinement. The refinement was conducted before any model training and is independent of our method design. Guided by the DARPA reports and REAPr, we manually examined the corresponding audit logs to identify and verify missing entities using multiple evidence sources—timestamp alignment, filename or command-line matching, file-path consistency, and causal dependencies—with cross-review among annotators. This process required extensive manual inspection and verification effort, and it only supplements verifiable malicious entities without altering any original Orthrus labels. In total, 21 additional malicious nodes were confirmed across three datasets—THEIA-E3 (+11), CLEARSCOPEE3 (+8), and CLEARSCOPE-E5 (+2). We use this refined ground truth as the primary reference for evaluation, as it more comprehensively reflects all attack activities documented in the official engagements. For completeness, results under the original Orthrus labels are also provided in Appendix M. The newly added malicious nodes (including UUID, attributes, and timestamps) are listed in Appendix J for transparency and will be publicly released to support community benchmarking and future research on PIDSs. 5.2.2. Attack Campaign Coverage. An effective detection system should detect every attack campaign without omission. We consider an attack “detected” if at least one node directly involved in the campaign is flagged as malicious. Table 1 summarizes the results. P ROV F USION successfully detects all attacks across nine datasets except for a single campaign in CLEARSCOPE-E5, achieving near-complete coverage across diverse environments. This result highlights P ROV F USION’s ability to capture diverse attack campaigns. Analysis of the missed campaign. The missed case, “Lockwatch APK Java APT,” was exceptionally stealthy. Inspection of its anomaly scores shows that the involved nodes were already ranked near the decision boundary. For instance, one node yielded (Sattr = 0.401, Sstruc = 0.994, Scausal = 0.390), suggesting that the attack effectively mimicked benign attributes and causal flows while exhibiting a distinct structural deviation. Another node in the same campaign had (Sattr = 0.983, Sstruc = 0.984, Scausal = 0.988), with its aggregated score slightly below the detection threshold. These results indicate that P ROV F USION was responsive to this campaign’s abnormal structure, and a slightly relaxed threshold could have captured it. This single miss reflects a conservative thresholding choice made to reduce FPs. 5.2.3. Node-level Performance on the Refined Ground Truth. We compare P ROV F USION with six provenancebased intrusion detection systems across all six TC datasets. As summarized in Table 2, P ROV F USION consistently achieves higher detection performance than all baselines,

CADETS THEIA CLEARSCOPE OpTC E3 E5 E3 E5 E3 E5 PIDS↓ A1A2A3 A1A2 A1A2A3 A1 A1 A1A2A3A4 A1A2A3 Kairos ✗ ✓ ✗ ✗ ✗ ✓ ✓ ✗ ✗ ✓ ✓ ✗ ✗ ✗ ✓ ✓ ✓ Magic ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✗ ✓ ✓ ✓ ✗ ✓ ✓ ✓ ✓ ✓ NodLink ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✗ ✓ ✓ ✓ ✗ ✓ ✗ ✓ ✓ ✓ Flash ✗ ✓ ✓ ✓ ✓ ✓ ✓ ✗ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ Orthrus ✗ ✓ ✓ ✗ ✓ ✓ ✗ ✓ ✓ ✓ ✓ ✗ ✗ ✗ ✓ ✓ ✓ Velox ✓ ✓ ✓ ✗ ✗ ✓ ✓ ✓ ✓ ✓ ✓ ✗ ✗ ✓ ✓ ✓ ✓ M+V ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✗ ✗ ✓ ✓ ✓ ✓ ProvFusion ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✗ ✓ ✓ ✓ ✓

Dataset→

TABLE 1: Attack coverage of P ROV F USION and state-ofthe-art baselines. ✓ denotes a detected campaign (at least one malicious node flagged), while ✗ denotes a miss. P ROVF USION successfully detects 16 out of 17 campaigns. We provide the details of each attack in Appendix D. identifying more true positives while generating fewer false positives—24.8 TPs and 6.2 FPs on average, compared with 7.2 TPs and 136.5 FPs for the next-best system (i.e., Velox, the strongest baseline by ADP and also competitive in TP/FP). This result demonstrates P ROV F USION’s high precision in identifying individual malicious entities. To test whether trivial detector combinations suffice, Table 2 also reports a score-level fusion of MAGIC and Velox; it inherits MAGIC’s FP cost without matching P ROVF USION’s TPs, confirming the value of multi-view fusion over naive score combination. Many of the newly verified malicious nodes in the refined ground truth were successfully detected not only by P ROV F USION but also by some strong baselines (e.g., Velox or Orthrus), which further supports the validity of these entities and the quality of the refined annotations. We further verified that the same performance trends hold under the original Orthrus ground truth (Appendix M), with only minor differences in absolute values due to the previously missing attacks. A qualitative example in Appendix H illustrates how P ROV F USION’s alerts align with the recorded attack sequences, helping analysts reconstruct intrusion chains and better interpret attack progression. Generalization to OpTC. Table 3 shows that P ROV F U SION ’s advantage transfers to the DARPA OpTC dataset: node-centric baselines accumulate 104 –105 FPs while edgecentric baselines collapse to < 3 TPs, whereas P ROV F U SION sustains a usable detection profile (average of 7 TPs / 13 FPs across the three hosts), confirming that multiview fusion remains effective under different logging infrastructure and attack distributions. Furthermore, a benignshift stability study (Appendix G) shows that rotating the validation day has minimal impact on detection. FP/FN Analysis. To illustrate how multi-view fusion resolves node/edge-specific FP/FN patterns, we compare P ROV F USION with MAGIC and Velox, two SOTA node- and edge-centric baselines. The two baselines exhibit complementary false negatives: MAGIC misses /tmp/mozilla_admin0/jAG_iSHt.bin.part, whose subgraph resembles common temporary-file activity, whereas Velox detects it from its anomalous edge context. Conversely, Velox misses the network endpoint

System Dataset TP↑ FP↓ TN↑ FN↓ F1↑ ADP↑ MCC↑ Dataset TP↑ FP↓ TN↑ FN↓ F1↑ ADP↑ MCC↑ Kairos 1 959 280.6K 67 0.00 0.00 0.00 0 6 3.14M 123 0 0.01 0 Magic 22 16.5K 265.0K 46 0.00 0.01 0.02 28 245.3K 2.89M 95 0.00 0.17 0.00 NodLink CADETS 18 34.3K 247.2K 50 0.00 0.49 0.01 CADETS 73 755.9K 2.38M 50 0.00 0.03 0.01 Flash 3 4.5K 277.0K 65 0.00 0.04 0.01 6 34.9K 3.10M 117 0.00 0.02 0.01 E3 E5 Orthrus 7 1 281.5K 61 0.18 0.81 0.30 1 8 3.14M 122 0.00 0.34 0.03 Velox 9 1 281.5K 59 0.23 0.97 0.35 0 2 3.14M 123 0.00 0.01 0.00 M+V 17 173 281.3K 51 0.13 0.92 0.15 7 810 3.14M 116 0.02 0.09 0.02 ProvFusion 24 1 281.5K 44 0.52 1.00 0.58 7 9 3.14M 116 0.10 0.92 0.16 Kairos 2 21 701.5K 127 0.03 0.15 0.04 0 7 1.86M 69 0.00 0.00 0.00 Magic 20 97.3K 604.2K 109 0.00 0.00 0.00 61 737.3K 1.12M 8 0.00 0.00 0.00 NodLink THEIA 26 258.3K 443.2K 103 0.00 0.01 0.00 THEIA 20 175.3K 1.68M 49 0.00 0.00 0.00 Flash 21 251.2K 450.3K 108 0.00 0.01 0.00 41 316.2K 1.54M 28 0.00 0.01 0.01 E3 E5 Orthrus 4 0 701.5K 125 0.06 0.67 0.18 1 31 1.86M 68 0.02 0.30 0.02 Velox 21 117 701.4K 108 0.16 1.00 0.16 2 63 1.86M 67 0.03 0.33 0.03 M+V 16 44 701.4K 113 0.17 0.49 0.18 2 306 1.86M 67 0.01 0.1 0.01 ProvFusion 91 2 701.5K 38 0.82 1.00 0.83 11 2 1.86M 58 0.27 1.00 0.37 Kairos 14 8.5K 102.9K 35 0.00 0.01 0.02 1 1 150.9K 52 0.04 0.33 0.10 Magic 42 8.4K 102.9K 7 0.01 0.02 0.06 7 8.3K 142.7K 46 0.00 0.00 0.01 NodLink CLRSCP 39 22.8K 88.5K 10 0.00 0.06 0.03 CLRSCP 3 27.0K 123.9K 50 0.00 0.00 0.00 Flash 36 11.1K 100.2K 13 0.01 0.01 0.04 0 0 150.9K 53 0.00 0.00 0.00 E3 E5 Orthrus 2 5 111.3K 47 0.07 0.40 0.11 2 8 150.9K 51 0.06 0.17 0.09 Velox 3 623 110.7K 46 0.01 0.50 0.02 8 13 150.9K 45 0.22 0.45 0.24 M+V 1 150 112.2K 48 0.01 0.16 0.01 6 58 150.9K 47 0.10 0.17 0.10 ProvFusion 6 7 111.3K 43 0.19 0.86 0.24 10 16 150.9K 43 0.25 0.77 0.27

TABLE 2: Node-level detection performance on the refined ground truth. P ROV F USION significantly outperforms all baselines in terms of F-1, ADP, and MCC, while maintaining a much lower number of false positives (FP). Best results are in bold. 128.55.12.110:44354→8:0, whose surrounding interactions appear syntactically ordinary, whereas MAGIC flags it due to its atypical structural footprint. P ROV F USION recovers both through cross-view voting. False positives, however, do not accumulate under fusion: a node that appears anomalous in one view is often supported as benign by the others, causing its combined evidence to remain below the voting threshold. True attack nodes typically lack such conflicting benign evidence; even when no single view is decisive, all views assign at least moderate suspicion, allowing consensus to emerge. This difference explains why P ROV F USION improves recall while suppressing false positives, rather than trading one for the other. RQ1 Answer: Across six datasets, P ROV F USION achieves comprehensive attack coverage and strong node-level precision while maintaining a low falsepositive rate. The refined ground truth—constructed independently from method design and validated through official DARPA records—provides a more complete benchmark for future research. Overall, the results confirm that P ROV F USION effectively detects diverse and stealthy attacks in provenance graphs under realistic conditions.

5.3. In-depth Analysis (RQ2) Having established P ROV F USION’s overall detection performance in Section §5.2, we now conduct ablation studies to understand the roles of its core design components and the mechanisms driving its effectiveness. This section addresses our RQ2: How do P ROV F USION’s core design choices contribute to its detection capability?

Dataset Method TP↑ FP↓ F1↑ ADP↑ MCC↑ Kairos 1 3 0.02 0.50 0.05 MAGIC 104 106K 0.00 0.20 0.06 NodLink 36 96.4K 0.00 0.06 0.02 FLASH 27 513K 0.00 0.14 0.00 H051 Orthrus 1 8 0.02 0.33 0.03 Velox 2 4 0.03 0.43 0.08 M+V 5 254 0.03 0.05 0.03 P ROV F USION 16 30 0.20 0.94 0.22 Kairos 2 3 0.00 1.00 0.02 MAGIC 2.48K 594K 0.01 0.33 0.65 NodLink 1.40K 77.7K 0.03 0.07 0.76 FLASH 751 168K 0.01 0.45 0.21 H201 Orthrus 1 7 0.0 0.50 0.01 Velox 1 9 0.0 0.33 0.01 M+V 10 1.17K 0.01 0.89 0.10 P ROV F USION 3 5 0.00 1.000 0.02 Kairos 1 5 0.00 0.20 0.01 MAGIC 433 510K 0.00 1.00 0.14 NodLink 365 102K 0.07 0.14 0.04 FLASH 31 26.6K 0.00 0.01 0.01 H501 Orthrus 1 4 0.00 0.25 0.02 Velox 1 7 0.00 0.33 0.01 M+V 3 7.93K 0.00 0.03 0.00 P ROV F USION 2 4 0.01 1.00 0.03

TABLE 3: Detection results on the DARPA OpTC dataset. 5.3.1. Validation of Normalization Strategy. We first validate our a priori choice of percentile normalization (discussed in §4.3). We compared its performance within the full P ROV F USION framework against three common alternatives: min-max, z-score [64], and robust scaling [65]. The detailed results are in Appendix F. The evaluation confirms our choice rationale: percentile normalization consistently outperforms all alternatives. Min-max proved highly sensitive to outliers in the validation data, incorrectly compressing the score range of some views (e.g., scaling a 98th-percentile anomaly to only

CADETS-E3 THEIA-E3

CLEARSCOPE-E3 CADETS-E5

1.0

THEIA-E5 CLEARSCOPE-E5

0.8

F1

FP

6000 4000

CADETS-E3 THEIA-E3

CLEARSCOPE-E3 CADETS-E5

0.6 0.4

2000

0.2

0

0.0

1

2

3

4

Threshold

5

6

7

1.0

THEIA-E5 CLEARSCOPE-E5

0.8

MCC

8000

CADETS-E3 THEIA-E3

CLEARSCOPE-E3 CADETS-E5

THEIA-E5 CLEARSCOPE-E5

0.6 0.4 0.2

1

2

3

4

Threshold

5

6

7

0.0

1

2

3

4

Threshold

5

6

7

(a) False Positive Count (b) F1 Score (c) MCC Figure 3: Impact of the voting threshold (Tv ) on performance. A threshold of Tv = 4 provides the optimal balance, maximizing the (b) F1 and (c) MCC scores while dramatically reducing the (a) False Positives seen at lower thresholds. MCC ↑ ADP ↑ Detect E3 E5 Avg E3 E5 Avg Individual Model Families (Best over 4 normalizations) CADETS .571 .100 .336 1.00 .917 .958 5/5 AE CLRSCP .117 .291 .204 1.00 .553 .776 3/5 THEIA .735 .295 .515 1.00 1.00 1.00 4/4 CADETS .396 .136 .266 .000 .000 .000 5/5 GMM CLRSCP .150 .284 .217 .000 .007 .004 4/5 THEIA .475 .273 .374 .000 .000 .000 3/4 CADETS .243 .101 .172 1.00 .917 .958 5/5 Isolation CLRSCP .076 .189 .132 .184 .310 .247 4/5 Forest THEIA .215 .000 .108 .670 .005 .338 2/4 CADETS .523 .114 .319 1.00 .750 .875 4/5 KNN CLRSCP .116 .330 .223 .033 .646 .339 4/5 THEIA .481 .295 .388 1.00 1.00 1.00 4/4 CADETS .184 .136 .160 .720 .950 .835 5/5 MVG CLRSCP .157 .284 .220 .333 .612 .473 4/5 THEIA .496 .273 .384 1.00 1.00 1.00 4/4 CADETS .582 .116 .349 .000 .015 .008 5/5 OCSVM CLRSCP .247 .320 .284 .000 .073 .036 4/5 THEIA .319 .367 .343 .000 .001 .000 4/4 CADETS .594 .122 .358 1.00 1.00 1.00 5/5 Linear CLRSCP .303 .261 .282 1.00 1.00 1.00 5/5 THEIA .845 .289 .567 1.00 1.00 1.00 4/4 CADETS .582 .158 .370 1.00 .925 .962 5/5 Voting CLRSCP .238 .269 .254 .857 .774 .815 4/5 (Ours) THEIA .831 .367 .599 1.00 1.00 1.00 4/4 Method

Dataset

TABLE 4: Best performance of different fusion/detection paradigms vs. our voting mechanism. Bold: best per dataset; underline: second best. All results use the best normalization found in hyperparameter search. 0.2). This silenced their contribution, leading to missed detections (8/14). Z-score and Robust scaling failed on the heterogeneous, non-Gaussian distributions, as their underlying extreme numerical expansion (e.g., normalized scores > 1000) led to different instabilities. For Z-Score, this expansion forced an inflated detection threshold, which silenced true attack signals and caused a catastrophic loss of detection (5/14). Robust scaling was the most erratic: its expansion mechanism resulted in a dual failure, allowing some views to improperly dominate the fusion (144 FPs) while simultaneously silencing others (1/14 detections). 5.3.2. Ablation: Evaluating the Rationale of the MultiDimensional Anomaly Fusion and Detection Framework. As described in §4.3, our fusion and detection framework integrates multi-view anomaly evidence through two stages: (i) aggregating view-specific scores using monotonic linear detectors, and (ii) combining their outputs through a vot-

ing rule. This ablation examines whether such a structured score–decision integration contributes to stable and reliable detection across datasets. We replace the proposed framework with representative alternatives from major anomaly detection paradigms and compare their performance under the same experimental setting. Comparative setup. We evaluate five model families: (1) linear combination models (weighted summations of the three scores under different weighting schemes); (2) probabilistic models (Gaussian Mixture Models [66] and multivariate Gaussian density estimation [67]); (3) distance/density-based models (K-Nearest Neighbor [43]); (4) boundary-based models (One-Class SVM [68] and Isolation Forest [69]); and (5) reconstruction-based deep models (a lightweight autoencoder using reconstruction error [70]). Each family was explored over a broad parameter space (including all four normalization methods from §5.3.1) and we report its best observed result per dataset (Table 4). For instance, the linear combination models enumerated all integer weight triplets (wattr , wstruc , wcausal ) ∈ {0, . . . , 19}3 \ {(0, 0, 0)}—a total of 7,999 combinations. Other families varied core hyperparameters such as neighborhood sizes and learning rates (see details in Appendix N). All models used identical normalized inputs and data partitions for fairness. Finding 1: linear fusion best supports the monotonicity principle. Across all datasets, the linear combination models achieved the most stable and overall highest performance, whereas other paradigms performed well only on certain datasets. Methods relying on statistical rarity often assigned high anomaly scores to rare but benign nodes while overlooking nodes whose three scores were all high. For example, under Isolation Forest with Percentile normalization, a node with a skewed score vector (e.g., [0.98, 0.50, 0.11]) can receive a higher anomaly score than a node with uniformly high scores (e.g., [0.99, 0.99, 0.98]). This behavior suggests that when individual scores already convey strong anomaly evidence, monotonic linear fusion provides a more natural and consistent integration. Finding 2: stability of the proposed fusion across datasets. Although linear fusion performed best overall, the optimal weight combinations varied by dataset—some favored two dominant views, others preferred balanced weighting. A fixed-weight model would thus require datasetspecific tuning to remain competitive. Because P ROV F U SION employs seven monotonic detectors that collectively cover representative linear fusion patterns—from singleview to all-view combinations—the voting stage maintains

w/o Group 1 w/o Group 2 w/o Group 3 Final TP FP TP FP TP FP TP FP CADETS-E3 24 1 24 1 24 1 24 1 THEIA-E3 91 2 96 214 94 5 91 2 CLRSCP-E3 6 7 44 1367 6 8 6 7 CADETS-E5 4 3 7 23 8 16 7 9 THEIA-E5 11 2 11 32 14 41 11 2 CLRSCP-E5 10 16 13 80 10 18 10 16 Dataset

TABLE 5: An ablation study of different detector groups. While individual detector groups yield high False Positives or low True Positives, our ensemble method achieves a more effective balance. performance comparable to the best observed weights, and consistently stable across datasets. This alignment between theoretical coverage and observed consistency supports the soundness of the overall design. Summary. Under our uniform experimental setting and datasets: (i) monotonic linear fusion provides a reliable basis for integrating multi-view anomaly scores; and (ii) the voting-based decision integration mitigates dataset-specific sensitivity, yielding stable results without further tuning. 5.3.3. Ablation: Detector Group Effectiveness and Voting Sensitivity. Within the multi-dimensional anomaly fusion and voting-based detection framework, P ROV F USION employs three detector groups corresponding to single-, two, and three-view combinations. To assess the contribution of each group and the sensitivity of the voting threshold, we conduct two ablation studies analyzing their respective impacts on detection performance. Ablation of Detector Groups. We assess each group’s contribution by disabling it in turn and measuring performance changes across datasets (Table 5). The results indicate that removing any group degrades performance. Excluding Group 2 or Group 3 increases false positives, while removing Group 1 causes a missed attack on CADETS-E5, reducing campaign coverage. These observations indicate that each group captures distinct anomaly patterns, and combining all three yields broader coverage with fewer false alarms. Detailed per-group and per-detector results are provided in Appendix I and Appendix E. Threshold sensitivity of the voting rule. An entity is flagged as anomalous if it receives at least Tv votes from the seven detectors. We vary Tv from 1 to 7 and record MCC, F1, and false-positive counts (Figure 3). When the threshold is small (Tv = 1–2), the detection criterion becomes overly permissive, flagging many benign nodes as anomalies and resulting in a sharp increase in false positives. Conversely, when the threshold is large (Tv = 6–7), the detection criterion becomes too strict: although false positives drop, true positives also decline noticeably, leading to lower F1 and MCC scores. Between these extremes, moderate thresholds (Tv = 3–5) maintain high F1 and MCC values while keeping false positives at a manageable level, indicating a balanced trade-off between recall and precision. The default Tv = 4 (a simple majority of seven) lies within this stable range and provides consistently strong results across datasets. Summary. Each detector group contributes distinct strengths, and the voting mechanism remains robust within

a wide threshold range (Tv = 3–5), showing consistent behavior across datasets. 5.3.4. Ablation: Enhancements to Node- and EdgeCentric Detectors. P ROV F USION incorporates two viewspecific enhancements as part of its architecture. (i) For the node-centric view, we decouple structural and attributebased scoring to mitigate false positives caused by noisy or infrequent attributes (e.g., rare file paths). (ii) For the edge-centric view, we apply a class-weighted training loss to address the imbalance between frequent and rare edge types in provenance graphs. To assess their contribution within P ROV F USION, we perform ablations by removing each enhancement in turn and measuring the impact on detection performance. Experimental setup. We compare the full system (Edgew/ E -W eight , Nodew/ SA-Split ) with two degraded variants: Nodew/o SA-Split , which concatenates one-hot node types with attribute embeddings as input features, causing the encoder to aggregate structural and attribute information into a single representation; and Edgew/o E -W eight , which treats all edge types equally. Results on THEIA-E3, a representative dataset, are shown in Table 6. Findings. Introducing the structural–attribute split increases true positives, improving the system’s ability to identify real attack nodes while filtering attribute-level noise. Disabling edge weighting lowers ADP, indicating reduced ability to emphasize rare but critical edges relative to common benign ones. Both enhancements therefore strengthen the detectors’ discriminative power and stability. 5.3.5. Contribution of the Multi-View Architecture. To assess the contribution of each view, we disable one at a time and re-evaluate performance (Table 7). All three views contribute meaningfully but in distinct ways. Removing the Edge-centric Causal View mainly reduces recall, whereas removing either the Node-centric Attribute View or the Structure View increases false positives and lowers precision. These patterns indicate that each view captures distinct aspects of anomalous behavior, and together they contribute to balanced and stable detection across datasets. RQ2 Answer. Our in-depth analysis shows that P ROV F USION’s detection capability results from the coordinated contributions of multiple design components rather than any single factor. (1) Enhancements to the node- and edge-centric detectors reinforce the underlying anomaly signals; (2) the multidimensional anomaly fusion with voting-based detection helps achieve stable performance and robustness to moderate threshold variations; and (3) the multi-view architecture allows different types of anomalies to be captured across diverse scenarios. Overall, P ROV F USION’s effectiveness arises from the synergy of these components, which together provide reliable and consistent detection across datasets.

5.4. Computational Overhead (RQ3) In this section, we evaluate the runtime and scalability of P ROV F USION to answer our third research question: Is

Ablation Component Edge-centricw/ E -W eight Edge-centricw/o E -W eight Node-centricw/ SA-Split Node-centricw/o SA-Split

Single centric Performance Final Performance TP↑ FP↓ TN↑ FN↓ F1↑ ADP↑ MCC↑ TP↑ FP↓ TN↑ FN↓ F1↑ ADP↑ MCC↑ 95 36 701.5K 34 0.73 1.00 0.73 91 2 701.5K 38 0.82 1.00 0.83 19 10 701.4K 110 0.24 0.73 0.31 16 3 701.5K 113 0.22 0.64 0.32 37 19 701.5K 92 0.40 0.48 0.44 91 2 701.5K 38 0.82 1.00 0.83 0 19 701.5K 129 0.00 0.00 0.00 6 0 701.5K 123 0.09 1.00 0.22

TABLE 6: Impact of our core detector enhancements on the THEIA-E3 dataset. This ablation study shows that removing either the edge weighting or the structure-attribute (SA) split causes a significant drop in performance. Dataset System TP↑ FP↓ CADETS-E3 24 1 THEIA-E3 90 1 w/o CLRSCP-E3 6 288 attribute CADETS-E5 3 5 view THEIA-E5 14 35 CLRSCP-E5 1 11 CADETS-E3 0 1 THEIA-E3 1 15 w/o CLRSCP-E3 38 983 causal CADETS-E5 5 9 view THEIA-E5 0 12 CLRSCP-E5 4 20

TN↑ 281.5K 701.5K 111.1K 3.14M 1.86M 150.9K 281.5K 701.5K 110.4K 3.14M 1.86M 150.9K

FN↓ 44 39 43 120 55 52 68 128 11 118 69 49

F1↑ ADP↑ MCC↑ System TP↑ FP↓ 0.52 1.00 0.58 24 1 0.82 1.00 0.83 61 17 w/o 0.04 1.00 0.05 3 0 structure 0.05 1.00 0.10 2 3 view 0.24 1.00 0.24 1 10 0.03 0.83 0.04 11 23 0.00 0.85 0.00 24 1 0.01 0.34 0.02 91 2 0.07 0.97 0.17 original 6 7 0.07 1.00 0.12 ProvFusion 7 9 0.00 0.45 0.00 11 2 0.10 0.76 0.11 10 16

TN↑ FN↓ 281.5K 44 701.5K 68 111.4K 46 3.14M 121 1.86M 68 150.9K 42 281.5K 44 701.5K 38 111.3K 43 3.14M 116 1.86M 58 150.9K 43

F1↑ ADP↑ MCC↑ 0.52 1.00 0.58 0.59 1.00 0.61 0.12 1.00 0.25 0.03 0.90 0.08 0.03 1.00 0.04 0.25 0.82 0.26 0.52 1.00 0.58 0.82 1.00 0.83 0.19 0.86 0.24 0.10 0.92 0.16 0.27 1.00 0.37 0.25 0.77 0.27

TABLE 7: Ablation study quantifying the importance of each analysis view. Removing any single view results in a significant performance drop on multiple datasets, demonstrating effectiveness of our multi-view fusion.

103

TH E3 EIA CL -E3 RS C CA P-E3 DE TS TH E5 EIA CL -E5 RS CP -E5

102

DE CA

Seconds (log scale)

ProvFusion

103

Orthrus

Velox

ProvFusion

102 101

TS TH E3 EIA CL -E3 RS C CA P-E3 DE TS TH E5 EIA CL -E5 RS CP -E5

Velox

DE

Orthrus

CA

104

TS

Seconds (log scale)

(b) Inference Time. (a) Training Time. Figure 4: Runtime Overhead. 5.4.1. Training and Inference Efficiency. Figure 4 shows the training and inference times across datasets, with a detailed runtime breakdown in Appendix K. Overall, P ROVF USION achieves broadly comparable training efficiency to Velox and Orthrus, while exhibiting higher inference efficiency, as we detail below. First, during both training and inference, P ROV F USION represents multiple edges between two nodes as a single multi-hot edge, rather than processing each edge individually. This design avoids redundant computation and reduces the total number of edge-level operations, thereby mitigating part of the additional cost introduced by the multi-view framework. As a result, the overall training time is competitive. For instance, on the largest dataset, CADETS-E5, our training time is comparable to Velox (with only ∼2.6% overhead) and faster than Orthrus. Second, a critical component of P ROV F USION is the KNN-based scoring for the attribute and structure views.

1e7 1.0

1.0

0.5

0.5 0

10 20 30 GPU Memory (GB)

0.0

3

1e6

1e7 2

# of Edges

1e6

# of Edges # of Nodes

1.5 # of Nodes

P ROV F USION computationally efficient and scalable? We measure the training time, inference time, and scalability with respect to graph size, and compare P ROV F USION with Velox [24] and Orthrus [25], two SOTA systems that report higher efficiency than all other compared approaches. All experiments are conducted under the same hardware and software configuration for fair comparison.

2

1

1 300 400 500 600 Process Time (s)

0

(a) Peak GPU Memory. (b) Inference Throughput. Figure 5: Scalability Analysis. While this step, which compares test nodes against all training nodes, is often a computational bottleneck, we address this by integrating the Faiss library [44], [45] to accelerate the nearest-neighbor search. This optimization is highly effective, leading to a consistently strong inference performance. As a result, P ROV F USION is faster than both baselines on many datasets. For instance, on the largest dataset (CADETS-E5), P ROV F USION is ∼1.10× faster than Velox and ∼1.25× faster than Orthrus. On THEIA-E5, it is ∼2.5× faster than both baselines, demonstrating the practical efficiency of our optimized framework. 5.4.2. Scalability Analysis. To evaluate scalability, we construct graphs of increasing size from CADETS-E5 by randomly sampling nodes (300K to 3M) and their associated edges. We then measure GPU memory usage and inference processing time against the graph size. As shown in Figure 5, GPU memory consumption scales approximately linearly with the graph size, as expected. More importantly, the inference time exhibits a favorable linear growth along with graph size. Our analysis shows that while the number of nodes increases by 10× (from 3K to 3M) and the number of edges increases by ∼75× (from 330K to 24.7M), the total processing time increases by only ∼2.7× (from 249s to 670s). This result demonstrates that our Faiss-accelerated [44] KNN scoring step successfully avoids

becoming a quadratic bottleneck. RQ3 Answer. P ROV F USION achieves competitive training efficiency and relatively faster inference performance compared to SOTA baselines, particularly on large-scale datasets. This efficiency is driven by our multi-hot edge representation and Faissaccelerated KNN scoring. Scalability analysis confirms that the memory usage and the inference processing time grow linearly with graph size. This demonstrates that our framework is scalable and efficient for large-scale graphs.

5.5. Hyperparameter Study (RQ4) In this section, we answer RQ4 using the THEIA-E3 dataset, which contains a relatively large number of labeled nodes (128) and three distinct attack campaigns, making it suitable for controlled sensitivity analysis. Figure 6 summarizes the results, reporting true positives (TPs), false positives (FPs) under different parameter settings. Weighting coefficient α. As described in Section §4.3, α controls the adaptive weighting. Performance remains highly consistent for α between 4 and 7 (TPs stay at 91, FPs at 2–4), showing limited sensitivity to this parameter. Mask rate mr. A moderate rate of mr = 0.3 yields the better balance (91 TPs, 2 FPs), while lower (mr = 0.1) and higher (mr = 0.5) rates degrade performance (83 TPs / 4 FPs and 87 TPs / 3 FPs, respectively). Learning rate lr. A proper learning rate ensures smooth training dynamics. We find that lr = 1 × 10−3 provides the best performance (91 TPs, 2 FPs), as higher rates (e.g., 1e − 2) failed to converge (0 TPs), and lower rates produced more false positives (e.g., 5 FPs at 1e−4 or 17 FPs at 1e−6). Encoder layers L. The number of GNN encoder layers controls neighborhood aggregation. A two-layer (L = 2) encoder achieves the best balance (91 TPs, 2 FPs). A single layer (L = 1) provides insufficient representational capacity (1 TP, 33 FPs), while too many layers (e.g., L = 3 or L = 4) cause over-smoothing, degrading performance (13 TPs and 5 TPs, respectively). Hidden dimension h. The hidden dimension determines the representational granularity. A dimension of h = 64 provides an effective balance (91 TPs, 2 FPs). A smaller dimension (h = 32) underfits and performs poorly (13 TPs, 44 FPs), while a larger one (h = 128) slightly degrades performance (83 TPs, 4 FPs). RQ4 Answer. P ROV F USION shows stable performance across its core tuning parameters, with limited sensitivity to the fusion coefficient α and mask rate mr. Consistent with general GNN behavior, architectural parameters such as encoder layers (L) and hidden dimension (h) require careful tuning to balance representational capacity and avoid underfitting or over-smoothing.

6. Adversarial Robustness A representative example is the mimicry attack proposed by Goyal et al. [71], which targets subgraph-based detectors

(e.g., Unicorn [11]) by crafting malicious subgraphs that imitate benign topological motifs. These attacks rely on the assumption of global structural similarity, embedding malicious interactions within subgraphs that visually and statistically resemble benign patterns. However, this assumption does not directly apply to node- or edge-centric paradigms, which assess anomalies based on localized behavioral contexts—focusing on entity-level attributes and interaction plausibility rather than global subgraph resemblance. This distinction has been empirically supported by the benchmark analysis of Velox [24], which evaluated multiple recent PIDSs [17], [18], [23], [24], [25], [26], [72] under subgraphlevel mimicry attacks and found that node- and edge-centric systems were largely unaffected, as localized behavioral features remained stable under such global perturbations. Consequently, mimicry attacks designed for subgraph-based detectors are not directly suitable for evaluating systems following node- or edge-centric paradigms such as ours. To empirically assess the reasoning, we reproduced the mimicry attack experiment of Goyal et al. [71] using their publicly released implementation [24], [71]. Detailed results are in Appendix §L. Consistent with prior observations, the attack failed to induce false negatives in P ROV F USION, and all original malicious entities were correctly identified. However, we observed a subtle secondary effect: the structural perturbations introduced by the attack slightly increased false positives. This likely occurred because the edges inserted by the attacker—intended to make the subgraph appear benign—introduced local behavioral inconsistencies that our fine-grained detectors recognized as anomalous. This finding provides empirical confirmation that subgraph-level mimicry attacks are largely ineffective against node- and edge-centric paradigms, whose decisions rely on localized rather than global behavioral evidence. It also points to a next-step challenge—designing adaptive evasion strategies that can deceive fine-grained detectors without introducing new detectable inconsistencies, thereby enabling a more rigorous evaluation of PIDSs.

7. Related Work We position P ROV F USION with respect to two established paradigms in provenance-based intrusion detection: node-centric and edge-centric analysis. Each paradigm provides a distinct analytical focus. Reviewing them helps clarify the scope of prior progress and the rationale for combining multiple analytical perspectives. Node-centric approaches. Node-centric methods analyze the properties of individual entities and their immediate neighborhoods. They generally operate under the notion that malicious behavior alters the attributes or structural context of affected nodes. Representative works such as ThreaTrace [72] and Flash [18] train GNNs to predict node types (e.g., process, file) based on contextual information, with classification inconsistency used as an indicator of potential abnormality. Other methods adopt self-supervised objectives to model regularities of benign behavior without relying on predefined labels. SIGL [22] employs a Graph

0

1 2 3 4 5 6 7 8 9 10

0

(a) Weighting Coefficient

50

2

25 0

4

0.1

0.3

0.5

(b) Mask Rate mr

100 75 50 25 0

0

20

TP FP

15 10 5

1e-2 1e-3 1e-4 1e-5 1e-6

(c) Learning Rate lr

0

100 75 50 25 0

TP FP

100

40 20

32

64

128

256

0

(d) Hidden Dimension h

TP FP

75

40

False Positives

2

25

TP FP

75

False Positives True Positives

50

100

4

False Positives True Positives

TP FP

False Positives True Positives

75

False Positives True Positives

True Positives

100

30

50

20

25

10

0

1

2

3

4

0

(e) Encoder Layers L

Figure 6: Hyperparameter sensitivity on THEIA-E3 dataset. LSTM-based [73] autoencoder to reconstruct process features, while MAGIC [17] uses a GAT-based masked reconstruction framework and measures deviations via searchbased metrics such as K-D tree distance [43]. R-CAID [21] further combines GNNs with root-cause reasoning to associate detected deviations with their potential origins. Overall, these approaches provide detailed per-entity assessment but remain centered on local reasoning. When coordinated attack behaviors span multiple entities that each appear individually plausible, purely node-centric detectors may provide limited visibility into the broader execution context. Edge-centric approaches. Edge-centric methods [9] shift the analytical focus from entities to their interactions, modeling whether interactions between entities conform to typical causal patterns. Kairos [26] leverages Temporal Graph Networks (TGNs) [74] to represent long-range temporal dependencies in system event streams. To improve scalability, Orthrus [25] and Velox [24] introduce more lightweight formulations emphasizing efficiency and interpretability. Such systems are effective at identifying irregular causal flows but may miss anomalies residing within entity states. For instance, a compromised process may continue to perform syntactically valid operations, generating causal links that appear ordinary when viewed in isolation. In such cases, evaluating edge plausibility alone may not sufficiently expose the underlying compromise. Synthesis. Existing work collectively highlights a persistent granularity gap. Node-centric detectors focus primarily on entity states and tend to place limited emphasis on interaction semantics, whereas edge-centric detectors concentrate on causal relationships while offering only partial visibility into the internal conditions of participating entities. Both perspectives contribute valuable insights yet capture different aspects of system behavior. As modern attack campaigns often involve subtle changes across entities and their interactions, reliance on a single detection view can reduce coverage in complex scenarios. P ROV F USION is designed to address this practical limitation by jointly analyzing entitylevel and interaction-level anomalies within a unified multiview formulation, providing a more context-aware basis for provenance-based intrusion detection.

modeling such as temporal graph networks [74] into the multi-view pipeline is a promising direction. Single-host scope. P ROV F USION, like existing PIDSs [17], [18], [23], [24], [25], [26], analyzes provenance from a single host and does not directly capture lateral movement across hosts. Extending multi-view fusion to cross-host settings requires addressing heterogeneous logging formats, cross-boundary entity resolution, and substantially larger graph scales; we view federated or hierarchical fusion across host-level detectors as a natural next step. Noisy and incomplete logging. We assume faithful provenance capture (Section §3); in practice, log pruning failures can break causal chains and degrade the structural and causal views. Developing uncertainty-aware scoring or imputation under degraded logging is a valuable future direction.

9. Conclusion In this work, we examined the different focuses of node-centric and edge-centric approaches in provenancebased intrusion detection. To bridge these gaps, we introduced P ROV F USION, a multi-view detection framework that integrates anomalies from three detection views: entity attributes, structure pattern, and causal interactions. By systematically fusing these views and employing a votingbased detection mechanism that relies solely on validation data, P ROV F USION provides a broader and more consistent assessment of system behavior. Experiments on six benchmark datasets show improved detection quality over representative state-of-the-art systems, with higher detection rates and lower false positives. Overall, this work offers an effective design for provenance-based defense systems.

References [1]

[2]

[3]

8. Limitations and Future Work Static graph and temporal dynamics. P ROV F USION encodes the events between an entity pair as a single multi-hot edge, which discards their ordering and timing. Like other static-graph PIDSs, it may therefore underestimate attacks that are anomalous only in sequence (e.g., slow exfiltration interleaved with legitimate accesses). Integrating temporal

[4]

[5]

W. U. Hassan, S. Guo, D. Li, Z. Chen, K. Jee, Z. Li, and A. Bates, “Nodoze: Combatting threat alert fatigue with automated provenance triage,” in 26th Annual Network and Distributed System Security Symposium, NDSS 2019, San Diego, California, USA, February 2427, 2019. The Internet Society, 2019. W. U. Hassan, A. Bates, and D. Marino, “Tactical provenance analysis for endpoint detection and response systems,” in 2020 IEEE Symposium on Security and Privacy, SP 2020, San Francisco, CA, USA, May 18-21, 2020. IEEE, 2020, pp. 1172–1189. Y. Liu, M. Zhang, D. Li, K. Jee, Z. Li, Z. Wu, J. Rhee, and P. Mittal, “Towards a timely causality analysis for enterprise security,” in 25th Annual Network and Distributed System Security Symposium, NDSS 2018, San Diego, California, USA, February 18-21, 2018. The Internet Society, 2018. M. Inam, Y. Chen, A. Goyal, J. Liu, J. Mink, N. Michael, S. Gaur, A. Bates, and W. Hassan, “Sok: History is a vast early warning system: Auditing the provenance of system intrusions,” in 2023 IEEE Symposium on Security and Privacy (SP), 2023. K. Pei, Z. Gu, B. Saltaformaggio, S. Ma, F. Wang, Z. Zhang, L. Si, X. Zhang, and D. Xu, “HERCULE: attack story reconstruction via community discovery on correlated log graph,” in Proceedings of the

32nd Annual Conference on Computer Security Applications, ACSAC 2016, Los Angeles, CA, USA, December 5-9, 2016, S. Schwab, W. K. Robertson, and D. Balzarotti, Eds. ACM, 2016, pp. 583–595. [6] A. Alshamrani, S. Myneni, A. Chowdhary, and D. Huang, “A survey on advanced persistent threats: Techniques, solutions, challenges, and research opportunities,” IEEE Commun. Surv. Tutorials, vol. 21, no. 2, pp. 1851–1877, 2019. [7] S. M. Milajerdi, R. Gjomemo, B. Eshete, R. Sekar, and V. N. Venkatakrishnan, “HOLMES: real-time APT detection through correlation of suspicious information flows,” in 2019 IEEE Symposium on Security and Privacy, SP 2019, San Francisco, CA, USA, May 19-23, 2019. IEEE, 2019, pp. 1137–1152. [8] W. U. Hassan, M. Lemay, N. Aguse, A. Bates, and T. Moyer, “Towards scalable cluster auditing through grammatical inference over provenance graphs,” in 25th Annual Network and Distributed System Security Symposium, NDSS 2018, San Diego, California, USA, February 18-21, 2018. The Internet Society, 2018. [9] J. Zeng, X. Wang, J. Liu, Y. Chen, Z. Liang, T. Chua, and Z. L. Chua, “SHADEWATCHER: recommendation-guided cyber threat analysis using system audit records,” in 43rd IEEE Symposium on Security and Privacy, SP 2022, San Francisco, CA, USA, May 22-26, 2022. IEEE, 2022, pp. 489–506. [10] Q. Wang, W. U. Hassan, D. Li, K. Jee, X. Yu, K. Zou, J. Rhee, Z. Chen, W. Cheng, C. A. Gunter, and H. Chen, “You are what you do: Hunting stealthy malware via data provenance analysis,” in 27th Annual Network and Distributed System Security Symposium, NDSS 2020, San Diego, California, USA, February 23-26, 2020. The Internet Society, 2020. [11] X. Han, T. F. J. Pasquier, A. Bates, J. Mickens, and M. I. Seltzer, “Unicorn: Runtime provenance-based detector for advanced persistent threats,” in 27th Annual Network and Distributed System Security Symposium, NDSS 2020, San Diego, California, USA, February 2326, 2020. The Internet Society, 2020. [12] I. J. King, X. Shu, J. Jang, K. Eykholt, T. Lee, and H. H. Huang, “Edgetorrent: Real-time temporal graph representations for intrusion detection,” in Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses, ser. RAID ’23. Association for Computing Machinery, 2023, p. 77–91. [13] K. Xu, W. Hu, J. Leskovec, and S. Jegelka, “How powerful are graph neural networks?” arXiv preprint arXiv:1810.00826, 2018. [14] E. Altinisik, F. Deniz, and H. T. Sencar, “Provg-searcher: A graph representation learning approach for efficient provenance graph search,” in Proceedings of the 2023 ACM SIGSAC conference on computer and communications security, 2023, pp. 2247–2261. [15] F. Yang, J. Xu, C. Xiong, Z. Li, and K. Zhang, “PROGRAPHER: an anomaly detection system based on provenance graph embedding,” in 32nd USENIX Security Symposium, USENIX Security 2023, Anaheim, CA, USA, August 9-11, 2023, J. A. Calandrino and C. Troncoso, Eds. USENIX Association, 2023, pp. 4355–4372. [16] X. Han, T. Pasquier, and M. Seltzer, “Provenance-based intrusion detection: opportunities and challenges,” in Proceedings of the 10th USENIX Conference on Theory and Practice of Provenance, ser. TaPP’18. USA: USENIX Association, 2018, p. 3. [17] Z. Jia, Y. Xiong, Y. Nan, Y. Zhang, J. Zhao, and M. Wen, “Magic: Detecting advanced persistent threats via masked graph representation learning,” 2023. [18] M. Rehman, H. Ahmadi, and W. Hassan, “Flash: A comprehensive approach to intrusion detection via provenance graph representation learning,” in 2024 IEEE Symposium on Security and Privacy (SP). [19] L. Akoglu, H. Tong, and D. Koutra, “Graph based anomaly detection and description: a survey,” Data mining and knowledge discovery, vol. 29, no. 3, pp. 626–688, 2015. [20] X. Ma, J. Wu, S. Xue, J. Yang, C. Zhou, Q. Z. Sheng, H. Xiong, and L. Akoglu, “A comprehensive survey on graph anomaly detection with deep learning,” IEEE Transactions on Knowledge and Data Engineering, vol. 35, no. 12, p. 12012–12038, Dec. 2023. [21] A. Goyal, G. Wang, and A. Bates, “R-caid: Embedding root cause analysis within provenance-based intrusion detection,” in 2024 IEEE Symposium on Security and Privacy (SP), 2024. [22] X. Han, X. Yu, T. Pasquier, D. Li, J. Rhee, J. Mickens, M. Seltzer, and H. Chen, “SIGL: Securing software installations through deep graph

learning,” in 30th USENIX Security Symposium (USENIX Security 21), 2021. [23] S. Li, F. Dong, X. Xiao, H. Wang, F. Shao, J. Chen, Y. Guo, X. Chen, and D. Li, “Nodlink: An online system for fine-grained apt attack detection and investigation,” in Proceedings 2024 Network and Distributed System Security Symposium. Internet Society, 2024. [24] T. Bilot, B. Jiang, Z. Li, N. El Madhoun, K. Al Agha, A. Zouaoui, and T. Pasquier, “Sometimes Simpler is Better: A Comprehensive Analysis of State-of-the-Art Provenance-Based Intrusion Detection Systems,” in Security Symposium (USENIX Sec’25). USENIX, 2025. [25] B. Jiang, T. Bilot, N. El Madhoun, K. Al Agha, A. Zouaoui, S. Iqbal, X. Han, and T. Pasquier, “ORTHRUS: Achieving High Quality of Attribution in Provenance-based Intrusion Detection Systems,” in Security Symposium (USENIX Sec’25). USENIX, 2025. [26] Z. Cheng, Q. Lv, J. Liang, Y. Wang, D. Sun, T. Pasquier, and X. Han, “Kairos: Practical intrusion detection and investigation using wholesystem provenance,” 2023. [27] DARPA I2O, “Transparent computing engagement 3 data release accessed 29th january 2025.” https://github.com/darpa-i2o/ Transparent-Computing/blob/master/README-E3.md, 2018. [28] DARPA 5, “Transparent computing engagement 5,” https://github. com/darpa-i2o/Transparent-Computing, 2019. [29] DARPA OpTc, “Operationally transparent cyber (optc) data,” https: //github.com/FiveDirections/OpTC-data, 2019. [30] T. N. Kipf and M. Welling, “Semi-supervised classification with graph convolutional networks,” 2017. [31] P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Liò, and Y. Bengio, “Graph attention networks,” 2018. [32] Don Marshall, “Winodws event tracing,” https://docs. microsoft.com/en-us/windows-hardware/drivers/devtest/ event-tracing-for-windows--etw-, 2021. [33] Steve Grubb, “Linux audit,” https://linux.die.net/man/8/auditd, 2021. [34] George V. Neville-Neil, “Dtrace,” https://wiki.freebsd.org/DTrace, 2018. [35] R. Paccagnella, P. Datta, W. U. Hassan, A. Bates, C. W. Fletcher, A. Miller, and D. Tian, “Custos: Practical tamper-evident auditing of operating systems using trusted execution,” in 27th Annual Network and Distributed System Security Symposium, NDSS 2020, San Diego, California, USA, February 23-26, 2020. The Internet Society, 2020. [36] A. Bates, D. J. Tian, K. R. Butler, and T. Moyer, “Trustworthy WholeSystem provenance for the linux kernel,” in 24th USENIX Security Symposium (USENIX Security 15). Washington, D.C.: USENIX Association, Aug. 2015, pp. 319–334. [37] R. Paccagnella, K. Liao, D. Tian, and A. Bates, “Logging to the danger zone: Race condition attacks and defenses on system audit frameworks,” in CCS ’20: 2020 ACM SIGSAC Conference on Computer and Communications Security, Virtual Event, USA, November 9-13, 2020, J. Ligatti, X. Ou, J. Katz, and G. Vigna, Eds. ACM, 2020, pp. 1551–1574. [38] T. F. J. Pasquier, X. Han, M. Goldstein, T. Moyer, D. M. Eyers, M. I. Seltzer, and J. Bacon, “Practical whole-system provenance capture,” in Proceedings of the 2017 Symposium on Cloud Computing, SoCC 2017, Santa Clara, CA, USA, September 24-27, 2017. ACM, 2017, pp. 405–418. [39] Z. Zhang, P. Qi, and W. Wang, “Dynamic malware analysis with feature engineering and feature learning,” in Proceedings of the AAAI conference on artificial intelligence, vol. 34, no. 01, 2020, pp. 1210– 1217. [40] T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” arXiv preprint arXiv:1301.3781, 2013. [41] M. Leimeister and B. J. Wilson, “Skip-gram word embeddings in hyperbolic space,” 2019. [42] J. Gilmer, S. S. Schoenholz, P. F. Riley, O. Vinyals, and G. E. Dahl, “Neural message passing for quantum chemistry,” CoRR, vol. abs/1704.01212, 2017. [43] G. Guo, H. Wang, D. Bell, Y. Bi, and K. Greer, “Knn model-based approach in classification,” in On The Move to Meaningful Internet Systems 2003: CoopIS, DOA, and ODBASE, R. Meersman, Z. Tari, and D. C. Schmidt, Eds., 2003. [44] M. Douze, A. Guzhva, C. Deng, J. Johnson, G. Szilvasy, P.-E. Mazaré,

M. Lomeli, L. Hosseini, and H. Jégou, “The faiss library,” 2024. [45] J. Johnson, M. Douze, and H. Jégou, “Billion-scale similarity search with GPUs,” IEEE Transactions on Big Data, vol. 7, no. 3, pp. 535– 547, 2019. [46] Z. Hou, X. Liu, Y. Cen, Y. Dong, H. Yang, C. Wang, and J. Tang, “Graphmae: Self-supervised masked graph autoencoders,” 2022. [47] S. Brody, U. Alon, and E. Yahav, “How attentive are graph attention networks?” arXiv preprint arXiv:2105.14491, 2021. [48] M.-C. Popescu, V. E. Balas, L. Perescu-Popescu, and N. Mastorakis, “Multilayer perceptron and neural networks,” WSEAS Trans. Cir. and Sys., vol. 8, no. 7, p. 579–588, Jul. 2009. [49] D. Arp, E. Quiring, F. Pendlebury, A. Warnecke, F. Pierazzi, C. Wressnegger, L. Cavallaro, and K. Rieck, “Dos and don’ts of machine learning in computer security,” 2021. [50] L. John Wiley & Sons, “Learning from data: concepts, theory, and methods,” 2007. [51] E. A. Manzoor, S. M. Milajerdi, and L. Akoglu, “Fast memoryefficient anomaly detection in streaming heterogeneous graphs,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, August 13-17, 2016, B. Krishnapuram, M. Shah, A. J. Smola, C. C. Aggarwal, D. Shen, and R. Rastogi, Eds. ACM, 2016, pp. 1035– 1044. [52] A. Alsaheel, Y. Nan, S. Ma, L. Yu, G. Walkup, Z. B. Celik, X. Zhang, and D. Xu, “ATLAS: A sequence-based learning approach for attack investigation,” in 30th USENIX Security Symposium (USENIX Security 21), 2021. [53] J. Zeng, Z. L. Chua, Y. Chen, K. Ji, Z. Liang, and J. Mao, “WATSON: abstracting behaviors from audit logs via aggregation of contextual semantics,” in 28th Annual Network and Distributed System Security Symposium, NDSS 2021, virtually, February 21-25, 2021. The Internet Society, 2021. [54] “The freebsd project,” https://www.freebsd.org/, 2025. [55] “The linux kernel archives,” https://www.kernel.org/, 2025. [56] “Android open source project,” https://source.android.com/, 2025. [57] W. Qiao, Y. Feng, T. Li, Z. Ma, Y. Shen, J. Ma, and Y. Liu, “Slot: Provenance-driven apt detection through graph reinforcement learning,” 2025. [58] A. Aly, E. Mansour, and A. Youssef, “Ocr-apt: Reconstructing apt stories from audit logs using subgraph anomaly detection and llms,” ser. CCS ’25. [59] L. Wang, X. Shen, W. Li, Z. Li, R. Sekar, H. Liu, and Y. Chen, “Incorporating gradients to rules: Towards lightweight, adaptive provenancebased intrusion detection,” arXiv preprint arXiv:2404.14720, 2024. [60] C. X. Ling, J. Huang, and H. Zhang, “Auc: a statistically consistent and more discriminating measure than accuracy,” in Proceedings of the 18th International Joint Conference on Artificial Intelligence, ser. IJCAI’03. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc., 2003, p. 519–524. [61] D. Chicco, M. J. Warrens, and G. Jurman, “The matthews correlation coefficient (mcc) is more informative than cohen’s kappa and brier score in binary classification assessment,” IEEE Access, vol. 9, pp. 78 368–78 381, 2021. [62] D. T. Computing, “Ground truth file,” https://drive.google.com/file/d/ 1mrs4LWkGk-3zA7t7v8zrhm0yEDHe57QU/view, 2018. [63] REAPr, “REAPr label set,” https://bitbucket.org/sts-lab/ reapr-ground-truth/src/master/, 2024. [64] J. Han, M. Kamber, and J. Pei, Data mining: Concepts and techniques. Elsevier, 2011. [65] P. Röchner, H. O. Marques, R. J. G. B. Campello, A. Zimek, and F. Rothlauf, Robust Statistical Scaling of Outlier Scores: Improving the Quality of Outlier Probabilities for Outliers. Springer Nature Switzerland, Oct. 2024, p. 215–222. [66] J. Liu, D. Cai, and X. He, “Gaussian mixture model with local consistency,” in Proceedings of the Twenty-Fourth AAAI Conference on Artificial Intelligence, ser. AAAI’10. AAAI Press, 2010, p. 512–517. [67] H. Otneim and D. TjØstheim, “The locally gaussian density estimator for multivariate data,” Statistics and Computing, vol. 27, no. 6, p. 1595–1616, Nov. 2017. [68] B. Schölkopf, R. Williamson, A. Smola, J. Shawe-Taylor, and J. Platt,

“Support vector method for novelty detection,” ser. NIPS’99. Cambridge, MA, USA: MIT Press, 1999, p. 582–588. [69] F. T. Liu, K. M. Ting, and Z.-H. Zhou, “Isolation forest,” in 2008 Eighth IEEE International Conference on Data Mining, 2008, pp. 413–422. [70] R. Chalapathy and S. Chawla, “Deep learning for anomaly detection: A survey,” 2019. [Online]. Available: https://arxiv.org/abs/1901.03407 [71] A. Goyal, X. Han, G. Wang, and A. Bates, “Sometimes, you aren’t what you do: Mimicry attacks against provenance graph host intrusion detection systems,” in 30th Annual Network and Distributed System Security Symposium, NDSS 2023, San Diego, California, USA, February 27 - March 3, 2023. The Internet Society, 2023. [72] S. Wang, Z. Wang, T. Zhou, H. Sun, X. Yin, D. Han, H. Zhang, X. Shi, and J. Yang, “THREATRACE: detecting and tracing hostbased threats in node level through provenance graph learning,” IEEE Trans. Inf. Forensics Secur., vol. 17, pp. 3972–3987, 2022. [73] N. Peng, H. Poon, C. Quirk, K. Toutanova, and W. tau Yih, “Crosssentence n-ary relation extraction with graph lstms,” 2017. [74] E. Rossi, B. Chamberlain, F. Frasca, D. Eynard, F. Monti, and M. M. Bronstein, “Temporal graph networks for deep learning on dynamic graphs,” CoRR, vol. abs/2006.10637, 2020.

A. Entity and Event Type Specification Table 8 enumerates the entity and event types used in constructing the provenance graph G = (V, E). Category

Types Considered in P ROV F USION

Entity Types Event Types

process, file, netflow CONNECT, EXECUTE, OPEN, READ, RECVFROM, RECVMSG, SENDMSG, SENDTO, WRITE, CLONE

TABLE 8: Summary of entity and event types.

B. Dataset Statistics Table 9 presents detailed statistics for each of the six processed DARPA TC datasets. Dataset CADETS-E3 THEIA-E3 ClEARSCOPE-E3 CADETS-E5 THEIA-E5 ClEARSCOPE-E5 OpTC-H201 OpTC-H501 OpTC-H051

Number of Nodes 1,103,333 1,308,080 260,700 8,045,589 4,390,381 366,025 3,839,000 3,884,805 3,089,126

Number of Edges 4,476,834 5,969,280 838,876 69,319,401 46,035,572 4,523,379 8,289,319 8,454,821 8,182,521

TABLE 9: Statistics of the evaluation datasets.

C. Dataset Splitting Methodology We adopted the identical data splitting methodology used in recent works [24], [25]. Table 10 provides the specific date ranges used to partition each dataset. The training sets consist exclusively of benign data, while the test sets contain a mix of benign and malicious activities.

D. Timestamp and Name of Each attack This section illustrates the detailed name and the timestamp of their occurrence. Table 11 shows the details.

E. The performance of each detector group. We further evaluate the performance of each of the three detector groups in isolation. Table 12 presents their standalone performance on all datasets. The results starkly illustrate the necessity of an ensemble approach. No single

Datasets Train CADETS-E3 2018-04-03 ∼ 10 THEIA-E3 2018-04-02 ∼ 08 CLRSCP-E3 2018-04-03 ∼ 10 CADETS-E5 2019-05-08/09/11 THEIA-E5 2019-05-08/09/10 CLRSCP-E5 2019-05-08/09 OpTC-H201 2019-09-19 ∼ 21 OpTC-H501 2019-09-19 ∼ 21 OpTC-H051 2019-09-19 ∼ 21

Valid Test 2018-04-02 2018-04-06/11/12/13 2018-04-09 2018-04-10/12/13 2018-04-02 2018-04-11/12 2019-05-12 2019-05-16/17 2019-05-11 2019-05-14/15 2019-05-11 2019-05-14/15/17 2019-09-22 2019-09-23 ∼ 25 2019-09-22 2019-09-23 ∼ 25 2019-09-22 2019-09-23 ∼ 25

TABLE 10: Date-based partitioning of each dataset in (yyyymm-dd). Dates containing malicious attack scenarios in the test set are marked in bold. Dataset Name CADETS A1 A2 E3 A3 A1 THEIA E3 A2 A3 CLRSCP A1 E3 CADETS A1 E5 A2 THEIA A1 E5 A1 CLRSCP A2 E5 A3 A4 OpTC A1 OpTC A2 OpTC A3

Date(yymm-dd) Nginx Backdoor w/ Drakon In-Memory 18-04-06 Nginx Backdoor w/ Drakon In-Memory 18-04-12 Nginx Backdoor w/ Drakon In-Memory 18-04-13 Firefox Backdoor w/ Drakon 18-04-10 In-Memory Browser Extension w/ Drakon Dropper 18-04-12 Phishing E-mail w/ EXE Attachment 18-04-13 Firefox Backdoo w/ Drakon In-Memory 18-04-11

To evaluate robustness to a benign distribution shift, we rotate which day serves as the validation split on THEIA-E3 (four days, each taking one turn as validation). Results are in Table 13. P ROV F USION detects all three attack campaigns in every split, confirming that percentile-based fusion is not sensitive to the choice of a benign holdout. Training Days 3,4,5 Days 4,5,9 Days 3,5,9 Days 3,4,9

Validation TP FP F1 ADP MCC Det. Day 9 91 2 0.82 1.00 0.83 3/3 Day 3 110 12 0.88 1.00 0.88 3/3 Day 4 86 14 0.73 1.00 0.74 3/3 Day 5 95 3 0.84 1.00 0.84 3/3

TABLE 13: Benign-shift study on THEIA-E3.

Description

Nginx Drakon APT Nginx Drakon APT Firefox Drakon APT BinFmt-Elevate Inject Appstarter APK Micro APT Elevate Firefox Drakon APT Lockwatch APK Java APT Tester Micro APT BinFmt-Elevate Attack on OpTC-H201 Attack on OpTC-H501 Attack on OpTC-H051

19-05-16 19-05-17 19-05-15 18-05-15 19-05-17 19-05-17 19-05-17 19-09-23 19-09-24 19-09-25

TABLE 11: Mapping of shorthand attack IDs (A1, A2, . . . ) to their full descriptions. detector group achieves a satisfactory balance between detection rate and false positives. For instance, while Group 2 and Group 3 (i.e., Pairwise Corroboration detector group and Holistic Fusion detector group) identify most attacks, they suffer from an increased number of False alarms. Conversely, Group 1 (i.e., a Specialist detector group) generates fewer false alarms but fails to detect several stealthy attacks entirely. This demonstrates that relying on any single detector group, or a simple fusion node-centric and edge-centric method, is insufficient, motivating our ensemble design.

F. Ablation Study on Normalization Strategy The aggregate results are presented in Table 14. Across both the TC and OpTC families, only Percentile Normalization provides the necessary balance of sensitivity and precision required by our framework. Groups 1 2 3 ✓ ✓ ✓

G. Benign-Shift Stability Study

CADETS E3 E5 0/0 0/0 24/1 94/5 24/1 104/36

THEIA E3 E5 3/7 5/8 6/8 8/16 46/5894 4/17

CLRSCP E3 E5 5/0 6/1 14/41 10/18 11/32 13/85

TABLE 12: Analysis of detector groups with cell of (TP/FP).

H. Case Study: Reconstructing an Attack from Sparse Signals Org.mozilla.fennec-firefox-dev _287651 …

…

…

…

Org.mozilla.fennec-firefox-dev _287344 …

…

…

…

Org.mozilla.fennec-firefox-dev _279719 …

…

…

…

1. Connect 2. Sendto

166.199.230.185.80 _ 195543

3. Recvfrom 1. Connect 2. Sendto

166.199.230.185.80 _198077

3. Recvfrom 1. Connect 2. Sendto

166.199.230.185.80 _198786

3. Recvfrom

Figure 7: The attack subgraph generated by P ROV F USION for the CLEARSCOPE-E3 dataset, highlighting the key malicious events and their causal relationships. In this graph, circles represent FILE, squares represent PROCESS, and diamonds represent NETWORK nodes. The entity marked as red denotes true positives detected by P ROV F USION. Beyond aggregate metrics, we examine whether P ROVF USION produces actionable alerts for analysts. We use the CLEARSCOPE-E3 scenario because its attack surface is small: in our run, P ROV F USION surfaced only six truepositive nodes. A 1-hop expansion around them yields a single concise subgraph (Figure 7) centered on process org.mozilla.fennec_firefox_dev and its connection to 166.199.230.185:80, already suggesting a plausible exfiltration path. Per-Node Scoring Walkthrough. Table 16 traces the six true positives in Figure 7, plus two false alarms for contrast, through raw view scores, percentile-normalized scores, and the final seven-detector vote vector. To keep decisions inspectable, P ROV F U SION exposes a per-entity vote vector. For org.mozilla.fennec_firefox_dev, the vector is [1,0,1,1,1,0,1] with normalized scores [1.0, 0.3, 1.0], showing which internal detectors contributed to the alert. This allows analysts to verify cross-detector agreement (“5 of 7 voted anomalous”) and prioritize the associated edges and processes for follow-up. The resulting subgraph reveals a repeated pattern of file access followed by network transmission, observed three times, consistent with the ground truth that the adversary

Dataset Normalization Detection TPs TC Percentile 13/14 149 TC Min-Max Scaling 8/14 110 TC Z-Score Scaling 5/14 18 TC Robust Scaling 1/14 6 OpTC Percentile 3/3 21 OpTC Min-Max Scaling 2/3 2 OpTC Z-Score Scaling 0/3 0 OpTC Robust Scaling 2/3 5

ADP 0.926 0.848 0.791 0.650 0.980 0.786 0.223 0.897

Total FPs 37 2378 27 144 39 25 19 144

TABLE 14: Aggregate Performance of Normalization Strategies on TC and OpTC Datasets. Detectors D1 D2 D3 D4 D5 D6 D7 All

CADETS E3 E5 0/67 14/249 0/2 6/314 24/1 0/15 24/1 8/16 25/2 8/16 24/1 4/17 24/1 4/17 24/1 7/9

THEIA E3 E5 0/25 7/21 1/3882 0/30 95/36 12/47 91/2 14/41 94/5 14/41 91/360 1/31 94/7 11/32 91/2 11/2

CLRSCP E3 E5 38/2240 0/23 3/522 9/23 6/8 8/14 6/7 10/18 6/8 10/18 43/2445 6/79 9/3929 13/85 6/7 10/16

TABLE 15: A comparison of individual detectors (D1–D7) and the final fused output with each cell TP/FP. attempted exfiltration three times under unstable connectivity. Together, precise node selection and transparent voting provide a small, context-rich view that supports efficient attack reconstruction without producing many disconnected or low-information alerts.

I. The performance of Each Detector. The standalone performance of each of the seven detectors is presented in Table 15. A key observation is that no single detector strikes an acceptable trade-off between detection rate and the number of false positives, which highlights the limitations of individual heuristics and underscores the need for an ensemble method. For example, the D6 detector, a simple fusion of three views, produces a high volume of false positives. Conversely, the D4 detector achieves greater precision but at the cost of more false alarms. These results confirm that any single method is inadequate on its own, providing the primary motivation for developing our ensemble model.

J. Added New Ground-Truth Entities. As illustrated in Section §5.2.3, we found and added new true positives into the ground truth file. Here we proved a Table 18 that details the UUID, attribute, and timestamp for each added entity.

K. Detailed Runtime Performance Breakdown To support our analysis of computational overhead in Section 5.4, this section provides a detailed breakdown of the runtime components for P ROV F USION across all six datasets. All measurements were conducted on the same hardware (NVIDIA A100 GPU) and averaged over five runs. Table 17 provides definitive empirical evidence for our claim in Section 5.4.

L. Mimicry Attack Evaluation We reproduced the mimicry attack of Goyal et al. [71] to assess whether subgraph-level perturbations affect P ROV F U SION . The attacker inserted 1K–5K camouflage edges into

Raw UUID 287651 287344 TP 279719 195543 198077 198786 FP 115182 284778

SS 3.14 1.39 1.04 0.92 0.90 0.80 0.84 4.37

SA 0.10 0.10 0.10 0.41 0.41 0.41 0.05 0.43

Norm SC 1.56 1.59 1.56 1.56 1.59 1.56 1.57 0.93

SS′ 1.00 1.00 1.00 1.00 1.00 0.99 1.00 1.00

′ SA 0.30 0.30 0.30 1.00 1.00 1.00 0.09 1.00

′ SC 1.00 1.00 1.00 1.00 1.00 1.00 1.00 0.99

Votes 1011101 1011101 1011101 0011111 0011111 0011111 1011101 1101111

V 5 5 5 5 5 5 5 6

TABLE 16: Scoring trace. SS /SA /SC denote the attribute, structural, and causal view scores; primed symbols denote percentile-normalized scores. The vote vector lists D1 –D7 ; an alert fires when V ≥ Tv = 4. Training Time (s) Inference Time (s) Dataset GMAE Causal Total GNN Causal KNN Total Train MLP Train Embed Pred. Score Infer CA-E3 9.00 394.0 403.0 1.3 14.2 14.8 30.3 TH-E3 4.5 407.7 412.2 1.3 17.7 15.0 33.0 CL-E3 4.0 157.2 161.2 0.4 6.2 3.2 9.8 CA-E5 3190.5 797.4 3987.8 9.7 86.4 574.0 670.1 TH-E5 2100.1 416.0 2516.1 7.5 72.9 169.0 249.4 CL-E5 118.5 587.9 706.5 0.97 18.6 5.0 24.5

TABLE 17: Runtime breakdown for P ROV F USION, highlighting the computational cost of each component. malicious subgraphs to imitate benign motifs. As shown in Table 20, all original malicious entities remained correctly detected, while false positives slightly increased with higher perturbation levels. Overall, mimicry perturbations reduced MCC modestly but did not conceal true malicious behaviors, indicating that localized detectors in node/edge-centric paradigms remain resilient to subgraph-level manipulations.

M. Results under the Orthrus Ground Truth We repeated all evaluations using the original Orthrus ground truth to verify consistency (Table 19). The overall performance trends of P ROV F USION and all baselines remain unchanged, confirming that the additional ground-truth refinements do not alter the relative ranking or main conclusions. Minor variations in absolute metrics are observed, primarily due to the previously missing attacks.

N. Hyperparameter Settings for Comparative Anomaly Detection Algorithms Here we summarize the hyperparameter ranges explored for each model family mentioned in Section §5.3.2. All settings were applied under identical data partitions and four normalization schemes (min-max, z-score, robust, and percentile) for strict comparability. Linear Combination Models. All integer weight triplets (wattr , wstruc , wcausal )! ∈!0, . . . , 193 !\!(0, 0, 0) were enumerated, yielding 7,999 combinations. Probabilistic Models. Multivariate Gaussian (MVG): regularization coefficient added to the covariance diagonal regularizationcov ∈ 10−2 , 10−3 , 10−4 , 10−5 , 10−6 . Gaussian Mixture Model (GMM): number of mixture components ncomponents ∈ 1, 3, 5, 7, 10, 12, 15. Distance/Density-based Models. K-Nearest Neighbor (KNN): neighborhood size nneighbors ∈ 5, 10, 20, 30, 40, 50.

Dataset

UUID Attribute ED35A9B7-0200-0000-0000-000020 /home/admin/profile 273847BC-0200-0000-0000-000020 /home/admin/profile D037D3BA-0200-0000-0000-000020 /home/admin/profile D237D6BA-0200-0000-0000-000020 /home/admin/profile EE35ABB7-0200-0000-0000-000020 /home/admin/profile THEIA-E3 1D38E3BB-0200-0000-0000-000020 /home/admin/profile EC3598B7-0200-0000-0000-000020 /home/admin/profile 54387BBE-0200-0000-0000-000020 /home/admin/profile 223838BC-0200-0000-0000-000020 /home/admin/profile 62175519-0400-0000-0000-000020 /bin/bash (./tcexec) 0100D00F-2925-2E00-0000-1889CA01 /tmp/mozilla admin0/jAG iSHt.bin.part 00000000-0000-0000-000028A44E org.mozilla.fennec firefox dev 00000000-0000-0000-00002740D6 org.mozilla.fennec firefox dev 00000000-0000-0000-0000279530 128.55.12.166 | 45525 | 166.199.230.185 | 80 128.55.12.166 | 46162 | 166.199.230.185 | 80 Clearscope-E3 00000000-0000-0000-00002974E9 00000000-0000-0000-00002B1ADE 128.55.12.166 | 46316 | 166.199.230.185 | 80 00000000-0000-0000-000027871D 128.55.12.166 | 55331 | 111.82.111.27 | 80 00000000-0000-0000-00002974D0 128.55.12.166 | 55968 | 111.82.111.27 | 80 00000000-0000-0000-00002B1AB3 128.55.12.166 | 56122 | 111.82.111.27 | 80 00000000-0000-0000-0000F3D696 org.mozilla.fennec vagrant Clearscope-E5 516DD5D2-8D47-CB75-4EC4/data/data/org.mozilla.fennec vagrant/files/ 82FF9B5054BB mozilla/profiles.ini

Timestamp 2018-04-12 12:57 2018-04-12 12:57 2018-04-12 13:10 2018-04-12 13:10 2018-04-12 12:57 2018-04-12 13:15 2018-04-12 12:57 2018-04-12 13:26 2018-04-12 13:16 2018-04-13 14:06 2018-04-13 14:06 2018-04-11 14:12 2018-04-11 13:54 2018-04-11 13:54 2018-04-11 14:12 2018-04-11 14:20 2018-04-11 13:54 2018-04-11 14:11 2018-04-11 14:20 2019-05-17 11:57 2019-05-17 11:57

TABLE 18: The newly added True Positives. System Dataset TP↑ FP↓ TN↑ FN↓ F1↑ ADP↑ MCC↑ Dataset TP↑ FP↓ TN↑ FN↓ F1↑ ADP↑ MCC↑ Kairos 1 959 280.6K 67 0.00 0.00 0.00 0 6 3.14M 123 0 0.01 0 Magic 22 16.5K 265.0K 46 0.00 0.01 0.02 28 245.3K 2.89M 95 0.00 0.17 0.00 NodLink CADETS 18 34.3K 247.2K 50 0.00 0.49 0.01 CADETS 73 756.0K 2.38M 50 0.00 0.03 0.01 Flash 3 4.5K 277.0K 65 0.00 0.04 0.01 6 34961 3.10M 117 0.00 0.02 0.01 E3 E5 Orthrus 7 1 281.5K 61 0.18 0.81 0.30 1 8 3.14M 122 0.00 0.34 0.03 Velox 9 1 281.5K 59 0.23 0.97 0.35 0 2 3.14M 123 0.00 0.01 0.00 ProvFusion 24 1 281.5K 44 0.52 1.00 0.58 7 9 3.14M 116 0.10 0.92 0.16 Kairos 1 22 701.4K 117 0.01 0.11 0.02 0 7 1.86M 69 0.00 0.00 0.00 Magic 19 97.2K 604.2K 99 0.00 0.00 0.00 61 737.3K 1.12M 8 0.00 0.00 0.00 NodLink THEIA 26 258.3K 443.2K 92 0.00 0.01 0.00 THEIA 20 175.3K 1.68M 49 0.00 0.00 0.00 Flash 20 251.2K 450.2K 98 0.00 0.01 0.00 41 316.2K 1.54M 28 0.00 0.01 0.01 E3 E5 Orthrus 2 2 701.4K 116 0.03 0.33 0.09 1 31 1.86M 68 0.02 0.30 0.02 Velox 18 120 701.3K 100 0.14 0.89 0.14 2 63 1.86M 67 0.03 0.33 0.03 ProvFusion 82 11 701.4K 36 0.78 1.00 0.78 11 2 1.86M 58 0.27 1.00 0.37 Kairos 6 8.4K 102.9K 35 0.00 0.00 0.01 1 1 150.9K 50 0.04 0.37 0.10 Magic 37 8.4K 102.9K 4 0.01 0.01 0.06 7 8.2K 142.6K 44 0.00 0.00 0.01 NodLink CLRSCP 38 22.8K 88.5K 3 0.00 0.06 0.03 CLRSCP 3 27.0K 123.9K 48 0.00 0.00 0.00 Flash 32 11.1K 100.2K 9 0.01 0.03 0.04 17 76.0K 74.9K 34 0.00 0.01 0.00 E3 E5 Orthrus 1 9 111.3K 40 0.00 0.17 0.05 2 8 150.9K 49 0.07 0.21 0.09 Velox 1 625 110.7K 40 0.00 0.10 0.00 8 13 150.9K 43 0.22 0.50 0.24 ProvFusion 1 12 111.3K 40 0.04 0.50 0.04 9 17 150.9K 42 0.23 0.71 0.25

TABLE 19: Results under the Original Orthrus Ground Truth Added Edges 1000 3000 5000 True Positives (TP) 24 23 22 False Positives (FP) 9 31 34 MCC 0.51 0.38 0.36 ADP 0.92 1.00 1.00

TABLE 20: Performance of P ROV F USION under mimicry attacks with different numbers of added edges. Boundary-based Models. One-Class SVM (OC-SVM): error upper bound parameter ν ∈ 10−3 , 10−4 , 10−5 , 10−6 , 10−7 . Isolation Forest (iForest): number of estimators nestimators ∈ 50, 100, 150, 200, 300, 500, 700, 900. Reconstruction-based Model. Autoencoder: learning rate lr ∈ 10−2 , 10−3 , 10−4 , 10−5 , 10−6 . The best-performing configuration per dataset was reported in the main results (Table 4).

O. Multi-Hot vs. Per-Edge Encoding To verify that collapsing parallel events between the same entity pair into a single multi-hot edge does not degrade detection, we compare both encodings on THEIAE3 under identical settings. We choose THEIA-E3 because its larger pool of true positives makes any encoding-induced differences clearly observable. Encoding

TP

FP

TN

FN

F1

ADP

MCC

Multi-hot (Ours) Per-edge

91 108

2 19

701K 701K

38 21

0.82 0.84

1.00 1.00

0.83 0.84

TABLE 21: Multi-hot vs. per-edge encoding on THEIA-E3. The two encodings achieve essentially equivalent MCC; per-edge recovers a few extra TPs but at nearly 10× the FP cost, confirming that multi-hot is a sound efficiencymotivated choice rather than a detection compromise.

P. Meta-Review The following meta-review was prepared by the program committee for the 2026 IEEE Symposium on Security and Privacy (S&P) as part of the review process as detailed in the call for papers.

P.1. Summary P ROV F USION is a multi-view provenance-based intrusion detection framework that fuses anomaly signals from three views (attribute, structure, causality) using percentile normalization, monotonic linear detectors, and voting-based aggregation. It is evaluated on DARPA TC datasets and outperforms single-view baselines in detection accuracy and false-positive rates.

P.2. Scientific Contributions • •

Creates a New Tool to Enable Future Science. Provides a Valuable Step Forward in an Established Field.

P.3. Reasons for Acceptance 1) 2) 3)

Well-motivated problem with a compelling diagnosis of why single-view detectors fail in complementary ways. Clear empirical gains over both individual baselines and naive ensemble combinations, validated through extensive ablations. The system is lightweight and practical. The fusion and voting mechanism adds minimal overhead, and the evaluation shows improvements in runtime efficiency.

P.4. Noteworthy Concerns 1)

Reviewers flagged an area where finer detail would strengthen the paper: a per-campaign FP/FN breakdown.

Record · ID 18957 · SHA-256 4d8022a17acfe719
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.