arXiv:2609.08623v1 [cs.CR] 8 Sep 2026
Navigating the Latent Manifold: Proactive Concept Drift Adaptation for Resilient NIDS Chao Zha
Zifeng Kang
Tian Liu
[email protected] Zhejiang University Hangzhou, China Institute of Computing Technology, Chinese Academy of Sciences Beijing, China
[email protected] Beijing University of Posts and Telecommunications Beijing, China
[email protected] Institute of Agricultural Equipment, Zhejiang Academy of Agricultural Sciences Hangzhou, China
Dakun Shen
Ruyun Zhang∗
[email protected] Zhejiang University Hangzhou, China
[email protected] Shanghai AI Laboratory Shanghai, China
Abstract
1
Network intrusion detection systems (NIDS) are critical for cybersecurity, safeguarding services and data from potential attacks. However, existing AI-based NIDS often assume static data distributions and fail to handle concept drift, leading to degraded performance and increased false positives in dynamic network environments. To address this issue, we propose DriftXpert† , a novel NIDS for drift-adaptive detection. Specifically, we propose a decoupled twostage offline adaptive framework. In Phase 1, we introduce an unsupervised anomaly metric based on latent manifold deviation. By performing outlier analysis within the latent space, the framework achieves high-sensitivity detection of network traffic concept drift. In Phase 2, to mitigate catastrophic forgetting under non-stationary distributions, we design a representation consistency alignment strategy. This strategy constrains the feature mapping between the legacy model and the drifted distribution, ensuring the model captures emerging attack characteristics while retaining discriminative power over known patterns. Furthermore, we incorporate crossepoch neuron weight aggregation and selective freezing mechanisms to enable fine-grained knowledge transfer in the parameter space, effectively balancing model plasticity and stability. Extensive experiments on public datasets demonstrate that DriftXpert effectively adapts to drifted data without catastrophic forgetting. Furthermore, real-world evaluations on enterprise network further confirm its robustness and practical applicability, contributing to improved security protection for millions of users.
Network intrusion detection systems are essential for modern network security, protecting environments and providing reliable digital spaces. NIDS monitors system traffic or activity in real time, detecting threats such as malware, DDoS attacks, and data breaches. Traditional anomaly or signature-based methods struggle with unknown threats [12, 24, 33, 44, 53], while AI-based approaches have shown superior performance [23, 28, 71, 74, 78–80]. Classical machine learning methods, such as clustering, decision trees, and random forests, have been applied to classify malicious traffic [19– 21, 36, 65, 67, 75], while reconstruction mechanisms and transfer learning help detect unknown attacks [37, 54, 62, 73, 77, 82]. Despite strong performance in traditional classification tasks [29, 45, 60], AI algorithms often underperform in real-world NIDS due to unique domain challenges [64]. Security research should emphasize understanding the nature of security problem over merely improving benchmark scores. Many AI-NIDS studies rely on static traffic assumptions [69, 73, 75], which are suitable for controlled experiments but do not reflect practical scenarios [16, 59]. In reality, network traffic evolves over time with changing statistical properties and inter-feature relationships, which is called concept drift [22]. Concept drift severely degrades detection performance, affecting the model’s ability to recognize emerging threats and evolving patterns [3]. Furthermore, concept drift can bias decision making [40], increasing false positives and false negatives. To assess the impact of concept drift, we conducted experiments on the CICIDS-2017 dataset, as presented in Figure 1.
CCS Concepts • Do Not Use This Code → Generate the Correct Terms for Your Paper.
Keywords Intrusion detection, Concept drift, Proactive sensing, Evolutionary alignment, Adaption, Catastrophic forgetting.
∗ Corresponding author. † Github repository: https://github.com/cc-sec-hub/DriftXpert
Introduction
Addressing concept drift should include two key components: drift detection and model adaptation. Model adaption should enable the updated model to fit new distributions while mitigating catastrophic forgetting, which is critical as historical patterns often reoccur and forgetting may increase false positives. Drift detection acts as the trigger for model adaption, ensuring that adaptation is performed only after drift is identified. Concept drift in NIDS arises from both temporal and spatial factors [81]. Temporally, evolving services and emerging traffic patterns induce distribution shifts, while storing historical data is often impractical due to resource
Trovato et al.
(a) Local drift.
(b) Global drift.
Figure 1: Evaluation of concept drift. We select four different existing methods, train them on the historical data, and then test them on the drifted data on the CICIDS-2017 dataset.
constraints. Spatially, deploying models across different environments further exacerbates drift, while reusing or sharing historical network traffic for adaptation is often infeasible, as such data may contain sensitive user and organizational information and introduce risks of privacy leakage and unauthorized exposure. Consequently, the drift detection and the model adaption should be studied under the assumption of no access to historical data. On the other hand, as real-world network traffic is massive, rendering manual labeling costly and impractical. Moreover, labeling inherently requires data analysis, during which drift can already be observed. Therefore, unsupervised drift detection is both necessary and inevitable in practical scenarios. However, existing drift detection methods typically rely on labeled drift data or directly incorporate historical data. For example, statistical approaches [7, 14, 30, 47] based on KL divergence or local density changes inherently assume such access, violating the aforementioned practical constraints and limiting their applicability in real-world settings. Moreover, hyperparameter threshold selection remains challenging and often requires careful tuning. From the perspective of model adaptation, existing mask-perturbation-based approaches heavily rely on historical data [76], limiting their practicality in real-world settings. Meanwhile, model update methods [27, 48, 70, 81, 83] that do not require historical data (e.g., constrained fine-tuning and generative replay) suffer from unstable training dynamics and suboptimal performance or introduce additional security risks, such as data poisoning. Our work. We propose DriftXpert, a unified and problemdriven framework for autonomous NIDS adaptation under evolving traffic distributions. Unlike prior work that separately studies drift detection, incremental learning, or knowledge preservation, DriftXpert tightly integrates drift awareness with model evolution. Specifically, we introduce two clustering-based drift indicators, CI and OR, with theoretical analysis to enable label-free drift detection and use detected drift as an adaptation criterion to trigger model updates only under distribution shifts. To address catastrophic forgetting, we further design representation consistency alignment, together with weight aggregation and selective freezing, to balance adaptation and knowledge preservation without requiring historical traffic replay or storage. Extensive experiments on public benchmarks and real-world enterprise traffic demonstrate the effectiveness of DriftXpert in long-term adaptive NIDS deployment. Contributions. This study makes the following contributions.
• We propose a unified drift-aware NIDS adaptation framework that integrates unsupervised drift detection and autonomous model evolution, rather than treating them as independent problems. • We propose an unsupervised drift detection mechanism based on latent representation deviation, where two clusteringbased indicators, Confidence Interval and Outlier Ratio, are introduced with theoretical analysis to enable effective and interpretable concept drift identification. • We design a drift-driven model adaptation framework that performs representation consistency alignment under evolving distributions, together with weight aggregation and selective freezing, to balance adaptation and knowledge preservation without relying on historical traffic replay.
2 Overview 2.1 Background Concept drift in NIDS refers to the evolution of the statistical properties of network traffic over time, leading to a mismatch between training and test data. Most NIDS models assume a stationary distribution, which rarely holds in real-world environments due to changing user behaviors, service updates, and evolving attack strategies. Consequently, learned decision boundaries become outdated, degrading model performance and increasing false positives or false negatives. A common manifestation of concept drift is shifts in benign traffic distributions, where even subtle changes can significantly affect prediction results by altering feature distributions or inter-feature relationships, often leading to a sharp increase in false positives in practical deployments. Handling concept drift typically involves two components: drift detection and model adaptation. Drift detection identifies distribution changes and triggers model updates. Model adaptation then updates the model to fit new data while preserving historical knowledge. Therefore, an effective approach must balance adaptability and stability. Specifically, concept drift is defined as a change in the joint probability distribution of the samples 𝑥 (from the feature space 𝑋 ) and labels 𝑦 (from the label space 𝑌 ), denoted 𝑃 (𝑥 ∈ 𝑋, 𝑦 ∈ 𝑌 ) = 𝑃 (𝑥, 𝑦). As in 𝑃 (𝑥, 𝑦) = 𝑃 (𝑥) · 𝑃 (𝑦|𝑥), two drift paradigms can be defined as follows. DEFINITION 1 (LOCAL DRIFT). The local drift is the change of 𝑃 (𝑥 ∈ 𝑋𝐿 , 𝑦) = 𝑃 (𝑥) · 𝑃 (𝑦|𝑥 ∈ 𝑋𝐿 ), 𝑋𝐿 represents the subspace of characteristics where the drift occurs (e.g., a specific attack). 𝑃 his (𝑦|𝑥) ≠ 𝑃drift (𝑦|𝑥), ∀𝑥 ∈ 𝑋𝐿 𝑃 his (𝑦|𝑥) = 𝑃drift (𝑦|𝑥), ∀𝑥 ∈ 𝑋
(1)
where 𝑃his and 𝑃drift (𝑦|𝑥) denote the conditional probability distributions in the historical and drifted distributions, respectively. DEFINITION 2 (GLOBAL DRIFT). The global drift is the change of 𝑃 (𝑥 ∈ 𝑋, 𝑦) = 𝑃 (𝑥) · 𝑃 (𝑦|𝑥 ∈ 𝑋 ), denoted 𝑃his (𝑥, 𝑦) ≠ 𝑃 drift (𝑥, 𝑦), ∀𝑥 ∈ 𝑋 . There can be two sources of drift: • Global Concept Drift: 𝑃his (𝑦|𝑥) ≠ 𝑃drift (𝑦|𝑥), ∀𝑥 ∈ 𝑋 (e.g., novel attack emerges, the original decision boundary to classify attacks becomes invalid, requiring a relearning of attack patterns.) • Global Covariate Shift: 𝑃his (𝑥) ≠ 𝑃drift (𝑥), ∀𝑥 ∈ 𝑋 (e.g., the statistical characteristics of traffic undergo a global change, leading to a distributional difference between historical and drifted data.)
Navigating the Latent Manifold: Proactive Concept Drift Adaptation for Resilient NIDS
2.2
Challenges
2.2.1 C1: Limited Capability for Unsupervised Drift Detection. The occurrence of concept drift indicates that the model requires updating to adapt to new network environments. Supervised drift detection methods depend on extensive manual inspection of prediction results, making them costly and impractical. Consequently, unsupervised approaches offer a more viable solution in real-world settings [25]. Detecting concept drift serves as a critical initial step for model updates. This should be considered a qualitative issue that can be addressed through quantitative analysis. Furthermore, concept drift detection does not need to be performed in real-time; instead, it should involve periodic monitoring of current network traffic conditions to identify potential drift. 2.2.2 C2: Inadequate Adaptation to Drifted Data. When concept drift is detected, model updating becomes necessary to adapt to new data. Retraining is often considered the ideal approach, as it typically achieves optimal performance when real data are available. However, in resource-constrained or privacy-sensitive scenarios, storing or accessing historical data is often impractical. Fine-tuning [26] provides an alternative, enabling model updates without relying on historical data. Although it can handle local drift effectively, it may struggle under global drift and lead to performance degradation. Therefore, designing efficient model update methods without historical data remains a key challenge for applying AI to NIDS. 2.2.3 C3: Catastrophic Forgetting of Historical Data. Model updates can lead to a critical issue: achieving accurate detection of recent data at the cost of significantly reduced detection performance in historical data, a phenomenon known as catastrophic forgetting [34, 39]. Although concept drift occurs within the current network traffic space, historical data does not completely vanish. Within a certain timeframe, it remains essential to retain the ability to detect historical data. Balancing improved detection performance for drifted data while preserving detection capabilities for historical data without access to historical datasets represents another key challenge in applying AI to NIDS.
2.3
• Evasion via distribution shift. An adversary may craft traffic patterns that mimic benign distribution changes, causing the model to misclassify malicious traffic as benign. • False positive amplification. Drift in benign traffic may significantly increase false alarms, degrading system usability in real deployments. • Forgetting-induced vulnerability. Improper model updates may erase previously learned attack patterns, allowing known attacks to bypass detection. Note that effective drift detection is indispensable in this context. On the one hand, a system may not currently exhibit significant drift but could become vulnerable as traffic evolves over time. On the other hand, an adversary may implicitly exploit existing drift patterns to bypass detection. Therefore, robust NIDS requires both reliable drift detection and stable model adaptation mechanisms.
3
Design
In this section, we delineate the underlying motivation behind the design of DriftXpert and provide a comprehensive exposition of its methodological details.
3.1
Design Motivation
3.1.1 Problem Statement. The application of AI to NIDS scenarios can be conceptualized as a classification task. Let 𝑝 (𝑋 ) represent the predicted distribution of the model, 𝑞(𝑋 ) denote the true label distribution, and optimize using the loss of cross entropy, which can be expressed as 𝐻 (𝑝, 𝑞) = 𝐻 (𝑝) + 𝐷 𝐾𝐿 (𝑝 ||𝑞), where 𝐻 (𝑝, 𝑞) represents the cross-entropy, 𝐻 (𝑝) denotes the entropy of 𝑝 (𝑋 ), 𝐷 𝐾𝐿 (𝑝 ||𝑞) refers to the Kullback-Leibler (KL) divergence.
Threat Model
Our threat model assumes a dynamic network environment in which data drift naturally occurs (excluding adversarial drift), and adversaries may exploit such distribution shifts to degrade a deployed NIDS. We consider a setting where the model is trained on historical data, while incoming traffic gradually deviates from the original distribution due to evolving benign behaviors or adaptive attack strategies. If this drift is not properly addressed, the model may produce incorrect predictions, leading to increased false positives or missed detections. We define two types of risks in this setting. First, undetected drift refers to concept drift that occurs without being identified, leading directly to performance degradation. Second, catastrophic forgetting arises when model updates fail to preserve historical knowledge, potentially amplifying detection errors. When both occur simultaneously, it constitutes a critical failure scenario, where the system loses both adaptability and stability. 2.3.1 In-scope Threats. We consider three in-scope threats of this study as follows.
(a) 𝑝 (𝑋 ).
(b) 𝑧.
Figure 2: Trends in the variations of predicted probability distribution and latent variables. The predict probability density function 𝑝 (𝑋 ) can be expressed as a Gaussian distribution G(𝜇𝑝 , 𝜎𝑝 ) (more detailed theoretical analysis can be found in Appendix D). If the intrusion detection task is considered a binary classification problem, the positive label distribution 𝑞(𝑋 ) should be treated as a set 𝑞(𝑋 ) = {1, ..., 1}. Based on above classification optimization objective, where 𝐷 𝐾𝐿 represents the KL divergence [68] between 𝑝 (𝑋 ) and 𝑞(𝑋 ), which quantifies the closeness between the two distributions, and 𝑝 (𝑋 ) follows a Gaussian distribution G, the optimal result is that the predictive distribution of the classifier is given by: 1 1 exp − 2 (𝑋 − 𝜇𝑝 ) 2, 𝜇𝑝 → 1, 𝜎𝑝 → 0. (2) 𝑝 (𝑋 ) = √︃ 2𝜎 2 𝑝 2𝜋𝜎 𝑝
Trovato et al.
We further consider the impact of concept drift, where the prediction probability for positive labels decreases, leading to misclassification as the negative class. This results in an erroneous classification outcome. In this case, the predicted probability distribution 𝑝 ′ (𝑋 ) for the drifted data can be expressed as: 𝑝 ′ (𝑋 ) = √︃
1 2𝜋𝜎𝑝2′
exp −
1 (𝑋 − 𝜇𝑝 ′ ) 2, 𝜇𝑝 ′ → 0, 𝜎𝑝 ′ → 1. 2𝜎𝑝2′
(3)
Compared to Equation (2), Equation (3) illustrates that the impact of concept drift manifests itself as a degradation in predictive performance, with changes in the probability distribution depicted in Figure 2a. Based on this analysis, our goal is to attribute and interpret the effect of concept drift on AI models for intrusion detection, identifying patterns of drift and the corresponding feature changes in the model. However, the inherent nonlinearity and opacity of deep learning models make them challenging to interpret, a well-known difficulty in the academic community. To address this, we shift our focus to reverse attribution - inferring possible changes in latent variables by analyzing variations in prediction outcomes. The model architecture widely used in the current study consists of an encoder with multiple fully connected layers and a classifier with a single fully connected layer [42, 61]. By examining the changes in the predicted probability distributions, our aim is to perform inverse causal inference on the behavior of the latent variables 𝑧 produced by the encoder. Since 𝑧 and 𝑝 (𝑋 ) exhibit a linear relationship, this approach significantly reduces the complexity of the attribution process. The predicted probability and the latent variables are connected through a single fully connected layer, and their relationship can be expressed as: 𝑝 (𝑋 ) = 𝑊 · 𝑧 (𝑋 ) + 𝑏, (4) 𝑝 ′ (𝑋 ) = 𝑊 · 𝑧 ′ (𝑋 ) + 𝑏, where 𝑝 (𝑋 ) and 𝑧 (𝑋 ) represent the predicted probability and latent variables of historical data, and 𝑝 ′ (𝑋 ) and 𝑧 ′ (𝑋 ) represent the predicted probability and latent variables of the drifted data. Since the relationship between the predicted probability 𝑝 (𝑋 ) and the latent variables 𝑧 is linear, combining Equation (2) and Equation (3), we can infer that the variation pattern of 𝑧 and 𝑧 ′ should be approximately similar to that of 𝑝 and 𝑝 ′ . The pattern of variation of latent variables 𝑧 is also depicted graphically in Figure 2b. From the perspective of the encoder, the drifted data is not included in its training process. As the encoder extracts features from historical data to achieve classification, this results in differences in the distribution of latent variables, as illustrated in Figure 2b. Similarly, when a new model is trained using drifted data to identify new samples, the historical data may exhibit drift relative to the recently collected data. The attribution results remain consistent with the aforementioned analysis. 3.1.2 Heuristic Insights. The analytical results mentioned above provide crucial insight that informs the architectural design of DriftXpert. Specifically, these findings lead to the following design principles: • As illustrated in Figure 2b, the evolution of drifted data distributions within the latent space exhibits strong congruence with the fluctuations in predictive probability distributions. Since the
representation process in the latent space is label-independent, this distributional heterogeneity provides a pivotal entry point: we can achieve high-sensitivity, unsupervised drift detection by monitoring distribution shifts directly within the latent representation space. • Figure 2b suggests that aligning the latent distributions of drifted and historical data via distribution overlap or subspace embedding enables the model to generalize across temporal domains. Since the output layer is typically linear, if a weighted combination of drifted samples yields a prediction close to 1, historical features in the same subspace will similarly produce near-unity outputs, providing a mathematical guarantee of consistent detection during model evolution.
3.2
System Architecture
We illustrate the proposed DriftXpert in Figure 3, a lightweight and robust real-time NIDS. The system is structured into the following components: • Online Task (Intrusion Detection). Given the resource constraints and the real-time requirements, we designed a lightweight model structure. The structure consists of a multi-layer fully connected stacked encoder and a classifier made of a single fully connected layer. The classifier outputs a binary result: Benign or Attack. • Phase 1 (Proactive Sensing). As shown in Figure 3, phase 1 continuously monitors traffic within a fixed window at regular intervals (C1). Once drift is detected, the system triggers a model update to mitigate its impact. This process runs in the offline mode without real-time constraints, as detailed in § 3.3. • Phase 2 (Evolutionary Alignment). As shown in Figure 3, upon detecting concept drift, phase 2 employs representation alignment, cross-epoch weight aggregation, and selective freezing to facilitate new knowledge acquisition (C2) while mitigating catastrophic forgetting (C3). This task also runs in the offline mode without real-time requirements, as detailed in § 3.4. From a holistic perspective, DriftXpert employs two phases to accomplish model updates: drift detection continuously monitors network traffic for data drift, and once detected, triggers model adaptation. Through model alignment, a new NIDS model is trained and replaces the historical one, achieving adaptability while preserving robustness to historical data.
3.3
Phase 1: Proactive Sensing
As discussed in § 3.1.2, using the consistency between latent variable changes and predicted probability changes (see § 3.1.1 and Figure 2), we propose detecting drift through latent space dynamics. Specifically, we develop a deep clustering-based detection framework [5, 6, 81] as shown in Figure 3. By reusing the historical model’s encoder and performing outlier analysis on latent variables, this model achieves robust unsupervised drift detection. To construct the deep clustering model, we first freeze all encoder parameters and then perform clustering in the latent variable space 𝑍 . Various approaches can be used to implement the clustering model [72], we choose the K-Means algorithm [49]. This algorithm iteratively optimizes an objective function by assigning data points to 𝐾 clusters, minimizing the sum of squared distances between the data points and their respective cluster centers. The
Navigating the Latent Manifold: Proactive Concept Drift Adaptation for Resilient NIDS
Figure 3: The overall architecture of DriftXpert. The system operates across three main components: an Online Task for real-time intrusion detection using a lightweight encoder-classifier structure. Phase 1 for continuous drift detection and deep 1 and clustering within traffic windows, and Phase 2 for model adaptation. Phase 2 employs a dual-objective optimization (○) 2 to integrate new knowledge while mitigating catastrophic forgetting. model aggregation (○) clustering model can be represented as 𝑓𝐶 : 𝑧 → [0, 1] 𝐾 , 𝐶 denotes the cluster centers, and 𝑌 = {1, 2, ..., 𝐾 } represents the set of labels. Let 𝐷 (𝑧, 𝐶𝑧 ) = ||𝑧 − 𝐶𝑧 || 2 denote the distance of latent variables 𝑧 to their corresponding cluster centers 𝐶𝑧 . Based on 𝐷 (𝑧, 𝐶𝑧 ), we define two critical metrics, the confidence interval and the outlier ratio, to detect concept drift. Confidence Interval (𝐶𝐼 ). 𝐶𝐼 defines a reliable distance range. When the distance 𝐷 (𝑧, 𝐶𝑧 ) is less than 𝐶𝐼 , the data point is considered reliable, indicating that no drift has occurred. In contrast, when 𝐷 (𝑧, 𝐶𝑧 ) exceeds 𝐶𝐼 , the data point is classified as an outlier, suggesting that drift has taken place. The mathematical definition of 𝐶𝐼 is as follows: 𝐶𝐼 = E𝑧∼𝑍ℎ [𝐷 (𝑧, 𝐶𝑧 )] + 𝛼 · S𝑧∼𝑍ℎ [𝐷 (𝑧, 𝐶𝑧 )],
(5)
where 𝑍ℎ represents the latent variable space of historical data, E𝑧∼𝑍ℎ and S𝑧∼𝑍ℎ denote the expectation and standard deviation of the distances from each sample point in 𝑍ℎ to 𝐶, and 𝛼 denotes an adjustable hyperparameter. Outlier Ratio (𝑂𝑅). 𝑂𝑅 is defined as the ratio of outliers in 𝑍𝑟 within the recent data window to outliers in 𝑍ℎ within the historical data window, under the same 𝐶𝐼 . The mathematical definition of 𝑂𝑅 is as follows: 1 Í𝑁 𝑍𝑟 𝑖=0 I{𝐷 (𝑧, 𝐶𝑧 ) ∉ [0, 𝐶𝐼 ]} 𝑁 𝑍𝑟 𝑂𝑅 = , (6) 1 Í𝑁𝑍ℎ 𝑗=0 I{𝐷 (𝑧, 𝐶𝑧 ) ∉ [0, 𝐶𝐼 ]} 𝑁𝑍 ℎ
where 𝑁𝑍𝑟 and 𝑁𝑍ℎ denote the number of samples in 𝑍𝑟 and 𝑍ℎ , respectively. 𝐼 {𝐴} = 1 if and only if condition 𝐴 is true. When the denominator is 0, we set it to 1 for numerical stability. This only decreases the calculated 𝑂𝑅 and does not affect drift detection, since a smaller 𝑂𝑅 already satisfies the drift criterion under our detection rule. Once 𝑂𝑅 is calculated, we can determine whether drift has occurred by comparing it with a threshold 𝑇 . The threshold 𝑇 is not treated as a model hyperparameter requiring sensitivity analysis. Since 𝑂𝑅 = 1 naturally indicates equal outlier ratios in the recent
and historical windows, 𝑂𝑅 > 1 already provides evidence of drift. The choice of 𝑇 only determines the detector’s operating sensitivity and should therefore be specified according to the application’s tolerance for drift and the relative costs of false alarms and missed detections. This module also involves two key hyperparameters, 𝐾 and 𝛼. Since the encoder is optimized for the classification task, we recommend that 𝐾 be within the range [𝐿, 3𝐿], where 𝐿 represents the total number of classification labels. A more detailed theoretical analysis of 𝛼 will be provided in Appendix E. Moreover, our method operates in the latent variable rather than in the original data, where the encoder incorporates knowledge from historical and drift data. This enables the drift detection model to be updated using only drift data, preventing catastrophic forgetting.
3.4
Phase 2: Evolutionary Alignment
As shown in Figure 2b, latent variable variations exhibit trends consistent with predicted probability distributions. Consistent with § 3.1.2, our method is based on latent space synchronization. By embedding historical latent space into the drifted subspace, DriftXpert enables adaptive learning while preserving prior knowledge. Based on this, we propose the following design strategies. Model Alignment. Figure 3 illustrates our proposed model update process. To align the distribution of latent variables of the drifted data with that of the historical data, we employ contrastive learning [35, 43]. We maintain three models: the current epoch model, the previous epoch model, and the historical model, producing latent variables 𝑧, 𝑧𝑝𝑟𝑒𝑣 , and 𝑧ℎ𝑖𝑠 , respectively. The goal is to align the current latent variables with the historical ones while distancing them from the previous epoch, formalized through the following contrastive loss. L𝑐𝑜𝑛 = − log
exp (sim(𝑧, 𝑧ℎ𝑖𝑠 )/T ) , (7) exp (sim(𝑧, 𝑧𝑝𝑟𝑒𝑣 )/T ) + exp (sim(𝑧, 𝑧ℎ𝑖𝑠 )/T )
where sim(𝑎, 𝑏) represents the cosine similarity between two vectors 𝑎 and 𝑏, and T denotes the temperature coefficient.
Trovato et al.
Model Aggregation. Our design philosophy aims to approximate the latent variable distribution of historical data to that of drifted data, thereby addressing the catastrophic forgetting problem while simultaneously learning new knowledge. However, during model updates, the training data consists primarily of the most recent data. As a result, the output layer focuses not only on the important features of the latent space corresponding to historical data, but also on those of the recent data. To mitigate this, we aggregate the historical model with the updated model, allowing the model to converge more effectively and minimize the loss of knowledge due to forgetting. The aggregation process of the models is as follows. Wℎ𝑖𝑠 = 𝛽 · Wℎ𝑖𝑠 + (1 − 𝛽) · W𝑐𝑢𝑟 ,
(8)
where Wℎ𝑖𝑠 and W𝑐𝑢𝑟 represent the weights of the historical model and the model of the current epoch, respectively. 𝛽 is a hyperparameter. Input-Layer Weight Freezing. To mitigate catastrophic forgetting, we employ a parameter-freezing strategy that keeps the weights of the initial encoder layers unchanged during adaptation. Since early layers typically encode general and transferable features, freezing them helps preserve previously learned representations. Meanwhile, the higher layers remain trainable to capture new patterns, enabling effective adaptation while reducing the risk of forgetting. Optimizing Objective. Beyond aligning the latent space, classification on drifted data is performed using cross-entropy loss [51] to ensure convergence. The loss function is defined as follows. 𝑛 ∑︁ L𝑐𝑒 = − 𝑞(𝑋 ) log 𝑝 (𝑋 ), (9) 𝑖=1
where 𝑞(𝑋 ) represents the true label distribution and 𝑝 (𝑥) denotes the predicted label distribution. Finally, we optimize the model update objective using two loss functions (Equation (7) and Equation (9). Specifically, our goal is to learn new knowledge while ensuring that the latent variable distribution of the drifted data approximates the latent variable distribution of the historical model, thus preventing catastrophic forgetting. Our optimization objective can be expressed as follows. L𝑜𝑏 𝑗 = L𝑐𝑒 + 𝛾 · L𝑐𝑜𝑛 ,
(10)
where 𝛾 is a hyperparameter.
4 Experimental Setup 4.1 Implementation 4.1.1 Environment. We deployed DriftXpert on PC with the following main configuration: 13th Gen Intel® Core™ i5-13500 CPU, Windows 11, 16GB DDR5 5200Hz RAM, and 1TB SSD. Our prototype is implemented in Python, the phase 1 is built using scikit-learn [57], and the phase 2 is built using PyTorch [56]. 4.1.2 Public Datasets. We used two widely public datasets, CICIDS2017 [17, 63] and CICIDS-2018 [18, 63], to simulate real-world scenarios. Using CICFlowMeter [41], we extracted 80 features, and retained 50 after evaluation, removing around 30 for the following reasons: i) Features causing "cheat effects" (e.g., 5-tuples). ii) Features with excessive outliers that could mislead results. iii) Zerodominant binary features that are confirmed to have negligible impact on performance through experiments. Furthermore, we also
addressed common issues, such as duplicate packets and labeling errors [15]. Our evaluation encompasses two data drift paradigms, local drift and global drift, to rigorously assess the system’s adaptability across diverse distributional shift scenarios. Local drift affects only specific regions of the data, while others remain stable, typically caused by distribution shifts in certain subsets. In contrast, global drift alters the entire distribution, often due to environmental changes or the emergence of new attacks. We partitioned the data sets to simulate global and local drift scenarios (more details on Table 15 in Appendix G.2). Since drift is distribution-related rather than time-related, we applied distributionbased rather than time-based partitioning. • For local drift, we partition the data with the same label into historical and drift data by distribution-based sampling. • For global drift, several new attack traffic differing from historical data is considered drift data. Historical and current data were divided into training and test sets, and the experiments used only the training set to keep the test set unseen. 4.1.3 Real-world Datasets. We collected real-world network traffic over several consecutive months in early 2026 from the external network of a large-scale enterprise. True labels are generated by the existing security system and then manually verified by domain experts to ensure high data quality and minimal labeling noise. The dataset includes diverse and representative attack types, such as command injection, directory traversal, and SQL injection, ensuring strong realism and coverage. For evaluation, the data are partitioned monthly: earlier traffic is treated as historical data, while later traffic is regarded as newly observed data. This setup enables systematic assessment of performance degradation under dynamic environments and the effectiveness of the proposed update mechanism. Due to the continuous evolution of real-world systems, the two stages exhibit not only significant shifts in benign traffic distributions but also previously unseen attack patterns, forming a more challenging concept drift scenario. Moreover, the enterprise serves millions of users, resulting in highly concurrent and complex traffic patterns. Consequently, the dataset faithfully reflects real operational environments, providing a solid foundation to evaluate the robustness, adaptability, and deployability. 4.1.4 Baselines. We selected five methods for comparison to validate the effectiveness of our method, as detailed below. • Mlp-FineTuning (M.F.). No parameter freezing, and the new model is fine-tuned over 5 epochs using the drifted dataset. • Mlp-LayerFreezed-FT (M.LF.FT). The first encoder layer is frozen, and the new model is fine-tuned for 5 epochs. • OWAD @NDSS23. Han et al. [27] proposed weighting model parameters to restrict updates on important ones, preserving old knowledge while adapting new distributions. • A-NIDS @TIFS25. Zha et al. [81] proposed a method that uses CTGAN to replay historical data, combining it with the latest data to update the model. • SSF @INFOCOM25. Zhang et al. [83] proposed SSF, an incremental method that captures concept drift by selecting representative new samples and discarding outdated ones.
Navigating the Latent Manifold: Proactive Concept Drift Adaptation for Resilient NIDS
4.1.5 Metrics. We used five metrics to evaluate DriftXpert, including five commonly used metrics: Accuracy (AC), Precision, Recall, F1-Score, and True Positive Rate (TPR) [31].
5
Research Questions
Our research questions (RQs) in this section are as follows. • RQ1: How effective is DriftXpert in accurately identifying concept drift within dynamic network traffic streams? • RQ2: What extent does DriftXpert outperform SOTA methodologies in terms of model adaptation efficiency? • RQ3: What are the individual contributions of the key components within DriftXpert to the overall detection performance? • RQ4: Can DriftXpert effectively align latent representation spaces across different temporal stages? • RQ5: How sensitive is the performance of DriftXpert to variations in key hyperparameters, such as 𝛼, 𝛽, 𝛾?
5.2
RQ1: Proactive Sensing Results by DriftXpert
To evaluate the efficacy of the drift detection module, we conducted a suite of evaluation experiments in four distinct scenarios derived from two benchmark datasets. We quantified the presence of outliers for both historical and drifted telemetry within each cluster and calculated the corresponding outlier ratio. The empirical results are detailed in Table 1 and Table 2. Experimental results indicate that outlier divergence is more pronounced within isolated data clusters. Across four distinct scenarios, our proposed deep clustering-based drift detection methodology effectively identifies concept drift. Crucially, as the primary objective is to serve as a high-fidelity trigger for phase 2, these findings empirically corroborate the module’s functional integrity. Furthermore, this component incorporates two critical hyperparameters: the cluster count 𝐾 and the outlier threshold coefficient 𝛼. While 𝐾 can be precisely estimated leveraging historical data and clustering heuristics, the sensitivity of 𝛼 is systematically analyzed in § 5.6, with its optimal value derivation and corresponding mathematical simulations provided in Appendix E.
5.3
Outliers / Samples of i-th Cluster
Datasets 1
Benchmark Evaluation on Public Datasets
In this section, we evaluate the performance of DriftXpert through comparative experiments on two publicly available datasets, complemented by a series of analytical studies, including ablation experiments and hyperparameter sensitivity analyses.
5.1
Table 1: Local Drift Detection Results on Public Datasets.
RQ2: Performance of Evolutionary Alignment vs. Baselines
To evaluate our proposed method, we conducted comparative experiments with two fine-tuning-based methods and three SOTA methods [27, 81, 83] under two different drift scenarios in the CICIDS2017 and CICIDS-2018 dataset. As shown in Table 3, Table 4, Table 5 and Table 6, Model-FineTuning achieves high recall under local drift on CICIDS-2017 (99.86% benign, 86.22% attack) without forgetting, but its attack recall drops sharply to 10.37% under global drift. Model-LayerFreezed-FT improves attack recall by 14% under local drift and increases it to 24.40% under global drift, but remains overall inferior. In CICIDS-2018, both fine-tuning methods adapt
2
3
4
5
2017-Set1 2017-Set2 OR
686/4,277 84/551 15,727/21,331 860/860 4.60 6.60
99/692 15/259 766/3,195 356/3,528 1.68 1.74
0/1 1/2 -
2018-Set1 2018-Set2 OR
22/2,049 773/12,928 5.57
5/988 1/277 16/15,269 10/3,100 -
0/629 1/3,106 -
19/3,791 23/23 -
- OR is not reported for clusters with negligible outlier counts due to lack of representative significance.
Table 2: Global Drift Detection Results on Public Datasets. Outliers / Samples of i-th Cluster
Datasets 1
2
3 4/407 35/178 -
4
2017-Set1 2017-Set2 OR
169/6,936 1/1,948 91,90/45,676 50/7,187 8.26 13.55
2018-Set1 2018-Set2 OR
37/7,095 1/2,696 100/3,717 1/97 4,988/8,184 4,298/38,978 843/13,486 9/1,325 116.87 297.28 2.32 -
5
12/1,044 0/1,670 0/2,238 0/741 0/1,486 0/1,466 -
well to recent data under local drift (recall >98%) but suffer severe catastrophic forgetting, with only slight mitigation from layer freezing; similar trends persist under global drift. Among SOTA methods, OWAD shows limited adaptability (64.30% recall for drifted anomalies in CICIDS-2017), A-NIDS improves detection via generative replay but still lags behind, and SSF preserves historical knowledge yet fails on drifted anomalies (30.41% recall), further degrading under global drift. On CICIDS-2018, OWAD struggles with both forgetting and adaptation, A-NIDS sacrifices historical knowledge for recent performance, and SSF exhibits the opposite behavior. Overall, DriftXpert consistently achieves a better balance between adaptability and stability, outperforming all baselines across datasets. To further address the severe class-imbalance issue, we additionally report the Macro-F1 score on both the historical and drifted datasets. As illustrated in Figure 4, most baseline methods exhibit a pronounced adaptation–retention gap, indicating an inherent tradeoff between learning new traffic patterns and preserving historical knowledge. Conventional fine-tuning methods fail to adequately adapt to the drifted distribution, whereas SSF achieves stronger adaptation at the cost of substantial degradation on historical data, revealing severe catastrophic forgetting. In contrast, DriftXpert consistently achieves the smallest adaptation–retention gap across all four experimental settings, demonstrating its ability to effectively accommodate concept drift while maintaining balanced recognition performance on previously learned traffic.
Trovato et al.
Table 3: Comparative Experiments (Benign%/ Malicious%) in the CICIDS-2017 Dataset on the Local Drift Scenario. Set 1 (historical data)
Methods Model-FineTuning Model-LayerFreezed-FT OWAD @NDSS23 A-NIDS @TIFS25 SSF @INFOCOM25 DriftXpert (Ours)
Set 2 (recent data)
Precision.
Recall.
F1-Score.
Precision.
Recall.
F1-Score.
98.67%/ 98.31% 99.99%/ 94.87% 99.91%/ 33.31% 99.97%/ 98.78% 99.77%/ 99.40% 99.87%/ 97.09%
99.86%/ 86.22% 99.47%/ 99.99% 82.27%/ 99.16% 99.89%/ 99.73% 99.94%/ 97.65% 99.71%/ 98.72%
99.26%/ 91.87% 99.74%/ 97.37% 90.23%/ 49.47% 99.93%/ 99.25% 99.86%/ 98.52% 99.79%/ 97.90%
99.66%/ 98.22% 98.87%/ 96.27% 96.01%/ 26.10% 99.16%/ 97.04% 93.32%/ 81.91% 99.70%/ 98.88%
99.83%/ 96.51% 99.68%/ 87.98% 82.50%/ 64.30% 99.73%/ 91.14% 99.31%/ 30.41% 99.89%/ 97.06%
99.74%/ 97.36% 99.27%/ 91.94% 88.74%/ 37.13% 99.45%/ 94.00% 96.22%/ 44.35% 99.79%/ 97.96%
Table 4: Comparative Experiments (Benign%/Malicious%) in the CICIDS-2017 Dataset on the Global Drift Scenario. Set 1 (historical data)
Methods Model-FineTuning Model-LayerFreezed-FT OWAD @NDSS23 A-NIDS @TIFS25 SSF @INFOCOM25 DriftXpert (Ours)
Set 2 (recent data)
Precision.
Recall.
F1-Score.
Precision.
Recall.
F1-Score.
50.13%/ 75.39% 51.60%/ 65.45% 82.10%/ 97.02% 72.11%/ 96.21% 98.89%/ 91.86% 78.84%/ 97.22%
96.38%/ 10.37% 86.22%/ 24.40% 97.33%/ 80.42 97.33%/ 64.29% 90.81%/ 99.02% 97.70%/ 75.35%
65.95%/ 18.23% 64.56%/ 35.55% 89.07%/ 87.94% 82.85%/ 77.07% 94.67%/ 95.31% 87.20%/ 84.90%
98.40%/ 95.06% 97.12%/ 93.72% 51.31%/ 47.86% 98.56%/ 98.61% 60.94%/ 56.32% 98.89%/ 97.64%
95.34%/ 98.30% 94.05%/ 96.96% 63.58%/ 35.66% 98.74%/ 98.42% 60.32%/ 56.96% 97.81%/ 98.80%
96.85%/ 96.65% 95.56%/ 95.31% 56.79%/ 40.87% 98.65%/ 98.52% 60.63%/ 56.64% 98.35%/ 98.22%
Table 5: Comparative Experiments (Benign%/Malicious%) in the CICIDS-2018 Dataset on the Local Drift Scenario. Set 1 (historical data)
Methods Model-FineTuning Model-LayerFreezed-FT OWAD @NDSS23 A-NIDS @TIFS25 SSF @INFOCOM25 DriftXpert (Ours)
Set 2 (recent data)
Precision.
Recall.
F1-Score.
Precision.
Recall.
F1-Score.
50.70%/ 0.01% 50.69%/ 0.01% 50.96%/ 0.01% 51.48%/ 48.48% 99.99%/ 99.99% 87.47%/ 97.70%
98.21%/ 0.01% 98.17%/ 0.01% 98.24%/ 0.01% 98.87%/ 1.14% 99.99%/ 99.99% 98.08%/ 85.29%
66.88%/ 0.01% 66.86%/ 0.01% 67.11%/ 0.01% 67.98%/ 2.24% 99.99%/ 99.99% 92.47%/ 91.07%
99.95%/ 97.91% 99.97%/ 97.50% 56.51%/ 11.27% 99.97%/ 97.77% 99.97%/ 0.00% 99.99%/ 97.54%
98.44%/ 99.93% 98.15%/ 99.97% 97.86%/ 0.36% 98.29%/ 99.98% 98.29%/ 0.00% 98.08%/ 99.99%
99.19%/ 98.91% 99.05%/ 98.72% 71.65%/ 0.70% 99.12%/ 98.85% 73.10%/ 0.00% 99.03%/ 98.75%
Table 6: Comparative Experiments (Benign%/Malicious%) in the CICIDS-2018 Dataset on the Global Drift Scenario. Set 1 (historical data)
Methods Model-FineTuning Model-LayerFreezed-FT OWAD @NDSS23 A-NIDS @TIFS25 SSF @INFOCOM25 DriftXpert (Ours)
5.4
Set 2 (recent data)
Precision.
Recall.
F1-Score.
Precision.
Recall.
F1-Score.
45.68%/ 83.89% 81.08%/ 83.94% 43.84%/ 51.90% 70.29%/ 92.28% 98.73%/ 92.69% 79.98%/ 99.07%
97.58%/ 9.82% 78.91%/ 85.69% 98.34%/ 1.40% 92.34%/ 70.11% 91.59%/ 98.91% 99.03%/ 80.73%
62.23%/ 17.58% 79.98%/ 84.80% 60.64%/ 2.73% 79.82%/ 79.68% 95.02%/ 95.70% 88.49%/ 88.97%
96.78%/ 98.84% 96.31%/ 92.18% 51.57%/ 28.02% 98.21%/ 99.60% 62.51%/ 58.60% 97.11%/ 99.11%
98.95%/ 96.44% 92.52%/ 96.13% 97.35%/ 1.12% 99.64%/ 98.02% 61.52%/ 59.62% 99.19%/ 96.83%
97.85%/ 97.63% 94.38%/ 94.11% 67.43%/ 2.15% 98.92%/ 98.80% 62.01%/ 59.10% 98.14%/ 97.96%
RQ3: Ablation Study
We conducted ablation experiments in the CICIDS-2017 [17, 63] and CICIDS-2018 [18, 63], and presented in Table 7, Table 8, Table 9 and Table 10. Specifically, the following nomenclature defines the ablation variants, delineating the correspondence between the model configurations, their sub-modules, and the data scopes.
• D.X-With-H . Consistent with the DriftXpert model architecture, the training dataset used is Set1 (historical data). • D.X-With-D. Consistent with the DriftXpert model architecture, the training dataset used is Set2 (drifted data). • D.X-Only-C. Consistent with the DriftXpert model architecture, only contrastive learning is retained as the adaptation strategy.
Navigating the Latent Manifold: Proactive Concept Drift Adaptation for Resilient NIDS
Table 8: Ablation Studies (Benign%/Malicious%) in the CICIDS-2017 Dataset on the Global Drift Scenario.
Methods
(a) CICIDS-2017 (Local Drift).
(b) CICIDS-2017 (Global Drift).
D.X-With-H D.X-With-D D.X-Only-C D.X-EnFre D.X-NoFre D.X-NoAgg DriftXpert
5.5 (c) CICIDS-2018 (Local Drift).
(d) CICIDS-2018 (Global Drift).
Figure 4: Comparative Results (Macro-F1) vs. Baselines. • D.X-EnFre. All encoder parameters are frozen, the new model is trained in Set2 with contrastive loss and aggregation. • D.X-NoFre. Without parameters freezing, the new model is trained in Set2 using contrastive loss with aggregation. • D.X-NoAgg. The first layer parameters are frozen, the new model is trained in Set2 with contrastive loss, without aggregation. In the ablation study, D.X-Static-H and D.X-Static-R exhibit high intra-distribution accuracy but suffer from catastrophic forgetting or poor adaptability. D.X-Only-C demonstrates strong adaptability to new knowledge, however, its ability to mitigate catastrophic forgetting remains limited. D.X-EncFreezed partially preserves historical knowledge by freezing the encoder, yet its adaptability remains limited. Further comparisons with D.X-NoFreezed and D.XNoAggregation confirm that first-layer freezing is essential for feature retention, while the aggregation strategy is vital for drifted traffic detection. Consequently, DriftXpert optimizes the trade-off between adaptability and stability. It maintains approximately 98% precision and recall on drifted data and high historical recall on the CICIDS-2017 dataset, effectively mitigating catastrophic forgetting.
Methods D.X-With-H D.X-With-D D.X-Only-C D.X-EnFre D.X-NoFre D.X-NoAgg DriftXpert
Set 1 (historical data)
Set 2 (recent data)
Recall.
F1-Score.
Recall.
F1-Score.
99.99/99.99 99.99/0.01 99.89/72.48 99.97/99.69 99.92/80.15 98.24/99.99 99.71/98.72
99.99/99.99 95.34/0.01 99.91/83.50 99.97/99.69 99.00/88.60 99.11/91.72 99.79/97.90
99.97/29.27 99.98/97.95 99.98/98.58 99.90/32.06 99.90/98.69 99.79/96.04 99.89/97.06
96.64/45.18 99.89/98.87 99.91/99.19 96.68/48.21 99.89/98.88 99.70/96.95 99.79/97.96
Set 2 (recent data)
Recall.
F1-Score.
Recall.
F1-Score.
98.59/99.98 99.83/0.57 92.13/19.27 76.21/81.62 92.49/31.74 97.90/69.35 97.70/75.35
99.28/99.34 65.21/1.14 66.16/30.44 77.62/80.08 69.67/45.75 84.87/80.96 87.20/84.90
79.40/40.34 98.53/97.83 99.29/99.24 72.04/78.37 99.29/99.17 97.96/98.81 97.81/98.80
67.73/49.60 98.29/98.09 99.26/99.20 74.86/75.61 99.27/99.19 98.44/98.27 98.35/98.22
RQ4: Alignment Analysis of Latent Space
To quantify the contribution of the alignment mechanism during phase 2, we performed a quantitative analysis and visualized the latent space distributions of both drifted and historical traffic. By comparing these distributions across the historical and updated model states, we aim to elucidate the pivotal role of this mechanism in mitigating distribution shifts. The results are illustrated in Figure 5 and Figure 6. In the latent space of the historical model, the distributions of historical data and drifted data exhibit a certain degree of separation, which is particularly pronounced in the local drift scenario. In contrast, although the separation under global drift is less evident than that of local drift, observable distribution differences can still be identified from the box plots. After model adaption, the historical and drifted data demonstrate a more consistent alignment. From the perspective of density distributions, the proportion of aligned samples between the two data types can be observed more intuitively. It should be noted that the two distributions are unlikely to fully overlap in visualization, however, achieving effective alignment within certain subspaces is sufficient to support the adaptive capacity of the model. Table 9: Ablation Studies (Benign%/Malicious%) in the CICIDS-2018 Dataset on the Local Drift Scenario.
Methods Table 7: Ablation Studies (Benign%/Malicious%) in the CICIDS-2017 Dataset on the Local Drift Scenario.
Set 1 (historical data)
D.X-With-H D.X-With-D D.X-Only-C D.X-EnFre D.X-NoFre D.X-NoAgg DriftXpert
5.6
Set 1 (historical data)
Set 2 (recent data)
Recall.
F1-Score.
Recall.
F1-Score.
99.52/99.97 99.94/ 0.01 98.56/0.01 91.83/ 0.01 98.56/ 0.01 98.08/85.56 98.08/85.29
99.75/99.74 67.66/ 0.01 67.04/0.01 63.92/ 0.01 67.04/ 0.01 92.59/91.23 92.47/91.07
99.15/0.01 99.97/99.99 98.37/99.97 90.51/98.10 98.48/99.99 97.60/99.97 98.08/99.99
72.58/ 0.01 99.99/99.98 99.16/98.88 94.26/92.71 99.23/98.99 98.77/98.42 99.03/98.75
RQ5: Sensitivity Analysis of Hyperparameter
In this section, we conducted a comprehensive sensitivity analysis on three key hyperparameters, 𝛼, 𝛽, and 𝛾, to evaluate the model’s
Trovato et al.
Table 10: Ablation Studies (Benign%/Malicious%) in the CICIDS-2018 Dataset on the Global Drift Scenario.
Methods D.X-With-H D.X-With-D D.X-Only-C D.X-EnFre D.X-NoFre D.X-NoAgg DriftXpert
Set 1 (historical data)
Set 2 (recent data)
Recall.
F1-Score.
Recall.
F1-Score.
99.82/99.80 61.97/72.00 92.58/37.63 4.71/99.99 97.04/25.70 99.71/69.01 99.03/80.73
99.78/99.83 62.60/71.44 67.87/52.49 8.99/72.97 66.32/40.15 83.24/81.55 88.49/88.97
94.43/25.77 99.88/97.33 99.94/97.42 64.28/73.11 99.76/97.65 99.26/96.69 99.19/96.83
71.84/39.10 98.71/98.58 98.82/98.66 68.21/68.71 98.81/98.68 98.12/97.92 98.14/97.96
(a) D.X-Static-H (Local Drift).
5.6.1 Sensitivity Analysis of 𝛼. We conducted sensitivity analysis on 𝛼 using CICIDS-2017 [17, 63] and CICIDS-2018 [18, 63]. The clustering model uses 5 clusters, and a minimum number of outliers is required to ensure a reliable OR, since drift detection is qualitative and OR alone is insufficient. As 𝛼 increases from 0, historical outliers decrease while the OR first increases and then decreases, as shown in Figure 7. The maximum occurs when the derivative of outliers with respect to 𝛼 is zero, consistent with Appendix E, indicating the inflection point as an appropriate choice of 𝛼.
(a) Local Drift (CICIDS-2017).
(b) Global Drift (CICIDS-2017).
(c) Local Drift (CICIDS-2018).
(d) Global Drift (CICIDS-2018).
(b) DriftXpert (Local Drift).
Figure 7: Sensitivity Analysis of Hyperparameter 𝛼. (c) D.X-Static-H (Global Drift).
(d) DriftXpert (Global Drift).
Figure 5: Alignment Analysis on the CICIDS-2017 Dataset.
(a) D.X-Static-H (Local Drift).
(b) DriftXpert (Local Drift).
(c) D.X-Static-H (Global Drift).
(d) DriftXpert (Global Drift).
Figure 6: Alignment Analysis on the CICIDS-2018 Dataset.
performance stability and robustness. Regarding the hyperparameter 𝐾, we omit a separate sensitivity evaluation as its optimal value can be empirically estimated from the clustering profiles of historical data, providing a principled basis for its selection.
5.6.2 Sensitivity Analysis of 𝛽. The hyperparameter 𝛽 is used to control the aggregation weight between the historical model and the updated model, where 𝛽 = 0.5 indicates equal contributions from both. Based on this, we conducted a sensitivity analysis by selecting seven values within the range of 0.35 to 0.65, as shown in Figure 8. The experimental results demonstrate that, under different values, the model performance remains relatively stable in both local and global drift scenarios. This indicates that setting 𝛽 around 0.5 can effectively ensure the robustness of the model. 5.6.3 Sensitivity Analysis of 𝛾. The hyperparameter 𝛾 controls the weight of the alignment regularization term. We performed a sensitivity analysis by selecting seven values within the range of 0.7 to 1.3, as shown in Figure 9. The experimental results indicate that, under different values, the model performance remains relatively stable in two drift scenarios. This suggests that setting 𝛾 around 1.0 effectively ensure the robustness of the model.
6
Real-world Test on Enterprise Network
For real-world validation, we deployed DriftXpert within a major corporation’s infrastructure, utilizing empirical datasets from their external-facing services for training and performance monitoring.
6.1
Research Questions
Our RQs in this section are as follows. • RQ1: Does DriftXpert maintain its efficacy in identifying drift when deployed within real-world enterprise infrastructures?
Navigating the Latent Manifold: Proactive Concept Drift Adaptation for Resilient NIDS
distributions across individual clusters. Specifically, in the cluster exhibiting the most significant drift, the outlier count for drifted telemetry reached 1,926 out of 10,326 samples, compared to only 39 out of 4,265 for historical data. This yields a localized OR of 20.4, providing strong empirical evidence of localized distribution shifts and highlighting the imperative for adaptive model updates. (a) Local Drift (Historical).
(b) Local Drift (Drifted).
(c) Global Drift (Historical).
(d) Global Drift (Drifted).
Figure 8: Sensitivity Analysis of Hyperparameter 𝛽 on the CICIDS-2017 Dataset. Figure 10: Proactive Sensing in the Real-world Environment.
6.3 (a) Local Drift (Historical).
(b) Local Drift (Drifted).
(c) Global Drift (Historical).
(d) Global Drift (Drifted).
Figure 9: Sensitivity Analysis of Hyperparameter 𝛾 on the CICIDS-2017 Dataset. • RQ2: Can DriftXpert demonstrate robust adaptability to drifted telemetry while concurrently mitigating catastrophic forgetting? • RQ3: Can DriftXpert achieve robust performance under encrypted traffic conditions, while still outperforming baseline methods?
6.2
RQ1: Proactive Sensing in the Real-world Environment
Initially, we conducted drift detection experiments within an enterprise network environment, with the results summarized in Figure 10. The experimental curves indicate that as 𝛼 increases, the outlier count in historical telemetry progressively declines. The optimal configuration is identified at the inflection point (𝛼 ≈ 3.5), where the outlier ratio reaches its peak. This selection process is entirely deterministic, relying solely on historical data for parameter calibration. With 𝛼 fixed at 3.5, we further analyzed the outlier
RQ2: Evolutionary Alignment in the Real-world Environment vs. Baselines
In our real-world enterprise network evaluation, we benchmarked DriftXpert against three SOTA baselines, including OWAD [27], A-NIDS [81], and SSF [83], as illustrated in Table 11. The results demonstrate that DriftXpert achieves a superior F1-Score of 99.43% and 97.10% on the drifted data set, significantly outperforming all baselines and manifesting an exceptional capacity for adaptation of new knowledge. Crucially, while existing models such as OWAD suffer from acute performance degradation due to catastrophic forgetting (with an F1-Score that plummets to 15.22%), DriftXpert maintains a robust recall of 96.12% on historical data. This underscores our model’s ability to concurrently capture emerging threat patterns and preserve foundational knowledge, providing a decisive advantage in continuous learning for dynamic environments.
6.4
RQ3: Evaluation on Encrypted Environment vs. Baselines
Given that encrypted payloads eliminate plaintext visibility, assessing detection integrity under encryption is critical for real-world deployment. As shown in Table 12, our comparative analysis reveals that DriftXpert consistently outperforms OWAD, A-NIDS, and SSF in encrypted scenarios. While DriftXpert secures a 0.88 ACC and 0.94 TPR on baseline distributions and preserves an 0.85 ACC post-drift, competing models suffer from severe performance degradation. Specifically, the accuracy of A-NIDS falls to 0.71 after drift, while OWAD and SSF do not provide reliable detection. These findings underscore the efficacy of our latent alignment strategy, which enables DriftXpert to navigate complex encrypted traffic evolution with a superior balance between plasticity and stability compared to main baselines.
Trovato et al.
Table 11: Real-world Test on the Enterprise Network (Benign% / Malicious%). Set 1 (historical data)
Methods D.X-With-H D.X-With-D OWAD @NDSS23 A-NIDS @TIFS25 SSF @INFOCOM25 DriftXpert (Ours)
Set 2 (recent data)
Precision.
Recall.
F1-Score.
Precision.
Recall.
F1-Score.
98.33%/ 90.49% 93.11%/ 48.37% 75.59%/ 26.74% 97.41%/ 48.34% 90.72%/ 75.65% 98.45%/ 61.16%
96.75%/ 94.96% 70.88%/ 83.88% 90.48%/ 10.63% 67.43%/ 94.44% 92.96%/ 69.69% 80.14%/ 96.12%
97.53%/ 92.67% 80.49%/ 61.36% 82.37%/ 15.22% 79.69%/ 63.95% 91.82%/ 72.55% 88.36%/ 74.75%
97.80%/ 91.98% 99.29%/ 95.94% 83.12%/ 13.89% 99.37%/ 94.80% 92.80%/ 88.18% 99.64%/ 96.08%
98.43%/ 89.06% 99.21%/ 96.35% 93.06%/ 5.59% 98.94%/ 96.88% 98.38%/ 61.26% 99.22%/ 98.14%
98.11%/ 90.49% 99.25%/ 96.15% 87.81%/ 7.97% 99.15%/ 95.83% 95.51%/ 72.29% 99.43%/ 97.10%
Table 12: Evaluation on Encrypted Environment. Datasets
Methods
ACC
F1
TPR
Historical
OWAD @NDSS23 A-NIDS @TIFS25 SSF @INFOCOM25 DriftXpert (Ours)
0.29 0.87 0.64 0.88
0.39 0.93 0.77 0.93
0.24 0.93 0.64 0.94
Drifted
OWAD @NDSS23 A-NIDS @TIFS25 SSF @INFOCOM25 DriftXpert (Ours)
0.32 0.71 0.54 0.85
0.21 0.82 0.60 0.89
0.13 0.92 0.49 0.90
7 Discussion 7.1 Foundation Models and Drift Adaptation As Large Language Models (LLMs) demonstrate remarkable potential in zero-shot task [4, 8, 9, 55], a fundamental question naturally arises: Can the inherent robustness of general-purpose foundation models entirely obviate the need for explicit concept drift adaptation? In this section, we provide a critical discussion. Protocol v.s. Linguistic Semantics. Network traffic consists of unstructured bitstreams governed by rigorous state-machine logic and spatio-temporal dependencies, fundamentally diverging from natural language semantics. Although LLMs capture macroscopic patterns, they lack the perceptual precision to sense fine-grained protocol shifts, such as subtle variations in proprietary industrial protocols. By quantifying latent manifold deviation, DriftXpert models the underlying geometry of traffic representations, capturing essential distribution shifts more acutely than the probabilistic heuristics of LLMs. The Plasticity-Stability Dilemma. Even foundation models require fine-tuning to maintain precision when encountering new attacks or environmental shifts. However, full-parameter tuning is not only computationally prohibitive but also suffers from catastrophic forgetting, where adapting to new patterns inevitably compromises the discriminative power over legacy attack signatures. DriftXpert addresses this via a decoupled two-stage framework, employing representation consistency alignment and selective neuron freezing. This approach provides a surgical and controllable path for model evolution that is far more precise than vanilla fine-tuning, thereby ensuring the long-term integrity of the defensive system.
Operational Bottlenecks and Deployment Overhead. NIDS are typically deployed on high-throughput backbone links, requiring microsecond-level inference latency. The prohibitive computational overhead and hardware redundancy necessitated by LLMs are impractical for most production environments, such as the power grid infrastructure evaluated in this study. In contrast, DriftXpert maintains a lightweight architecture, significantly lowering deployment thresholds and energy consumption while meeting the stringent real-time requirements of industrial networks. Crucially, our approach is not antithetical to the rise of LLMs. We envision a profound synergistic potential between foundation models and task-specific frameworks like DriftXpert. In future autonomous security ecosystems, LLMs can function as high-level commanders that leverage reasoning capabilities for cross-domain threat modeling and strategic orchestration. Conversely, DriftXpert acts as a front-line scout residing at the high-speed network edge. Its primary role is to perform sub-millisecond real-time sensing and robust distribution alignment. This hierarchical integration ensures a defensive posture that is both strategically comprehensive and operationally agile.
7.2
Drift Detection and Annotation Efficiency
We first evaluate the drift detection performance under both local and global drift scenarios using Macro-Precision (Macro-Prc), Macro-Recall (Macro-Rec), and Accuracy, as summarized in Table 13. The results demonstrate that DriftXpert achieves effective drift identification in both scenarios, with Macro-Recall values of 84.29% and 81.77% for local and global drift, respectively, indicating that most drifted samples can be successfully detected. The relatively lower Macro-Precision (50.07%) under local drift is mainly attributed to the inherent class imbalance, where only a small fraction of samples undergo distribution changes while the majority remain consistent with the historical distribution. Table 13: Evaluation of Drift Detection. Scenarios
Macro-Prc
Macro-Rec
Accuracy
Local Drift Global Drift
0.5007 0.8181
0.8429 0.8177
0.8107 0.7869
With accurate drift detection, DriftXpert can substantially reduce the annotation overhead by requiring manual labeling only for
Navigating the Latent Manifold: Proactive Concept Drift Adaptation for Resilient NIDS
detected drift samples, rather than continuously annotating all incoming traffic. For example, in our enterprise network deployment, we collected two months of traffic data and observed that only 1,357 out of 42,875 sampled flows in the second month exhibited distribution shifts compared with the previous month. This observation indicates that, in practical scenarios, drifted samples typically constitute only a small fraction. In real-world networks, traffic evolution is generally gradual, and drifted samples rarely dominate the data stream within a short period unless major events such as largescale business changes or system migrations occur. Such events are usually foreseeable, allowing operators to proactively adjust data collection and annotation strategies. Therefore, with reliable drift detection, DriftXpert effectively minimizes unnecessary annotation efforts during continuous model maintenance.
8
Related Work
In this section, we summarize previous related research efforts, with detailed information provided as below. Detection Methods of Concept Drift. Statistics-based methods are widely used for concept drift detection [25]. Examples include using density changes in active learning as a drift trigger [7], Treap-based Kolmogorov-Smirnov tests for distribution comparison [14], density estimation for local or global drift [47], pseudo-error rates from dynamic classifiers [58], distance measures like Kullback-Leibler (KL) divergence [30], and trainable detectors based on Restricted Boltzmann Machines [40]. However, in NIDS scenarios, since historical data is not retained, direct comparison of historical data with drifted data is not applicable [14, 30, 47]. Furthermore, model updates must be considered, as continuous stacking of classifiers can introduce additional problems [58]. Methods based on distance, among others, require specific hyperparameters, and selecting appropriate hyperparameters is a key challenge [40]. In contrast, our method can be implemented without access to historical data and does not rely on labeled data. Moreover, the selection of hyperparameters is supported by solid theoretical foundations and an automated analysis strategy. ML-Based Intrusion Detection Methods. Machine learning algorithms have been extensively applied to intrusion detection. For example, Yang et al. [73] proposed a conditional probability model for detecting known traffic, complemented by a reconstruction mechanism for zero-day attacks. Fu et al. [19] employed discrete Fourier transform to extract frequency-domain features, combined with clustering models for detection. Yang et al. [75] introduced an ensemble tree model for known attack detection alongside a clustering-based approach for zero-day scenarios. Wang et al. [69] developed a cloud-based method integrating a Stacked Contractive Auto-Encoder with Support Vector Machine (SVM). Fu et al. [21] proposed an unsupervised point cloud analysis method that models traffic feature vectors as high-dimensional points, aggregates them into voxels, and leverages voxel density to reduce false positives. Mirsky et al. [54] combined stacked autoencoders with reconstruction techniques to identify malicious traffic, while King et al. [38] utilized temporal graph representations and messagepassing neural networks for anomaly detection. Ding et al. [10] further introduced generative adversarial networks to mitigate class imbalance. Despite achieving strong performance on benchmark
datasets, these approaches largely overlook concept drift in realworld environments. Consequently, they implicitly assume a static traffic distribution, under which model effectiveness is only temporary. Once drift occurs, performance deteriorates significantly, often resulting in a substantial increase in false positives. Adaptation Methods for Concept Drift. The objective of model adaptation is to update models upon drift detection, enabling them to accommodate new data while mitigating catastrophic forgetting. Existing approaches can be broadly categorized into ensemble learning [13] and incremental learning [52]. For instance, Wang et al. [70] proposed constructing new subclassifiers and adjusting their weights to adapt to dynamic data; however, this method lacks explicit drift detection and depends on continuous labeling. Liu et al. [48] introduced an instance-based ensemble approach that dynamically selects classifiers under different drift scenarios. Andresini et al. [2] leveraged clustering for drift detection and employed a neighbor-based strategy to generate pseudo-labels for model updates, though these pseudo-labels are susceptible to noise and may induce self-poisoning. Yang et al. [76] generated drift-aware perturbation vectors via masking, aligned them with historical representations, and fine-tuned a constrained classifier using selected historical samples; however, this approach relies on access to historical data and lacks a robust drift detection mechanism. Han et al. [27] assigned importance weights to model parameters, restricting updates to critical parameters to preserve prior knowledge while adapting less significant ones to new distributions. Zhang et al. [83] proposed SSF (Strategic Selection and Forgetting), an incremental NIDS method that captures drift by selecting representative new samples and discarding outdated ones. Zha et al. [81] utilized CTGAN to replay historical data and combine it with recent data for model updates, though such generation inevitably incurs information loss, potentially degrading performance. Lian et al. [46] proposed leveraging uncertainty to identify drift-induced anomalies. However, because their method cannot learn new benign patterns, it may generate a large number of false positives during natural drift, such as that caused by business upgrades. Unlike most prior work that separately focuses on drift detection, incremental learning, or knowledge preservation, we propose a unified and problem-driven adaptation framework rather than a simple combination of existing techniques. Specifically, we introduce two clustering-based drift indicators, CI and OR, with theoretical analysis to validate their effectiveness and guide hyperparameter selection, enabling unsupervised drift detection. The detected drift serves as an adaptation criterion to trigger updates only under significant distribution shifts. Motivated by the analysis in § 3.1, we further introduce contrastive learning to align representations between historical and evolving traffic distributions. To balance adaptation and catastrophic forgetting, we employ weight aggregation and first-layer freezing to preserve stable knowledge. Unlike existing methods that rely on historical data replay or storage, our framework requires no access to historical traffic and performs label-free drift detection, reducing storage and annotation costs in long-term deployment.
Trovato et al.
9
Conclusion
We propose a robust and lightweight NIDS, DriftXpert. We begin with a backward-attribution analysis of concept drift, common in IDS research, to identify variation patterns in the latent space. Based on this, we introduce a deep clustering method for drift detection and a model contrastive approach to align latent spaces of historical and drifted data, mitigating catastrophic forgetting during updates. We provide a theoretical analysis of drift detection hyperparameters and propose a reliable selection method using image-based simulations. Finally, we validate the performance of drift and intrusion detection in two public datasets, analyze latent space distributions to confirm our attribution, and conduct latency and throughput tests on a low-end PC, showing the suitability of DriftXpert for resource-constrained deployment.
Acknowledgments This work was supported by the National Key R&D Program of China under Grant No. 2023YFB2904000, No. 2023YFB2904001 and No. 2024YFE0203800.
References [1] Shun-Ichi Amari. 1998. Natural gradient works efficiently in learning. Neural computation 10, 2 (1998), 251–276. [2] Giuseppina Andresini, Feargus Pendlebury, Fabio Pierazzi, Corrado Loglisci, Annalisa Appice, and Lorenzo Cavallaro. 2021. Insomnia: Towards conceptdrift robustness in network intrusion detection. In Proceedings of the 14th ACM workshop on artificial intelligence and security. 111–122. [3] Firas Bayram, Bestoun S Ahmed, and Andreas Kassler. 2022. From concept drift to model degradation: An overview on performance-aware drift detectors. Knowledge-Based Systems 245 (2022), 108632. [4] Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901. [5] Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze. 2018. Deep clustering for unsupervised learning of visual features. In Proceedings of the European conference on computer vision (ECCV). 132–149. [6] Jianlong Chang, Lingfeng Wang, Gaofeng Meng, Shiming Xiang, and Chunhong Pan. 2017. Deep adaptive image clustering. In Proceedings of the IEEE international conference on computer vision. 5879–5887. [7] Albert França Josuá Costa, Régis Ant&nio Saraiva Albuquerque, and Eulanda Miranda Dos Santos. 2018. A drift detection method based on active learning. In 2018 international joint conference on neural networks (IJCNN). IEEE, 1–8. [8] Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Shuming Ma, Zhifang Sui, and Furu Wei. 2023. Why can GPT learn in-context? language models secretly perform gradient descent as meta-optimizers. In Findings of the Association for Computational Linguistics: ACL 2023. 4005–4019. [9] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers). 4171–4186. [10] Hongwei Ding, Yu Sun, Nana Huang, Zhidong Shen, and Xiaohui Cui. 2023. TMGGAN: Generative adversarial networks-based imbalanced learning for network intrusion detection. IEEE Transactions on Information Forensics and Security 19 (2023), 1156–1167. [11] Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio. 2017. Sharp minima can generalize for deep nets. In International Conference on Machine Learning. PMLR, 1019–1028. [12] Bo Dong and Xue Wang. 2016. Comparison deep learning method to traditional methods using for network intrusion detection. In 2016 8th IEEE international conference on communication software and networks (ICCSN). IEEE, 581–585. [13] Xibin Dong, Zhiwen Yu, Wenming Cao, Yifan Shi, and Qianli Ma. 2020. A survey on ensemble learning. Frontiers of Computer Science 14 (2020), 241–258. [14] Denis Moreira Dos Reis, Peter Flach, Stan Matwin, and Gustavo Batista. 2016. Fast unsupervised online drift detection using incremental kolmogorov-smirnov test. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining. 1545–1554.
[15] Gints Engelen, Vera Rimmer, and Wouter Joosen. 2021. Troubleshooting an intrusion detection dataset: the CICIDS2017 case study. In 2021 IEEE Security and Privacy Workshops (SPW). IEEE, 7–12. [16] Prahlad Fogla, Monirul Islam Sharif, Roberto Perdisci, Oleg M Kolesnikov, and Wenke Lee. 2006. Polymorphic Blending Attacks.. In USENIX security symposium. 241–256. [17] Canadian Institute for Cybersecurity. 2017. CICIDS 2017 Dataset. https://www. unb.ca/cic/datasets/ids-2017.html. [18] Canadian Institute for Cybersecurity (CIC). 2018. CICIDS 2018 Dataset. https: //www.unb.ca/cic/datasets/ids-2018.html. [19] Chuanpu Fu, Qi Li, Meng Shen, and Ke Xu. 2021. Realtime robust malicious traffic detection via frequency domain analysis. In Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security. 3431–3446. [20] Chuanpu Fu, Qi Li, and Ke Xu. 2024. Flow Interaction Graph Analysis: Unknown Encrypted Malicious Traffic Detection. IEEE/ACM Transactions on Networking (2024). [21] Chuanpu Fu, Qi Li, Ke Xu, and Jianping Wu. 2023. Point cloud analysis for MLbased malicious traffic detection: Reducing majorities of false positive alarms. In Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security. 1005–1019. [22] João Gama, Indrė Žliobaitė, Albert Bifet, Mykola Pechenizkiy, and Abdelhamid Bouchachia. 2014. A survey on concept drift adaptation. ACM computing surveys (CSUR) 46, 4 (2014), 1–37. [23] Luíza Caetano Garaffa, Maik Basso, Andréa Aparecida Konzen, and Edison Pignaton de Freitas. 2021. Reinforcement learning for mobile robotics exploration: A survey. IEEE Transactions on Neural Networks and Learning Systems 34, 8 (2021), 3796–3810. [24] Pedro Garcia-Teodoro, Jesus Diaz-Verdejo, Gabriel Maciá-Fernández, and Enrique Vázquez. 2009. Anomaly-based network intrusion detection: Techniques, systems and challenges. computers & security 28, 1-2 (2009), 18–28. [25] Rosana Noronha Gemaque, Albert França Josuá Costa, Rafael Giusti, and Eulanda Miranda Dos Santos. 2020. An overview of unsupervised drift detection methods. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery 10, 6 (2020), e1381. [26] Yunhui Guo, Honghui Shi, Abhishek Kumar, Kristen Grauman, Tajana Rosing, and Rogerio Feris. 2019. Spottune: transfer learning through adaptive finetuning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 4805–4814. [27] Dongqi Han, Zhiliang Wang, Wenqi Chen, Kai Wang, Rui Yu, Su Wang, Han Zhang, Zhihua Wang, Minghui Jin, Jiahai Yang, et al. 2023. Anomaly Detection in the Open World: Normality Shift Detection, Explanation, and Adaptation.. In NDSS. [28] Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, et al. 2022. A survey on vision transformer. IEEE transactions on pattern analysis and machine intelligence 45, 1 (2022), 87–110. [29] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE international conference on computer vision. 1026–1034. [30] Meenal Jain, Gagandeep Kaur, and Vikas Saxena. 2022. A K-Means clustering and SVM based hybrid concept drift detection technique for network anomaly detection. Expert Systems with Applications 193 (2022), 116510. [31] Nathalie Japkowicz and Mohak Shah. 2011. Evaluating learning algorithms: a classification perspective. Cambridge University Press. [32] Edwin T Jaynes. 1957. Information theory and statistical mechanics. Physical review 106, 4 (1957), 620. [33] Soo-Yeon Ji, Bong-Keun Jeong, Seonho Choi, and Dong Hyun Jeong. 2016. A multi-level intrusion detection method for abnormal network behaviors. Journal of Network and Computer Applications 62 (2016), 9–17. [34] Ronald Kemker, Marc McClure, Angelina Abitino, Tyler Hayes, and Christopher Kanan. 2018. Measuring catastrophic forgetting in neural networks. In Proceedings of the AAAI conference on artificial intelligence, Vol. 32. [35] Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. 2020. Supervised contrastive learning. Advances in neural information processing systems 33 (2020), 18661–18673. [36] Ilhan Firat Kilincer, Fatih Ertam, and Abdulkadir Sengur. 2021. Machine learning methods for cyber security intrusion detection: Datasets and comparative study. Computer Networks 188 (2021), 107840. [37] Jin-Young Kim, Seok-Jun Bu, and Sung-Bae Cho. 2018. Zero-day malware detection using transferred generative adversarial networks based on deep autoencoders. Information Sciences 460 (2018), 83–102. [38] Isaiah J King, Xiaokui Shu, Jiyong Jang, Kevin Eykholt, Taesung Lee, and H Howie Huang. 2023. EdgeTorrent: Real-time Temporal Graph Representations for Intrusion Detection. In Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses. 77–91.
Navigating the Latent Manifold: Proactive Concept Drift Adaptation for Resilient NIDS
[39] James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. 2017. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences 114, 13 (2017), 3521– 3526. [40] Łukasz Korycki and Bartosz Krawczyk. 2021. Concept drift detection from multiclass imbalanced data streams. In 2021 IEEE 37th International Conference on Data Engineering (ICDE). IEEE, 1068–1079. [41] Arash Habibi Lashkari, Gerard Draper Gil, Mohammad Saiful Islam Mamun, and Ali A Ghorbani. 2017. Characterization of tor traffic using time based features. In International Conference on Information Systems Security and Privacy, Vol. 2. SciTePress, 253–262. [42] Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015. Deep learning. nature 521, 7553 (2015), 436–444. [43] Qinbin Li, Bingsheng He, and Dawn Song. 2021. Model-contrastive federated learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 10713–10722. [44] Wenjuan Li, Steven Tug, Weizhi Meng, and Yu Wang. 2019. Designing collaborative blockchained signature-based intrusion detection in IoT environments. Future Generation Computer Systems 96 (2019), 481–489. [45] Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer. 2022. Mvitv2: Improved multiscale vision transformers for classification and detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 4804–4814. [46] Xinglin Lian, Chengtai Cao, Yan Liu, Xovee Xu, Yu Zheng, and Fan Zhou. 2025. Facing Anomalies Head-On: Network traffic anomaly detection via uncertaintyinspired inter-sample differences. In Proceedings of the ACM on Web Conference 2025. 3908–3917. [47] Anjin Liu, Jie Lu, Feng Liu, and Guangquan Zhang. 2018. Accumulating regional density dissimilarity for concept drift detection in data streams. Pattern Recognition 76 (2018), 256–272. [48] Anjin Liu, Jie Lu, and Guangquan Zhang. 2020. Diverse instance-weighting ensemble based on region drift disagreement for concept drift adaptation. IEEE transactions on neural networks and learning systems 32, 1 (2020), 293–307. [49] Hongfu Liu, Junxiang Chen, Jennifer Dy, and Yun Fu. 2023. Transforming complex problems into K-means solutions. IEEE transactions on pattern analysis and machine intelligence 45, 7 (2023), 9149–9168. [50] Scott Lundberg. 2017. A unified approach to interpreting model predictions. arXiv preprint arXiv:1705.07874 (2017). [51] Anqi Mao, Mehryar Mohri, and Yutao Zhong. 2023. Cross-entropy loss functions: Theoretical analysis and applications. In International conference on Machine learning. PMLR, 23803–23828. [52] Marc Masana, Xialei Liu, Bartłomiej Twardowski, Mikel Menta, Andrew D Bagdanov, and Joost Van De Weijer. 2022. Class-incremental learning: survey and performance evaluation on image classification. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 5 (2022), 5513–5533. [53] Weizhi Meng, Wenjuan Li, and Lam-For Kwok. 2014. EFM: enhancing the performance of signature-based network intrusion detection systems using enhanced filter mechanism. computers & security 43 (2014), 189–204. [54] Yisroel Mirsky, Tomer Doitshman, Yuval Elovici, and Asaf Shabtai. 2018. Kitsune: An Ensemble of Autoencoders for Online Network Intrusion Detection. In Network and Distributed Systems Security (NDSS) Symposium. [55] Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems 35 (2022), 27730–27744. [56] Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems 32 (2019). [57] Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al. 2011. Scikit-learn: Machine learning in Python. the Journal of machine Learning research 12 (2011), 2825–2830. [58] Felipe Pinagé, Eulanda M dos Santos, and João Gama. 2020. A drift detection method based on dynamic classifier selection. Data Mining and Knowledge Discovery 34, 1 (2020), 50–74. [59] Yuqi Qing, Qilei Yin, Xinhao Deng, Xiaoli Zhang, Peiyang Li, Zhuotao Liu, Kun Sun, Ke Xu, and Qi Li. 2025. Training Robust Classifiers for Classifying Encrypted Traffic under Dynamic Network Conditions. In Proceedings of the 2025 ACM SIGSAC Conference on Computer and Communications Security. 3564–3578. [60] Yongming Rao, Wenliang Zhao, Zheng Zhu, Jiwen Lu, and Jie Zhou. 2021. Global filter networks for image classification. Advances in neural information processing systems 34 (2021), 980–993. [61] David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams. 1986. Learning representations by back-propagating errors. nature 323, 6088 (1986), 533–536. [62] Nerella Sameera and M Shashi. 2020. Deep transductive transfer learning framework for zero-day attack detection. ICT Express 6, 4 (2020), 361–367.
[63] Iman Sharafaldin, Arash Habibi Lashkari, Ali A Ghorbani, et al. 2018. Toward generating a new intrusion detection dataset and intrusion traffic characterization. ICISSp 1 (2018), 108–116. [64] Robin Sommer and Vern Paxson. 2010. Outside the closed world: On using machine learning for network intrusion detection. In 2010 IEEE symposium on security and privacy. IEEE, 305–316. [65] Vincent F Taylor, Riccardo Spolaor, Mauro Conti, and Ivan Martinovic. 2016. Appscanner: Automatic fingerprinting of smartphone apps from encrypted network traffic. In 2016 IEEE European Symposium on Security and Privacy (EuroS&P). IEEE, 439–454. [66] MTCAJ Thomas and A Thomas Joy. 2006. Elements of information theory. WileyInterscience. [67] Thijs Van Ede, Riccardo Bortolameotti, Andrea Continella, Jingjing Ren, Daniel J Dubois, Martina Lindorfer, David Choffnes, Maarten Van Steen, and Andreas Peter. 2020. Flowprint: Semi-supervised mobile-app fingerprinting on encrypted network traffic. In Network and distributed system security symposium (NDSS), Vol. 27. [68] Tim Van Erven and Peter Harremos. 2014. Rényi divergence and Kullback-Leibler divergence. IEEE Transactions on Information Theory 60, 7 (2014), 3797–3820. [69] Wenjuan Wang, Xuehui Du, Dibin Shan, Ruoxi Qin, and Na Wang. 2020. Cloud intrusion detection method based on stacked contractive auto-encoder and support vector machine. IEEE transactions on cloud computing 10, 3 (2020), 1634–1646. [70] Xian Wang. 2022. Enidrift: A fast and adaptive ensemble system for network intrusion detection under real-world drift. In Proceedings of the 38th Annual Computer Security Applications Conference. 785–798. [71] Lingfei Wu, Yu Chen, Heng Ji, and Bang Liu. 2021. Deep learning on graphs for natural language processing. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2651–2653. [72] Dongkuan Xu and Yingjie Tian. 2015. A comprehensive survey of clustering algorithms. Annals of data science 2 (2015), 165–193. [73] Jian Yang, Xiang Chen, Shuangwu Chen, Xiaofeng Jiang, and Xiaobin Tan. 2021. Conditional variational auto-encoder and extreme value theory aided two-stage learning approach for intelligent fine-grained known/unknown intrusion detection. IEEE Transactions on Information Forensics and Security 16 (2021), 3538–3553. [74] Luming Yang, Lin Liu, JunJie Huang, Zhuotao Liu, Shiyu Liang, Shaojing Fu, and Yongjun Wang. 2025. MM4flow: A Pre-trained Multi-modal Model for Versatile Network Traffic Analysis. In Proceedings of the 2025 ACM SIGSAC Conference on Computer and Communications Security. 1664–1678. [75] Li Yang, Abdallah Moubayed, and Abdallah Shami. 2021. MTH-IDS: A multitiered hybrid intrusion detection system for internet of vehicles. IEEE Internet of Things Journal 9, 1 (2021), 616–632. [76] Shuo Yang, Xinran Zheng, Jinze Li, Jinfeng Xu, Xingjun Wang, and Edith CH Ngai. 2024. Recda: Concept drift adaptation with representation enhancement for network intrusion detection. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 3818–3828. [77] Selim Yılmaz, Emre Aydogan, and Sevil Sen. 2021. A transfer learning approach for securing resource-constrained IoT devices. IEEE Transactions on Information Forensics and Security 16 (2021), 4405–4418. [78] Chao Zha, Tian Liu, Chungang Lin, Bing Bai, and Ruyun Zhang. 2026. Sparse gaussian markov modeling for robust and trustworthy unknown cyber defense. IEEE Transactions on Cognitive Communications and Networking (2026). [79] Chao Zha, Haolin Pan, Bing Bai, Jiangxing Wu, and Ruyun Zhang. 2026. FlowXpert: Context-Aware Flow Embedding for Enhanced Traffic Detection in IoT Network. IEEE Transactions on Mobile Computing (2026). [80] Chao Zha, Dakun Shen, Ruyun Zhang, and Kui Ren. 2026. Normality in Anomaly: Rethinking Traffic Labels. IEEE Transactions on Dependable and Secure Computing (2026). [81] Chao Zha, Zhiyu Wang, Yifei Fan, Bing Bai, Yinjie Zhang, Sainan Shi, and Ruyun Zhang. 2025. A-NIDS: Adaptive Network Intrusion Detection System Based on Clustering and Stacked CTGAN. IEEE Transactions on Information Forensics and Security (2025). [82] Chao Zha, Zhiyu Wang, Yifei Fan, Xingming Zhang, Bing Bai, Yinjie Zhang, Sainan Shi, and Ruyun Zhang. 2024. SKT-IDS: Unknown attack detection method based on Sigmoid Kernel Transformation and encoder–decoder architecture. Computers & Security 146 (2024), 104056. [83] Xinchen Zhang, Running Zhao, Zhihan Jiang, Handi Chen, Yulong Ding, Edith C.H. Ngai, and Shuang-Hua Yang. 2025. Continual Learning with Strategic Selection and Forgetting for Network Intrusion Detection. In IEEE INFOCOM 2025 - IEEE Conference on Computer Communications. 1–10. doi:10.1109/ INFOCOM55648.2025.11044615
A
Open Science
All artifacts of this work are available at our GitHub repository (https://github.com/cc-sec-hub/DriftXpert), including the complete
Trovato et al.
implementation of our adaptation framework, hyperparameter configurations, and scripts to reproduce the main results. The public benchmarks used in this study are accessible via their respective official websites.
B
Ethical Considerations
This study focuses on enhancing the robustness of NIDS against data drift. Our ethical statements are as follows: • Data Provenance and Privacy: We used public benchmarks and a dataset provided by an industrial partner. All data was strictly authorized for research and underwent thorough anonymization. No personally identifiable information or sensitive network configurations are included. • No Subject Involvement: This work involves no human or animal subjects. Thus, IRB approval was not required.
C
Generative AI Usage
Generative AI tools, such as ChatGPT and Gemini, are used only for polishing the language and improving the readability of the manuscript. These tools are not used for data analysis, code implementation, or generating research ideas. All contents are carefully reviewed and verified by the authors to ensure their technical accuracy and academic integrity.
D
Theoretical Analysis of Predict Distribution
DEFINITION 3. Let 𝑝 (𝑋 ) represent the predicted distribution of the model and 𝑞(𝑋 ) denote the true distribution of the labels. THEOREM 1. The predict probability density function 𝑝 (𝑋 ) can be expressed as: 𝑝 (𝑋 ) = √
1 2𝜋𝜎 2
exp −
1 (𝑋 − 𝜇) 2, 2𝜎 2
(11)
PROOF. The classification results are optimized through convergence using cross-entropy loss, which can be expressed as follows:
For constrained optimization problems, we solve them using the Lagrangian function. Thus, our objective function can be formulated as: ∫ ∞ L =− 𝑝 (𝑋 ) ln 𝑝 (𝑋 ) 𝑑𝑋 −∞ ∫ ∞ + 𝜆0 𝑝 (𝑋 ) 𝑑𝑋 − 1 −∞ ∫ ∞ + 𝜆1 𝑋 𝑝 (𝑋 ) 𝑑𝑋 − 𝜇 −∞ ∫ ∞ + 𝜆2 (𝑋 − 𝜇) 2 𝑝 (𝑋 ) 𝑑𝑋 − 𝜎 2 . (17) −∞
By solving the above objective function, we can derive 𝑝 (𝑋 ) = 𝐶 exp (𝜆1𝑋 + 𝜆2 (𝑋 − 𝜇) 2 ). Further analysis leads to the following parameter values: 𝐶 = √ 1 2 , 𝜆1 = 0, 𝜆2 = − 2𝜎1 2 . 2𝜋𝜎 Consequently, the predict probability density function 𝑝 (𝑋 ) can be expressed as: 𝑝 (𝑋 ) = √
1 2𝜋𝜎 2
exp −
1 (𝑋 − 𝜇) 2, 2𝜎 2
(18)
which follows a Gaussian distribution.
E
Theoretical Analysis of Hyperparameter 𝛼 for Drift Detection
The hyperparameters 𝛼 directly determine the value of 𝐶𝐼 , thus influencing the number of outliers from both historical data and drifted data in the drift detection model. This, in turn, affects the magnitude of 𝑂𝑅. Following the theoretical analysis in [81], we examine the relationship among 𝑂𝑢𝑡𝑙𝑖𝑒𝑟𝑠, 𝑂𝑅, and 𝛼, with the goal of understanding their interdependencies and possible correlations. ASSUMPTION 1. Assume the distribution of i-th cluster follows the Gaussian distribution, that is, 1 1 (19) 𝑃𝑖 (𝑧) = √︃ exp − 2 (𝑧 − 𝜇𝑖 ) 2, 2𝜎𝑖 2𝜋𝜎 2 𝑖
𝐻 (𝑝, 𝑞) = 𝐻 (𝑝) + 𝐷𝐾𝐿 (𝑝 ||𝑞),
(12)
where 𝐻 (𝑝, 𝑞) represents the cross-entropy, 𝐻 (𝑝) denotes the entropy of 𝑝 (𝑋 ), 𝐷 𝐾𝐿 (𝑝 ||𝑞) refers to the Kullback-Leibler (KL) divergence. 𝐻 (𝑝) represents the intrinsic uncertainty of the predicted distribution 𝑝 (𝑋 ), expressed as follows: ∫ ∞ 𝐻 (𝑝) = − 𝑝 (𝑋 ) ln 𝑝 (𝑋 ) 𝑑𝑋 . (13) −∞
To derive the probability density function of the predicted distribution 𝑝 (𝑋 ) from the perspective of maximizing entropy [32, 66], the following constraints must first be satisfied: ∫ ∞ 𝑝 (𝑋 ) 𝑑𝑋 = 1, (14) −∞
𝑂𝑢𝑡𝑙𝑖𝑒𝑟𝑠 =
𝐾 ∑︁
𝑁𝑖 ·
𝑖=1
exp (−(𝐴𝛼 + 𝐵) 2 ) . √ (𝐴𝛼 + 𝐵) 𝜋
(20)
where 𝐾 is the total number of clusters, 𝑁𝑖 is the sample number of i-th cluster, 𝐴 and 𝐵 are constants. PROOF. According to |𝑥 | > 𝑉𝑖 (𝑉𝑖 > 𝜇𝑖 ), the number of outliers can be estimated by the probability distribution: ∫ +∞ 𝐾 ∑︁ 𝑂𝑢𝑡𝑙𝑖𝑒𝑟𝑠 = 2𝑁𝑖 · 𝑃𝑖 (𝑧)𝑑𝑧. (21) 𝑖=1
𝑉𝑖 𝑧−𝜇
∫ ∞ 𝑋 𝑝 (𝑋 ) 𝑑𝑋 = 𝜇,
(15)
(𝑋 − 𝜇) 2 𝑝 (𝑋 ) 𝑑𝑋 = 𝜎 2,
(16)
−∞
∫ ∞
where 𝜇𝑖 and 𝜎𝑖 are the mean and standard deviation of 𝑃𝑖 (𝑧), respectively. THEOREM 2. When |𝑥 | > 𝑉𝑖 (𝑉𝑖 > 𝜇𝑖 ), samples are outliers, such that the number of outliers is:
−∞
where 𝜇 represents the mean, 𝜎 2 denotes the variance.
A standard transformation is applied as 𝑡 = 𝜎𝑖 𝑖 ; then ∫ +∞ ∫ +∞ 1 𝑡2 𝑃𝑖 (𝑧)𝑑𝑧 = 𝑉 −𝜇 √ exp (− )𝑑𝑡 𝑖 𝑖 2 2𝜋 𝑉𝑖 𝜎𝑖 = 1 − Φ(
𝑉𝑖 − 𝜇𝑖 ), 𝜎𝑖
(22)
Navigating the Latent Manifold: Proactive Concept Drift Adaptation for Resilient NIDS
∫ 𝑉𝑖 −𝜇𝑖 𝜎𝑖 𝑉𝑖 − 𝜇𝑖 1 𝑡2 Φ( )= (23) √ exp (− )𝑑𝑡 . 𝜎𝑖 2 2𝜋 −∞ Based on Plackett’s approximation and the Gaussian error func𝑧−𝜇 tion, Φ( 𝜎𝑖 𝑖 ) is approximated as follows: 𝑉𝑖 − 𝜇𝑖 𝑉𝑖 − 𝜇𝑖 1 ) , (24) Φ( )= 1 + erf( √ 𝜎𝑖 2 2𝜎𝑖 ∫ 𝑉𝑖 −𝜇𝑖 𝜎𝑖 𝑉𝑖 − 𝜇𝑖 2 erf( √ )=√ exp (−𝑡 2 )𝑑𝑡 . (25) 𝜋 0 2𝜎𝑖 Applying the Taylor series expansion gives the following. 2 𝑒 −𝑥 1 3 1 erf(𝑥) = 1 − √ 1 − 2 + 4 + 𝑂 ( 6 ) . (26) 2𝑥 4𝑥 𝑥 𝑥 𝜋 Let 𝑉𝑖 = 𝐶𝑖 + 𝐶𝐼 , 𝐶𝑖 be the center of the i-th cluster, according 𝑉 −𝜇 𝐶 +𝐶𝐼 −𝜇 to Equation (5), 𝑖𝜎𝑖 𝑖 = 𝑖 𝜎𝑖 𝑖 = 𝐴𝛼 + 𝐵. Consequently, combining Equation (19), Equation (21) - Equation (26), 𝑂𝑢𝑡𝑙𝑖𝑒𝑟𝑠 can be represented as: 𝑂𝑢𝑡𝑙𝑖𝑒𝑟𝑠 =
𝐾 ∑︁
𝑁𝑖 ·
𝑖=1
exp (−(𝐴𝛼 + 𝐵) 2 ) . √ (𝐴𝛼 + 𝐵) 𝜋
(27)
Figure 11: Graphical simulation of 𝑂𝑢𝑡𝑙𝑖𝑒𝑟𝑠 (𝛼) and 𝑂𝑅(𝛼). DEFINITION 4. Given a parameter 𝜃 ∗ as a local minimum point of the loss function L (𝜃 ), such that ∇L (𝜃 ∗ ) = 0 and ∇2 L (𝜃 ∗ ) > 0. THEOREM 3. Given a parameter transformation method 𝜂 ∗ = 𝑔 −1 (𝜃 ∗ ), the parameter 𝜂 ∗ is also a local minimum point of the loss function L (𝜃 ). PROOF. We define L𝜂 ∗ = L ◦ 𝑔 as the loss function with respect to parameters 𝜂 ∗ . We extend the derivation as follows: L𝜂 ∗ (𝜂 ∗ ) = L (𝑔(𝜂 ∗ ))
Considering the relationship between 𝑂𝑢𝑡𝑙𝑖𝑒𝑟𝑠 and 𝛼, the connection between 𝑂𝑅 and 𝛼 becomes clearer. To simplify the analysis, two scenarios are considered to illustrate their relationship. • As 𝛼 gradually increases from 0, the drifted data distribution completely consists of outliers, whereas the number of outliers in the historical data distribution gradually decreases. The outlier ratio 𝑂𝑅1 can be expressed as: 1
𝑂𝑅1 =
𝑁 𝑍𝑟
.
𝑂𝑅2 =
𝑁 𝑍𝑟
· 𝑂𝑢𝑡𝑙𝑖𝑒𝑟𝑠 (𝛼) 1 𝑁 𝑍ℎ · 1
.
(29)
To provide a clearer illustration of the relationship among 𝑂𝑢𝑡𝑙𝑖𝑒𝑟𝑠, 𝑂𝑅, and 𝛼, a graphical simulation was conducted based on Equation (27) – Equation (29), as shown in Figure 11.
F
Theoretical Analysis of Contrastive Loss on Convergence
An important aspect of neural networks is understanding the nonEuclidean geometry of the parameter space of neural structures [1]. Our interest does not lie in the parameters themselves but in the mappings they represent. The non-Euclidean geometry of the parameter space, combined with the manifold of models exhibiting equivalent behavior under observation, allows movement from one region of the parameter space to another. This enables changes in the curvature of the model without altering the actual calculation of the function [11].
(30) ∗
= ∇L (𝜃 ∗ )(∇𝑔)(𝜂 ∗ )
(31)
(∇2 L𝜂 ∗ )(𝜂 ∗ ) = (∇𝑔)(𝜂 ∗ )𝑇 (∇2 L)(𝑔(𝜂 ∗ ))(∇𝑔) (𝜂 ∗ ) + (∇L)(𝑔(𝜂 ∗ ))(∇2𝑔)(𝜂 ∗ ) = (∇𝑔)(𝜂 ∗ )𝑇 ∇2 L (𝜃 ∗ )(∇𝑔)(𝜂 ∗ )
(32)
+ (∇L)(𝜃 ∗ )(∇2𝑔)(𝜂 ∗ )
(28)
• As 𝛼 continues to increase, there are no outliers in the historical data (to ensure that the denominator is not zero, we set it to 1). Meanwhile, the number of outliers in the drifted data distribution gradually decreases, following a trend similar to that of 𝑂𝑢𝑡𝑙𝑖𝑒𝑟𝑠 (𝛼). The outlier ratio 𝑂𝑅2 can be expressed as:
∗
(∇L𝜂 ∗ )(𝜂 ) = (∇L)(𝑔(𝜂 ))(∇𝑔)(𝜂 )
· 𝑁 𝑍𝑟
𝑁𝑍ℎ · 𝑂𝑢𝑡𝑙𝑖𝑒𝑟𝑠 (𝛼)
1
∗
By DEFINITION 4, ∇L (𝜃 ∗ ) = 0 and ∇2 L (𝜃 ∗ ) > 0, we can thus derive: (∇L𝜂 ∗ )(𝜂 ∗ ) = 0 (33) (∇2 L𝜂 ∗ )(𝜂 ∗ ) = (∇𝑔)(𝜂 ∗ )𝑇 ∇2 L (𝜃 ∗ )(∇𝑔)(𝜂 ∗ ) > 0
(34)
Thus, 𝜂 ∗ is also a local minimum point of the loss function L (𝜃 ). It means that if the drifted data can converge on the classification model, then by employing contrastive learning to shift the model parameters within the non-Euclidean geometry to another parameter space, the new model will still converge on the drifted data.
G More Details of Experiments G.1 Features Used in Our Experiments We use CICFlowMeter [41] to extract approximately 80 relevant features and remove features such as source IP, destination IP, source port and destination port that could introduce "cheating effects." [15] After experimental comparisons, we also eliminated features with numerous outliers (e.g. extremely large values or negative numbers) and features that had no impact on the experimental results. Finally, we selected 50 features used in our experiments, as shown in Table 14.
G.2
Settings of Drift Scenario
Table 15 illustrates details of two drift scenario settings in CICDS2017 and CICIDS-2018 datasets.
Trovato et al.
Table 14: Features Used in Our Experiments.
G.3
Type
Features
Time-Realted
Flow Duration, Flow IAT Mean, Flow IAT Std, Flow IAT Max, Flow IAT Min, Fwd IAT Tot, Fwd IAT Mean, Fwd IAT Std, Fwd IAT Max, Fwd IAT Min, Bwd IAT Tot, Bwd IAT Mean, Bwd IAT Std, Bwd IAT Max, Bwd IAT Min
Table 16 presents details of our model architecture and parameter counts, and Table 17 provides the hyperparameters used in our experiments.
Packet-Realted
Length-Realted
Tot Fwd Pkts, Tot Bwd Pkts, Flow Byts/s, Flow Pkts/s, Fwd Pkts/s, Bwd Pkts/s, Subflow Fwd Pkts, Subflow Fwd Byts, Subflow Bwd Pkts, Subflow Bwd Byts, Init Fwd Win Byts, Init Bwd Win Byts, Fwd Act Data Pkts TotLen Fwd Pkts, TotLen Bwd Pkts, Fwd Pkt Len Max, Fwd Pkt Len Min, Fwd Pkt Len Mean, Fwd Pkt Len Std, Bwd Pkt Len Max, Bwd Pkt Len Min, Bwd Pkt Len Mean, Bwd Pkt Len Std, Fwd Header Len, Bwd Header Len, Pkt Len Min, Pkt Len Max, Pkt Len Mean, Pkt Len Std, Pkt Len Var, Pkt Size Avg, Fwd Seg Size Avg, Bwd Seg Size Avg, Fwd Seg Size Min
Protocol-Related
Details of Model and Training
Table 16: Details of Model Structure.
Blocks
Dims
Layers
Output
Encoder
Linear ReLU Linear ReLU Linear
50 96 96 64 64
96 96 64 64 32
4896 0 6208 0 2080
Classifier
Linear
32
2
66
* Total parameters: 11550.
Table 17: Recommended Hyperparameters. Type
Fwd PSH Flags
Datasets
Hyper-Paras
Value
CICIDS-2017
learning rate batch size epoch T 𝛽 𝛾
1e-4 256 500 2.0 0.5 1.0
CICIDS-2018
learning rate batch size epoch T 𝛽 𝛾
1e-4 256 120 2.0 0.3 0.1
CICIDS-2017
learning rate batch size epoch T 𝛽 𝛾
1e-4 256 200 2.0 0.5 1.0
CICIDS-2018
learning rate batch size epoch T 𝛽 𝛾
1e-5 256 100 3.0 0.5 0.1
Table 15: Datasets Settings Global Drift Dataset
Type
Set1
Set2
Local Drift
Benign, DoSGoldenEye
Benign, DoSGoldenEye
Benign, DoSGoldenEye, Dos-Hulk, FTP-Patator, SSH-Patator.
Benign, Botnet, Portscan, DoSSlowhttptest, DoS-Slowloris, XSS, DDoS, Bruteforce.
Local Drift
Benign, DoSGoldenEye
Benign, DoSGoldenEye
Global Drift
Benign, DoS- Benign, Botnet, Slowhttptest, FTP-Patator, DoS-Hulk, DosSSH-Patator, GoldenEye, DDoS-LOICDoS-Slowloris. HTTP, DDoS-LOICUDP, XSS, Web-Attacks, SQL-Injection.
CICIDS-2017 Global Drift
CICIDS-2018
Parameters
Input
Local Drift
H Supplementary Experiment H.1 Feature Importance Analysis under Concept Drift We analyze feature influence using SHAP (SHapley Additive Explanations) [50] for both historical and drifted models (Figure 12).
Navigating the Latent Manifold: Proactive Concept Drift Adaptation for Resilient NIDS
The results show that key features differ, indicating drift beyond latent distributions. While initialization from the historical model alleviates catastrophic forgetting, SHAP reveals evolving feature importance. Therefore, freezing the first layer of the encoder preserves the initial feature influence while allowing subsequent fully connected layers to adapt, providing a well-justified balance between stability and adaptability. (a) Local Drift.
(b) Global Drift.
Figure 13: Contrastive Loss During Training.
(a) Historical Model.
(a) Local Drift (Historical).
(b) Local Drift (Drifted).
(c) Global Drift (Historical).
(d) Global Drift (Drifted).
(b) Drifted Model.
Figure 12: SHAP summary plots of the historical and drifted models. Higher-ranked features on the left indicate greater impact on predictions.
H.2
Analysis of Contrastive Loss
During the model update process, both the aggregation model and the training model are initialized using the historical model parameters. According to contrastive loss (as shown in Equation (7), when both sim(𝑧, 𝑧ℎ𝑖𝑠 ) and sim(𝑧, 𝑧𝑝𝑟𝑒𝑣 ) are 1.0 (completely similar), contrastive loss should be 0.693, indicating the optimal state of distribution alignment. However, as the model is trained on drifted data, it tends to fit the new data, causing parameter changes that shift the latent variable distribution away from the historical model’s latent variable distribution. As a result, the contrastive loss gradually increases, meaning that during the initial epochs, the contrastive loss does not contribute to aligning the model. As training progresses, the contrastive loss begins to align the latent variable distribution of the new model with that of the aggregated model, and the contrastive loss should decrease steadily, but it should not fall below 0.693. The graph of contrastive loss during our experiments is shown in Figure 13.
H.3
Sensitivity Analysis of Hyperparameter on the CICIDS-2018 Dataset
Figure 14: Sensitivity Analysis of Hyperparameter 𝛽.
(a) Local Drift (Historical).
(b) Local Drift (Drifted).
(c) Global Drift (Historical).
(d) Global Drift (Drifted).
We conducted sensitivity analyses of the hyperparameters 𝛽 and 𝛾 on the CICIDS-2018 dataset. The results are presented in Figure 14 and Figure 15, with the detailed analysis provided as follows:
Figure 15: Sensitivity Analysis of Hyperparameter 𝛾.
H.3.1 Sensiticity Analysis of 𝛽. We investigated the sensitivity of the hyperparameter 𝛽 on the CICIDS-2018 dataset under both local drift and global drift scenarios. DriftXpert maintains stable performance across different values of 𝛽 in both scenarios, with only
marginal performance variations. This indicates that the proposed framework is not sensitive to the selection of 𝛽, and a wide range of values can effectively preserve the balance between catastrophic forgetting and model adaptation.
Trovato et al.
H.3.2 Sensitivity Analysis of 𝛾. We further analyzed the impact of the hyperparameter 𝛾 under local and global drift scenarios. The performance remains relatively stable as 𝛾 varies, without significant degradation in either scenario. This observation suggests that DriftXpert is insensitive to the choice of 𝛾 and can achieve reliable adaptation across different parameter settings.
H.4
Sensitivity Analysis of Hyperparameter on the Real-World Dataset
We further conducted sensitivity analyses of the hyperparameters 𝑎 and 𝑏 on the real-world dataset, with the results presented in Figure 16 and Figure 17. The detailed analysis is provided as follows:
(a) Historical.
(b) Drifted.
H.5
Performance Test
H.5.1 Memory Usage. We estimated the memory usage of the model. DriftXpert is lightweight, with 11,550 parameters (see Appendix D), which requires approximately 46 KB. Taking into account the storage of gradient and momentum, the optimizer adds an estimated 92 KB. Contrastive learning further requires storing the previous and aggregated models, adding another 92 KB. In total, the training memory footprint of DriftXpert is around 230 KB, making it suitable for deployment on devices with limited resources. H.5.2 Latency. We measured per sample detection latency of the Encoder and Classifier, as illustrated in Figure 18a and Figure 18b. The Encoder ranged from 0.1 to 2.1 µs (average 0.7 µs), while the simpler Classifier ranged from 0.0 to 0.24 µs (average 0.1 µs). Overall latency did not exceed 2.3 µs, demonstrating the strong real-time performance of DriftXpert. H.5.3 Throughput. We evaluated the throughput of the Encoder and Classifier, as illustrated in Figure 18c and Figure 18d. The Encoder ranged from 0.4 to 2.8 MQPS, averaging 1.6 MQPS, while the Classifier reached up to 40 MQPS, averaging 16 MQPS. In general, DriftXpert achieves an average throughput of 1.6 MQPS, highlighting its strong real-time capability.
Figure 16: Sensitivity Analysis of Hyperparameter 𝛽.
(a) Historical.
(a) Latency of Encoder.
(b) Latency of Classifier.
(c) Throughput of Encoder.
(d) Throughput of Classifier.
(b) Drifted.
Figure 17: Sensitivity Analysis of Hyperparameter 𝛾.
H.4.1 Sensiticity Analysis of 𝛽. Firstly, we evaluated the sensitivity of the hyperparameter 𝛽 on the real-world enterprise traffic dataset. DriftXpert achieves consistently stable performance across different values, with only minor fluctuations. This indicates that the proposed drift detection mechanism is robust to the variation of 𝛽 in practical deployment scenarios. Although 𝛽 controls the sensitivity of drift identification, DriftXpert can effectively distinguish evolving traffic patterns from normal variations across a wide range of settings, demonstrating its robustness in real-world environments. H.4.2 Sensitivity Analysis of 𝛾. Secondly, we investigated the impact of 𝛾 on the real-world enterprise traffic dataset. The results show that DriftXpert maintains stable performance under different configurations, without noticeable performance degradation. This demonstrates that the proposed adaptation strategy is not highly dependent on the exact choice of 𝛾, and can consistently preserve previous knowledge while adapting to emerging traffic patterns. These results further validate the practicality and robustness of DriftXpert for long-term deployment in dynamic network environments.
Figure 18: Performance results of DriftXpert.