PACT: Reducing Alert Fatigue in Low-Prevalence SOC Streams with Triggered Active Learning Samuel Ndichu1 , Tao Ban1 , Seiichi Ozawa2 , Takeshi Takahashi1 , Daisuke Inoue1
arXiv:2605.22324v1 [cs.CR] 21 May 2026
1
National Institute of Information and Communications Technology, Tokyo, Japan 2 Kobe University, Kobe, Japan {ndichu, bantao, takeshi takahashi, dai}@nict.go.jp, [email protected]
Abstract—Security operations centers face persistent alert fatigue: in low-prevalence streams, even low false-positive rates generate substantial investigation load, while aggregate F1 scores obscure analyst burden. We introduce PACT, a Paretoaware controller for triggered active learning, which wraps an already-deployed frozen XGBoost-Focal screener with an adaptive windowing score-shift trigger and a hybrid acquisition rule combining threshold-relative uncertainty with high-score sampling. On two public low-prevalence benchmarks, AIT-ADS (AIT Alert Data Set), and BOTSv1 (Boss of the SOC version 1), PACT attains the lowest benign-normalized false-positive (FP) burden among the adaptive methods tested. It reduces burden by 43% and 21%, respectively, relative to a frozen baseline, while using 3.8× and 5.2× fewer analyst queries than periodic uniform-random updating. A matched-trigger ablation controls trigger timing and shows that acquisition contributes beyond timing alone, at the cost of approximately ten percentage points of positive-window recall under freerunning triggers. A frozen threshold-only baseline pushes FP lower still but collapses BOTSv1 recall by 55 percentage points. Under the evaluated workload assumptions, pure FP minimization trades unacceptable recall for that lower burden. Index Terms—Security operations centers, alert fatigue, alert screening, intrusion detection systems, active learning, ADWIN, class imbalance, analyst workload, false-positive burden, operating-point trade-offs.
1. Introduction Security Operations Centers (SOCs) rely on intrusion detection systems (IDSs) and network security monitoring (NSM) tools to surface potentially malicious activity [1]–[4]. In practice, these systems generate large alert volumes, and most alerts are benign, redundant, or low-priority. Analysts must therefore perform downstream alert screening: deciding which alerts deserve attention, which can be suppressed, and how to maintain sensitivity to rare but consequential attacks [5]. Excessive false-positives (FPs) contribute to alert fatigue, reduce analyst efficiency, and increase the risk that important incidents are missed [6]–[10]. IDS-derived alert streams are low-prevalence: malicious events typically constitute well under 1% of traffic [11]–
[13]. At that scale, even a 0.1% false-positive rate (FPR) over a one-million-event ingestion day produces a thousand false alerts, enough to materially affect analyst workload. The streams we study illustrate this asymmetry. Extreme Gradient Boosting (XGBoost)-Focal screeners trained under the same protocol produce nearly three times as many stream-time false alerts per million benign events on the Boss of the SOC version 1 (BOTSv1) as on the AIT Alert Data Set (AIT-ADS) (Table 6), even though the offline focalloss F1 scores differ by fewer than ten points (Table 4). Because positives are rare, FPR movements can dominate investigation volume, while recall losses remain safetycritical and must be reported explicitly. Alert distributions also evolve as attacker tactics change, environments shift, and tools are reconfigured. Together, prevalence asymmetry and drift motivate adaptive alert screening that uses bounded analyst feedback to control FP burden. Adaptive screening is a design problem, not a free upgrade. Updating a detector from small queried batches can destabilize calibration, overfit recent samples, or trade recall for lower FPR. Drift triggers based on unlabeled score streams detect changes in model confidence, not necessarily true concept drift. Active-learning policies that acquire useful positive examples can also introduce sampling bias or update the model on unrepresentative data. We therefore ask when adaptation actually improves a deployed screener, rather than merely shifting errors from FPs to missed positives or consuming additional analyst labels. Scope. We focus deliberately on low-prevalence security alert streams derived from intrusion-detection and security-event telemetry, of which AIT-ADS and BOTSv1 are two representative public benchmarks. High-prevalence security information and event management (SIEM) incident streams, where positives are widespread throughout the stream, raise distinct calibration and threshold-management questions and are outside the scope of this paper. Contributions. This paper makes three focused contributions. •
Operational evaluation framing. We evaluate adaptive alert screening on SOC-facing endpoints, including FPs per million benign events (FP/1M benign), positive-window recall, missed positives, and realized query rate, rather than on aggregate F1
•
•
alone. PACT controller and matched-trigger ablation. We introduce PACT, a Pareto-aware controller for triggered active learning,1 and compare four streaming strategies on AIT-ADS and BOTSv1; a matchedtrigger ablation controls trigger timing and shows that acquisition contributes beyond timing alone. Pareto trade-off reporting. PACT attains the lowest benign-normalized FP burden among the adaptive strategies on both datasets, at the cost of roughly ten percentage points of positive-window recall; under the evaluated workload assumptions, periodic and adaptive windowing (ADWIN)-random preserve recall on BOTSv1 only at FP volumes operationally impractical for this SOC scenario.
The contribution is not a new drift detector or activelearning primitive. Rather, PACT combines three existing ingredients in a downstream SOC screening setting: a controller around an already-deployed frozen screener, an unsupervised score-shift trigger over the predicted-probability stream, and an analyst-efficient hybrid query rule. Prior IDS active-learning work retrains online classifiers on raw traffic [14]–[16], and prior SOC alert-prioritization work uses analyst feedback without a drift-triggered query loop [17]– [20]. We instead target downstream alert screening on public alert and log benchmarks and evaluate the resulting trade-off in analyst burden explicitly. Research questions. The paper is organized around five research questions: •
•
•
•
•
RQ1. How separable is the AIT-ADS alert stream under offline static baselines after leakage-mitigation filtering, and how do XGBoost loss variants compare on both datasets? RQ2. Under a low deployment prior, how does the frozen backbone translate into analyst-facing daily false-alert volume? RQ3. When does triggered active learning reduce benign-normalized FP burden relative to frozen deployment, and at what cost in positive-window recall and missed positives? RQ4. Under a matched ADWIN trigger schedule, which acquisition rule (random, uncertainty-only, high-score-only, or hybrid) yields the lowest FP burden, and how does the choice interact with recall and missed positives? RQ5. What is the realized analyst-labeling cost of each streaming strategy across multiple seeds, and how does the minimum-update-batch safeguard affect comparisons across nominal budgets?
The remainder covers related work (Section 2), the triggered active-learning design (Section 3), the experimental setup (Section 4), results (Section 5), discussion (Section 6), limitations (Section 7), and the conclusion (Section 8). 1. We adopt PACT as a pronounceable contraction of Pareto-Aware Controller for Triggered active learning; the active-learning component is the controller’s mechanism rather than part of the initialism.
2. Related Work Triggered active learning for low-prevalence SOC alert screening draws on five bodies of prior work. These are alert screening versus upstream detection, class-imbalance handling, drift-aware online intrusion detection, hybrid acquisition under drift and imbalance, and active learning with human-in-the-loop SOC support. Alert screening versus upstream detection. Most classical IDS and NSM literature targets detection: deciding whether raw traffic, events, or logs are malicious [8], [21], [22]. The complementary problem we address is downstream alert screening: given alerts already produced by IDSs or NSM tools, how should a SOC prioritize and filter them so analysts focus on consequential incidents while controlling FP burden and analyst workload [6], [7]. Class imbalance. Cybersecurity datasets are persistently imbalanced [23], and prior work has studied costsensitive learning, resampling methods such as the synthetic minority over-sampling technique [24]–[27], algorithm-level reweighting [28], and focal loss (FL) [29] for static IDS or alert-screening tasks. We use focal loss as the frozen backbone and focus on how stream-time updates change analyst-facing burden. Drift-aware online intrusion detection. Supervised, unsupervised, and neural methods have been applied to intrusion detection [30]–[35]. Drift-aware online IDS work often uses incremental updates, and ADWIN-style detectors [36], [37] provide a common change-detection mechanism that is sometimes coupled with active learning at the raw-traffic layer [14], [15]. ADWIN-U [37] shows how ADWIN-style windowing can be adapted for unsupervised drift monitoring; we apply ADWIN directly to the predictedprobability stream as a heuristic score-shift trigger and pair it with hybrid acquisition. Online and semi-supervised log analyzers [38] and isolation-forest-based online filters [39], [40] reduce labeling demand, but few jointly address lowprevalence imbalance, score-stream drift, and analyst-facing screening quality. Hybrid acquisition under drift and imbalance. Hybrid acquisition rules in streaming classification [41], [42] address the imbalance-drift interaction but typically target upstream traffic classification [43], [44] and not downstream alert screening. The method of Liu et al. [42] biases queries toward minority classes through an asymmetric-margin uncertainty rule, whereas our acquisition rule explicitly partitions the query budget between threshold-relative boundary uncertainty and high-positive-class-score exploitation. Outside security, Das et al. [45] combine a streaming drift detector with batch active learning over tree-based anomaly ensembles. This is the closest generic precedent for the “drift-trigger plus active-learning plus tree-ensemble” pattern. We instead target downstream alert screening with a frozen gradient-boosted screener and a hybrid (uncertainty plus high-score) rule, not diversity-aware sampling. Adaptive-XGBoost [46] provides the methodological precedent for tree-append updates under drift, which we use in the warm-start step. We treat ADWIN as a heuristic trigger
over predicted probabilities, not as a validated concept-drift oracle. Active learning and human-in-the-loop SOC support. Analyst time is the binding labeling resource [47], [48]. Active-learning theory provides uncertainty- and committeebased query rules [49], [50]. In security specifically, interactive labeling and situation-awareness frameworks have used active learning to reduce expert workload [51], [52]. Recent SOC-specific work has applied active learning to network-IDS alert classification [16], [53], false-alert filtering [54], and alert prioritization via reinforcement learning from analyst feedback [18]–[20], [55]. Provenance-based triage adds correlation-aware prioritization [17]; AI-assisted aggregation, anomaly detection, and post-correlation analysis further reduce alert volume [9], [13], [56]–[60]. However, these efforts rarely report the realized query rate or the composition of labels actually used to update a deployed screener, two quantities that, as we show, are essential for evaluating whether adaptation truly reduces analyst burden. Gap. Imbalance, drift, and analyst workload are typically studied separately, and many alert-management systems report aggregate accuracy or F1 without normalizing FPs by benign event volume, the operational denominator most directly tied to analyst-facing alert volume. We therefore compare triggered update policies on two lowprevalence SOC benchmarks and report the quantities that determine whether adaptation is operationally useful: FP burden, recall, and analyst labeling cost.
3. Triggered Active-Learning Design Fig. 1 shows the four operational stages of PACT, the streaming alert-screening controller used in the experiments: frozen-core scoring of the incoming alert stream, ADWINbased score-shift triggering, hybrid query selection with bounded analyst feedback, and warm-start booster update via hot-swap. The frozen, imbalance-aware classification core is trained offline; the streaming loop wraps it with a triggered active-learning controller. The “Pareto-aware” label describes the evaluation, not an internal optimizer. PACT produces one operating point per configuration. The Pareto view emerges when we compare configurations along FP burden, recall, and analyst-query cost.
3.1. Problem setting and decision rule t Let Dt = {(xi , yi )}ni=1 denote alerts observed at time window t, with xi a feature vector and yi ∈ {0, 1} a benign or malicious label. Given a screening model that outputs a probability p(x), the screening decision uses an operating threshold θ: ( 1, p(x) ≥ θ, ŷ(x) = (1) 0, otherwise.
Because focal-loss training can shift probability calibration relative to cross-entropy [29], θ is selected on a held-out validation tail using a maximum-F1 criterion rather than fixed at 0.5.
3.2. Preprocessing and leakage mitigation Alerts are sorted by timestamp, categorical fields onehot encoded, and numerical features standardized using ′ training-set statistics only: Xnum = (Xnum − µnum )/σnum . Missing values are imputed using training-set statistics: the median for numerical features and the mode for categorical features. To reduce post-decision leakage, the feature space is filtered using a conservative global denylist that removes labels, attack names, downstream verdicts, analyst annotations, incident identifiers, split metadata, and other post-decision fields. After filtering, both datasets are mapped into a compact five-feature alert-screening schema: alert category, severity, source port, destination port, and a causal time-since-last-alert/event feature. In the alert-screening setting we target, upstream IDS and security-event tools have already emitted the alerts; alert category and severity are upstream alert metadata available before downstream analyst triage, not analyst verdicts or post-investigation labels. Severity is therefore treated as pre-decision with respect to downstream analyst screening, and the time-since feature is computed only from prior events, never from future stream information. The exact retained columns per dataset are listed in Appendix A. This filter mitigates known leakage risks but does not, by itself, prove that every retained feature is observable pre-decision in every SOC deployment.
3.3. Imbalance-aware backbone (XGBoost-Focal) We use XGBoost [61] as the screening backbone because gradient-boosted trees handle heterogeneous tabular alert features well and admit custom objectives. Focal loss [29], [62], [63] emphasizes hard examples; for trueclass probability pt , FL(pt ) = −αt (1 − pt )γ log(pt ),
(2)
where αt controls class weighting and γ controls the focusing effect. We compare focal against plain cross-entropy and class-weighting (scale_pos_weight) under chronological or stratified validation, then deploy the selected variant frozen as the streaming core (Section 4).
3.4. Score-shift triggering with ADWIN We use ADWIN [36] to flag candidate change points during stream operation. ADWIN maintains a dynamic window and signals a change when mean(W1 ) − mean(W2 ) > ϵ,
(3)
with ϵ depending on confidence level and effective window size. We apply ADWIN to the predicted probability stream rather than to delayed ground-truth error, because labels are not available before analyst review. ADWIN alarms are therefore best interpreted as heuristic evidence of scoredistribution shift, not as validated concept-drift detections; the validity of any update is judged empirically through downstream FP burden, recall, realized query rate, and label yield.
Alert stream {𝑥ₜ} (low-prevalence regime) Deployed pipeline (frozen)
Train XGBoost-Focal
Frozen scorer p(𝑥), ŷ = 𝟙[p( 𝑥) ≥ θ]
Score stream {p(𝑥ₜ)}
ŷₜ
SOC analysts triage
Hot-swap update (append-only)
PACT – active-learning controller
ADWIN trigger Mean shift in p(𝑥)
Hybrid query ½ near θ + ½ highest p(𝑥)
Q
Analyst oracle Labels → pending 𝑃
|𝑃| ≥ 𝐵min
Warm-start + 𝑘w trees, cooldown 𝑐
Figure 1. PACT Architecture.
3.5. Hybrid acquisition rule On a trigger event, the controller selects a bounded query batch from the recent buffer B . The hybrid rule splits the batch: half the slots go to samples closest to the operating threshold, ranked by their threshold-relative score distance u(x) = |p(x) − θ|,
(4)
which we use as a proxy for boundary uncertainty under threshold shift. The other half goes to the highest predicted positive-class probabilities, representing high-score acquisition that prioritizes likely positives. After deduplication, the batch is topped up from the highest remaining scores. The threshold-relative quantity u(x) is used in place of binary entropy because θ ̸= 0.5 in general after F1 -based threshold selection, so binary entropy and threshold-relative score distance are not equivalent. This hybrid avoids a known failure mode of pure uncertainty sampling under severe imbalance: most lowconfidence samples are benign, so an entropy-only rule can starve the update of positive evidence. The high-score half supplies that positive evidence; the threshold-relative half preserves boundary exploration. We also evaluated a lightweight replay-buffer stabilizer for the warm-start update; it did not improve the operational profile and is reported as a negative robustness check in Appendix B.
3.6. Warm-start booster continuation Queried labels are treated as oracle/analyst feedback. Rather than retraining from scratch, we update the booster by appending a bounded number of new trees to the existing ensemble [61]. Two safeguards keep these updates stable. First, pending labels accumulate across triggers until the minimum update batch of Bmin = 32 labels is reached, which avoids single-sample updates that would destabilize the booster. Second, an event-based cooldown period after each update reduces repeated-label leakage from temporally correlated network segments. We fix Bmin = 32 as a practical stability safeguard; sensitivity to this setting is left
Algorithm 1 PACT streaming loop. 1: Train frozen XGBoost-Focal core on historical split 2: Select operating threshold θ on train-tail data 3: Initialize buffer B , pending set P ← ∅, ADWIN, cooldown c ← 0 4: for each stream batch Xt do 5: Predict p(x) and ŷ(x) = 1[p(x) ≥ θ] 6: Update rolling metrics; append Xt to B 7: Q←∅ 8: if strategy is ADWIN-random or ADWIN-hybrid then 9: Update ADWIN with predicted probabilities 10: if ADWIN triggers and c = 0 then 11: Q ← S ELECT Q UERY BATCH(B ; random or hybrid) 12: end if 13: else if strategy is periodic and update interval reached then 14: if c = 0 then 15: Q ← S ELECT Q UERY BATCH(B ; uniform random) 16: end if 17: end if 18: if Q ̸= ∅ then 19: Obtain oracle labels for Q, add to P 20: end if 21: if |P| ≥ Bmin then 22: Warm-start XGBoost: append bounded trees using P 23: P ← ∅; c ← cooldown duration 24: end if 25: c ← max(c − |Xt |, 0) 26: end for 27: Report final metrics over the full stream
to future work. Algorithm 1 gives the full streaming loop; pending labels accumulate across consecutive triggers until |P| ≥ Bmin , at which point the booster is updated. A consequence of the minimum-batch safeguard is that
the per-trigger budget is best read as a nominal per-trigger buffer budget, not as a strict global stream-labeling percentage; in low-budget settings, the realized stream-level rate can exceed the nominal percentage. We therefore report the realized query rate alongside every result.
4. Experimental Setup This section describes the two public alert benchmarks and their chronological splits, the backbone selection protocol used to choose the frozen XGBoost-Focal core, the streaming simulator parameters and strategies, and the metrics used to characterize endpoint behavior.
definitions. Hyperparameters are tuned with Optuna (treestructured Parzen estimator multivariate sampler, MedianPruner, 5 startup trials, 15 trials/variant/fold, inner-validation proportion 0.20, selection metric validation F1 ). Tuned parameters span learning rate, max depth, estimators, and minimum child weight, with focal γ and α additionally tuned for the focal variant.2 For AIT-ADS we use a 2-fold chronological expanding-window evaluation. For BOTSv1, strict chronological folding produces too few valid positivesupport folds, so we use 5-fold stratified holdout for imbalance stress testing and do not draw temporal claims from BOTSv1 offline folds.
4.3. Streaming simulator parameters 4.1. Datasets and chronological splits We use two public low-prevalence security alert benchmarks. The AIT-ADS is an alert/log dataset combining host-side AMiner and Wazuh outputs with Suricata IDS alerts [64]. BOTSv1 is a Splunk security-event dataset spanning network and endpoint telemetry [65]. We treat both as alert/security-event streams, not as modality-pure host-only or network-only sources. Our evaluation targets operational alert-screening behavior on these mixed telemetry-derived streams. Table 1 summarizes their volumetric properties,
Table 3 reports the streaming simulator parameters. Rolling metrics are computed over a trailing window of 10,000 events. The recent buffer holds 5,000 events. ADWIN uses δ = 0.002. The minimum update batch is 32 labels; warm-start updates append 10 boosting rounds, capped at a 500-tree maximum. Cooldown after each update is 2,000 events. The periodic-update interval is 10,000 events. Operating thresholds are selected on a 0.20 train-tail using a maximum-F1 criterion over a 101-point grid; no post-update threshold recalibration is performed in the main runs.
TABLE 1. DATASET VOLUMES AND MALICIOUS PREVALENCE .
TABLE 3. S TREAMING SIMULATOR CONFIGURATION .
Dataset
Source telemetry
Total Malicious Prev. (%)
AIT-ADS Multi-source IDS alerts 587,943 BOTSv1 Splunk security events 5,078,376
5,821 32,686
0.99 0.64
TABLE 2. C HRONOLOGICAL TRAIN AND STREAM SPLITS .
Dataset
Train Train Rows Pos.
AIT-ADS 100,787 BOTSv1 1,707,436
Stream Stream Str. Prev. Rows Pos. (%)
100 487,156 5,721 100 3,370,940 32,586
1.17 0.97
and Table 2 reports the exact chronological train and stream splits used in the streaming experiments. Both datasets exhibit extended peacetime intervals: AIT-ADS contains recurrent attack spikes across 23 days, while BOTSv1 contains a concentrated 47-minute attack burst following 28 dormant days within its 29-day span. Both streams are initialized with low-positive warm starts (100 training positives) so the streaming evaluation begins primarily from benigndominated training data.
4.2. Backbone selection protocol We compare three XGBoost variants, namely plain cross-entropy, class-weighted (scale_pos_weight), and focal loss, under each dataset’s chosen split protocol, sharing identical preprocessing, hyperparameter search budgets, validation policy, threshold-selection rule, and metric
Parameter
Value
Batch size Rolling metric window Recent buffer size ADWIN δ Minimum update batch (Bmin ) Boosting rounds per update Initial boosting rounds Maximum total trees Cooldown period Periodic update interval Random seed XGBoost objective XGBoost max depth / lr / tree Threshold policy
1,000 events 10,000 events 5,000 events 0.002 32 labels 10 100 500 2,000 events 10,000 events 42 binary:logistic 6 / 0.10 / hist Train-tail max-F1
4.4. Streaming strategies We compare four strategies in the streaming loop: 1) 2)
Frozen: the trained core deployed continuously with zero updates. Periodic: warm-start updates at fixed temporal intervals using uniform random sampling, bypassing ADWIN.
2. Search spaces: plain/weighted: lr ∈ [0.05, 0.20] log-uniform, depth ∈ [4, 8], estimators ∈ [100, 250] (step 50), min-child-weight ∈ [1, 5]; focal: lr ∈ [0.02, 0.20], depth ∈ [4, 9], estimators ∈ [100, 350] (step 50), minchild-weight ∈ [1, 6], γ ∈ [0.50, 2.50], α ∈ [0.20, 0.80]. Fixed: subsample 0.90, colsample 0.90, L2 reg 1.00.
3) 4)
ADWIN-random: ADWIN-triggered warm-start updates, but querying the oracle by uniform random sampling rather than the hybrid rule. ADWIN-hybrid: ADWIN-triggered warm-start updates with the hybrid uncertainty-plus-high-score acquisition rule.
ADWIN-hybrid is the full PACT configuration; the other three strategies are baselines and ablations. Frozen disables all updates; Periodic and ADWIN-random each omit one of the two PACT components (the score-shift trigger or the hybrid query rule, respectively). In the results that follow, we refer to the full configuration as PACT when reporting headline outcomes and as ADWIN-hybrid when contrasting it against the other strategies in the same table.
4.5. Metrics We report rolling F1 , precision, recall, and FPR over the trailing 10,000-event window, but treat them as undefined when the window contains no positive samples (which occurs at the end of both streams because the attack spikes are concentrated mid-stream). The primary operational endpoint is the cumulative benign-normalized FP burden, i.e. FPs per million benign events (FP/1M benign). This metric is more directly tied to analyst-facing alert volume than the all-events FPR because its denominator is peacetime traffic, and it approximates the false-alert volume that drives analyst queues under routine, low-prevalence conditions. To bound the safety cost of any FP reduction, we report two attack-window measures. Let W + denote the set of trailing windows of size W = 10,000 events that contain at least one positive sample, and let TPw and FNw denote the true positives and false negatives within window w. The average positive-window recall is the mean of rolling recall over W + , X TPw 1 , (5) Rec+ = + |W | TP + FNw w + w∈W
and remains well defined even when the final window contains no positives. We also report cumulative missed positives, the total number of stream events with y = 1 and ŷ = 0. For the matched-trigger ablation in Section 5.3.2, we additionally report the maximum missed-positive streak (the longest run of consecutive missed positives) and the mean burst-detection delay (the mean number of events between a positive’s arrival and its first within-burst detection). For each run we report cumulative queries, applied positive and negative labels, total updates, and the realized query rate (cumulative queries divided by total stream events).
5. Results Results proceed in three stages. We first characterize offline separability and select the frozen core (Section 5.1), then project that core onto a low-prevalence deployment scenario (Section 5.2). Section 5.3 reports endpoint metrics for the four streaming strategies and includes two ablations:
a threshold-only operating-point check (Section 5.3.1) and a matched-trigger acquisition comparison (Section 5.3.2). Sections 5.4 and 5.5 then report realized labeling cost and per-seed dispersion.
5.1. Backbone characterization (RQ1) To characterize the offline separability of the two streams, Table 4 reports XGBoost loss-variant performance. We additionally ran broader offline static baselines on AIT-ADS: XGBoost, random-forest, and an unsupervised Isolation Forest under five-fold evaluation. XGBoost and random-forest variants reach F1 between 95.98% and 96.24% at FPR between 0.0350% and 0.0371%. Isolation Forest collapses to F1 = 68.15 ± 16.77% at FPR = 0.70±0.11% despite comparable recall (98.24±1.94%).3 This confirms that the AIT-ADS feature representation is highly separable for supervised classifiers, and that an unsupervised detector is not competitive on burden under this prevalence regime. TABLE 4. O FFLINE XGB OOST LOSS - VARIANT COMPARISON . Dataset
Variant
Recall (%)
FPR (%)
Precision (%)
F1 (%)
Plain 76.02±31.90 0.0145±0.0096 98.31±0.42 83.89±20.79 AIT-ADS Weighted 76.05±32.15 0.0145±0.0096 98.31±0.42 83.88±20.95 (2-fold chron.) Focal 76.34±31.44 0.0150±0.0088 98.20±0.58 84.13±20.45 BOTSv1 (5-fold strat.)
Plain 95.63±5.72 0.0654±0.0486 90.57±6.74 Weighted 95.64±5.71 0.0654±0.0487 90.57±6.74 Focal 95.66±5.70 0.0654±0.0486 90.57±6.73
93.01±6.13 93.01±6.13 93.02±6.11
Table 4 compares the three XGBoost loss variants on both datasets. On AIT-ADS (2-fold chronological), focal loss is marginally above plain and weighted (F1 84.13% vs. 83.89% and 83.88%), with FPR between 0.0145% and 0.0150% across variants. The high cross-fold standard deviation reflects the small number of valid temporal folds and the recurrent attack-spike topology. On BOTSv1 (5-fold stratified holdout), the variants converge tightly: F1 ≈ 93.0%, FPR 0.0654%, with focal at 93.02 ± 6.11%. The crossvariant gaps in Table 4 are smaller than the cross-fold standard deviation, so the loss-variant choice is not statistically distinguishable on these splits. We select XGBoost-Focal as the streaming core on the imbalance-handling motivation of the framework rather than on a statistical preference: it is consistently competitive across both protocols, attains the highest mean F1 on each dataset, and avoids dataset-specific loss choices. Implication. Both datasets are well separable offline by a strong tabular classifier under the leakage-mitigation filter. The streaming evaluation must therefore be interpreted relative to an already strong frozen detector, not relative to a weak baseline. The static supervised baselines show high separability, while the chronological streaming setting remains challenging for the same backbones. 3. Full per-family offline tables are included in the artifact package.
5.2. Offline pre-deployment projection (RQ2) As a pre-deployment stress test, we project each focal core to a 0.10% deployment prior and 1,000,000 daily events using the cross-validated recall and FPR from Table 4. TABLE 5. O FFLINE PRE - DEPLOYMENT PROJECTION AT A 0.10% PRIOR .
Dataset
Core True Alerts False Alerts Precision (%)
AIT-ADS Focal BOTSv1 Focal
763 956
149 653
83.66 59.42
Table 5 4 shows that even sub-0.1% FPRs translate into hundreds of daily false alerts (149 for AIT-ADS, 653 for BOTSv1) at deployment precisions of 83.66% and 59.42% respectively, an asymmetric alert-level signal-to-noise ratio of 5.1:1 versus 1.5:1. This projection is not the measured stream-time burden; the chronological streaming split is harder and temporally shifted, so Section 5.3 uses realized cumulative FP/1M benign as the primary operational endpoint.
5.3. Streaming-strategy comparison (RQ3) The main empirical results are reported in two stages. We first report single-seed endpoint metrics (Table 6), which characterize the stream-level trajectory of each strategy. We then report a multi-seed Pareto summary (Table 7) that adds two safety-critical measures, average positive-window recall and cumulative missed positives, which jointly bound the operational cost of any FP reduction. Because endpoint F1 , precision, and recall are undefined at the final rolling window of both streams (Section 4.5), the primary endpoint is cumulative benign-normalized FP burden paired with positive-window recall. Hybrid acquisition is the lowest-FP adaptive strategy in these streams. On BOTSv1, PACT reduces FP/1M benign from 27,023 (frozen) to 21,271, a roughly 21% reduction; on AIT-ADS, the reduction is larger, from 9,532 to 5,467 (roughly 43%). Among the adaptive update strategies in Table 7, PACT attains the lowest FP/1M benign on both datasets, by substantial margins: 5,467 versus 9,532 for periodic and 16,590 for ADWIN-random on AIT-ADS, and 21,271 versus 316,540 for periodic and 999,455 for ADWIN-random on BOTSv1. The rolling benign-FPR trajectories (Fig. 2a,b) show that PACT stays below the frozen line for most of both streams. FP reduction trades off with positive-window recall. The positive-window recall column of Table 7 makes the trade-off explicit. On AIT-ADS, PACT attains 72.24% average positive-window recall versus 82.37% for periodic and 82.40% for ADWIN-random; cumulative missed positives are comparable across the three strategies (45, 41, and 37 4. True/False Alert counts are projected from Table 4 as TP = ⌊Recall × N+ ⌋ and FP = ⌊FPR × N− ⌋ with N+ = 1,000 and N− = 999,000 at the 0.10% prior. Precision is computed from the integer counts as TP/(TP + FP).
respectively). On BOTSv1, PACT attains 89.09% positivewindow recall and misses 3,280 positives, whereas periodic and ADWIN-random both attain 100% window recall and miss none. On AIT-ADS the two metrics tell slightly different stories: cumulative missed positives differ by fewer than ten across strategies, but average positive-window recall weights each positive-support window equally, so a few sparse low-recall windows pull the mean down. We report both: missed positives count the absolute safety cost; positive-window recall captures whether attack detection is uniform across the stream. PACT is not a dominance claim; it is an operating point for SOCs that constrain analyst workload and can tolerate the measured recall cost. An operator prioritizing recall over workload could prefer periodic or ADWIN-random on AITADS (roughly 82% versus 72% recall). On BOTSv1, however, those two strategies preserve recall only at FP volumes more than an order of magnitude above frozen (over 300,000 and nearly one million FP/1M benign, respectively). Mechanism observation. ADWIN-random isolates the trigger from the query rule. On BOTSv1, its rolling FPR reaches 100% during sustained portions of the stream (Fig. 2a), meaning every benign event in the window is being flagged, which suggests that trigger timing alone is insufficient. Table 6 and Fig. 3 show that PACT yields positive-enriched update batches, but ADWIN-random also obtains many positives on BOTSv1 and still fails. Positive density alone therefore does not explain the hybrid gain; alignment with the operating threshold also matters. Periodic updating, lacking a score-shift trigger, applies batches that are only 7.0% and 0.7% positive, respectively, dominated by benign noise. 5.3.1. Threshold-only operating-point check. To distinguish adaptation from operating-point movement, we compare PACT with a frozen threshold-only baseline. The baseline uses the same XGBoost-Focal core but selects a stricter (higher) operating threshold from the train-tail validation curve (the highest threshold that retains ≥ 95% of traintail recall), with no streaming updates. The original maxF1 thresholds and recall-constrained thresholds were θ = 0.0027 → 0.1854 on AIT-ADS and θ = 5.95×10−6 → 0.3622 on BOTSv1. Table 8 reports the comparison. Pure threshold movement is not an acceptable alertfatigue remedy on these streams under the evaluated recall constraints. On AIT-ADS it cuts FP/1M benign by 94% (9,532 to 536) but raises missed positives from 41 to 319 and lowers positive-window recall from 82.37% to 69.72%. On BOTSv1 the collapse is starker: FP/1M benign falls to 28 (about 1,000× below frozen), but positive-window recall drops from 89.09% to 33.73% and missed positives jump from 3,280 to 22,273. On BOTSv1, Frozen and PACT shares identical positive-window recall (89.09%) and missed-positive count (3,280) with threshold-only because PACT’s updates fire predominantly outside the 47-minute attack burst and therefore do not change within-burst classification; only stream-time FP burden differs. PACT does not reach the threshold-only FP floor on either dataset; its role
TABLE 6. S INGLE - SEED STREAMING ENDPOINTS AT A 1.00% NOMINAL BUDGET.
Rolling Benign FPR (%)
Dataset
Strategy
BOTSv1
Frozen Periodic ADWIN-random ADWIN-hybrid
Cum. FP / 1M Applied Queries Updates FP Benign Pos / Neg
2.65 90,213 31.59 1,091,055 100.00 3,336,536 0.56 71,010
Frozen Periodic AIT-ADS ADWIN-random ADWIN-hybrid
1.19 1.19 2.40 1.03
4,589 4,589 10,598 2,632
27,023 326,824 999,455 21,271
0 2,000 782 382
0 40 16 8
0/0 140 / 1,860 352 / 430 266 / 116
9,532 9,532 22,013 5,467
0 2,000 782 532
0 40 16 11
0/0 14 / 1,986 76 / 706 216 / 316
AIT-ADS Rolling Benign False-Positive Rate
Rolling Benign False-Positive Rate
BOTSv1 1.0 0.8 Frozen Periodic ADWIN-random ADWIN-hybrid
0.6 0.4 0.2 0.0
0K
500K
1,000K
1,500K
2,000K
Events Processed
2,500K
3,000K
3,500K
Frozen Periodic ADWIN-random ADWIN-hybrid
0.08 0.06 0.04 0.02 0.00
0K
100K
(a) BOTSv1 rolling benign FPR
200K
300K
Events Processed
AIT-ADS: Cumulative FP Burden
BOTSv1: Cumulative FP Burden FPs per 1M Benign Events
FPs per 1M Benign Events
800,000 Frozen Periodic ADWIN-random ADWIN-hybrid
400,000 200,000 0
0K
500K
500K
(b) AIT-ADS rolling benign FPR
1,000,000
600,000
400K
40,000 30,000 20,000 10,000 0
1,000K 1,500K 2,000K 2,500K 3,000K 3,500K
Events Processed
(c) BOTSv1 cumulative FP/1M benign
Frozen Periodic ADWIN-random ADWIN-hybrid
50,000
0K
100K
200K
300K
Events Processed
400K
500K
(d) AIT-ADS cumulative FP/1M benign
Figure 2. FP trajectories at a 1.00% nominal per-trigger budget.
Update Batch Label Composition: Applied Labels AIT-ADS BOTSv1 2,000
Applied Update Labels
2,000
Neg (benign) Pos (positive/malicious)
1,500 1,000
782
782 532
500
382
id Nhy WI AD
nd Nra WI AD
Figure 3. Applied update-label composition.
br
om
ic od Pe ri
br Nhy WI AD
AD
WI
Nra
nd
od
ic
om
id
0
Pe ri
Applied Update Labels
2,000
is to occupy a middle operating point with lower FP burden than frozen, far less recall collapse than threshold-only, and lower analyst-query cost than periodic updating. This reframes the headline claim: PACT is not the absolute lowestFP operating point on these streams. It is a less destructive operating point on the FP, recall, and query Pareto frontier than frozen, periodic, ADWIN-random, or threshold-only FP minimization. 5.3.2. Matched-trigger acquisition ablation (RQ4). The single-seed and multi-seed comparisons above co-vary trigger timing, update count, and acquisition rule. To isolate the effect of the acquisition rule, we replay the same ADWIN trigger schedule used by ADWIN-hybrid and substitute four
TABLE 9. M ATCHED - TRIGGER ACQUISITION ABLATION .
TABLE 7. M ULTI - SEED PARETO SUMMARY. Q U ↓
91.03 82.37 79.91 71.90
BOTSv1
Random 999,822 Uncertainty 27,023 High-score 21,739 Hybrid 21,271
100.00 89.09 89.09 89.09
Median over three seeds. Pos-win. rec: average positivewindow recall (mean of rolling recall restricted to windows that contain at least one positive sample). Q: queries. U: updates. Missed positives count stream events with y = 1 and ŷ = 0. The recall column reports the cost incurred at the lowest-FP operating point. Superscripts denote per-strategy interquartile range (IQR) on FP/1M benign; strategies without a superscript had zero IQR across the audited seeds. a FP/1M benign IQR 2,781. b FP/1M benign IQR 76,299. c FP/1M benign IQR 221,477.
FP/1M Pos-win. Missed Queries Ben. ↓ rec. % ↑ pos. ↓
Because updates require at least Bmin = 32 accumulated labels and triggers do not scale linearly with stream length, the realized stream-level query rate can differ substantially from the nominal per-trigger budget. We therefore report the realized query rate directly. Fig. 4 visualizes realized
query policies in turn: random, uncertainty-only, high-scoreonly, and hybrid. Table 9 reports the result.
Missed-positive characterization. Missed-positive diagnostics (Table 9) show no extended outage: AIT-ADS has maximum missed-positive streaks of 2 events under all four policies, while BOTSv1 has a maximum streak of 32 events with near-zero mean burst-detection delay, indicating prompt burst detection despite within-burst drops.
BOTSv1 Strategies
u=40
0.400%
Frozen Periodic ADWIN-random ADWIN-hybrid
0.300% 0.200%
u=16 u=11
0.100%
u=40
AD
WI
N-
ran
u=8
br id
od ic Pe ri
en Fr
u=16
u=0
Nhy
u=0
0.000%
oz
Realized Query Rate (% of Stream Events)
Hybrid acquisition is the lowest-FP policy on both datasets, supporting the view that the gain reported in Table 6 is at least partly attributable to acquisition rather than to trigger timing alone. On BOTSv1, hybrid, uncertainty, and high-score acquisition all attain identical positive-window recall and missed-positive counts, leaving FP burden as the only differentiator; hybrid is lowest on that endpoint. On AIT-ADS, where positives are not concentrated in a single burst, the four policies instead trace a clean Pareto front. Random has the highest recall but the highest FP burden, hybrid has the lowest FP burden but the lowest recall, and uncertainty and high-score sit in between. The clean front is therefore AIT-ADS-specific; on BOTSv1, the attack-burst topology collapses three of the four policies onto a shared recall plateau, consistent with Table 7.
Realized Annotation Burden: AIT-ADS, BOTSv1 AIT-ADS
WI
0 0 382
AD
89.09 3,280 33.73 22,273 89.09 3,280
5.4. Realized labeling cost (RQ5)
do m
0 0 532
ran
41 319 45
N-
Frozen 27,023 BOTSv1 Frozen threshold-only 28 ADWIN-hybrid 21,271
82.37 69.72 72.24
WI
9,532 536 5,467
0 1,222 25 3,280 532 11 3,280 482 10 3,280 482 10
AD
Frozen AIT-ADS Frozen threshold-only ADWIN-hybrid
872 18 722 15 672 14 672 14
od ic
Strategy
34 41 152 101
Q: queries; U: updates. Trigger timestamps were recorded from a reference ADWIN-hybrid run; update counts may differ slightly across rows because of post-query deduplication and minimumbatch eligibility. The AIT-ADS hybrid row differs from the free-running ADWIN-hybrid endpoint in Table 6 (5,467 FP/1M benign, 11 updates, 532 queries, 45 missed positives) because forced-schedule replay changes update eligibility after deduplication, pending-label accumulation, and cooldown; on BOTSv1 the two runs reach identical FP/1M benign, positive-window recall, and missed-positive endpoints, though query and update counts differ slightly (382 vs. 482 queries, 8 vs. 10 updates) for the same eligibility reasons. The ablation’s purpose is to compare acquisition policies under identical trigger timestamps.
TABLE 8. T HRESHOLD - ONLY OPERATING - POINT ABLATION . Dataset
Q U ↓
Pe ri
0 2,000 40 0 782 16 3,280 382 8
22,013 9,532 8,701 6,148
en
100.00 100.00 89.09
Random Uncertainty High-score Hybrid
oz
Periodic 316,540b BOTSv1 ADWIN-random 999,455c ADWIN-hybrid 21,271
AIT-ADS
Fr
41 2,000 40 37 682 14 45 532 11
FP/1M Pos-win. Missed Ben. ↓ rec. % ↑ pos. ↓
br id
82.37 82.40 72.24
Acquisition
Nhy
9,532 16,590a 5,467
Dataset
WI
Periodic AIT-ADS ADWIN-random ADWIN-hybrid
FP/1M Pos-win. Missed Ben. ↓ rec. % ↑ pos. ↓
AD
Strategy
do m
Dataset
Figure 4. Realized annotation burden.
rates and update counts. On AIT-ADS, the periodic strategy at a 1.00% nominal budget yields a realized rate of 0.41% over the full stream (2,000 queries across 487,156 stream events), because triggers fire at fixed intervals rather than scaling with stream length. On BOTSv1, the same nominal budget yields only 0.06% realized rate (2,000 queries across 3,370,940 events). Among adaptive strategies, PACT produces the lowest realized rate on both datasets: 0.01% on BOTSv1 (382 queries) and 0.11% on AIT-ADS (532 queries).
Operational summary. On a 3.4M-event BOTSv1 stream, PACT attains the lowest FP burden among the adaptive strategies using 382 queries and 8 updates. This corresponds to roughly one analyst review per 9,000 events on average, with strongly positive-enriched query batches concentrated in the attack-burst window rather than spread uniformly across the stream. On the smaller 487K-event AIT-ADS stream, the corresponding cost is 532 queries and 11 updates. Compared to periodic updating, PACT uses 5.2× fewer queries on BOTSv1 and 3.8× fewer on AITADS, at the recall cost reported in Section 5.3.
5.5. Per-seed dispersion The three-seed audit is intended as a dispersion check, not a claim of broad statistical robustness. PACT showed zero IQR on FP/1M benign, positive-window recall, queries, and updates across the audited seeds, while BOTSv1 periodic and ADWIN-random showed substantial FP/1M dispersion (IQR values reported in Table 7). In all cases the seed-level minimum of periodic and random remains well above the PACT median.
6. Discussion Strength of the frozen baseline. A well-trained imbalance-aware backbone with a calibrated threshold already achieves low FP burden on both datasets (1.19% rolling benign FPR on AIT-ADS; 2.65% on BOTSv1) before any adaptation. The offline AIT-ADS baselines (Section 5.1) show that AIT-ADS data is highly separable, leaving limited room for adaptation to improve endpoint FP burden. Adaptive updates therefore must improve on a strong frozen detector through targeted label acquisition rather than through update volume alone. Inference cost is bounded by design. The 500-tree ensemble cap fixes the per-event prediction footprint, and each warm-start trigger appends at most ten trees within the event-based cooldown horizon. PACT therefore has a bounded prediction footprint even on a one-million-event-per-day stream. Role of acquisition policy. The matched-trigger ablation (Table 9) holds the ADWIN trigger schedule fixed and varies only the query rule. Hybrid acquisition attains the lowest FP/1M benign on both datasets while random sampling at the same trigger times produces roughly 47× more FPs on BOTSv1 (999,822 FP/1M benign). The mechanism is best read not as “random batches are mostly benign” (Table 6 shows BOTSv1 ADWIN-random batches are about 45% positive) but as trigger timing alone is insufficient: even when triggered batches contain positives, uniformly sampled feedback can be poorly aligned with the operating threshold. The hybrid rule’s high-score half supplies positive-enriched evidence, while its threshold-relative half preserves boundary exploration. The cost, visible in Table 9 when the trigger schedule is held fixed, is roughly 19 percentage points of average positive-window recall on AIT-ADS relative to matched-schedule random (71.90% vs. 91.03%). Under freerunning triggers (Table 7), the cost is closer to ten percentage
points because each policy sets its own update schedule. The precision-recall trade-off pattern is itself consistent with classical active-learning theory [49], [50]; what we add is its quantification on SOC alert streams under benignnormalized FP burden. Failure modes of periodic updating. On AIT-ADS, periodic updating exactly reproduces the frozen endpoint because uniform-random batches are dominated by benign samples (14 positives in 2,000 queries). On BOTSv1, the failure is severe because the attack burst is concentrated in a 47-minute window of the 29-day stream, so periodic sampling is schedule-correlated with stream topology in a way no unsupervised query rule can correct. Relation to production-SOC alert triage. Automated Alert Classification and Triage (AACT) [19] reports alert reduction on production SOC telemetry using supervised triage-action labels, AlertPro [18] uses reinforcement learning over multi-step attack-graph structures, and the L2DHF learning-to-defer framework [20] requires analyst-deferral feedback signals. None of these inputs is available in the public AIT-ADS/BOTSv1 benchmarks, so we do not run direct empirical comparisons; the protocols are not commensurable. PACT instead studies bounded triggered querying around a deployed screener under the labels that public alert benchmarks expose. All four lines of work share the underlying point that alert reduction must be interpreted together with false-negative or recall cost. Operational implications. On these streams, the results suggest that no single-score selection captures the full operating-point picture: FP/1M benign captures investigation load, positive-window recall captures attack-window safety, and realized query rate captures analyst labeling cost. PACT is the most attractive of the tested strategies when labeling capacity and false-alert volume are constrained, but it is not universally preferable. Threshold-only tuning lowers FP burden further at unacceptable recall cost on BOTSv1, while periodic and ADWIN-random preserve recall there only with FP volumes more than an order of magnitude above frozen. The reported operating points together expose the trade-offs a SOC must weigh against its analyst capacity and missed-positive tolerance.
7. Limitations and Scope The headline claims apply to AIT-ADS/BOTSv1-like low-prevalence alert streams, with stream prevalences under 1.5% and attack patterns concentrated mid-stream, rather than to SOC streams in general. High-prevalence SIEM incident streams raise distinct calibration and thresholdmanagement issues outside this paper’s scope. ADWIN here is a score-shift trigger over the predicted-probability stream, not a labeled drift detector. Warm-start updates append bounded numbers of trees rather than retraining from scratch. The minimum-update-batch (Bmin = 32) and event-based cooldown safeguards reduce but do not eliminate biased small-batch updates. A replay-buffer variant (Appendix B) did not improve the operational profile, and
lifecycle policies after the maximum-tree cap are out of scope. The three-seed audit (seeds 40–42) is a limited sweep, not evidence of broad statistical robustness; broader sweeps will be included in the artifact package. With only two benchmarks the reported pattern is a two-point existence proof, not a generalization; broader dataset coverage would clarify whether the AIT-ADS-style clean Pareto and the BOTSv1-style three-way recall collapse are systematic or dataset-specific. The matched-trigger ablation in Section 5.3.2 controls trigger timing but not update count exactly, because deduplication and minimum-batch eligibility can change update timing; query and update counts are reported alongside FP burden so the residual difference between policies is visible. The hybrid rule uses a fixed 50/50 split, ADWIN uses δ = 0.002, and Bmin = 32 throughout, with sensitivity sweeps left to future work. The streaming evaluation runs in an offline simulator, so perevent scoring latency and warm-start update throughput on production telemetry are not characterized here. Finally, the feature filter uses a conservative denylist plus datasetspecific cleanup (Appendix A); per-dataset allowlists derived from operational alert schemas would strengthen deployability claims.
Ethical Considerations This study uses previously released benchmark datasets and an offline streaming simulator. It does not involve human subjects, live attacks, or intervention on production systems. Reported trade-offs concern alert-screening behavior under simulated deployment and are not used to automate real incident-response decisions. Per-dataset feature manifests are released so that operators can verify, before any deployment, that retained features are observable predecision in their environment.
References [1]
S. Bhatt, P. K. Manadhata, and L. Zomlot, “The operational role of security information and event management systems,” IEEE Security & Privacy, vol. 12, no. 5, pp. 35–41, 2014.
[2]
R. Gupta, S. Tanwar, S. Tyagi, and N. Kumar, “Machine learning models for secure data analytics: A taxonomy and threat model,” Computer Communications, vol. 153, pp. 406–440, 2020.
[3]
J. Zhu, S. He, J. Liu, P. He, Q. Xie, Z. Zheng, and M. R. Lyu, “Tools and benchmarks for automated log parsing,” in 2019 IEEE/ACM 41st International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 2019, pp. 121–130.
[4]
H. Debar and A. Wespi, “Aggregation and correlation of intrusiondetection alerts,” in International Workshop on Recent Advances in Intrusion Detection, Springer. Berlin Heidelberg: Springer-Verlag, 2001, pp. 85–103.
[5]
S. Ndichu, T. Ban, S. Ozawa, T. Takahashi, and D. Inoue, “AI-driven security alert screening and alert fatigue mitigation in security operations centers: A survey,” 2026. [Online]. Available: https://arxiv.org/abs/2605.08316
[6]
R. G. Bace, Intrusion detection.
[7]
H.-J. Liao, C.-H. R. Lin, Y.-C. Lin, and K.-Y. Tung, “Intrusion detection system: A comprehensive review,” Journal of Network and Computer Applications, vol. 36, no. 1, pp. 16–24, 2013. [Online]. Available: https://doi.org/10.1016/j.jnca.2012.09.004
[8]
Z. Ahmad, A. Shahid Khan, C. Wai Shiang, J. Abdullah, and F. Ahmad, “Network intrusion detection system: A systematic study of machine learning and deep learning approaches,” Transactions on Emerging Telecommunications Technologies, vol. 32, no. 1, p. e4150, 2021. [Online]. Available: https://doi.org/10.1002/ett.4150
[9]
T. Ban, T. Takahashi, S. Ndichu, and D. Inoue, “Breaking alert fatigue: AI-assisted SIEM framework for effective incident response,” Applied Sciences, vol. 13, no. 11, p. 6610, 2023. [Online]. Available: https://doi.org/10.3390/app13116610
8. Conclusion PACT recasts adaptive alert screening on low-prevalence streams as an operating-point choice. Two ingredients drive the result on the streams we study: applying ADWIN to the score stream rather than to labeled error, and splitting the query budget between threshold-relative uncertainty and high-score sampling. A matched-trigger ablation controls trigger timing and shows that the acquisition rule contributes beyond timing alone. The threshold-only baseline makes the limits of single-metric optimization explicit on these streams: it reaches a lower FP burden than every adaptive setting tested but collapses BOTSv1 recall by 55 percentage points. Reporting on AIT-ADS/BOTSv1-like low-prevalence streams should therefore surface FP burden, positive-window recall, and analyst labeling cost together. Among the strategies tested, an unsupervised score-shift trigger paired with positive-blind querying produced the least attractive operating point.
Data Availability BOTSv1 and AIT-ADS are public benchmark datasets available from their original providers. The artifact package will include split-generation scripts, the preprocessing and leakage-mitigation pipeline, the retained-feature manifest in machine-readable form, all run configuration files, raw perseed metric traces, and the figure and table generation scripts used to produce the reported results. Dataset files are not redistributed when prohibited by their original licenses; the artifact instead documents the exact download and preprocessing steps required to reproduce the splits used here.
Sams Publishing, 2000.
[10] L. Yang, Z. Chen, C. Wang, Z. Zhang, S. Booma, P. Cao, C. A. Withers, R. Iyer, D. Estrada, and G. Wang, “True attacks, attack attempts, or benign triggers? an empirical measurement of network alerts in a security operations center,” in Proceedings of the 33rd USENIX Security Symposium (USENIX Security 24). USENIX Association, 2024. [Online]. Available: https://www.usenix. org/conference/usenixsecurity24/presentation/yang-limin [11] A. Khraisat, I. Gondal, P. Vamplew, and J. Kamruzzaman, “Survey of intrusion detection systems: techniques, datasets and challenges,” Cybersecurity, vol. 2, no. 1, pp. 1–22, 2019. [Online]. Available: https://doi.org/10.1186/s42400-019-0038-7 [12] P. De and I. Nath, “Machine learning approaches on intrusion detection system: A holistic review,” in Advances in Communication, Devices and Networking, S. Dhar, D.-T. Do, S. N. Sur, and H. C.-M. Liu, Eds. Singapore: Springer Nature Singapore, 2023, pp. 387–400.
[13] T. Ban, S. Ndichu, T. Takahashi, and D. Inoue, “Combat security alert fatigue with AI-assisted techniques,” in Proceedings of the 14th Cyber Security Experimentation and Test Workshop, 2021, pp. 9–16. [14] G. Andresini, F. Pendlebury, F. Pierazzi, C. Loglisci, A. Appice, and L. Cavallaro, “INSOMNIA: Towards concept-drift robustness in network intrusion detection,” in Proceedings of the 14th ACM Workshop on Artificial Intelligence and Security (AISec). ACM, 2021, pp. 111–122. [Online]. Available: https://doi.org/10.1145/ 3474369.3486864 [15] F. Camarda, A. De Paola, S. Drago, P. Ferraro, and G. Lo Re, “Managing concept drift in online intrusion detection systems with active learning,” in Joint National Conference on Cybersecurity (ITASEC & SERICS 2025), ser. CEUR Workshop Proceedings, vol. 3962, 2025. [Online]. Available: https://ceur-ws.org/Vol-3962/ paper42.pdf
[29] T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in Proceedings of the IEEE international conference on computer vision. Venice, Italy: IEEE, 2017, pp. 2980– 2988. [30] M. Zaman and C.-H. Lung, “Evaluation of machine learning techniques for network intrusion detection,” in 2018 IEEE/IFIP Network Operations and Management Symposium (NOMS), IEEE. Taipei, Taiwan: IEEE, 2018, pp. 1–5. [31] Y. Chang, W. Li, and Z. Yang, “Network intrusion detection based on random forest and support vector machine,” in 2017 IEEE international conference on computational science and engineering (CSE) and IEEE international conference on embedded and ubiquitous computing (EUC), vol. 1, IEEE. Guangzhou, China: IEEE, 2017, pp. 635–638.
[16] R. Vaarandi and A. Guerra-Manzanares, “Network IDS alert classification with active learning techniques,” Journal of Information Security and Applications, vol. 81, p. 103687, 2024. [Online]. Available: https://doi.org/10.1016/j.jisa.2023.103687
[32] L. Zomlot, S. Chandran, D. Caragea, and X. Ou, “Aiding intrusion analysis using machine learning,” in 2013 12th International Conference on Machine Learning and Applications, vol. 2, IEEE. Miami, FL, USA: IEEE, 2013, pp. 40–47.
[17] W. U. Hassan, S. Guo, D. Li, Z. Chen, K. Jee, Z. Li, and A. Bates, “NoDoze: Combatting threat alert fatigue with automated provenance triage,” in Network and Distributed System Security Symposium (NDSS), 2019.
[33] R. S. S. Kumar, A. Wicker, and M. Swann, “Practical machine learning for cloud intrusion detection: Challenges and the way forward,” in Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security. Dallas Texas USA: ACM, 2017, pp. 81–90.
[18] X. Wang, X. Yang, X. Liang, X. Zhang, W. Zhang, and X. Gong, “Combating alert fatigue with AlertPro: Context-aware alert prioritization using reinforcement learning for multi-step attack detection,” Computers & Security, vol. 137, p. 103583, 2024. [Online]. Available: https://doi.org/10.1016/j.cose.2023.103583
[34] K. Narayana Rao, K. Venkata Rao, and P. R. P.V.G.D., “A hybrid intrusion detection system based on sparse autoencoder and deep neural network,” Computer Communications, vol. 180, pp. 77–88, 2021. [Online]. Available: https://doi.org/10.1016/j.comcom.2021.08. 026
[19] M. Turcotte, F. Labreche, and S.-O. Paquette, “Automated alert classification and triage (AACT): An intelligent system for the prioritisation of cybersecurity alerts,” 2025. [Online]. Available: https://arxiv.org/abs/2505.09843
[35] N. Shone, T. N. Ngoc, V. D. Phai, and Q. Shi, “A deep learning approach to network intrusion detection,” IEEE Transactions on Emerging Topics in Computational Intelligence, vol. 2, no. 1, pp. 41–50, 2018. [Online]. Available: https://doi.org/10.1109/TETCI. 2017.2772792
[20] F. Jalalvand, M. Baruwal Chhetri, S. Nepal, and C. Paris, “Adaptive alert prioritisation in security operations centres via learning to defer with human feedback,” 2025. [Online]. Available: https://arxiv.org/abs/2506.18462 [21] A. Milenkoski, M. Vieira, S. Kounev, A. Avritzer, and B. D. Payne, “Evaluating computer intrusion detection systems: A survey of common practices,” ACM Computing Surveys (CSUR), vol. 48, no. 1, pp. 1–41, 2015. [Online]. Available: https://doi.org/10.1145/2808691 [22] A. L. Buczak and E. Guven, “A survey of data mining and machine learning methods for cyber security intrusion detection,” IEEE Communications surveys & tutorials, vol. 18, no. 2, pp. 1153– 1176, 2015. [Online]. Available: https://doi.org/10.1109/COMST. 2015.2494502 [23] G. M. Weiss and F. Provost, “The effect of class distribution on classifier learning: an empirical study,” Rutgers University, Tech. Rep., 2001. [24] Y. Sun, A. K. Wong, and M. S. Kamel, “Classification of imbalanced data: A review,” International journal of pattern recognition and artificial intelligence, vol. 23, no. 04, pp. 687–719, 2009. [25] H. He and E. A. Garcia, “Learning from imbalanced data,” IEEE Transactions on knowledge and data engineering, vol. 21, no. 9, pp. 1263–1284, 2009. [26] N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, “SMOTE: synthetic minority over-sampling technique,” Journal of artificial intelligence research, vol. 16, pp. 321–357, 2002. [Online]. Available: https://doi.org/10.1613/jair.953 [27] X.-Y. Liu, J. Wu, and Z.-H. Zhou, “Exploratory undersampling for class-imbalance learning,” IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), vol. 39, no. 2, pp. 539–550, 2008. [28] K. R. M. Fernando and C. P. Tsokos, “Dynamically weighted balanced loss: Class imbalanced learning and confidence calibration of deep neural networks,” IEEE Transactions on Neural Networks and Learning Systems, vol. 33, no. 7, pp. 2940–2951, 2022. [Online]. Available: https://doi.org/10.1109/TNNLS.2020.3047335
[36] A. Bifet and R. Gavalda, “Learning from time-changing data with adaptive windowing,” in Proceedings of the 2007 SIAM international conference on data mining. SIAM, 2007, pp. 443–448. [37] D. N. Assis and V. M. A. Souza, “ADWIN-U: adaptive windowing for unsupervised drift detection on data streams,” Knowledge and Information Systems, vol. 67, pp. 10 005–10 034, 2025. [Online]. Available: https://doi.org/10.1007/s10115-025-02523-1 [38] X. Wang, J. Song, X. Zhang, J. Tang, W. Gao, and Q. Lin, “LogOnline: A semi-supervised log-based anomaly detector aided with online learning mechanism,” in 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE). IEEE, 2023, pp. 141–152. [39] M. E. Aminanto, L. Zhu, T. Ban, R. Isawa, T. Takahashi, and D. Inoue, “Combating threat-alert fatigue with online anomaly detection using isolation forest,” in Neural Information Processing: 26th International Conference, ICONIP 2019, Sydney, NSW, Australia, December 12–15, 2019, Proceedings, Part I 26. Springer, 2019, pp. 756–765. [40] M. E. Aminanto, T. Ban, R. Isawa, T. Takahashi, and D. Inoue, “Threat alert prioritization using isolation forest and stacked auto encoder with day-forward-chaining analysis,” IEEE Access, vol. 8, pp. 217 977–217 986, 2020. [Online]. Available: https: //doi.org/10.1109/ACCESS.2020.3041837 [41] W. Liu, C. Zhu, Z. Ding, H. Zhang, and Q. Liu, “Multiclass imbalanced and concept drift network traffic classification framework based on online active learning,” Engineering Applications of Artificial Intelligence, vol. 117, p. 105632, 2023. [Online]. Available: https://doi.org/10.1016/j.engappai.2022.105632 [42] W. Liu, H. Zhang, Z. Ding, Q. Liu, and C. Zhu, “A comprehensive active learning method for multiclass imbalanced data streams with concept drift,” Knowledge-Based Systems, vol. 215, p. 106778, 2021. [Online]. Available: https://doi.org/10.1016/j.knosys.2021.106778
[43] J. Pesek, D. Soukup, and T. Čejka, “Active learning framework to automate network traffic classification,” 2022. [Online]. Available: https://arxiv.org/abs/2211.08399
[59] R. Shittu, A. Healing, R. Ghanea-Hercock, R. Bloomfield, and M. Rajarajan, “Intrusion alert prioritisation and attack detection using postcorrelation analysis,” Computers & security, vol. 50, pp. 1–15, 2015.
[44] A. Shahraki, M. Abbasi, A. Taherkordi, and A. D. Jurcut, “Active learning for network traffic classification: A technical study,” 2021. [Online]. Available: https://arxiv.org/abs/2106.06933
[60] Y. Wang, Y. Guo, and C. Fang, “An end-to-end method for advanced persistent threats reconstruction in large-scale networks based on alert and log correlation,” Journal of Information Security and Applications, vol. 71, p. 103373, 2022.
[45] S. Das, M. R. Islam, N. Kannappan Jayakodi, and J. R. Doppa, “Effectiveness of tree-based ensembles for anomaly discovery: Insights, batch and streaming active learning,” 2019. [Online]. Available: https://arxiv.org/abs/1901.08930 [46] J. Montiel, R. Mitchell, E. Frank, B. Pfahringer, T. Abdessalem, and A. Bifet, “Adaptive XGBoost for evolving data streams,” in Proceedings of the International Joint Conference on Neural Networks (IJCNN). IEEE, 2020. [Online]. Available: https: //doi.org/10.1109/IJCNN48605.2020.9207555 [47] B. A. Alahmadi, L. Axon, and I. Martinovic, “99% false positives: A qualitative study of SOC analysts’ perspectives on security alarms,” in 31st USENIX Security Symposium (USENIX Security 22), 2022, pp. 2783–2800. [48] A. Madani, S. Rezayi, and H. Gharaee, “Log management comprehensive architecture in security operation center (SOC),” in 2011 International Conference on Computational Aspects of Social Networks (CASoN). IEEE, 2011, pp. 284–289. [49] R. M. Monarch, Human-in-the-Loop Machine Learning: Active learning and annotation for human-centered AI. Simon and Schuster, 2021. [50] M. Bilgic, L. Mihalkova, and L. Getoor, “Active learning for networked data,” in Proceedings of the 27th international conference on machine learning (ICML-10), 2010, pp. 79–86.
[61] T. Chen and C. Guestrin, “XGBoost: A scalable tree boosting system,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM, 2016, pp. 785– 794. [Online]. Available: https://doi.org/10.1145/2939672.2939785 [62] Mulyanto, S. W. Prakosa, M. Faisal, and J.-S. Leu, “Using optimized focal loss for imbalanced dataset on network intrusion detection system,” in 2022 IEEE 95th Vehicular Technology Conference: (VTC2022-Spring), 2022, pp. 1–7. [Online]. Available: https://doi.org/10.1109/VTC2022-Spring54318.2022.9861034 [63] A. S. Dina, A. Siddique, and D. Manivannan, “A deep learning approach for intrusion detection in internet of things using focal loss function,” Internet of Things, vol. 22, p. 100699, 2023. [Online]. Available: https://doi.org/10.1016/j.iot.2023.100699 [64] M. Landauer, F. Skopik, and M. Wurzenberger, “Introducing a new alert data set for multi-step attack analysis,” in Proceedings of the 17th Cyber Security Experimentation and Test Workshop (CSET ’24). Association for Computing Machinery, 2024, pp. 41–53. [Online]. Available: https://doi.org/10.1145/3675741.3675748 [65] Splunk Inc., “Boss of the SOC (BOTS) v1 Dataset,” Public dataset repository. https://github.com/splunk/botsv1, 2018, accessed: 202604-27.
[51] A. Beaugnon, P. Chifflier, and F. Bach, “ILAB: An interactive labelling strategy for intrusion detection,” in Research in Attacks, Intrusions, and Defenses (RAID 2017), ser. Lecture Notes in Computer Science, vol. 10453. Springer, 2017, pp. 120–140. [Online]. Available: https://doi.org/10.1007/978-3-319-66332-6 6
Appendix A. Feature-retention and leakage-mitigation manifest
[52] S. McElwee and J. Cannady, “Cyber situation awareness with active learning for intrusion detection,” in Proceedings of the IEEE International Conference on Big Data (Big Data). IEEE, 2019, pp. 3540–3549. [Online]. Available: https://doi.org/10.1109/ BigData47090.2019.9020599
To support reviewer-facing reproducibility for the leakage-mitigation filter described in Section 3, this appendix lists the exact features retained in the streaming experiments. The filter applies (i) an explicit drop-column set covering label, timestamp, split, and known post-decision attribute names, and (ii) a case-insensitive regex denylist on column names. Only columns present in both the train and stream parquet files that pass both filters are retained. No stream labels are used during retention. Explicit drop columns. label, timestamp, ts, datetime, date, split, fold_id, attack_type, dataset_name, time_group. Regex denylist (case-insensitive). Match (caseinsensitive) on column names containing any of: attack, verdict, malicious, suspicious, incriminated, dataset_name, time_group, split, fold. Retained features.
[53] R. Vaarandi and A. Guerra-Manzanares, “Stream clustering guided supervised learning for classifying NIDS alerts,” Future Generation Computer Systems, vol. 155, pp. 231–244, 2024. [Online]. Available: https://doi.org/10.1016/j.future.2024.01.032 [54] D. Du, Y. Li, Y. Cao, Y. Liu, G. Meng, N. Li, D. Han, and H. Feng, “FAF-BM: An approach for false alerts filtering using BERT model with semi-supervised active learning,” in Science of Cyber Security (SciSec 2024), ser. Lecture Notes in Computer Science, vol. 15441. Springer, 2024, pp. 295–312. [Online]. Available: https://doi.org/10.1007/978-981-96-2417-1 16 [55] L. Tong, A. Laszka, C. Yan, N. Zhang, and Y. Vorobeychik, “Finding needles in a moving haystack: Prioritizing alerts with adversarial reinforcement learning,” 2020. [Online]. Available: https://arxiv.org/abs/1906.08805 [56] M. Landauer, F. Skopik, M. Wurzenberger, and A. Rauber, “Dealing with security alert flooding: using machine learning for domainindependent alert aggregation,” ACM Transactions on Privacy and Security, vol. 25, no. 3, pp. 1–36, 2022. [57] M. Du, F. Li, G. Zheng, and V. Srikumar, “DeepLog: Anomaly detection and diagnosis from system logs through deep learning,” in Proceedings of the 2017 ACM SIGSAC conference on computer and communications security, 2017, pp. 1285–1298. [58] A. Hofmann and B. Sick, “Online intrusion alert aggregation with generative data stream modeling,” IEEE transactions on dependable and secure computing, vol. 8, no. 2, pp. 282–294, 2009.
•
•
AIT-ADS (5 features): feat_alert_category, feat_severity, feat_src_port, feat_dest_port, feat_time_since_last_alert. BOTSv1 (5 features): feat_severity, feat_src_port, feat_dest_port, feat_alert_category, feat_time_since_last_event.
A machine-readable copy of this manifest, together with the preprocessing-pipeline source, is included in the artifact package. Per-dataset allowlists derived from operational alert schemas remain a deployment-time activity.
Appendix B. Replay-buffer stabilizer: negative result To check whether a lightweight replay buffer would stabilize warm-start updates, we ran an additional ADWINhybrid configuration with a replay buffer of the most recent 512 labeled examples. Each queried batch was mixed with a random sample from the buffer at a replay ratio of r = 0.5. Table 10 reports the comparison against the no-replay configuration used in the main paper. On AITADS the replay variant increased FP/1M benign from 5,467 to 15,186 and increased cumulative missed positives from 45 to 107, although it slightly improved positive-window recall from 72.24% to 80.23%. On BOTSv1 the replay variant increased FP/1M benign from 21,271 to 28,046 with positive-window recall and missed-positive count unchanged but a small rise in query count (from 382 to 432) from
deduplication interactions with the replay buffer. We treat this as a negative robustness check; the simpler no-replay update used in the main paper is retained. TABLE 10. R EPLAY- BUFFER ABLATION UNDER ADWIN- HYBRID ( NO REPLAY VS . REPLAY BUFFER OF 512 EXAMPLES , r = 0.5 ). B OLD MARKS THE COLUMN - WISE BEST WITHIN EACH DATASET. FP / 1M Pos.-win. Missed Queries Benign ↓ rec. (%) ↑ pos. ↓
Dataset
Variant
AIT-ADS
No replay Replay512, r=0.5
5,467 15,186
72.24 80.23
45 107
532 532
BOTSv1
No replay Replay512, r=0.5
21,271 28,046
89.09 89.09
3,280 3,280
382 432
LLM Usage Statement LLMs were used for editorial purposes in this manuscript, and all outputs were inspected by the authors to ensure accuracy and originality. LLMs were not used to generate experimental results, run analyses, or write code that produced reported numbers.