Adversarial Evasion in Non-Stationary Malware Detection: Minimizing Drift Signals through Similarity-Constrained Perturbations
arXiv:2604.21310v1 [cs.CR] 23 Apr 2026
Pawan Acharya and Lan Zhang Northern Arizona University, Flagstaff, AZ, USA [email protected], [email protected]
Abstract. Deep learning has emerged as a powerful approach for malware detection, demonstrating impressive accuracy across various data representations. However, these models face critical limitations in realworld, non-stationary environments where both malware characteristics and detection systems continuously evolve. Our research investigates a fundamental security question: Can an attacker generate adversarial malware samples that simultaneously evade classification and remain inconspicuous to drift monitoring mechanisms? We propose a novel approach that generates targeted adversarial examples in the classifier’s standardized feature space, augmented with sophisticated similarity regularizers. By carefully constraining perturbations to maintain distributional similarity with clean malware, we create an optimization objective that balances targeted misclassification with drift signal minimization. We quantify the effectiveness of this approach by comprehensively comparing classifier output probabilities using multiple drift metrics. Our experiments demonstrate that similarity constraints can reduce output drift signals, with ℓ2 regularization showing the most promising results. We observe that perturbation budget significantly influences the evasiondetectability trade-off, with increased budget leading to higher attack success rates and more substantial drift indicators. Keywords: Adversarial malware · Dataset shift · Drift monitoring · Tabular malware detection · Similarity constraints
1
Introduction
In real-world cybersecurity environments, malware detection systems face a fundamental challenge: the inherently dynamic nature of malicious software [1]. Modern malware detection models predominantly operate on static, tabular feature vectors such as byte-frequency statistics and structural metadata which can be efficiently extracted at scale and integrated into standard classification pipelines. However, these systems are continuously challenged by the rapid evolution of malware, where attackers constantly modify their techniques through obfuscation, polymorphism, packing, and build-chain alterations [20].
2
Acharya and Zhang
To address this perpetual transformation, organizations have developed nonstationary malware detection systems with sophisticated drift monitoring mechanisms. These systems aim to detect and respond to significant changes in data distributions or model performance. Classic drift detection approaches include methods like DDM and EDDM, which monitor error statistics and trigger alerts when misclassification patterns deviate from historical baselines [4,3]. When ground truth labels are delayed or unavailable, unsupervised drift monitors employ statistical techniques such as Kolmogorov-Smirnov(KS) tests, JensenShannon(JS) divergence, Wasserstein distance, and Maximum Mean Discrepancy(MMD) to compare data distributions across different time windows [8,19]. The landscape of adversarial malware generation has evolved to address detection challenges through various sophisticated approaches. Researchers have developed gradient-based attacks specifically adapted to malware models [12,20], creating methods that can manipulate feature representations to evade detection. Search-based approaches utilizing evolutionary algorithms have emerged, allowing for more nuanced modifications of malware samples [2]. Reinforcement learning frameworks have further advanced these techniques [20]. Existing adversarial malware generation methods remain limited by key constraints. Most target detection models in a static setting, assuming models do not change over time. This overlooks concept drift in real-world malware detection, where malware and defenses continually co-evolve. Many approaches struggle to preserve malware semantics while achieving evasion, and they rarely evaluate whether generated adversarial samples remain similar to genuine malware. This context gives rise to a pivotal research question: In a non-stationary malware detection environment, can an attacker generate adversarial samples that not only evade classification but also remain inconspicuous to drift monitoring mechanisms? This challenge is especially acute in adversarial security, where uncoordinated distribution shifts can be deliberately exploited. To address this challenge, we propose a novel approach that generates targeted adversarial examples in the classifier’s standardized feature space, augmented with sophisticated similarity regularizers including Kullback-Leibler (KL) divergence, ℓ2 distance, and MMD. By carefully constraining perturbations to maintain distributional similarity with clean malware, we create an optimization objective that balances targeted misclassification with drift signal minimization. We quantify the effectiveness of this approach by comprehensively comparing classifier output probabilities using multiple drift metrics, including JS Divergence, Hellinger Distance, Wasserstein Metric, and others, thereby systematically probing the vulnerabilities of current concept drift detection strategies. Our experimental findings reveal three key insights into adversarial malware generation. Under a constrained low-budget setting, different similarity constraints demonstrate varying drift mitigation capabilities, with the ℓ2 regularization approach effectively minimizing drift signals. The iterative Fast Gradient Sign Method (i-FGSM) with ℓ2 constraints lowered KS and Population Stability Index (PSI) metrics most effectively, while MMD regularization introduced substantially larger distributional shifts. These findings underscore the
Adversarial Evasion in Non-Stationary Malware Detection
3
nuanced challenge of generating adversarial malware samples that can simultaneously evade detection and maintain distributional similarity, providing critical insights into the vulnerabilities of current concept drift monitoring mechanisms. Our contributions are summarized as follows: – We formulate similarity-constrained targeted feature-space attacks for tabular malware detection by augmenting targeted cross-entropy with KL, ℓ2 , and MMD regularizers. – We propose an drift evaluation protocol comparing clean vs. adversarial malware using JSD, Hellinger, Wasserstein, output-MMD, KS, and PSI. – We characterize evasion–detectability trade-offs across objectives, optimizers, and hyperparameters, showing perturbation budget as the dominant driver of ASR and drift.
2
Background and Related Work
2.1
Non-stationary Malware Detection and Concept Drift
The cybersecurity landscape is defined by the continuous evolution of malware, with attackers constantly developing new obfuscation and modification techniques [7,6,20]. This perpetual transformation challenges traditional static machine learning models, which quickly become obsolete as malware characteristics shift. Concept drift represents the fundamental challenge in maintaining effective malware detection systems [4,3]. It describes how statistical properties of target variables change over time, rendering previously trained models increasingly ineffective. Researchers have developed two primary strategies: retraining, which replaces existing models with new ones, and incremental algorithms that continuously update models as new information emerges [14]. Drift detection mechanisms have become crucial in managing this non-stationary environment. These mechanisms employ sophisticated statistical techniques to identify and flag anomalous inputs that deviate from the model’s expected behavior [9]. Classic methods like Drift Detection Method (DDM) and Early Drift Detection Method (EDDM) monitor error rates and error spacing, providing early warnings about potential distributional shifts [4,3]. The ultimate objective of non-stationary malware detection is to create adaptive systems capable of maintaining high detection performance in the face of continuous technological evolution [6]. By implementing drift monitoring techniques and flexible model strategies, researchers aim to develop more resilient detection mechanisms that can effectively respond to ever-changing landscape of malware threats. 2.2
Adversarial Malware Generation Methods
Adversarial malware generation has emerged as a critical research domain, demonstrating the vulnerabilities of learning-based malware detection systems [12,2]. Existing approaches can be categorized into three methodological families: gradientbased attacks, generative methods, and reinforcement learning techniques.
4
Acharya and Zhang
Gradient-based attacks leverage loss gradients to identify precise modifications that alter model predictions, strategically manipulating executable files to bypass detection with minimal changes. Researchers have demonstrated the effectiveness of systematically appending or modifying bytes while maintaining the executable’s core functionality [12,13]. Generative approaches train models to produce adversarial feature vectors that can transfer across detection systems, focusing on byte-level insertion strategies that maximize evasion potential. These methods represent a more sophisticated approach to adversarial sample generation [10]. Reinforcement learning-based methods introduce a dynamic approach, framing adversarial malware generation as a Markov Decision Process [2,20]. RL agents learn to apply binary transformations that maximize evasion against malware detectors, using advanced optimization techniques [2]. However, a critical limitation persists: most existing approaches evaluate success primarily through misclassification and executability preservation, overlooking how adversarial samples interact with drift monitoring mechanisms [20,6].
3
Methodology
Our research addresses a critical challenge in non-stationary malware detection by developing adversarial techniques that can simultaneously evade classification and remain inconspicuous to drift monitoring mechanisms. We propose a novel approach generating targeted adversarial examples within the classifier’s standardized feature space, focusing on static, fixed-length tabular representations derived from byte-level statistics and structural metadata. By incorporating sophisticated similarity-constrained approaches with regularization techniques based on KL divergence, ℓ2 distance, and MMD, we develop a nuanced framework for creating adversarial samples that can potentially bypass detection while maintaining statistical similarity to original malware distributions. Our comprehensive methodology systematically evaluates the effectiveness of these adversarial samples through a multi-dimensional analysis, employing a sophisticated suite of drift metrics including JS divergence, Hellinger distance, Wasserstein distance, output-based MMD, KS statistic, and PSI. 3.1
Framework Overview
Let fθ : Rd → [0, 1]2 denote a trained binary classifier that outputs class probabilities for benign and malicious classes. We assume access to feature vectors x ∈ Rd corresponding to malware samples (y = 1). Given a set of clean malware features X = {xi }m i=1 , our framework proceeds in three steps: 1. Adversarial generation. An attack procedure A takes the classifier fθ and the clean malware features X as input and produces a perturbed set Xadv = {xi,adv }m i=1 intended to be misclassified as benign while respecting a bounded perturbation budget. We consider both unconstrained and similarity-constrained objectives, implemented with iterative FGSM (i-FGSM) and projected gradient descent (PGD).
Adversarial Evasion in Non-Stationary Malware Detection Adversarial generation
Classifier evaluation
Clean malware features X
Drift monitoring
Drift metrics
Classifier fθ(·)
KS, PSI, JSD, Hellinger Wasserstein, output-MMD
Attack A: CE → benign + αR R ∈ {KL, L2, MMD} i-FGSM / PGD ε, δ, T
5
Classifier evaluation
Drift detection p_adv = fθ(X_adv)[:,1]
p_clean = fθ(X)[:,1] Report: ASR + drift (trade-off)
Adversarial features X_adv
Attack Success Rate
Fig. 1: Architecture overview of the methodology components.
2. Classifier evaluation. The adversarial examples Xadv are fed into fθ to compute the attack success rate, defined as the fraction of malware samples classified as benign after perturbation. 3. Drift detection. We compare the distributions of classifier outputs for clean and adversarial malware using drift metrics that quantify the adversarial shift from a distributional perspective.
Threat Model. We assume a white-box adversary with access to the trained classifier fθ and its gradients with respect to the standardized input feature vector x ∈ Rd . The attacker aims to produce targeted adversarial examples from malware inputs (y = 1) that are misclassified as benign (ytarget = 0), subject to a per-feature perturbation budget. This models an attacker who can manipulate the observed feature representation within bounded limits. Our attacks are conducted in the classifier’s feature space. We use feature-space perturbations to quantify the sensitivity of (i) the classifier and (ii) output-based drift monitoring under similarity-aware adversarial objectives. 3.2
Adversarial Malware Generation
All attacks are carried out in the standardized feature space of a fixed classifier fθ . We restrict attention to malware samples (y = 1) and perform targeted attacks toward the benign label (ytarget = 0). For each malware feature vector x, we construct an adversarial counterpart xadv by minimizing a composite loss under a per-feature perturbation budget. Adversarial Objective with Similarity Constraints. Given a clean malware feature vector x and its adversarial counterpart xadv , we define the general adversarial loss Ladv (xadv ) = LCE fθ (xadv ), ytarget + α R(xadv , x, Xadv , X), (1) where LCE is the cross-entropy loss encouraging misclassification as benign, R is an optional similarity regularizer, and α ≥ 0 controls the strength of the similarity constraint. We consider four choices for R:
6
Acharya and Zhang
– Baseline (unconstrained): R = 0, so Ladv reduces to standard targeted cross-entropy. – KL-regularized attack. We normalize the original and adversarial feature adv vectors into per-sample probability vectors p = P xxj +ε , q = P xxadv,j +ε , and j
j
define the regularizer RKL (xadv , x) = KL(p∥q), averaged over the batch. – ℓ2 -regularized attack. We penalize the Euclidean distance between clean and adversarial features, Rℓ2 (xadv , x) = xadv − x 2 , encouraging small perturbations in ℓ2 norm. – MMD-regularized attack. We measure discrepancy between batches of clean and adversarial features via the empirical squared MMD (MMD) under a kernel k (linear, RBF, or polynomial). Given a batch of clean features X and adversarial features Xadv , we define RMMD (Xadv , X) = MMD2k (X, Xadv ), which encourages the adversarial batch to remain close to the clean malware batch in feature space. Iterative FGSM (i-FGSM). The i-FGSM attack optimizes (1) via repeated signed gradient updates under a per-feature box constraint. For a clean malware sample (0) x, we initialize xadv = x, and perform T iterations of the form (t) (t) (t+1) (2) xadv = Π[x−δ, x+δ] xadv − ϵ · sign ∇x(t) Ladv (xadv ) , adv
where ϵ is the step size, δ is the perturbation budget (maximum allowed absolute deviation per feature), and Π[x−δ, x+δ] denotes element-wise clipping to the interval [xj − δ, xj + δ] for each feature j. We apply i-FGSM to all four loss configurations. For each configuration and hyperparameter setting (ϵ, T, δ), we generate adversarial examples for all malware samples and compute the attack success rate, defined as the fraction of adversarial examples classified as benign by fθ . Projected Gradient Descent (PGD) with Random Start. We also implement a PGD-style attack that uses a random initialization within the perturbation set, followed by iterative signed-gradient updates and projection. For a clean sample (0) x, we initialize xadv = Π[x−δ, x+δ] (x + u), u ∼ U([−δ, δ]), and iterate (t+1) (t) (t) xadv = Π[x−δ, x+δ] xadv − η · sign ∇x(t) Ladv (xadv ) , (3) adv
for t = 0, . . . , T − 1, where η is the step size. The random start allows PGD to explore different regions of the feasible set compared to i-FGSM initialized at x. 3.3
Drift Detection Metrics
To quantify how distinguishable adversarial malware is from clean malware, we compare the classifier’s output distributions on clean and adversarial samples using a suite of drift metrics. Let X denote the clean malware features from a reference test split and Xadv the corresponding adversarial features produced by
Adversarial Evasion in Non-Stationary Malware Detection
7
a given attack configuration. We compute the predicted malware-class probability for each sample: pclean = fθ (X)[:, 1], padv = fθ (Xadv )[:, 1]. We measure the discrepancy between the one-dimensional distributions of pclean and padv using: – Population Stability Index (PSI) [11]: a standard feature drift measure, computed by binning pclean PB into B bins, using the same bin edges for padv , and aggregating PSI = b=1 (cb − rb ) log rcbb , where rb and cb are the proportions of clean and adversarial samples in bin b. – Kolmogorov–Smirnov (KS) statistic [16]: the maximum absolute difference between the empirical cumulative distribution functions (ECDFs) of pclean and padv , as in the two-sample KS test. – Jensen–Shannon divergence [15]: a symmetric, bounded divergence between binned probability mass functions derived from pclean and padv . – Hellinger distance: a metric on probability distributions, computed from the square-rooted bin frequencies of pclean and padv . Concretely, if p̂ and q̂ are √ 2 P √ the binned empirical distributions, we use H 2 (p̂, q̂) = 21 b p̂b − q̂b . – Wasserstein distance [17]: the one-dimensional Earth Mover’s distance between the empirical distributions of pclean and padv . – MMD [5]: an empirical squared MMD between pclean and padv using an RBF kernel with bandwidth chosen by a median heuristic.
4
Experimental Results
We implement the framework described in Section 3 to answer three research questions: (RQ1) How sensitive are output-based drift metrics (JSD, Hellinger, Wasserstein, and output-MMD) compared to KS and PSI under a fixed lowbudget operating point? (RQ2) How do attack design choices affect the evasion detectability trade-off? (RQ3) How do attack hyperparameters (step size ϵ, budget δ, and iterations T ) affect the trade-off between ASR and drift detectability? Data and Feature Representation. We investigate a binary malware detection task using the BODMAS dataset [18], which comprises 57,293 malware and 77,142 benign Windows PE files. Each sample is represented by a comprehensive 2,381-dimensional feature vector that integrates byte-level statistics and continuous auxiliary features derived from structural and metadata characteristics. To ensure robust analysis, features are standardized using training data statistics and carefully partitioned into distinct training, validation, and test splits. Classifier Architecture and Training. The classifier fθ is implemented as a sophisticated feed-forward neural network designed to handle high-dimensional tabular features. The architecture incorporates two fully connected hidden layers with ReLU activations, complemented by batch normalization and dropout techniques to enhance training stability and mitigate overfitting. An ℓ2 weight regularization is applied to dense layers, with a softmax output layer generating benign and malware class probabilities. Trained using the Adam optimizer with categorical cross-entropy loss, the model demonstrates exceptional performance, achieving over 95% accuracy on clean test data.
8
Acharya and Zhang
Table 1: Drift scores under a fixed operating point (ϵ=0.0001, δ=0.0005, T =100, λKL =λℓ2 =λMMD =1). Attack
JSD
Hell.
Wass.
MMD
KS
PSI Detected
i-FGSM 0.0212 0.0176 0.000750 -6.77E-05 0.0183 0.0025 No i-FGSM + KL 0.0197 0.0164 0.000550 -8.15E-05 0.0159 0.0021 No i-FGSM + ℓ2 0.0124 0.0103 0.000377 -9.92E-05 0.0121 0.0008 No i-FGSM + MMD (Lin) 0.1099 0.0916 0.008141 1.80E-03 0.0865 0.0673 Yes (MMD) i-FGSM + MMD (RBF) 0.1099 0.0916 0.008259 1.81E-03 0.0865 0.0673 Yes (MMD) i-FGSM + MMD (Poly) 0.0565 0.0471 0.004104 6.47E-04 0.0453 0.0177 Yes (MMD)
Adversarial Attack Configurations Our adversarial evaluation framework generates targeted adversarial examples through a comprehensive approach. We explore multiple objectives, including a baseline cross-entropy approach and three sophisticated similarity-constrained regularization techniques: KL divergence, ℓ2 , and MMD. These objectives are implemented using two advanced optimization methods: iterative FGSM and PGD with random initialization. The attack configurations systematically vary critical hyperparameters such as perturbation budget, step size, iteration count, and regularization strength. Success is quantitatively measured by the proportion of adversarial malware samples misclassified as benign by the target classifier.
4.1
RQ1: Sensitivity of Drift Metrics Under a Fixed Operating Point
Table 1 reports drift scores computed between the clean and adversarial malwareprobability output distributions. Across all statistical metrics, similarity constraints reduce measured drift relative to the unconstrained baseline. For example, the baseline i-FGSM attack yields JSD=0.0212 and Hellinger=0.0176, whereas i-FGSM+KL reduces these to JSD=0.0197 and Hellinger=0.0164, and i-FGSM+ℓ2 further reduces them to JSD=0.0124 and Hellinger=0.0103. Consistent reductions are also observed for Wasserstein (0.00075 → 0.00055 → 0.000377) and for KS statistic (0.0183 → 0.0159 → 0.0121), indicating that KL and especially ℓ2 regularization produce adversarial outputs that are closer to the clean output distribution under all evaluated drift scores. In contrast, MMD-regularized variants induce substantially larger distribution shifts across every metric. For i-FGSM+MMD with linear and RBF kernels, drift magnitudes increase to JSD=0.1099 and Hellinger=0.0916, with Wasserstein increasing to approximately 8 × 10−3 and the KS statistic increasing to 0.0865. The polynomial-kernel variant yields intermediate shifts (JSD=0.0565, Hellinger=0.0471, Wasserstein=0.004104, KS=0.0453). This consistent separation across statistical metrics indicates that, at this operating point, statistical detectors are sensitive to the stronger output shifts produced by MMDregularized objectives, while remaining responsive to the smaller shifts induced by KL and ℓ2 constraints. Comparing against traditional metrics, KS aligns with the statistical measures in terms of relative sensitivity: it decreases for KL/ℓ2 and increases markedly
Adversarial Evasion in Non-Stationary Malware Detection
9
Table 2: ASR and drift trade off across optimization methods and similarityregularized objectives. All runs use ϵ=0.0001, δ=0.0005, T =100, and λKL =λℓ2 =λMMD =1. i-FGSM Objective
Loss ASR (%)
CE 6.55556 CE + KL 12.568620 CE + ℓ2 6.57777 CE + MMD (Lin) 6.237870 CE + MMD (RBF) 6.228153 CE + MMD (Poly) -1403.280884
3.10 3.09 3.09 3.71 3.71 3.34
PGD KS / PSI
Loss ASR (%)
0.0183 / 0.0025 6.5555787 0.0159 / 0.0021 12.451903 0.0121 / 0.0008 6.5903187 0.0865 / 0.0673 6.55568 0.0865 / 0.0673 6.555345 0.0453 / 0.0177 -1409.275391
3.10 3.09 3.08 3.10 3.10 3.10
KS / PSI 0.0183 / 0.0025 0.0159 / 0.0021 0.0038 / 0.0001 0.0183 / 0.0025 0.0183 / 0.0025 0.0159 / 0.0010
for MMD-regularized attacks. PSI values remain low across all attacks (0.0008– 0.0673), which falls within the negligible-drift range under standard PSI interpretability bands. As a result, PSI is less sensitive than KS and the statistical metrics for distinguishing these attack variants in this low-budget regime. Under a fixed operating point, statistical drift metrics (JSD, Hellinger, Wasserstein, and MMD) provide a consistent and discriminative signal over similarityconstrained adversarial examples: KL and ℓ2 regularization reduce drift magnitudes relative to baseline i-FGSM, whereas MMD-regularized attacks induce substantially larger shifts, are also reflected by KS. PSI remains small for all evaluated attacks, suggesting limited sensitivity in this configuration.
4.2
RQ2: Effect of Loss Configuration and Optimization Method
Table 2 reports the trade-off between attack effectiveness and drift detectability when varying (i) the optimization procedure (i-FGSM vs. PGD with random initialization) and (ii) the attack objective (cross-entropy only vs. KL-, ℓ2 -, and MMD-regularized losses). Unless stated otherwise, all experiments in this subsection use the same perturbation parameters (ϵ = 0.0001, per-feature budget δ = 0.0005, T = 100) and regularization strengths (λKL = λℓ2 = λMMD = 1). Drift is quantified on the malware-class output probabilities using KS and PSI, where larger values indicate greater distributional deviation. Optimizer effect (i-FGSM vs. PGD). Under the fixed, low-budget setting considered here, replacing i-FGSM with PGD (random start) yields negligible changes in attack effectiveness: ASR remains approximately 3% across the baseline, KL-, and ℓ2 -regularized configurations. The drift magnitudes are also similar for these configurations such as baseline KS=0.0183 and PSI=0.0025 for both i-FGSM and PGD, at this perturbation scale, random initialization alone does not materially alter either evasion or output-distribution shift. Loss configuration effect (CE vs. KL/ℓ2 /MMD). The choice of regularization has a clearer influence on drift. Both KL and ℓ2 regularization reduce drift relative to the unconstrained objective while maintaining essentially the same ASR. For i-FGSM, KS decreases from 0.0183 (CE) to 0.0159 (CE+KL) and 0.0121 (CE+ℓ2 ), with PSI decreasing from 0.0025 to 0.0021 and 0.0008, re-
10
Acharya and Zhang
Table 3: FGSM-only: representative hyperparameter sweeps illustrating the ASR and drift trade-off. Budget sweep (ϵ = 0.01, T = 100) Step-size sweep (δ = 0.01, T = 100) Iterations sweep (ϵ = 0.005, δ = 0.1) 0.05
0.1
0.001 0.005
0.02
0.1
ϵ 0.01 0.01 0.01 0.01 δ 0.005 0.01 0.02 0.05 T 100 100 100 100 ASR (%) 3.10 3.50 5.24 22.37 Avg. PSI 0.103 0.125 0.175 0.328
0.005
0.01
0.02
0.01 0.1 100 72.04 0.511
0.001 0.005 0.01 0.02 0.01 0.01 0.01 0.01 100 100 100 100 3.50 3.50 3.50 3.49 0.128 0.126 0.125 0.129
0.01
0.1 0.01 100 3.49 0.129
150
200
0.005 0.005 0.005 0.005 0.1 0.1 0.1 0.1 10 50 100 150 21.71 71.76 72.04 72.04 0.307 0.502 0.510 0.512
10
50
100
0.005 0.1 200 72.04 0.512
Fig. 2: FGSM ASR Fig. 3: FGSM Avg. Fig. 4: ASR vs. Fig. 5: Avg. PSI vs. vs. budget δ. PSI vs. budget δ. budget δ(KL). budget δ(KL).
spectively. A similar pattern holds for PGD, where ℓ2 regularization yields the smallest drift among the evaluated settings (KS=0.0038, PSI=0.0001). MMD regularization. MMD-regularized objectives exhibit behavior that is more sensitive to the kernel choice. For i-FGSM, linear and RBF MMD produce substantially larger shifts (KS=0.0865, PSI=0.0673) while yielding only a modest ASR increase (3.71%), whereas the polynomial kernel shows smaller drift (KS=0.0453, PSI=0.0177) with ASR close to baseline (3.34%). For PGD, the tested MMD configurations do not improve ASR and produce drift scores comparable to the baseline/KL settings. Results suggest that, under the present hyperparameters, MMD regularization does not consistently improve the ASR stealth trade-off and may increase detectability depending on the kernel. 4.3
RQ3: Effect of Attack Hyperparameters
FGSM-only (CE). As shown in Table 3 and illustrated in Figure 2 and 3, we analyze the baseline iterative FGSM attack (targeted CE loss) to understand how the attack hyperparameters such as, step size ϵ, perturbation budget δ, and iteration count T affect the trade-off between attack success rate (ASR) and drift-based detectability (Avg. PSI computed on model outputs). Budget δ is the dominant driver of the ASR and drift trade-off. Increasing the budget produces the largest change in both ASR and drift. Under a representative controlled slice (ϵ = 0.01, T = 100), increasing δ from 0.005 to 0.1 increases ASR from 3.10% to 72.04%, while Avg. PSI increases from 0.103 to 0.511. This pattern reflects the core mechanism of box-constrained FGSM: larger δ expands the feasible perturbation set [x − δ, x + δ], making it more likely for the adversarial example to cross the classifier decision boundary, but also inducing larger shifts in the monitored output distribution.
Adversarial Evasion in Non-Stationary Malware Detection
11
Table 4: FGSM with KL (λKL = 1): representative hyperparameter sweeps illustrating ASR and detectability trends. Budget sweep (fix ϵ = 0.01, T = 100) ϵ 1 2 3 4 5
δ
T
0.01 0.005 100 0.01 0.01 100 0.01 0.02 100 0.01 0.05 100 0.01 0.1 100
2.92 3.23 4.53 17.69 52.73
Step-size sweep (fix δ = 0.01, T = 100) Iterations sweep (fix ϵ = 0.005, δ = 0.1) T
ASR (%)
Avg. PSI
ϵ
0.087 0.001 0.01 100 0.098 0.005 0.01 100 0.138 0.01 0.01 100 0.02 0.01 100 0.245 0.378 0.1 0.01 100
3.23 3.23 3.23 3.24 3.24
0.105 0.101 0.098 0.114 0.114
0.005 0.005 0.005 0.005 0.005
ASR (%) Avg. PSI
ϵ
δ
T
ASR (%)
Avg. PSI
0.1 10 0.1 50 0.1 100 0.1 150 0.1 200
δ
18.02 50.03 50.01 50.01 50.01
0.232 0.369 0.371 0.371 0.371
Step size ϵ has a secondary effect once δ is fixed. When the budget is small, varying ϵ has minimal impact on ASR and Avg. PSI because updates quickly saturate under the projection constraint. At δ = 0.01 and T = 100, increasing ϵ from 0.001 to 0.1 leaves ASR essentially unchanged (3.50% → 3.49%), with Avg. PSI staying near 0.125–0.129. At larger budgets, ϵ can mildly modulate the final solution, but its effect remains substantially smaller than the effect of δ. Iteration count T shows diminishing returns and depends on budget. At low budgets (δ = 0.01), increasing T provides negligible gains because the attack saturates early. At higher budgets, additional iterations can improve ASR initially (for δ = 0.1 and ϵ = 0.005, ASR increases sharply from 21.71% at T = 10 to 71.76% at T = 50), but improvements beyond T (50–100) are marginal and Avg. PSI changes only slightly. δ primarily controls feasibility (and detectability), while ϵ and T mainly affect optimization efficiency within the feasible set. FGSM with KL regularization. We analyze a similarity-constrained FGSM variant where the attack loss augments targeted cross-entropy with a KL penalty. As shown in Table 4 and Figure 4 and 5, we systematically explore how key hyperparameters impact the attack’s performance. Budget δ remains the dominant driver of the ASR and drift trade-off. Holding ϵ = 0.01 and T = 100 fixed, increasing the budget from δ = 0.005 to δ = 0.1 increases ASR from 2.92% to 52.73%, while Avg. PSI increases from 0.087 to 0.378 (Table 4). As in the baseline FGSM case, the budget directly expands the feasible perturbation set [x − δ, x + δ], improving the chance of crossing the decision boundary while increasing output-distribution shift. Step size ϵ has minimal impact in the low-budget regime. At δ = 0.01 and T = 100, sweeping ϵ from 0.001 to 0.1 leaves ASR nearly unchanged (3.23% → 3.24%), with only modest PSI variation. The attack saturates under projection, with step size primarily affecting optimization efficiency rather than feasibility. Iterations T yield diminishing returns, especially after early gains at high budget. At a larger budget (δ = 0.1) and ϵ = 0.005, increasing iterations from T = 10 to T = 50 substantially increases ASR (18.02% → 50.03%), while Avg. PSI also increases (0.232 → 0.369). Beyond T ≈ 50, both ASR and Avg. PSI largely saturate (Table 4), indicating diminishing returns from additional iterations once the attack has converged within the feasible set. FGSM with ℓ2 regularization. We analyze FGSM augmented with an ℓ2 similarity penalty between clean and perturbed feature vectors. As shown in Table 5 and Figure 6 and 7, explore how hyperparameters influence the attack’s performance.
12
Acharya and Zhang
Table 5: FGSM with ℓ2 (λℓ2 = 1): representative hyperparameter sweeps illustrating ASR and detectability trends. Budget sweep (fix ϵ = 0.01, T = 100) ϵ 1 2 3 4 5 6
δ
T
0.01 0.005 100 0.01 0.01 100 0.01 0.02 100 0.01 0.05 100 0.01 0.10 100 0.01 0.20 100
2.61 2.63 2.71 2.85 3.09 3.51
Step-size sweep (fix δ = 0.01, T = 100) Iterations sweep (fix ϵ = 0.005, δ = 0.20) T
ASR (%)
Avg. PSI
ϵ
0.048 0.001 0.01 100 0.056 0.005 0.01 100 0.059 0.01 0.01 100 0.02 0.01 100 0.067 0.076 0.10 0.01 100 0.20 0.01 100 0.090
2.63 2.63 2.63 2.62 2.62 2.62
0.055 0.056 0.056 0.052 0.052 0.052
0.005 0.005 0.005 0.005 0.005
ASR (%) Avg. PSI
ϵ
δ
T
ASR (%)
Avg. PSI
0.20 10 0.20 50 0.20 100 0.20 150 0.20 200
δ
2.84 3.51 3.51 3.51 3.51
0.068 0.090 0.090 0.090 0.090
Fig. 6: ASR vs. Fig. 7: Avg. PSI vs. Fig. 8: ASR Fig. 9: Avg. PSI budget δ (λℓ2 = 1). budget δ (λℓ2 = 1). vs. budget δ vs. budget δ (δMMD = 0.1). (δMMD = 0.1).
Budget δ increases ASR only modestly under ℓ2 regularization. Unlike the baseline FGSM and FGSM with KL settings, increasing δ yields only a small improvement in evasion under the ℓ2 penalty. Under a representative slice (ϵ = 0.01, T = 100), increasing δ from 0.005 to 0.2 raises ASR from 2.61% to 3.51% (Table 5). This indicates that the ℓ2 penalty strongly constrains the optimization to remain close to the original feature vector, limiting the ability to cross the decision boundary even as the feasible box expands. Step size ϵ has negligible effect when δ and T are fixed. At δ = 0.01 and T = 100, sweeping ϵ from 0.001 to 0.2 leaves ASR essentially unchanged (approximately 2.62–2.63%), with only minor changes in Avg. PSI (Table 5). This suggests that, in the low-budget regime, the attack rapidly saturates under the combined effect of projection and the similarity penalty. Iterations T show early gains and then saturate. At a larger budget (δ = 0.2) and ϵ = 0.005, ASR increases from 2.84% at T = 10 to 3.51% at T = 50, after which additional iterations yield negligible improvement (Table 5). Avg. PSI exhibits a similar pattern, rising from 0.068 to ≈ 0.090 by T = 50 and then stabilizing. FGSM with MMD regularization. As shown in Table 6 and Figure 8 and 9, we analyze FGSM augmented with an MMD similarity penalty on the targeted cross-entropy objective. Effect of budget δ. The perturbation budget produces the largest change in both ASR and Avg. PSI. Under a representative slice (ϵ = 0.01, T = 100), increasing δ from 0.005 to 0.2 increases ASR from 3.10% to 94.91%, while Avg. PSI increases from 0.096 to 0.333 (Table 6). The budget is the dominant control knob for evasion, while also increasing distributional shift in the monitored outputs. Effect of step size ϵ. At fixed δ = 0.01 and T = 100, sweeping ϵ from 0.001 to 0.2 yields negligible changes in ASR (approx. 3.48% throughout) and only small
Adversarial Evasion in Non-Stationary Malware Detection
13
Table 6: FGSM with MMD (λMMD = 1): representative hyperparameter sweeps illustrating ASR and detectability trends. Budget sweep (fix ϵ = 0.01, T = 100)
1 2 3 4 5 6 7
ϵ
δ
T
0.01 0.01 0.01 0.01 0.01 0.01 0.01
0.005 0.010 0.020 0.050 0.080 0.100 0.200
100 100 100 100 100 100 100
ASR (%) Avg. PSI 3.10 3.48 4.99 17.43 41.55 58.48 94.91
0.096 0.108 0.140 0.223 0.286 0.299 0.333
Step-size sweep (fix δ = 0.01, T = 100) Iterations sweep (fix ϵ = 0.005, δ = 0.20) ϵ
δ
T
ASR (%)
Avg. PSI
ϵ
0.001 0.005 0.010 0.020 0.050 0.100 0.200
0.01 0.01 0.01 0.01 0.01 0.01 0.01
100 100 100 100 100 100 100
3.49 3.48 3.48 3.48 3.48 3.48 3.48
0.110 0.109 0.108 0.111 0.111 0.111 0.111
0.005 0.005 0.005 0.005 0.005
T
ASR (%)
Avg. PSI
0.20 10 0.20 50 0.20 100 0.20 150 0.20 200
δ
19.91 92.99 95.65 95.96 96.04
0.200 0.324 0.371 0.377 0.379
variation in Avg. PSI (approx. 0.108–0.111; Table 6). In the low-budget regime, step size primarily affects optimization dynamics rather than feasibility. Effect of iterations T . At larger budget (δ = 0.2) and ϵ = 0.005, increasing T from 10 to 50 increases ASR from 19.91% to 92.99%, while Avg. PSI increases from 0.200 to 0.324. Beyond T ≈ 100, ASR saturates near 96% while Avg. PSI continues to increase slightly (Table 6), indicating diminishing returns in evasion and a gradual increase in detectability.
5
Conclusion
Our research explored the intricate dynamics of targeted, gradient-based adversarial malware perturbations, focusing on their ability to simultaneously evade deep learning classifiers while maintaining statistical similarity to clean malware across drift metrics. Similarity constraints using KL divergence and ℓ2 distance consistently reduced measured drift across statistical metrics, yet paradoxically did not significantly alter the attack success rate, which remained around 3%. MMD-regularized variants, in contrast, generated larger output-distribution shifts and proved more detectable. Our key insights reveal the nuanced relationship between similarity constraints and evasion potential. Constraining perturbations can reduce outputdistribution drift without enhancing evasion capabilities, while expanding the perturbation budget substantially increases evasion potential at the cost of increased detectability. These findings underscore the complex challenge of generating stealthy adversarial malware, emphasizing the critical importance of evaluating attacks through multiple metrics beyond simple misclassification rates. The research illuminates the delicate balance between attack effectiveness and detection probability in the evolving landscape of adversarial machine learning.
References 1. Akerele, S., Adebola, N., Fagbohun, O., et al.: Modern deep learning approaches for malware detection and classification. J Artif Intell Mach Learn & Data Sci 2025 3(3), 2761–2768 2. Anderson, H.S., Kharkar, A., Filar, B., Evans, D., Roth, P.: Learning to evade static pe machine learning malware models via reinforcement learning. arXiv preprint arXiv:1801.08917 (2018)
14
Acharya and Zhang
3. Baena-Garcıa, M., del Campo-Ávila, J., Fidalgo, R., Bifet, A., Gavalda, R., Morales-Bueno, R.: Early drift detection method. In: Fourth international workshop on knowledge discovery from data streams. vol. 6, pp. 77–86 (2006) 4. Gama, J., Medas, P., Castillo, G., Rodrigues, P.: Learning with drift detection. In: Brazilian symposium on artificial intelligence. pp. 286–295. Springer (2004) 5. Gretton, A., Borgwardt, K., Rasch, M., Schölkopf, B., Smola, A.: A kernel method for the two-sample-problem. Advances in neural information processing systems 19 (2006) 6. Guerra-Manzanares, A., Bahsi, H.: Experts still needed: boosting long-term android malware detection with active learning. Journal of Computer Virology and Hacking Techniques 20(4), 901–918 (2024) 7. Guerra-Manzanares, A., Luckner, M., Bahsi, H.: Android malware concept drift using system calls: detection, characterization and challenges. Expert Systems with Applications 206, 117200 (2022) 8. Hinder, F., Vaquet, V., Hammer, B.: Adversarial attacks for drift detection. arXiv preprint arXiv:2411.16591 (2024) 9. Hinder, F., Vaquet, V., Hammer, B.: One or two things we know about concept drift—a survey on monitoring in evolving environments. part a: detecting concept drift. Frontiers in Artificial Intelligence 7, 1330257 (2024) 10. Hu, W., Tan, Y.: Generating adversarial malware examples for black-box attacks based on gan. In: International Conference on Data Mining and Big Data. pp. 409–423. Springer (2022) 11. Khademi, A., Hopka, M., Upadhyay, D.: Model monitoring and robustness of inuse machine learning models: quantifying data distribution shifts using population stability index. arXiv preprint arXiv:2302.00775 (2023) 12. Kolosnjaji, B., Demontis, A., Biggio, B., Maiorca, D., Giacinto, G., Eckert, C., Roli, F.: Adversarial malware binaries: Evading deep learning for malware detection in executables. In: 2018 26th European signal processing conference (EUSIPCO). pp. 533–537. IEEE (2018) 13. Kreuk, F., Barak, A., Aviv-Reuven, S., Baruch, M., Pinkas, B., Keshet, J.: Deceiving end-to-end deep learning malware detectors using adversarial examples. arXiv preprint arXiv:1802.04528 (2018) 14. Li, A.S., Iyengar, A., Kundu, A., Bertino, E.: Revisiting concept drift in windows malware detection: Adaptation to real drifted malware with minimal samples. arXiv preprint arXiv:2407.13918 (2024) 15. Lin, J.: Divergence measures based on the shannon entropy. IEEE Transactions on Information theory 37(1), 145–151 (2002) 16. Massey Jr, F.J.: The kolmogorov-smirnov test for goodness of fit. Journal of the American statistical Association 46(253), 68–78 (1951) 17. Villani, C., et al.: Optimal transport: old and new, vol. 338. Springer (2008) 18. Yang, L., Ciptadi, A., Laziuk, I., Ahmadzadeh, A., Wang, G.: Bodmas: An open dataset for learning based temporal analysis of pe malware. In: 4th Deep Learning and Security Workshop (2021) 19. Yang, W., Su, R., Cheng, Y., Guo, J.: A concept drift detection approach based on jensen-shannon divergence for network traffic classification. In: Proceedings of the 2022 5th international conference on artificial intelligence and pattern recognition. pp. 982–987 (2022) 20. Zhang, L., Liu, P., Choi, Y.H., Chen, P.: Semantics-preserving reinforcement learning attack against graph neural networks for malware detection. IEEE Transactions on Dependable and Secure Computing 20(2), 1390–1402 (2022)