Conceptio › Archive › arXiv CS
arXiv CSopen access

Breaking Windows Malware Detection: A Comprehensive Evaluation of Problem-Space Adversarial Robustness

Mashal Zainab et al. · arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

arXiv:2609.34456v1 [cs.CR] 28 Sep 2026

Breaking Windows Malware Detection: A Comprehensive Evaluation of Problem-Space Adversarial Robustness Salijona Dyrmishi Mashal Zainab Hamid Bostani Independent Researcher University of Luxembourg University of Luxembourg [email protected] [email protected] [email protected] Maxime Cordy Lorenzo Cavallaro University of Luxembourg University College London [email protected] [email protected] Abstract—Problem-space evasion attacks have exposed critical weaknesses in machine learning-based malware detectors; yet, their evaluation remains fragmented across models, datasets, and attack methodologies, often neglecting domain-specific requirements such as executability and functionality preservation. We address this gap with a unified, large-scale evaluation of nine state-of-the-art evasion attacks against eight Windows malware detectors, including seven open-source models and one commercial detector, under executability-preserving conditions. Our study analyzes attack effectiveness, complementarity, transferability, and adversarial hardening to evaluate robustness along complementary dimensions. We show that detector vulnerability depends strongly on both model representation and attack type: raw-byte detectors are particularly susceptible to several classes of problem-space manipulation, but no detector family is uniformly robust across all attacks. Importantly, effectiveness is not explained by transformation-space size alone: the strongest attacks can achieve substantially higher success while using fewer distinct transformations and concentrating on a small set of highimpact manipulations. We further show that two complementary attacks are sufficient to cover ≈ 99% of the adversarial examples produced by the remaining evaluated attacks. Transferability exhibits a different pattern from direct attack success: attacks with low direct success can produce highly transferable evasions. Finally, adversarial hardening is highly attack- and modeldependent: robustness gains often fail to transfer across attacks and can even increase susceptibility to unseen attacks. These findings highlight important limitations in current malware robustness evaluation practices, establish a comprehensive empirical baseline for realistic evasion attacks, and clarify key relationships between effectiveness, transferability, and defense robustness.

I. I NTRODUCTION As Windows malware continues to pose a major threat to modern computing systems, machine learning (ML) has emerged as a promising defense mechanism, providing the ability to generalize beyond traditional signature-based detection, and analyze large volumes of files at scale. Despite their widespread adoption, these systems operate in a fundamentally adversarial setting. Malware authors actively develop problemspace evasion attacks, methods that generate adversarial malware variants by modifying real portable executables (PEs) files in a way that preserves functionality while making these variants incorrectly classified as benign. As a result, even well-

performing ML-based detectors can be vulnerable to these carefully crafted samples, raising serious concerns about their reliability and robustness in real-world scenarios. In response to the growing sophistication of adversarial malware, the security community has pursued two primary directions: devising new evasion attacks to probe classifier weaknesses [1]–[8], and constructing new defenses to fortify them [9]–[14]. Yet, the practical efficacy of these efforts remains unclear, as they are evaluated within a fragmented and inconsistent landscape [15], [16]. This lack of a standardized benchmark means that new evasion attacks are often validated against only a handful of detectors (e.g., [17]), while new defenses are tested against a narrow and often impractical set of evasion attacks (e.g., [14]). Consequently, this fragmentation hinders progress by making it impossible to identify the most potent evasion attacks or the most effective defenses, ultimately leaving systems vulnerable. This state of uncertainty forces us to confront a critical question: How accurately do current evaluations capture the adversarial robustness of malware classifiers, and how reliable are the conclusions drawn from them? Answering this question is nontrivial. Evaluating adversarial robustness in Windows malware detection requires more than measuring whether an attack changes a model prediction. Problem-space attacks must produce valid executable binaries. At the same time, large-scale evaluation remains difficult due to the limited availability of reproducible datasets and models, as well as the computational cost of problem-space attacks. These challenges have constrained the scope of prior evaluations and can lead to robustness conclusions that do not generalize across attacks or detectors, often resulting in narrow or over-optimistic robustness claims. To address these challenges, we conduct a large-scale empirical study of adversarial robustness in Windows malware detection under a common evaluation methodology. We develop an evaluation framework that jointly considers dataset construction, detector representation, attack strategy, runtime validity, and cross-model evaluation. Our experimental testbed comprises eight malware classifiers and nine problem-space evasion attacks, spanning raw-byte, image-based, and feature-

engineered detectors, as well as attack strategies with substantially different optimization and transformation mechanisms. Rather than evaluating attacks in isolation, we study how their effectiveness, overlap, transferability, and induced robustness interact across detectors. Building on this framework, we investigate four research questions (Sec. IV) concerning (i) the effectiveness of blackbox problem-space evasion attacks, (ii) the complementarity and coverage of different attack strategies, (iii) cross-model transferability, and (iv) the effectiveness and cross-attack generalization of adversarial hardening. Our results reveal four main findings. First, attack effectiveness is strongly dependent on both the target detector and attack strategy: no single attack uniformly characterizes the vulnerability of all evaluated detectors, and transformation-space size alone does not explain attack strength. Second, complementary attacks expose adversarial examples missed by the strongest individual attack, and a compact two-attack ensemble covers approximately 99% of the evasive examples observed within our benchmark. Third, direct attack effectiveness and transferability are distinct properties: attacks with limited direct success can still produce highly transferable evasions, while transfer success depends strongly on the surrogate–target pairing. Finally, adversarial hardening is highly model- and attack-dependent. Robustness learned against one attack often generalizes unevenly to others and can, in some cases, increase susceptibility to previously unseen attacks and does not necessarily provide comparable protection against other, even closely related, attacks. In summary, we make the following contributions:

II. BACKGROUND 1) Evasion Attacks: Malware classifiers are vulnerable to adversarial attacks [18], in which adversaries craft subtle modifications to malicious code to evade ML-based detectors at inference-time [16], [19], [20], which are trained to distinguish between benign and malicious samples. These manipulations, known as evasion attacks, deliberately obfuscate malware to bypass detection while preserving its malicious functionality and undermine the trustworthiness of malware defense systems. Problem-Space vs. Feature-Space Attacks. Evasion attacks are commonly categorized into feature-space and problemspace attacks. Feature-space attacks model inputs in abstract feature vectors and manipulate the abstract feature representations used by the classifier, often without ensuring that the resulting samples remain realistic and functional. In contrast, problem-space attacks operate on executable malware, applying realizable transformations that preserve functionality, semantics, and plausibility [21], and therefore provide a realistic threat model for deployed systems [1]. Because problem-space attacks correspond to actual manipulations that an adversary could perform, they provide a more realistic and securityrelevant evaluation of the robustness of Windows malware detectors. White-Box vs. Black-Box Attacks. Evasion attacks can be further divided into white-box [20], [22], [23] and black-box [17], [24]–[27] depending on their access to or knowledge of the internal structure or parameters of the target model. Whitebox attacks assume full model knowledge (gradients, weights) and enable gradient-guided methods, while black-box attacks operate with limited feedback (labels or scores) and better reflect real-world attacker capabilities. Accordingly, this work focuses on black-box problem-space attacks, which provide a realistic assessment of malware detectors under the limited knowledge and operational constraints typical of deployed systems. 2) Static Windows Malware Classifiers: Malware classifiers have been extensively studied using either static or dynamic features. Static analysis extracts information from an executable without running it, for example, by examining headers, strings, byte sequences, or control-flow structures. Dynamic analysis, in contrast, observes program behavior during execution, such as API calls, system interactions, or network activity, typically within a controlled sandbox environment. Static techniques are generally more efficient, provide broader code coverage [28], and can scale to millions of files without executing potentially harmful code, making them a practical first line of defense in real-world deployments. Static Windows malware classifiers can be broadly categorized into four groups based on their feature representation [29], [30]: engineered (tabular) features, visualization-based methods, end-to-end raw-byte models, and graph-based approaches. Engineered feature-based techniques rely on statically extracted tabular features, such as section entropy, imported functions, byte histograms, and metadata. A notable contribution in this category is the EMBER dataset [31], [32], which

1) We develop a unified evaluation methodology for studying adversarial robustness in Windows malware detection across direct attack, attack complementarity, cross-model transfer, and adversarial hardening settings. Our implementation is freely available and can be used by future research to increase the rigor of robustness evaluation. 1 2) We conduct a large-scale empirical evaluation of eight malware classifiers and nine problem-space evasion attacks, covering raw-byte, image-based, and featureengineered detectors together with diverse attack strategies. To the best of our knowledge, this is the first study to jointly examine such a broad range of evasion attacks and detection paradigms. In contrast to prior evaluations centered primarily on individual attacks or robustness dimensions, our study jointly examines attack effectiveness, overlap, transferability, and cross-attack defense generalization. 3) We identify four key empirical insights concerning attack effectiveness, transferability, attack complementarity, and adversarial hardening. These findings provide concrete recommendations for designing more reliable robustness evaluations and for interpreting adversarial vulnerability and robustness claims in Windows malware detection. 1 https://github.com/mashalzainab/windows-malware-adversarialrobustness/

2

Evolutionary or Population or Genetic algorithm-based attacks generate a population of candidate adversarial binaries and evolve them over generations using genetic operations such as selection, crossover, and mutation. GAMMA [1], AIMED [56], MDEA [57], and AMG-PDG [58] fall into this category. Generative adversarial network-based attacks leverage generative adversarial networks (GANs) to produce adversarial payloads or perturbations that can be appended to the malware. GAPGAN [26], MalFox [27] and [5] exemplify this approach. Heuristic or search-based attacks employ targeted modifications guided by hill climbing, occlusion analysis, or other deterministic strategies to increase the probability of evasion while maintaining malware functionality. Malware-makeover [59], MiniMal [6], MalPatch [7], Game-up [60], [55], [61], and [62] are examples of such approaches. All problem-space evasion attacks share the objective of generating functional malware variants that evade detection. Most guarantee format preservation, while fewer ensure full executability and retention of malicious behavior. Recent works increasingly incorporate sandbox or semantic verification to improve realism. In this study, we evaluated a representative set of problemspace black-box attacks, including AIMED, AIMED-RL, ARMED, GAMMA and its variants, MAB-malware, GameUp, and MiniMal against multiple static malware detectors with diverse feature representations. This setup enabled a systematic comparison of attack effectiveness and highlighted potential robustness limitations in current machine learning–based malware detection methods.

has become a widely used benchmark and feature extraction standard. These approaches are efficient, interpretable, and remain relevant, though they depend on the quality of handcrafted features. Visualization-based techniques convert binaries into images for classification. [33] proposed one of the first grayscale image representations of executables, a method that has inspired numerous follow-up studies [34]–[39]. By exposing spatial patterns in the byte sequences, these methods enable convolutional neural networks to identify malware-specific textures effectively. End-to-end learning-based techniques operate directly on raw executable bytes, avoiding manual feature engineering. [40] introduced MalConv, a CNN trained on raw PE files, and [41] later improved it to better handle the extreme length of executables. These methods leverage the representational power of deep neural networks, but must manage computational challenges due to sequence length. Graph-based techniques represent executables as structured graphs, such as Function Call Graphs (FCGs), Control Flow Graphs (CFGs), or Program Dependence Graphs (PDGs), and apply graph representation learning (GRL). For example, MAGIC [42] converts assembly code into CFGs and applies a Deep Graph Convolutional Neural Network (DGCNN) [43], while MalGraph [44] uses hierarchical FCGs and CFGs to capture structural semantics. By modeling relationships between functions and instructions, graph-based methods can identify malware patterns that may be invisible to simpler representations. All these approaches face robustness challenges. Engineered features, while efficient and interpretable, can be bypassed by attacks targeting specific features. Visualization and raw-byte methods reduce manual feature design but remain vulnerable to subtle binary perturbations. Graph-based techniques capture richer program semantics but can still be manipulated through structural changes. In this work, we focus on static malware detectors and systematically evaluate a diverse set of detectors spanning the major categories to assess their robustness against problemspace evasion attacks. 3) Black-Box Problem-Space Windows Malware Evasion Attacks: In the Windows ecosystem, problem-space evasion attacks operating under black-box settings can be categorized by their search strategies [45]: Reinforcement learning-based attacks iteratively select sequences of semantic-preserving modifications to evade detection. Examples include Gym-malware [24], [46], MABmalware [17], AIMED-RL [25], MalwareTotal [47], AMGVAC [48], DQEAF [49], SRL [50], MalInfo [51], gymmalware-mini [52], and [53] which employ RL agents or actorcritic variants to learn effective modification policies. Randomization-based attacks apply stochastic transformations to the malware binary, such as section injection, code padding, or embedding into a dropper. ARMED [54] and Dropper [55] are representative methods, using random modifications while preserving functionality.

III. R ELATED W ORK As research on evaluating the robustness of malware detectors against evasion attacks matures, including black-box problem-space attacks, we conducted a comprehensive literature survey to assess the current state of the field. Table V in the Appendix summarizes our findings. Our survey shows that comprehensive evaluation of malware detectors against problem-space evasion attacks remains lacking. We identify four recurring limitations in the existing literature: limited detector diversity, limited comparison across attacks, insufficient evaluation of realistic problemspace black-box attacks, and limited analysis of cross-model transferability and attack complementarity. Limited detector diversity: Most studies focus on just two model families, MalConv and LGBM, which are used in 21 and 19 of the 33 studies, respectively. In comparison, other and more recent Windows malware detectors receive much less attention. A similar pattern can be seen in the experimental setups: most studies test only one type of model and one attack. Only two studies evaluate more than three model types, while 12 consider multiple attacks. Commercial detectors are somewhat better represented, appearing in 14 studies, but there is still limited broad and systematic testing across different models and attack strategies. These findings suggest a need for more diverse evaluations that include modern malware

3

detectors and examine their robustness against a wider range of evasion attacks. Limited comparison across attacks: A small number of works have systematically compared existing methods. [63] evaluates Extend, Full DOS, Shift, FGSM padding plus slack, and GAMMA against MalConv and LGBM under both whitebox and black-box settings, and additionally examines attack transferability; however, its experimental scope is relatively narrow, with limited model diversity, restricted datasets, and with limited independent functional validation of the generated adversarial binaries. [64] instead compares a different set of generators, including Partial DOS, Full DOS, GAMMA (padding and section-injection variants), and Gym-malware, evaluated against selected antivirus products rather than opensource detectors, and shows that combining generators can improve evasion; however, the evaluation is centered on commercial AV products and a relatively small set of generators, limiting architectural analysis of the underlying detectors. [23] evaluates white-box and black-box attacks against MalConv, two DNN variants, and a GBDT model, explicitly scoping its analysis to a small set of detectors and, while proposing potential mitigations, does not evaluate its attacks against any adversarially hardened models. More recently, [65] proposes a broader benchmark spanning performance, temporal drift, adversarial robustness, and computational overhead across a wider range of detectors, including transfer attacks between models; however, EXE-Bench uses essentially three manipulation families in its adversarial evaluation: FullDOS, contentshift, and GAMMA. Building on these contributions, there remains an opportunity to broaden coverage to more recently proposed attack methods, incorporate defensive strategies, and evaluate a wider and more consistent range of detection models. Limited evaluation of realistic problem-space black-box attacks: Prior studies [12], [66], [67] have primarily focused on feature-space training, using static representations extracted from PE files (e.g., LIEF features) and crafting adversarial examples with gradient-based white-box methods such as FGSM, PGD, or DeepFool. Other works [10] have targeted raw-binary classifiers like MalConv and AvastNet, employing scalable adversarial generation pipelines to train models against inplace randomization, displacement, or gradient-guided attacks, showing that even lower-effort adversarial samples can improve robustness across multiple attack strategies. While these studies provide important insights into adversarial robustness, feature-space or white-box attacks do not necessarily capture the constraints faced by an attacker who must produce a valid, functionality-preserving executable without access to the detector internals. [68] has highlighted the limitations of current ML-based malware detection systems and the need for more realistic adversarial evaluation. Limited understanding of transferability and attack complementarity: The transferability of adversarial malware across models has also been studied [69], suggesting that adversarial examples often fail to generalize across detectors, contrasting with image classification. However, these analyses typically

involve single-step attacks or small datasets, leaving open questions about the cross-model transferability of iterative, functionality-preserving attacks. Moreover, existing studies typically focus on single attacks in isolation, against singular models, leaving the interplay and complementarity of multiple attacks and models largely unexplored. To address these limitations, we present a unified large scale evaluation that enables systematic cross-attack and crossmodel comparison. More broadly, we conduct a holistic evaluation of diverse attack strategies against malware detectors spanning both classical and deep learning representations. IV. R ESEARCH Q UESTIONS We next define the research questions that structure our analysis. RQ1: What are the key factors that influence the efficacy of state-of-the-art problem-space evasion attacks against diverse Windows malware classifiers? Our first question challenges established results extracted from narrow experimental settings (as detailed in Section III) and aims to identify which factors are most critical for evasion across a diverse ecosystem of classifiers. To address this question, we evaluate the Attack Success Rate (ASR, of nine recent problem-space evasion attacks applied to the test sets of eight representative static malware classifiers, including seven open-source models and one commercial detector. This diverse testbed allows us to directly investigate RQ1 and identify the key factors that truly influence attack efficacy. RQ2: What is the optimal ensemble of malware evasion attacks? Having shown the relative effectiveness (or lack thereof) of established attacks in fooling diverse models, we raise the question of whether an attacker could combine multiple attacks into an ensemble and get significant improvement in success rate. This strategy was used in the computer vision domain to form what is today the strongest baseline for adversarial attacks on images, i.e. AutoAttack [70]. In this scenario, an ideal ensemble would be able to produce (almost) all adversarial examples that individual attacks can produce taken together, thereby minimizing the value of integrating the excluded attacks. To achieve this, we study the overlap and complementarity of individual state-of-the-art attacks, in terms of which detected original malware they can turn into a successful adversarial example. More precisely, for an attack A and any other attack B, we want to determine how many unique adversarial examples A can produce, i.e., how many of the original malware that A can turn into successful adversarial examples that B cannot. Attacks with a sufficient number of unique malware samples are added to the ensemble. RQ3: Do malware evasion attacks transfer across malware detectors?

4

Transferability is an important property of adversarial examples that is practically relevant for malware detection [71]. In real-world settings, the attacker may not know any information about the target model or they have only a limited number (even a single) of query attempts to the target model, e.g., due to query costs or to keep the attack elusive. Transfer attacks are relevant in such scenarios because they generate adversarial malware using a surrogate model (built by the attacker), with no extra cost and no risk of alerting the target system. Furthermore, studying transferability enables the risk estimation of generalizable adversarial malware, able to bypass detection from multiple models and thus to infect a large number of systems. We, therefore, study the transferability of malware evasion attacks across models. Starting from a given malware sample, a surrogate model, and an evasion attack (all assumed to be owned and controlled by the attacker), we produce an adversarial example that successfully bypasses the surrogate model. We then test whether this example transfers to a given target model (assumed unknown to the attacker). We repeat this process across all combination of (surrogate and target) models, attacks and original samples. Through this extensive study, we aim to uncover whether attacks remain as effective in a transfer-based threat model as they were in direct attack settings, and which attacks have the highest transferability – highlighting the generalization ability of the examples they produce.

on the influential roles of evasion attacks and defense mechanisms. We translate the insights gained into concrete recommendations for the research community, thereby fostering more rigorous and profound evaluation practices for malware classifiers. V. E VALUATION F RAMEWORK AND S ETTINGS This study conducts a comprehensive and realistic evaluation of malware classifiers under problem-space evasion attacks using a framework that separates the experimental factors, adversarial setting, and evaluation outcomes. The experimental factors comprise data, representation, model, and attack. The data factor captures dataset provenance, composition, and scale, which can influence model generalization and robustness. The representation factor describes how malware is presented to the detector, ranging from raw binaries and images to hand-crafted PE features. The model factor captures differences in detector architectures and learning paradigms. The attack factor characterizes how adversarial malware is generated, including the problem-space transformations, search strategy, and functional constraints required to preserve malware validity and behavior [21], [74]. These experimental factors are considered under a blackbox threat model, which defines the adversary’s knowledge and capabilities. Specifically, the adversary has no access to the detector’s internal architecture, parameters, or training data and can only query the detector to observe its outputs. This setting reflects a realistic attack scenario while enabling executabilitypreserving adversarial malware generation through problemspace modifications [75], [76]. On this foundation, we evaluate robustness from three complementary perspectives: direct attack effectiveness, transferability, and adversarial hardening. Direct attack effectiveness measures evasion against the target detector, transferability captures cross-detector evasion, and adversarial hardening assesses whether adversarial exposure improves robustness to later attacks. Together, these dimensions enable systematic analysis of the interplay among data, representations, models, and attacks across evaluation settings, as illustrated in Figure 1.

RQ4: How effective is adversarial hardening in improving robustness in single-attack and cross-attack scenarios? We next investigate the effectiveness of defenses. In other domains, adversarial hardening has been recognized as the only defense that stood the test of time [72], [73] and was shown to be effective in the malware domain as well [10]. We thus naturally focus on this well-established mechanism. Accordingly, for a given model m and attack a, we use a generate adversarial examples from the training set of a. We then retrain m from scratch using the produced adversarial examples, with the same training parameters as the original model. We then attack each model with the same protocol as in RQ1 and compute the resulting success rates. Moreover, the fact that new problem-space attacks continually emerge raise the question of whether a model hardened against a given attack is also robust to different attacks, or if it “overfits” to the specific attack patterns it learned from. We, therefore, evaluate model robustness across attacks by hardening a given model m against a given attack a, and then evaluating the robustness of the hardened model to different attacks and compare it to the vanilla model’s. Thereby, we aim to reveal whether (or not) hardening can fundamentally strengthen the model’s understanding of fundamental malware representations, or if it merely mitigates specific obfuscation techniques. By investigating these research questions, we aim to advance the understanding of requirements for reliable robustness evaluation of malware classifiers, with a specific focus

Malware Classifier

Raw bytes

D

ec is

io

n

bo u

nd ary

Original Malware

Cuckoo Sandbox

Engineered features Adversarial perturbation

feature_1

feature_2

...

feature_n

val 1 ...

val 2 ...

... ...

val n ...

Model Retraining

Gray-scale images Valid Adversarial Malware

Adversarial malware

Adversarial Generation

Feature Extraction

Classification and Retraining

Corrupt Adversarial Malware

Sandbox Evaluation

Fig. 1: Overview of the adversarial malware generation, evaluation, and retraining pipeline. A. Models Our evaluation includes eight pre-trained models. The models were selected to cover a diverse set of feature represen-

5

tations and architectures that are widely used. These models are highly relevant because they provide pretrained publicly available artifacts that are routinely used as baselines in adversarial malware research, allowing fair comparison and reproducibility. EMBER GBDT 2017 [31] is a gradient-boosted decision tree trained on engineered PE features, called ember features (v1), containing 2351 features, including section metadata, imports, and byte histograms. EMBER GBDT 2018 [31] is an updated version of the EMBER GBDT model, trained on a newer dataset distribution, ember feature v2 containing 2381 features, while keeping the same model structure. EMBER GBDT 2024 [32] relies on EMBER feature version 3 ("thrember"), which enriches previous versions of EMBER with new features (e.g Authenticode signatures, warnings, etc.), increasing the total feature vector dimension from 2,381 to 2,568. We reuse their "All PE files" classifier for our evaluations, which was trained using the Win32, Win64, and .NET files in the training set. MalConv [40] processes executable binaries as sequences of raw bytes using a 1D convolutional neural network. In our evaluations, we used an alternative version of MalConv that we trained on a subset of files (100K binaries) from the original training set while keeping the architecture unchanged. This approach is adopted because the complete training set used for MalConv (i.e. EMBER 2017) is not publicly available, while our adversarial retraining requires mixing clean examples with adversarial examples crafted from them. Therefore, we have trained this version of MalConv using the available subset of EMBER 2017. MalConv2 [41] is an improved version of MalConv, designed to handle the extreme length of executable files efficiently. Like MalConv, it operates directly on raw byte sequences, with architectural modifications to address limitations of the original model. FFNN SOREL [77] is a feed-forward neural network trained on engineered PE features for malware classification. Multiple random seeds were evaluated, with the bestperforming model (seed 3) selected for experiments. LGBM SOREL [77] is a LightGBM model trained on engineered PE features for malware detection. The best-performing model (seed 0) among different seeds was selected. Commercial is a closed source, commercial malware detector that operates on grayscale image representations of executable files. It uses a CNN-based architecture fine-tuned for malware detection. All models are pre-trained on standard malware datasets: EMBER GBDT 2017 on EMBER 2017, MalConv2 and EMBER GBDT 2018 on EMBER 2018, FFNN SOREL and LGBM SOREL on SOREL-20M, EMBER GBDT 2024 on EMBER 2024, the commercial image-based model on its proprietary dataset, and MalConv on a subset of files from EMBER 2017 dataset. We report the evaluation metrics of the model on clean samples in Table IX in Appendix D. Following established practice in malware detection, we calibrate each

detector’s decision threshold at a fixed low false-positive operating point rather than using an arbitrary default threshold. We initially target an FPR of 0.1%, consistent with prior static malware detection studies emphasizing deployability under stringent false-positive constraints [78], [77]. Both 0.1% and 1% FPR are commonly used low-FPR operating points in malware-classifier evaluation, with the former representing a more conservative deployment regime and the latter providing a less restrictive but still operationally relevant trade-off [31], [32]. For models whose TPR falls below 70% at 0.1% FPR, we therefore relax the operating point to 1% FPR. Consequently, every evaluated detector achieves at least 70% TPR at its calibrated operating point. B. Attacks The models described in Section V-A were evaluated against nine black-box problem-space adversarial attacks: MAB-malware [17] is a reinforcement learning framework that models the adversarial process as a multi-armed bandit, balancing exploration and exploitation to identify effective transformations for evading malware detectors. Gamma-shift, Gamma-padding, Gamma-section. 3 variations of the GAMMA attack [1], [79] based on a genetic algorithm that applies functionality-preserving manipulations, including padding, section injection, and shifting, to generate adversarial malware. Game-up [60] constructs chains of problem-space transformations using a greedy search strategy to maximize a universal evasion rate across malware samples, effectively flipping their labels to benign. ARMED [54] applies random binary-level modifications to malware, followed by functionality testing in a sandbox, iterating until evasion is achieved. AIMED [56] uses a genetic algorithm to generate adversarial malware by evaluating perturbations for functionality, evasion success, similarity to the original, and diversity, with the fittest samples used to produce subsequent generations. AIMED-RL [25] is a reinforcement-learning based attack using a DQN agent to learn effective transformations, with mechanisms to encourage diversity and achieve high evasion rates with fewer modifications. MiniMal [6]: MiniMal is a black-box hard-label attack that performs hierarchical perturbation minimization via action reduction, binary search, and particle swarm optimization under functionality constraints. To select representative attacks, we conducted a thorough literature review and established a set of selection criteria: the selected attacks operate in the problem space; assume black-box access to the target model (or can be adapted from white-box settings); and provide open-source artifacts such as pseudo-code or code implementations. Additionally, we prioritized attacks that are well studied and considered stateof-the-art in the literature. The resulting set covers widely studied and practically relevant families of problem-space malware evasion techniques, including reinforcement learning methods (MAB-malware,

6

AIMED-RL), evolutionary and genetic algorithms (AIMED, GAMMA), greedy search (Game-up, MiniMal), and random mutation strategies (ARMED). These attacks operate through a commonly used set of functionality-preserving transformations for Windows executable files, detailed in Appendix B, and collectively capture the dominant strategies used in recent adversarial malware research. Other attacks identified in the literature were excluded because they relied on the same transformation sets or search algorithms as the selected methods, and therefore would not add meaningful diversity to the evaluation. This overlap arises because many attacks build upon the Gym-Malware environment [80]. Some attacks were excluded due to the lack of reproducibility, as discussed in Sec. VII. To ensure a fair comparison, all attacks were executed under comparable effort constraints. While different evasion methods encode the notion of configurations specific to the techniques they employ (e.g., the penalty term in GAMMA or the bandit related parameters in the MAB-Malware attack) the black-box settings common to all attacks are kept identical. Specifically, the query budget is set to 50, the transformation budget to 10, 5 optimization rounds, and all attacks operate in a hardlabel setting. Each attack also involves additional methodspecific parameters, which are configured according to the recommendations of their respective authors.

models by retraining. Additional details about the datasets are provided in Appendix C. D. Sandboxing environment and executability checks Our study recognizes the importance of ensuring that adversarial samples fooling detection are actually executable. We validate the executability of every adversarial sample to avoid overstating attack success. Although previous studies have recognized this need, they evaluate it only through spot checks on a limited set of examples (10 or 50) [7], [17], [59]; in contrast, we assess every generated candidate. A sample is counted as a successful evasion only if it both evades the detector and executes successfully. We validate all adversarial samples using the Cuckoo sandbox [85]. Several attacks already incorporate executability checking during generation: FAME-based [86] attacks (AIMED, AIMED-RL, ARMED, Game-up) natively filter broken samples as part of their attack pipeline. For MiniMal, we retain its static file checking, but replace the original functionality checking pipeline with Cuckoo Sandbox to ensure consistency across all evaluated attacks. This modification is necessary for two reasons: (1) to maintain a uniform executability evaluation framework across all methods, and (2) because the original MiniMal implementation relies on proprietary IDA Pro [87]-based analysis. The remaining attacks do not perform executability checking as part of their attack procedures; for these attacks, we post-validate the generated samples using Cuckoo Sandbox. Samples that execute without crashing are labeled executable evasive samples, whereas crashes or sandbox failures are labeled non-executable and excluded from subsequent analyses. We note that executability alone does not guarantee complete preservation of malicious behavior; accordingly, our executability check is intended as a validation of operational validity, rather than a direct measurement of behavioral equivalence with the original malware.

C. Datasets Used for robustness evaluation. To evaluate model robustness, we curated a set of 1,000 malware samples from a larger test pool of approximately 65k PE files collected from VirusShare 499 [81], the Dike Dataset [82], and the commercial test set (files collected from MalwareBazaar [83] and other sources). All samples in this pool were strictly disjoint from the training data used for any model and are valid functional PE files. From this pool, we sampled 1,000 files using the STAS sampling technique [84], consistently classified as malware by all evaluated models. STAS combines statistically representative sampling with domain-specific constraints to reduce sampling bias and preserve diversity along multiple axes. We adapt this principle to our PE malware pool so that the resulting subset is not dominated by a narrow group of temporally similar samples. This set will be referred as Unified Evaluation Set (UES). To ensure that our insights are general and not artifacts of dataset selection, we also performed additional experiments across datasets by sampling 1,000 malware instances from the test set of the dataset originally used for each model, which were correctly classified as malicious by the corresponding model. We refer to these as the Native Test Set (NTS). These results are included in the Appendix F. Used for adversarial hardening. To perform adversarial hardening, we randomly selected 15% malware of the training set corresponding to each model. Adversarial examples were then generated from these samples using the attacks described in Section V-B, and used to improve the robustness of the

E. Evaluation Metrics ASR. The effectiveness of each attack in RQ1 is measured using the Attack Success Rate (ASR), which quantifies the proportion of originally detected malware samples that are successfully modified to evade detection while remaining executable. Formally, for a set of malware samples M and an attack a, let m′ ∈ M′ denote the executable modified version of m ∈ M generated by a. Then,

ASR(a) =

|{m ∈ M | m′ ∈ M′ evades detection}| |M|

Coverage (C). Coverage is used in RQ2 to compare attacks and determine how much the evasions of one attack overlap with those of another. Let EA denote the set of executable malware samples evaded by the most successful attack on average (across models), or by an ensemble of attacks, and EB the set of executable malware samples evaded by another attack. The set of overlapping samples is O = EA ∩ EB . The

7

coverage of the B attack by the single or ensemble attacks A, C, is then: |O| C(EA , EB ) = |EB |

Moreover, in Appendix F (Table XI), we provide the ASR results for the models evaluated on the Native Test Set (NTS) dataset. The trends of the attacks remain largely consistent across both evaluation settings, yielding the same effective ranking of the attacks. Among all attacks methods, MAB-Malware is the most effective attack, achieving the highest average ASR (93.64%) and almost complete evasion on the commercial detector. Gamma-shift and MiniMal also perform strongly, with average ASRs of 65.79% and 65.74%, respectively, and all three attacks remain the most effective even against the newer EMBER-24 model, which otherwise shows more robustness against other attacks. In our experiments, MAB-Malware operated over seven PE transformations, all of which are contained in the ten-action FAME transformation space; thus, its action space is a strict subset of the FAME action space. Nevertheless, as shown in Table I, MAB-Malware achieves a substantially higher average ASR (93.64%) than the FAME-based attacks, whose average ASRs range from 8.35% to 19.19%. This advantage is not explained by a larger transformation set or longer perturbation sequences. As shown in Table II, successful MAB-Malware evasions use only 2.89 transformations on average and 1.55 distinct transformation types, compared with 10.00 and 6.40 for AIMED and 10.00 and 10.00 for ARMED. Although MAB-Malware requires slightly more target-model queries among successful evasions, these query counts are successconditioned and should therefore be interpreted jointly with ASR. We further analyze the composition of successful perturbation sequences. MAB-Malware exhibits a concentrated transformation profile: section addition and overlay append occur in 54.16% and 51.65% of successful evasions, respectively, and also account for 47.20% and 40.97% of the final actions preceding the first observed evasion. In contrast, AIMED distributes successful sequences across a broader portion of its action space. These results show that transformation-set breadth alone does not explain attack effectiveness and suggest that the mechanisms used to search, select, and compose transformations are important contributors to MAB-Malware’s higher evasion success. Notably, from X), the Gamma-section attack is heavily penalized by executability-preservation checks: its ASR drops by an average of 94%, compared with minor decreases of < 1% for Gamma-padding, MAB-malware and MiniMal, and no decrease for Gamma-shift (see Table X). This sharp reduction suggests that section-level modifications frequently violate execution constraints. [63] evaluates the Gamma-sections attack and reports a high number of evasions against MalConv (and transferability to other detectors such as EMBER17). However, they do not report any runtime execution or functionality checks on the generated adversarial PE files. Our executability validation shows that, although Gammasections initially yields many evasions, a substantial fraction of the generated binaries fail to preserve the original malware executability. We provide more details about this behavior in

This metric captures the fraction of samples evaded by a given attack that are also evaded by the another successful attack or ensemble of attacks on average, helping to establish a hierarchy of attack effectiveness. Transferability (TR). Transferability is relevant to RQ3, quantifying how adversarial examples generated for one model generalize to another. For a given attack, we denote by T RA⇒B its transferability rate from the surrogate model A to the target model B, i.e., the proportion of executable adversarial samples m′A successfully produced by the attack on A that also successfully evade detection by model B. We sur also denote by ASRA⇒B the attack success rate in transferbased settings, i.e. the proportion of eligible original samples whose adversarial variants evade both A and B. T RA⇒B =

sur ASRA⇒B =

|{m′A ∈ M′A | m′A evades B}| |M′A |

|{m ∈ M | m′A evades A ∧ m′A evades B}| |M|

T RA⇒B is conditioned on successful evasion of the surrosur measures end-to-end transfer success gate, whereas ASRA⇒B over all eligible samples. Delta ASR (∆ASR). To evaluate the impact of adversarial hardening and the transferability of defenses across attacks (RQ4), we measure the change in ASR before and after adversarial retraining. Let ASRbefore denote the ASR of a model prior to hardening, and ASRafter the ASR after adversarial retraining. We define two variants of ∆ASR, the absolute and relative change, defined as: ∆ASRabs = ASRafter −ASRbefore ,

∆ASRrel =

∆ASRabs ASRbefore

VI. R ESULTS RQ1: Attack Effectiveness Table I reports the ASR of nine black-box, problem-space adversarial attacks evaluated against eight Windows malware detectors on the Unified Evaluation Set (UES) dataset. Due to the stochastic nature of the attacks, each experiment is repeated three times, and we report the mean ASR across the three runs while the standard deviation remains below 1 percentage point in all cases. The results are stable across repetitions and are not driven by individual random runs, which supports the reproducibility of the observed attack and robustness trends. For completeness, the Appendix includes an additional Table (Tab. X) comparing ASR values before and after executability-preservation checks for the GAMMA attack variants MAB-malware attack and the MiniMal attack, since such checks are not an integrated component of the attacks themselves but are applied post-hoc to ensure their validity.

8

TABLE I: ASR against Windows malware detectors, reported as Mean over 3 runs on the UES dataset. Bold entries indicate the highest mean ASR for each attack, underlined entries indicate the second-highest mean ASR.

Model

MAB- Gamma- Gamma- GammaAIMEDGameMiniMal AIMED ARMED malware shift padding section RL up

Avg

FFNN-SOREL LGBM-SOREL EMBER-17 EMBER-18 EMBER-24 MalConv MalConv2 Commercial

85.63 95.30 89.03 98.33 83.40 98.83 98.70 99.90

15.97 70.50 48.83 91.93 12.73 100.00 88.53 97.80

14.17 70.40 48.87 91.00 7.33 98.53 81.97 91.13

3.27 3.93 5.43 4.60 0 4.53 2.20 3.20

65.20 81.50 64.80 77.40 30.60 98.20 11.00 97.23

5.05 16.23 0.93 56.30 1.77 22.40 35.15 15.70

0.57 7.77 0.20 30.53 0.03 14.30 22.27 7.70

0.60 5.20 0.13 19.50 0.23 13.60 19.23 8.30

0 4.40 0.30 26.00 0 12.50 25.80 9.50

21.16 39.47 28.72 55.07 15.12 51.43 42.76 47.83

Avg

93.64

65.79

62.92

3.39

65.74

19.19

10.42

8.35

9.81

37.70

TABLE II: Attack effectiveness and effort. Query and perturbation statistics are averaged over successful attacks; ASR is averaged over all models. Attack

Search strategy

MAB-Malware AIMED AIMED-RL ARMED Game-Up

Thompson-sampling MAB Genetic programming Deep reinforcement learning Random perturbation sequence Universal perturbation search

Transforms

Queries

Seq. length

Distinct transforms

Runtime (s)

ASR (%)

7 10 10 10 10

6.51 5.79 2.13 1.18 1.00

2.89 10.00 2.13 10.00 9.00

1.55 6.40 1.93 10.00 4.33

5.17 s 424.60 1.95 14.75 20.10

93.64 19.19 10.42 8.35 9.81

Appendix E. Insights from RQ1: The effectiveness of problem-space malware evasion attacks varies substantially across target models, with both ASR and attack ranking changing according to the detector being attacked. Our results indicate that this variability is associated with both the transformation space and the mechanism used to search and compose transformations. MAB-Malware achieves the highest average ASR despite operating over a strict subset of the FAME transformation space, while successful evasions typically rely on only a small number of distinct transformations. Section addition and overlay append dominate successful MAB-Malware sequences, and effective attack construction depends more on identifying and combining high-yield transformations than on simply enlarging the available action set. Detector architecture also plays an important role. Rawbyte end-to-end models are generally more susceptible to the evaluated problem-space manipulations, whereas detectors based on engineered PE features are more resistant to several attacks. This highlights the need for functionality-preserving transformations that meaningfully alter the representations used by feature-engineered detectors, rather than only modifying binary structure at a superficial level. More broadly, evaluating attacks across heterogeneous detector families is necessary to avoid conclusions that are specific to a narrow model class or experimental setting.

evaluate MAB-Malware individually, as it was the strongest single attack in RQ1, and measure how many adversarial examples generated by the remaining attacks are also evaded by MAB-Malware. The corresponding results are provided in Appendix G (Table XII). The observed coverage shows that MAB-Malware does not fully capture the evasive behavior space observed in our evaluation: a subset of samples evaded by other attacks remains uncovered. We therefore iteratively add Gamma-shift, the second most effective attack from RQ1, which complements MAB-Malware by targeting samples that the latter fails to evade, and evaluate the coverage of the resulting two-attack ensemble. Table III summarizes this analysis by reporting, for each attack considered in RQ1, the number of adversarial samples it generates and the proportion of those samples that are also covered by the MAB-Malware+Gamma-shift ensemble.The ensemble achieves complete or near-complete coverage across most attack-model combinations. Aggregated over the evaluated samples, coverage ranges from 98.89% for Gammasection to 100% for MiniMal, AIMED, and Game-Up, with similarly high values for the remaining attacks. We obtain similar results on the NTS dataset, with MABMalware and Gamma-shift again providing the strongest coverage among the evaluated combinations, and the overall coverage trends remaining consistent, implying that the complementarity observed on the primary dataset is not limited to a single sample set. We emphasize that this ensemble is not intended to be exhaustive or to represent the complete space of possible malware attacks. Rather, it provides a compact attack set that captures the adversarial examples observed in our

RQ2: Attack Complementarity We adopt an iterative coverage-based approach to identify a compact and effective ensemble of attacks. We first

9

TABLE III: Coverage of other attacks by the ensemble attack (MAB-malware+Gamma-shift) samples across models. Each attack column is divided into two sub-columns: Overlapping / Evasive and Coverage (%). Model

Gammapadding

Gammasection

MiniMal

AIMED

AIMED-RL

GameUp

ARMED

FFNN-SOREL LGBM-SOREL EMBER-17 EMBER-18 EMBER-24 MalConv MalConv2 Commercial

142 / 142 704 / 704 486 / 488 910 / 910 73 / 73 986 / 986 820 / 820 911 / 911

100% 100% 99.59% 100% 100% 100% 100% 100%

27 / 30 40 / 40 54 / 54 46 / 46 0/0 46 / 46 22 / 22 32 / 32

90% 100% 100% 100% – 100% 100% 100%

652 / 652 815 / 815 648 / 648 774 / 774 306 / 306 982 / 982 110 / 110 973 / 973

100% 100% 100% 100% 100% 100% 100% 100%

50 / 50 162 / 162 9/9 563 / 563 17 / 17 224 / 224 351 / 351 157 / 157

100% 100% 100% 100% 100% 100% 100% 100%

5/5 72 / 76 2/2 305 / 305 3/3 143 / 143 222 / 222 77 / 77

100% 94.74% 100% 100% 100% 100% 100% 100%

6/6 52 / 52 2/2 195 / 195 3/4 136 / 136 193 / 197 83 / 83

100% 100% 100% 100% 75% 100% 97.97% 100%

0/0 44 / 44 3/3 260 / 260 0/0 125 / 125 258 / 258 95 / 95

– 100% 100% 100% – 100% 100% 100%

Total

5032 / 5034

99.96%

267 / 270

98.89%

5260 / 5260

100%

1533 / 1533

100%

829 / 833

99.52%

670 / 675

99.26%

785 / 785

100%

Median

Mean

evaluation and can serve as a practical baseline for robustness assessment. Insights from RQ2: MAB-Malware and Gamma-shift form the compact attack ensemble identified in our evaluation that achieves near-complete coverage of the adversarial examples generated by the evaluated attacks. Their complementarity shows that high individual ASR does not imply complete coverage: attacks with lower overall success can still expose samples missed by the strongest attack. This supports the use of cumulative coverage, rather than ASR alone, when selecting attacks for robustness evaluation. Their complementary behavior allows the ensemble to capture adversarial examples that are missed by either attack individually, while avoiding the additional cost of evaluating a larger collection of attacks. We therefore propose this two-attack ensemble as a practical baseline for evaluating the robustness of malware classifiers. As new attacks are introduced, the same coveragebased procedure can be reapplied to determine whether they expose previously uncovered vulnerabilities and materially expand the evaluation set, while attacks that provide little additional coverage may be considered redundant for this evaluation objective. This iterative procedure provides a potential foundation for developing a standardized and evolving attack suite for malware robustness evaluation, analogous in spirit to established robustness benchmarks such as RobustBench [88].

Fig. 2: Distribution of average (T RA⇒B %) (over all surrogate models) for each attack, over all target models.

RQ3: Transferability

Fig. 3: Adversarial attack transferability across models and attacks.

Mode

100

TR (%)

80 60 40 20 0 MAB Gamma Gamma Gamma MiniMal AIMED AIMED ARMED Gameup malware shift padding section RL Attack

81.5

76.9

37.4

20.8

19.4

7.6

20.4

3.7

9.3

12.6

5.9

0.0

16.7

Gamma padding

7.5

20.4

3.7

6.7

11.8

5.9

0.6

14.9

Gamma section

0.0

0.2

0.1

0.0

1.1

0.2

0.1

0.9

MiniMal 74.9

23.1

63.0

1.1

12.9

70.4

86.1

10.1

AIMED 13.0

0.9

35.5

1.7

5.0

13.7

18.1

25.5

AIMED RL

6.6

0.2

21.3

0.0

0.0

6.6

11.3

19.0

ARMED

7.5

0.1

13.6

0.2

0.6

4.7

11.6

15.1

Gameup

0.0

0.0

0.0

0.0

0.0

0.0

0.0

0.0

Commercial EMBER 17

EMBER 18

EMBER 24

FFNN SOREL

LGBM SOREL

MalConv MalConv2

Target Model

100

80

60

40

20

0

MAB malware 73.9

87.7

85.6

83.4

85.1

77.6

61.2

64.4

Gamma 12.9 shift

25.6

8.6

9.6

13.3

8.9

0.0

24.0

Gamma padding

8.8

26.6

9.2

6.8

12.6

8.9

1.1

20.5

Gamma section

0.0

0.5

0.2

0.0

1.4

0.3

0.2

1.1

MiniMal 99.7

42.9

76.9

3.9

29.9

78.4

93.7

10.7

AIMED 15.2

0.9

46.7

1.8

5.0

15.3

21.2

30.3

AIMED RL

7.5

0.2

28.1

0.0

7.0

13.2

22.3

ARMED

8.3

0.1

19.5

0.2

0.6

5.2

13.6

19.2

Gameup

0.0

0.0

0.0

0.0

0.0

0.0

0.0

0.0

Commercial EMBER 17

EMBER 18

EMBER 24

FFNN SOREL

LGBM SOREL

MalConv MalConv2

Target Model

100

80

Best Surrogate ASR (%)

48.9

Attack

78.1

Gamma shift

Avg Surrogate ASR (%)

Attack

MAB malware 33.2

60

40

20

0

sur sur across attack– (b) Maximum ASRA⇒B under (a) Avg ASRA⇒B target model pairs. an oracle surrogate assumption.

To quantify transferability, we first take an attack-centric perspective and ask whether adversarial examples that successfully evade a surrogate model also evade a different target model. For each attack and target model, and for all successful adversarial samples produced for RQ1, we test these samples against the target model and compute transferability rate as the proportion of surrogate-model evasions that also evade the target, as formally defined in Section V-E. Figure 2 reports the conditional transferability of successful surrogate-model evasions. AIMED, AIMED-RL, and ARMED exhibit the highest transferability among their successful evasions, while MAB-Malware and MiniMal show more moderate transfer rates. This illustrates that direct attack effectiveness and transferability are distinct properties: attacks with low ASR can nevertheless produce adversarial examples that generalize well across models.

We next move from the attack-centric perspective to the target-model perspective. Rather than asking whether successful adversarial examples transfer, we ask: how vulnerable is a given target model when adversarial examples are generated on a different model? We consider two transfer settings: an arbitrary-surrogate setting, where transfer success is averaged over all non-target surrogates, and a best-surrogate setting, where the attacker selects the surrogate yielding the highest transferred ASR. Figure 3a reports the transferred ASR against each target model, averaged over all surrogate models different from the target. Transfer performance is generally lower than directattack ASR, but the magnitude of the reduction varies substantially across attacks and targets. MAB-Malware retains

10

TABLE IV: Absolute and relative change in mean ∆ASRabs and ∆ASRrel before and after adversarial hardening on the UES dataset, illustrating robustness and defense transfer across attacks.

comparatively high transferred ASR against several targets, reaching 81.5% on EMBER-24, 78.1% on EMBER-17, and 76.9% on FFNN-SOREL, whereas its average transfer to MalConv and MalConv2 drops to 20.8% and 19.4%, respectively. MiniMal exhibits a different pattern, transferring strongly to MalConv (86.1%), the commercial detector (74.9%), and LGBM-SOREL (70.4%), but only weakly to EMBER-24 (1.1%). These differences demonstrate that transferability depends strongly on the specific attack and surrogate–target model relationship rather than on direct attack effectiveness alone. Lastly, Figure 3b considers an oracle setting in which the attacker selects the surrogate yielding the highest transferred ASR for each target. Surrogate selection substantially increases transfer success. For MAB-Malware, for example, transferred ASR increases from 33.2% to 73.9% against the commercial detector, from 20.8% to 61.2% against MalConv, and from 19.4% to 64.4% against MalConv2. In some cases, the best-surrogate transfer approaches or matches directattack performance, such as MAB-Malware against EMBER24 (83.4%). These results show that surrogate choice can substantially alter the apparent transfer robustness of a detector Notably, transfer can occasionally exceed the corresponding direct-attack ASR. For example, MiniMal reaches 99.7% transferred ASR against the commercial detector with the best surrogate, compared with 97.23% under direct attack. This shows that surrogate-generated adversarial examples can expose target-model vulnerabilities that are not necessarily reached by directly optimizing against that target. Insights from RQ3: Our results confirm that transferability depends strongly on the choice of surrogate and target models, with carefully selected surrogates achieving substantially higher transfer rates than arbitrary ones. However, we also observe an important distinction between attack effectiveness and transferability: attacks with lower direct-attack success can nevertheless produce adversarial examples that transfer well, whereas highly effective attacks may produce more surrogate-specific examples. This highlights the importance of evaluating both attack effectiveness and transferability when assessing the threat posed by adversarial malware.

Model

Hardened Against

Test Attack

ASRbefore

ASRafter

∆ASRabs

∆ASRrel (%)

Gamma-padding

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

48.87 48.83 89.03 64.80 0.93

0.36 1.00 77.9 40.5 0

↓ -48.51 ↓ -47.83 ↓ -11.13 ↓ -24.30 ↓ -0.93

↓ -99.26% ↓ -97.95% ↓ -12.50% ↓ -37.50% ↓ -100.00%

Gamma-shift

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

48.87 48.83 89.03 64.80 0.93

0.26 0.43 77.7 52.0 0

↓ -48.61 ↓ -48.40 ↓ -11.33 ↓ -12.80 ↓ -0.93

↓ -99.47% ↓ -99.12% ↓ -12.73% ↓ -19.75% ↓ -100.00%

MAB-malware

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

48.87 48.83 89.03 64.80 0.93

0.66 1.4 3.6 1.4 0

↓ -48.21 ↓ -47.43 ↓ -85.43 ↓ -63.40 ↓ -0.93

↓ -98.65% ↓ -97.13% ↓ -95.96% ↓ -97.84% ↓ -100.00%

MiniMal

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

48.87 48.83 89.03 64.80 0.93

0.10 0.80 83.6 0 0

↓ -48.77 ↓ -48.03 ↓ -5.43 ↓ -64.80 ↓ -0.93

↓ -99.80% ↓ -98.36% ↓ -6.10% ↓ -100.00% ↓ -100.00%

AIMED

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

48.87 48.83 89.03 64.80 0.93

83.03 83.6 85.4 60.0 0

↑ 34.16 ↑ 34.77 ↓ -3.63 ↓ -4.80 ↓ -0.93

↑ 69.90% ↑ 71.21% ↓ -4.08% ↓ -7.41% ↓ -100.00%

Gamma-padding

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

98.53 100 98.83 98.20 22.40

91.54 99.92 98.53 82.9 20.03

↓ -6.99 ↓ -0.08 ↓ -0.30 ↓ -15.30 ↓ -2.37

↓ -7.09% ↓ -0.08% ↓ -0.30% ↓ -15.58% ↓ -10.58%

Gamma-shift

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

98.53 100 98.83 98.20 22.40

88.152 99.49 98.55 88.6 19.63

↓ -10.38 ↓ -0.51 ↓ -0.28 ↓ -9.60 ↓ -2.77

↓ -10.53% ↓ -0.51% ↓ -0.28% ↓ -9.78% ↓ -12.37%

MAB-malware

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

98.53 100 98.83 98.20 22.40

89.64 99.75 97.57 89.8 17.93

↓ -8.89 ↓ -0.25 ↓ -1.26 ↓ -8.40 ↓ -4.47

↓ -9.02% ↓ -0.25% ↓ -1.27% ↓ -8.55% ↓ -19.96%

MiniMal

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

98.53 100 98.83 98.20 22.40

92.45 99.65 98.11 75.7 15.8

↓ -6.08 ↓ -0.35 ↓ -0.72 ↓ -22.50 ↓ -6.60

↓ -6.17% ↓ -0.35% ↓ -0.73% ↓ -22.91% ↓ -29.46%

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

98.53 100 98.83 98.20 22.40

98.50 100 98.72 87.1 8.30

↓ -0.03

↓ -0.03%

AIMED

Gamma-padding

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

91.13 97.80 99.90 97.23 15.70

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

EMBER-17

MalConv

0.00

0.00%

↓ -0.11 ↓ -11.10 ↓ -14.10

↓ -0.11% ↓ -11.30% ↓ -62.95%

0.10 1.19 99.90 93.2 20.5

↓ -91.03 ↓ -96.61

↓ -99.89% ↓ -98.78%

0.00 ↓ -4.03 ↑ 4.80

↓ -4.14% ↑ 30.57%

91.13 97.80 99.90 97.23 15.70

3.18 0.30 99.90 88.3 16.1

↓ -87.95 ↓ -97.50

↓ -96.51% ↓ -99.69%

0.00 ↓ -8.93 ↑ 0.40

↓ -9.18% ↑ 2.55%

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

91.13 97.80 99.90 97.23 15.70

46.31 71.72 99.90 92.2 20.3

↓ -44.82 ↓ -26.08

↓ -49.18% ↓ -26.67%

0.00 ↓ -5.03 ↑ 4.60

↓ -5.17% ↑ 29.30%

MiniMal

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

91.13 97.80 99.90 97.23 15.70

92.24 99.69 100 92.1 10.5

↑ 1.11 ↑ 1.89 ↑ 0.10 ↓ -5.13 ↓ -5.20

↑ 1.22% ↑ 1.93% ↑ 0.10% ↓ -5.28% ↓ -33.12%

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

91.13 97.80 99.90 97.23 15.70

98.05 99.45 99.90 92.7 10.20

↑ 6.92 ↑ 1.65

↑ 7.59% ↑ 1.69%

AIMED

0.00 ↓ -4.53 ↓ -5.50

↓ -4.66% ↓ -35.03%

Gamma-shift

Commercial MAB-malware

RQ4: Adversarial Hardening We evaluate adversarial hardening on three representative models, EMBER-17, MalConv, and the Commercial model, which span diverse feature representations and have each shown vulnerability to problem-space attacks. Each model is retrained using adversarial examples from the five attacks with the highest average impact: MAB-malware, Gamma-shift, MiniMal, Gamma-padding, and AIMED, all of which achieve an average ASR above 15% and as shown in RQ2, MABmalware and Gamma-shift also achieve broad coverage across other attack types making this set both representative of highperforming attacks. Table IV and Figure 4 present the ASRs of hardened models against both the attack used for hardening and other unseen attacks. Each attack was executed 3 times and results are

0.00%

0.00%

0.00%

0.00%

averaged across runs, whereas the standard deviation remains minimal. For comparison, the corresponding pre-hardening ASRs are also shown, with results for the attack used in hardening highlighted to emphasize direct defense effectiveness. To evaluate the hardened models, we use the same test sets as for the original models, i.e., the UES dataset. Over 90% of files remain identical across the two evaluations; in the few cases where a file was misclassified, it was replaced with a correctly detected malware sample to maintain consistency. Ta-

11

EMBER-17

MalConv Gamma padding

-7.1

-0.1

50

Gamma -99.5 -99.1 -12.7 -19.8 -100.0 shift

Gamma -10.5 shift MAB malware

50

100

MAB malware -98.7 -97.1 -96.0 -97.8 -100.0 MiniMal -99.8 -98.4 AIMED

69.9 71.2

-6.1 -100.0 -100.0 -4.1

-7.4 -100.0

Gamma Gamma MAB MiniMal AIMED padding shift malware

Test Attack

Commercial

-0.3 -15.6 -10.6

Gamma padding -99.9 -98.8

0.0

-4.1

30.6

-0.5

-0.3

-9.8 -12.4

Gamma -96.5 -99.7 shift

0.0

-9.2

2.5

-9.0

-0.2

-1.3

-8.6 -20.0

MAB malware -49.2 -26.7

0.0

-5.2

29.3

MiniMal

-6.2

-0.3

-0.7 -22.9 -29.5

MiniMal

1.2

1.9

0.1

-5.3 -33.1

AIMED

-0.0

0.0

-0.1 -11.3 -63.0

AIMED

7.6

1.7

0.0

-4.7 -35.0

Gamma Gamma MAB MiniMal AIMED padding shift malware

Test Attack

Defense Attack

0

Defense Attack

Gamma padding -99.3 -98.0 -12.5 -37.5 -100.0

ASRrel (\%) Defense Attack

100

Gamma Gamma MAB MiniMal AIMED padding shift malware

Test Attack

Fig. 4: Relative change in ASR (∆ASRrel , %) after adversarial hardening on the UES dataset. Defense Attack denotes the attack used for training and Test Attack the attack used for evaluation. Negative values indicate reduced ASR (improved robustness), while positive values indicate increased susceptibility.

ble XIII (Appendix H) reports the performance metrics of the adversarially retrained models, with classification thresholds determined using the same procedure as for the clean models (Section V-A). We additionally report results on the NTS dataset in Appendix I (Table XIV). Notably, AIMED shows a low initial ASR on the EMBER-17 model when evaluated on the UES dataset, but a substantially higher ASR on the NTS dataset; despite this shift in initial ASR, the robustness gains from hardening remain consistent across both settings. When hardened models are attacked with the same attacks used for retraining, adversarial hardening generally reduces ASR substantially, although the magnitude of the improvement depends strongly on the model. EMBER-17 exhibits the largest and most consistent gains, with relative ASR reductions of 100% for MiniMal and AIMED, approximately 99% for Gamma-padding and Gamma-shift, and 95.96% for MAB-malware. This indicates that the feature-based EMBER17 detector can effectively adapt to the adversarial patterns represented during retraining. In contrast, MalConv benefits substantially less from attack-specific hardening: the relative reductions are only 7.09%, 0.51%, and 1.27% for Gamma-padding, Gamma-shift, and MAB-malware, respectively, although larger improvements are observed for MiniMal (22.91%) and particularly AIMED (62.95%). The Commercial model shows strong robustness gains against Gamma-padding and Gamma-shift, with relative ASR reductions of 99.89% and 99.69%, respectively, and a moderate improvement against AIMED (35.03%); however, hardening against MAB-malware produces no reduction in ASR. To assess defense transferability, we evaluate each hardened model against attacks not used during retraining. The strongest reciprocal transfer occurs between Gamma-padding and Gamma-shift: on EMBER-17, hardening against Gammapadding reduces Gamma-shift ASR by 97.95%, while Gammashift hardening reduces Gamma-padding ASR by 99.47%; similarly, the Commercial model shows reductions of 98.78% and 96.51%, respectively. In contrast, transfer is much weaker for MalConv, where Gamma-padding hardening reduces Gamma-shift ASR by only 0.08%, while Gamma-

shift hardening reduces Gamma-padding ASR by 10.53%. In contrast, transfer between MAB-Malware and AIMED is substantially less consistent despite their overlapping transformation spaces (Table VI), as discussed in RQ1. This difference is also reflected in how the two attacks use those transformations: successful MAB-Malware evasions are concentrated on a small subset of actions, while AIMED distributes successful sequences across a much broader portion of its action space. Shared transformation availability alone is insufficient to predict defense transferability; the way transformations are selected and composed during attack generation leads to the results. More generally, Figure 4 shows that adversarial hardening can also increase susceptibility to unseen attacks; for example, AIMED-hardened EMBER-17 increases ASR under Gamma-padding and Gamma-shift by 69.90% and 71.21%, respectively.

Insights from RQ4: Our findings suggest that adversarial hardening does not benefit all malware detection paradigms equally. Engineered feature-based detectors exhibit substantially stronger robustness gains than the end-to-end detector considered in our evaluation, indicating that the effectiveness of adversarial retraining depends strongly on the underlying feature representation and model architecture. We further observe that similarity between attacks is not, by itself, a reliable predictor of defense transferability. Although strong reciprocal transfer occurs between Gamma-padding and Gamma-shift for some detectors, the same effect is weak or asymmetric for MalConv, while attacks sharing overlapping functionality-preserving transformations, such as MAB-malware and AIMED, exhibit considerably less consistent transfer. These results motivate further investigation into which properties of adversarial examples determine whether robustness learned against one attack generalizes to others. In practice, adversarial hardening may therefore need to be complemented by defenses that are less dependent on the particular attack distribution observed during training.

12

VII. L IMITATIONS AND F UTURE W ORK Our findings across all research questions reveal that robustness in malware detection is not an inherent property of a model, but is shaped by architectural biases, data representation, and the diversity of adversarial manipulations. While this study provides a comprehensive analysis, some limitations must be acknowledged. First, our work inherits the inherent constraints of problemspace attacks in Windows malware, including computational complexity and data availability. Acquiring original malware and benign executables is challenging due to copyright and safety concerns; consequently, our models were trained on the EMBER dataset, which represents only a subset of the full distribution. Second, while we evaluated nine representative problem-space evasion attacks, other published strategies were not reproducible, meaning our analysis may not capture the entire threat landscape of problem-space vulnerabilities. In particular, the public implementation of [10], [59], [89] is either incomplete or relies on proprietary disassembly tools (e.g., IDA Pro), which challenges their reproducibility. Taken together, these limitations and our core findings advocate for a paradigm shift: from building attack-specific defenses toward fostering representation-level resilience. Our results indicate that many vulnerabilities stem not from the learning algorithm itself, but from feature representations that allow superficial, semantics-preserving modifications to dictate a model’s decision. Therefore, we propose that future research should prioritize the exploration of robust feature representations for Windows malware detection, which are inherently less susceptible to such adversarial manipulations. VIII. C ONCLUSION In this work, we systematically evaluated the vulnerability of Windows malware classifiers to realistic evasion attacks, examining attack effectiveness, cross-model transferability, and adversarial hardening. Our findings reveal significant variability in attack effectiveness and transferability across models and datasets, highlighting important limitations in the robustness of current malware detectors. We also show that evaluating attacks in isolation can underestimate real-world risk, and that adversarial hardening, while beneficial in some cases, does not consistently protect against diverse attack strategies. Overall, our study provides a reproducible benchmark and systematic evaluation methodology that can support the development of more robust malware classifier and more realistic evaluation practices.

13

O PEN S CIENCE

[9] S. Dyrmishi, S. Ghamizi, T. Simonetto, Y. Le Traon, and M. Cordy, “On the empirical effectiveness of unrealistic adversarial hardening against realistic adversarial attacks,” in 2023 IEEE symposium on security and privacy (SP). IEEE, 2023, pp. 1384–1400. [10] K. Lucas, S. Pai, W. Lin, L. Bauer, M. K. Reiter, and M. Sharif, “Adversarial training for {Raw-Binary} malware classifiers,” in 32nd USENIX security symposium (USENIX security 23), 2023, pp. 1163– 1180. [11] H. Bostani, Z. Zhao, Z. Liu, and V. Moonsamy, “Level up with ml vulnerability identification: leveraging domain constraints in feature space for robust android malware detection,” ACM Transactions on Privacy and Security, vol. 28, no. 2, pp. 1–32, 2025. [12] B. G. Doan, S. Yang, P. Montague, O. De Vel, T. Abraham, S. Camtepe, S. S. Kanhere, E. Abbasnejad, and D. C. Ranashinghe, “Feature-space bayesian adversarial learning improved malware detector robustness,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 12, 2023, pp. 14 783–14 791. [13] H. Bostani, J. Cortellazzi, D. Arp, F. Pierazzi, V. Moonsamy, and L. Cavallaro, “On the effectiveness of adversarial training on malware classifiers,” arXiv preprint arXiv:2412.18218, 2024. [14] D. Li, S. Cui, Y. Li, J. Xu, F. Xiao, and S. Xu, “Pad: Towards principled adversarial malware detection against evasion attacks,” IEEE Transactions on Dependable and Secure Computing, 2023. [15] A. Bensaoud, J. Kalita, and M. Bensaoud, “A survey of malware detection using deep learning,” Machine Learning With Applications, vol. 16, p. 100546, 2024. [16] K. Aryal, M. Gupta, M. Abdelsalam, P. Kunwar, and B. Thuraisingham, “A survey on adversarial attacks for malware analysis,” IEEE Access, 2024. [17] W. Song, X. Li, S. Afroz, D. Garg, D. Kuznetsov, and H. Yin, “Mabmalware: A reinforcement learning framework for attacking static malware classifiers,” arXiv preprint arXiv:2003.03100, 2020. [18] B. Biggio, I. Corona, D. Maiorca, B. Nelson, N. Šrndić, P. Laskov, G. Giacinto, and F. Roli, “Evasion attacks against machine learning at test time,” in Joint European conference on machine learning and knowledge discovery in databases. Springer, 2013, pp. 387–402. [19] I. J. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and harnessing adversarial examples,” arXiv preprint arXiv:1412.6572, 2014. [20] B. Kolosnjaji, A. Demontis, B. Biggio, D. Maiorca, G. Giacinto, C. Eckert, and F. Roli, “Adversarial malware binaries: Evading deep learning for malware detection in executables,” in 2018 26th European signal processing conference (EUSIPCO). IEEE, 2018, pp. 533–537. [21] F. Pierazzi, F. Pendlebury, J. Cortellazzi, and L. Cavallaro, “Intriguing properties of adversarial ml attacks in the problem space,” in 2020 IEEE symposium on security and privacy (SP). IEEE, 2020, pp. 1332–1349. [22] L. Demetrio, B. Biggio, G. Lagorio, F. Roli, and A. Armando, “Explaining vulnerabilities of deep learning to adversarial malware binaries,” arXiv preprint arXiv:1901.03583, 2019. [23] L. Demetrio, S. E. Coull, B. Biggio, G. Lagorio, A. Armando, and F. Roli, “Adversarial exemples: A survey and experimental evaluation of practical attacks on machine learning for windows malware detection,” ACM Transactions on Privacy and Security (TOPS), vol. 24, no. 4, pp. 1–31, 2021. [24] H. S. Anderson, A. Kharkar, B. Filar, D. Evans, and P. Roth, “Learning to evade static pe machine learning malware models via reinforcement learning,” arXiv preprint arXiv:1801.08917, Jan. 2018. [25] R. Labaca-Castro, S. Franz, and G. D. Rodosek, “Aimed-rl: Exploring adversarial malware examples with reinforcement learning,” in Joint European Conference on Machine Learning and Knowledge Discovery in Databases (ECML PKDD). Springer, 2021, pp. 37–52. [26] J. Yuan, S. Zhou, L. Lin, F. Wang, and J. Cui, “Black-box adversarial attacks against deep learning based malware binaries detection with gan,” in ECAI 2020. IOS Press, 2020, pp. 2536–2542. [27] F. Zhong, X. Cheng, D. Yu, B. Gong, S. Song, and J. Yu, “Malfox: Camouflaged adversarial malware example generation based on convgans against black-box detectors,” IEEE Transactions on Computers, vol. 73, no. 4, pp. 980–993, 2023. [28] Z. Chen, E. Brophy, and T. Ward, “Malware classification using static disassembly and machine learning,” arXiv preprint arXiv:2201.07649, 2021. [29] D. Ucci, L. Aniello, and R. Baldoni, “Survey of machine learning techniques for malware analysis,” Computers & Security, vol. 81, pp. 123–147, 2019.

To support reproducibility, we release the code, experimental configurations, and trained models required to reproduce the reported results. Raw malware binaries will not be redistributed due to safety and licensing considerations. LLM U SAGE C ONSIDERATIONS LLMs were used for editorial assistance and support with code development and debugging. All generated content was reviewed and validated by the authors. LLMs were not used to generate experimental data or results, and the authors remain fully responsible for the accuracy and integrity of the manuscript. E THICAL C ONSIDERATIONS This study analyzes Windows malware detection systems in fully isolated environments, using authorized or public samples with no exposure to live systems. All malware processing and execution is conducted within these isolated research environments, which are separated from production systems and configured to prevent interaction with third-party hosts. We do not deploy malware outside these environments, evaluate attacks against operational security products without authorization, or distribute live malware, generated adversarial binaries, or transformation payloads. The problem-space attacks used in our evaluation are existing techniques from prior work; we do not introduce new functionality-preserving binary transformations. Vulnerabilities discovered during this study were responsibly disclosed to the affected vendors, with the goal of improving adversarial robustness and strengthening malware detection systems. R EFERENCES [1] L. Demetrio, B. Biggio, G. Lagorio, F. Roli, and A. Armando, “Functionality-preserving black-box optimization of adversarial windows malware,” IEEE Transactions on Information Forensics and Security, vol. 16, pp. 3469–3478, 2021. [2] W. Song, X. Li, S. Afroz, D. Garg, D. Kuznetsov, and H. Yin, “Automatic generation of adversarial examples for interpreting malware classifiers,” arXiv preprint arXiv:2003.03100, 2020. [3] M. Sharif, K. Lucas, L. Bauer, M. K. Reiter, and S. Shintre, “Optimization-guided binary diversification to mislead neural networks for malware detection,” arXiv e-prints, pp. arXiv–1912, 2019. [4] A. Khormali, A. Abusnaina, S. Chen, D. Nyang, and A. Mohaisen, “Copycat: practical adversarial attacks on visualization-based malware detection,” arXiv preprint arXiv:1909.09735, 2019. [5] I. Rosenberg, A. Shabtai, Y. Elovici, and L. Rokach, “Query-efficient black-box attack against sequence-based malware classifiers,” in Proceedings of the 36th Annual Computer Security Applications Conference, 2020, pp. 611–626. [6] C. Li, Z. Jiang, Y. Wang, T. Xia, Y. Zhang, and Y. Mao, “Minimal: hardlabel adversarial attack against static malware detection with minimal perturbation,” in Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, 2025, pp. 5589–5597. [7] D. Zhan, Y. Duan, Y. Hu, W. Li, S. Guo, and Z. Pan, “Malpatch: Evading dnn-based malware detection with adversarial patches,” IEEE Transactions on Information Forensics and Security, vol. 19, pp. 1183– 1198, 2023. [8] Y. Liu, J. Ning, Q. Feng, Y. Zhang, Y. Huang, and L. Y. Zhang, “Grasp: Hard-label black-box malware evasion with higher success, fewer queries, and smaller perturbations.”

14

detection,” IEEE Transactions on Dependable and Secure Computing, vol. 20, no. 2, pp. 1390–1402, 2022. [51] F. Zhong, P. Hu, G. Zhang, H. Li, and X. Cheng, “Reinforcement learning based adversarial malware example generation against blackbox detectors,” Computers & Security, vol. 121, p. 102869, 2022. [52] J. Chen, J. Jiang, R. Li, and Y. Dou, “Generating adversarial examples for static pe malware detector based on deep reinforcement learning,” in Journal of Physics: Conference Series, vol. 1575, no. 1. IOP Publishing, 2020, p. 012011. [53] D. Gibert, M. Fredrikson, C. Mateu, J. Planes, and Q. Le, “Enhancing the insertion of nop instructions to obfuscate malware via deep reinforcement learning,” Computers & Security, vol. 113, p. 102543, 2022. [54] R. Labaca-Castro, C. Schmitt, and G. D. Rodosek, “Armed: How automatic malware modifications can evade static detection?” in 2019 5th International Conference on Information Management (ICIM). IEEE, 2019, pp. 20–27. [55] F. Ceschin, M. Botacin, H. M. Gomes, L. S. Oliveira, and A. Grégio, “Shallow security: On the creation of adversarial variants to evade machine learning-based malware detectors,” in Proceedings of the 3rd Reversing and Offensive-oriented Trends Symposium, 2019, pp. 1–9. [56] R. Labaca-Castro, C. Schmitt, and G. D. Rodosek, “Aimed: Evolving malware with genetic programming to evade detection,” in 2019 18th IEEE International Conference On Trust, Security And Privacy In Computing And Communications/13th IEEE International Conference On Big Data Science And Engineering (TrustCom/BigDataSE). IEEE, 2019, pp. 240–247. [57] X. Wang and R. Miikkulainen, “Mdea: Malware detection with evolutionary adversarial learning,” in 2020 IEEE Congress on Evolutionary Computation (CEC). IEEE, 2020, pp. 1–8. [58] S. Wang, Y. Fang, Y. Xu, and Y. Wang, “Black-box adversarial windows malware generation via united puppet-based dropper and genetic algorithm,” in 2022 IEEE 24th Int Conf on High Performance Computing & Communications; 8th Int Conf on Data Science & Systems; 20th Int Conf on Smart City; 8th Int Conf on Dependability in Sensor, Cloud & Big Data Systems & Application (HPCC/DSS/SmartCity/DependSys). IEEE, 2022, pp. 653–662. [59] K. Lucas, M. Sharif, L. Bauer, M. K. Reiter, and S. Shintre, “Malware makeover: Breaking ml-based static analysis by modifying executable bytes,” in Proceedings of the 2021 ACM Asia Conference on Computer and Communications Security, 2021, pp. 744–758. [60] R. Labaca-Castro, L. Muñoz-González, F. Pendlebury, G. D. Rodosek, F. Pierazzi, and L. Cavallaro, “Realizable universal adversarial perturbations for malware,” arXiv preprint arXiv:2102.06747, 2022. [61] F. Ceschin, M. Botacin, G. Lüders, H. M. Gomes, L. Oliveira, and A. Gregio, “No need to teach new tricks to old malware: Winning an evasion challenge with xor-based adversarial samples,” in Reversing and Offensive-oriented Trends Symposium, 2020, pp. 13–22. [62] W. Fleshman, E. Raff, R. Zak, M. McLean, and C. Nicholas, “Static malware detection & subterfuge: Quantifying the robustness of machine learning and current anti-virus,” in 2018 13th International Conference on Malicious and Unwanted Software (MALWARE). IEEE, 2018, pp. 1–10. [63] M. Imran, A. Appice, and D. Malerba, “Evaluating realistic adversarial attacks against machine learning models for windows pe malware detection,” Future Internet, vol. 16, no. 5, p. 168, 2024. [64] P. Louthánová, M. Kozák, M. Jureček, M. Stamp, and F. Di Troia, “A comparison of adversarial malware generators,” Journal of Computer Virology and Hacking Techniques, vol. 20, no. 4, pp. 623–639, 2024. [65] A. Ponte, D. Gibert, M. Kozak, D. Trizna, M. Pintor, B. Biggio, F. Roli, and L. Demetrio, “Exe-bench: Ranking the tradeoffs of ai-based windows malware detectors for real-world usability,” arXiv preprint arXiv:2607.24177, 2026. [66] A. Al-Dujaili, A. Huang, E. Hemberg, and U.-M. O’Reilly, “Adversarial deep learning for robust detection of binary encoded malware,” in 2018 IEEE Security and Privacy Workshops (SPW). IEEE, 2018, pp. 76–82. [67] L. Lobascio, G. Andresini, A. Appice, and D. Malerba, “Adversarial training to improve accuracy and robustness of a windows pe malware detection model,” 2025. [68] P. He, Y. Mao, C. Li, L. Cavallaro, T. Wang, and S. Ji, “On the security risks of ml-based malware detection systems: A survey,” arXiv preprint arXiv:2505.10903, 2025.

[30] D. Gibert, “Machine learning for windows malware detection and classification: methods, challenges, and ongoing research,” in Malware: Handbook of Prevention and Detection. Springer, 2024, pp. 143–173. [31] H. S. Anderson and P. Roth, “Ember: an open dataset for training static pe malware machine learning models,” arXiv preprint arXiv:1804.04637, 2018. [32] R. J. Joyce, G. Miller, P. Roth, R. Zak, E. Zaresky-Williams, H. Anderson, E. Raff, and J. Holt, “Ember2024-a benchmark dataset for holistic evaluation of malware classifiers,” in Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2, 2025, pp. 5516–5526. [33] L. Nataraj, S. Karthikeyan, G. Jacob, and B. S. Manjunath, “Malware images: visualization and automatic classification,” in Proceedings of the 8th international symposium on visualization for cyber security, 2011, pp. 1–7. [34] K. Kancherla and S. Mukkamala, “Image visualization based malware detection,” in 2013 IEEE symposium on computational intelligence in cyber security (CICS). IEEE, 2013, pp. 40–44. [35] J. H. Go, T. Jan, M. Mohanty, O. P. Patel, D. Puthal, and M. Prasad, “Visualization approach for malware classification with resnext,” in 2020 IEEE Congress on Evolutionary Computation (CEC). IEEE, 2020, pp. 1–7. [36] R. Vinayakumar, M. Alazab, K. Soman, P. Poornachandran, and S. Venkatraman, “Robust intelligent malware detection using deep learning,” IEEE access, vol. 7, pp. 46 717–46 738, 2019. [37] S. Seneviratne, R. Shariffdeen, S. Rasnayaka, and N. Kasthuriarachchi, “Self-supervised vision transformers for malware detection,” IEEE Access, vol. 10, pp. 103 121–103 135, 2022. [38] B. Prima and M. Bouhorma, “Using transfer learning for malware classification,” The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences, vol. 44, pp. 343– 349, 2020. [39] N. Bhodia, P. Prajapati, F. Di Troia, and M. Stamp, “Transfer learning for image-based malware classification,” arXiv preprint arXiv:1903.11551, 2019. [40] E. Raff, J. Barker, J. Sylvester, R. Brandon, B. Catanzaro, and C. Nicholas, “Malware detection by eating a whole exe,” arXiv preprint arXiv:1710.09435, 2017. [41] E. Raff, W. Fleshman, R. Zak, H. S. Anderson, B. Filar, and M. McLean, “Classifying sequences of extreme length with constant memory applied to malware detection,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 11, 2021, pp. 9386– 9394. [42] J. Yan, G. Yan, and D. Jin, “Classifying malware represented as control flow graphs using deep graph convolutional neural network,” in 2019 49th annual IEEE/IFIP international conference on dependable systems and networks (DSN). IEEE, 2019, pp. 52–63. [43] M. Zhang, Z. Cui, M. Neumann, and Y. Chen, “An end-to-end deep learning architecture for graph classification,” in Proceedings of the AAAI conference on artificial intelligence, vol. 32, no. 1, 2018. [44] X. Ling, L. Wu, W. Deng, Z. Qu, J. Zhang, S. Zhang, T. Ma, B. Wang, C. Wu, and S. Ji, “Malgraph: Hierarchical graph neural networks for robust windows malware detection,” in IEEE INFOCOM 2022-IEEE Conference on Computer Communications. IEEE, 2022, pp. 1998– 2007. [45] X. Ling, L. Wu, J. Zhang, Z. Qu, W. Deng, X. Chen, Y. Qian, C. Wu, S. Ji, T. Luo et al., “Adversarial attacks against windows pe malware detection: A survey of the state-of-the-art,” Computers & Security, vol. 128, p. 103134, 2023. [46] H. S. Anderson, A. Kharkar, B. Filar, and P. Roth, “Evading machine learning malware detection,” black Hat, vol. 2017, pp. 1–6, 2017. [47] S. He, C. Fu, H. Hu, J. Chen, J. Lv, and S. Jiang, “Malwaretotal: Multi-faceted and sequence-aware bypass tactics against static malware detection,” in Proceedings of the IEEE/ACM 46th International Conference on Software Engineering, 2024, pp. 1–12. [48] M. Ebrahimi, J. Pacheco, W. Li, J. L. Hu, and H. Chen, “Binary black-box attacks against static malware detectors with reinforcement learning in discrete action spaces,” in 2021 IEEE security and privacy workshops (SPW). IEEE, 2021, pp. 85–91. [49] Z. Fang, J. Wang, B. Li, S. Wu, Y. Zhou, and H. Huang, “Evading anti-malware engines with deep reinforcement learning,” IEEE Access, vol. 7, pp. 48 867–48 879, 2019. [50] L. Zhang, P. Liu, Y.-H. Choi, and P. Chen, “Semantics-preserving reinforcement learning attack against graph neural networks for malware

15

[92] B. Chen, Z. Ren, C. Yu, I. Hussain, and J. Liu, “Adversarial examples for cnn-based malware detectors,” IEEE Access, vol. 7, pp. 54 360– 54 371, 2019. [93] C. Wu, J. Shi, Y. Yang, and W. Li, “Enhancing machine learning based malware detection model by reinforcement learning,” in Proceedings of the 8th International Conference on Communication and Network Security, 2018, pp. 74–78. [94] M. Kozák, M. Jureček, M. Stamp, and F. D. Troia, “Creating valid adversarial examples of malware,” Journal of Computer Virology and Hacking Techniques, vol. 20, no. 4, pp. 607–621, 2024. [95] A. Sang, Z. Wang, L. Yang, L. Zhou, J. Jia, and H. Yang, “Generating adversarial malware examples against multiple machine learning detectors,” IEEE Transactions on Industrial Informatics, 2025. [96] L. De Rose, G. Andresini, A. Appice, and D. Malerba, “Olivander: a counterfactual-based method to generate adversarial windows pe malware,” Data Mining and Knowledge Discovery, vol. 39, no. 5, p. 46, 2025. [97] Y. Fang, Y. Zeng, B. Li, L. Liu, and L. Zhang, “Deepdetectnet vs rlattacknet: An adversarial method to improve deep learning-based static malware detection model,” Plos one, vol. 15, no. 4, p. e0231626, 2020. [98] T. Quertier, B. Marais, S. Morucci, and B. Fournel, “Merlin–malware evasion with reinforcement learning,” arXiv preprint arXiv:2203.12980, 2022. [99] M. Ebrahimi, N. Zhang, J. Hu, M. T. Raza, and H. Chen, “Binary black-box evasion attacks against deep learning-based static malware detectors with adversarial byte-level language model,” arXiv preprint arXiv:2012.07994, 2020. [100] [Online]. Available: https://www.hybrid-analysis.com [101] H. Aghakhani, F. Gritti, F. Mecca, M. Lindorfer, S. Ortolani, D. Balzarotti, G. Vigna, and C. Kruegel, “When malware is packin’heat; limits of machine learning classifiers based on static analysis features,” in Network and Distributed System Security Symposium. Internet Society, 2020.

[69] O. Suciu, S. E. Coull, and J. Johns, “Exploring adversarial examples in malware detection,” in 2019 IEEE Security and Privacy Workshops (SPW). IEEE, 2019, pp. 8–14. [70] F. Croce and M. Hein, “Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks,” in Proceedings of the 37th International Conference on Machine Learning, ser. ICML’20. JMLR.org, 2020. [71] M. Nasr, Y. Fratantonio, L. Invernizzi, A. Albertini, L. Farah, A. PetitBianco, A. Terzis, K. Thomas, E. Bursztein, and N. Carlini, “Evaluating the robustness of a production malware detection system to transferable adversarial attacks,” arXiv preprint arXiv:2510.01676, 2025. [72] A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu, “Towards deep learning models resistant to adversarial attacks,” arXiv preprint arXiv:1706.06083, 2017. [73] F. Yu, Z. Qin, C. Liu, L. Zhao, Y. Wang, and X. Chen, “Interpreting and evaluating neural network robustness,” arXiv preprint arXiv:1905.04270, 2019. [74] H. Bostani and V. Moonsamy, “Evadedroid: A practical evasion attack on machine learning for black-box android malware detection,” Computers & Security, vol. 139, p. 103676, 2024. [75] P. He, Y. Xia, X. Zhang, and S. Ji, “Efficient query-based attack against ml-based android malware detection under zero knowledge setting,” in Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security, 2023, pp. 90–104. [76] H. Li, Z. Cheng, B. Wu, L. Yuan, C. Gao, W. Yuan, and X. Luo, “Blackbox adversarial example attack towards {FCG} based android malware detection under incomplete feature information,” in 32nd USENIX Security Symposium (USENIX Security 23), 2023, pp. 1181–1198. [77] R. Harang and E. M. Rudd, “Sorel-20m: A large scale benchmark dataset for malicious pe detection,” arXiv preprint arXiv:2012.07634, 2020. [78] J. Saxe and K. Berlin, “Deep neural network based malware detection using two dimensional binary program features,” in 2015 10th international conference on malicious and unwanted software (MALWARE). IEEE, 2015, pp. 11–20. [79] L. Demetrio and B. Biggio, “secml-malware: A python library for adversarial robustness evaluation of windows malware classifiers,” 2021. [80] endgameinc, “Github - endgameinc/gym-malware,” 2017. [Online]. Available: https://github.com/endgameinc/gym-malware [81] 2020. [Online]. Available: https://virusshare.com/ [82] G.-A. Iosif, “Dikedataset,” Aug 2022. [Online]. Available: https: //github.com/iosifache/DikeDataset [83] [Online]. Available: https://bazaar.abuse.ch/ [84] T. Chow, M. D’Onghia, L. Linhardt, Z. Kan, D. Arp, L. Cavallaro, and F. Pierazzi, “Beyond the TESSERACT: Trustworthy Dataset Curation for Sound Evaluations of Android Malware Classifiers,” in IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), 2026. [85] Cuckoo, “Cuckoo sandbox - automated malware analysis,” 2024. [Online]. Available: https://cuckoosandbox.org/ [86] R. Labaca-Castro, Machine Learning under Malware Attack. Springer Nature, 2023. [87] hex rays, “ida-pro,” 2024. [Online]. Available: https://hex-rays.com/ ida-pro [88] F. Croce, M. Andriushchenko, V. Sehwag, E. Debenedetti, N. Flammarion, M. Chiang, P. Mittal, and M. Hein, “Robustbench: a standardized adversarial robustness benchmark,” in Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, J. Vanschoren and S. Yeung, Eds., vol. 1, 2021. [89] F. Kreuk, A. Barak, S. Aviv-Reuven, M. Baruch, B. Pinkas, and J. Keshet, “Adversarial examples on discrete sequences for beating whole-binary malware detection,” arXiv preprint arXiv:1802.04528, pp. 490–510, 2018. [90] K. Tran, F. Di Troia, and M. Stamp, “Robustness of image-based malware analysis,” in Silicon valley cybersecurity conference. Springer, 2022, pp. 3–21. [91] H. Do Thi Thu, T. D. Phan, H. Le Anh, L. Nguyen Duy, K. Nghi Hoang, and V.-H. Pham, “A method of mutating windows malwares using reinforcement learning with functionality preservation,” in Proceedings of the 11th international symposium on information and communication technology, 2022, pp. 142–149.

16

A PPENDIX

when available. For example, EMBER-2017 originally provides 300k malware and 300k benign samples in the training split; from this set we recovered 51,352 malware and 43,983 benign binaries for our training experiments. Similarly, for EMBER-2018 we recovered 2,500 malware and 2,500 benign samples from the original test split of 100k / 100k. In contrast, the proprietary Commercial dataset was available to us as full binaries. SOREL-20M presents a particular challenge: the public release includes a large collection of disarmed (nonexecutable) malware samples and extracted features; disarmed binaries are unsuitable for executability-preserving problemspace attacks because they cannot be executed in a sandbox environment. Therefore, for SOREL we again used the hashbased search strategy to recover runnable binaries and constructed a balanced test subset (reported in Table VIII) that preserves the temporal and class distributions recommended by the dataset authors [77].

A. Overview of Evaluated Models and Attack Methods in Prior Work We provide a structured overview of existing literature on evasion attacks in malware detection, highlighting the attack settings and models evaluated in each study in Table V. B. Problem-Space binary transformations and obfuscation techniques We define the problem-space transformations and obfuscation techniques used by the evaluated attacks. All transformations are applied in the problem space and, when used by an attack, are constrained to preserve runtime executability and functionality (see Section V-D). 1) Gamma-Padding. Appends benign or innocuous bytes (an overlay) to the end of the PE file so that the file size and byte-level distribution change while the original code and headers remain intact. 2) Gamma-Shift. Shifts section contents and inserts benign padding inside or between sections, adjusting offsets so that payload bytes are relocated rather than simply appended to the file end. 3) Gamma-Sections. Adds one or more new named PE sections containing benign data and updates the section table and headers accordingly, thereby modifying the binary’s structural metadata (section names, offsets, and sizes). 4) MiniMal. The MiniMal attack uses 4 transformations from Table VI including MD, AS, SA, and padding. 5) MAB-malware. The MAB-malware attack uses 7 transformations from Table VI including OA, SR, SA, SP, RC, RD and BC. 6) Other attacks. The remaining attacks select from a predefined set of 10 problem-space transformations OA, IP, SR, SA, SP, RS, RD, UP, UPD, and BC from (Table VI) according to each attack’s search strategy and constraints.

D. Model performance on clean samples We report threshold-dependent performance metrics for all evaluated malware detectors on clean (non-adversarial) samples in IX, including false positive rate (FPR), recall (true positive rate), precision, F1-score, accuracy, and AUC. Following common practice in the malware detection literature, we first calibrate all models at a fixed FPR of 0.1%. However, for models where this operating point results in a recall (TPR) below 70%, we relax the constraint and report results at an FPR of 1.0% to ensure a minimally functional detection regime. Overall, most models achieve high AUC values (>0.95), indicating strong ranking ability, while performance differences are primarily reflected in recall–precision tradeoffs induced by threshold selection. In particular, EMBERbased models and MalConv variants achieve the highest F1scores and recall at low FPR settings, whereas the commercial model exhibits comparatively lower accuracy despite strong precision, suggesting a more conservative detection strategy. E. Attack results (RQ1)

C. Datasets

Table X compares ASR computed on the raw set of generated evasive samples (“w/o exec.”) with ASR computed after we filter out non-executable samples using our sandboxbased checks (“w. exec.”). For each model–attack pair the “w/o exec.” row reports the ASR obtained by counting every generated sample that evaded the detector, whereas the “w. exec.” row reports the ASR after retaining only those evasive samples that also passed our executability tests. Two points are worth emphasizing. First, enforcing executability substantially reduces the measured effectiveness of some attacks. For example, the Gamma-section attack suffers a dramatic drop in ASR once non-executable samples are removed, indicating that many of its reported evasions produce non-operational binaries. In contrast, the other Gamma variants (shift and padding) exhibit only minor differences between ASR before and after executability checks. The reason is structural: Gamma-section modifies the PE file’s section table and other header fields, inserting new

Table VII summarizes the datasets used in our experiments, reporting the feature type, collection period, dataset size, train-test splits, and data availability, thereby highlighting differences in scale, temporal coverage, and representation. Additionally, Table VIII provides, for each dataset and split, the number of executable samples we were able to retrieve for our experiments alongside the original published counts (“collected / original”). Acquiring executable samples is a non-trivial step for problem-space evaluations because several public datasets only publish extracted features, images, or disarmed binaries rather than runnable PE files. Hence to create the NTS dataset and to obtain runnable binaries we used a hash-based search strategy: for each sample in the feature-only datasets (e.g., EMBER v1, v2, and v3 and SOREL-20M) we searched public malware repositories and archival services (e.g., Hybrid Analysis [100], VirusShare [81], [101]), and downloaded matching binaries

17

TABLE V: Overview of adversarial malware attacks and the models or platforms evaluated in prior work, including both academic and commercial detectors Paper [90]

Attacks Salting

CNN/DNN/MLP ✓

[26]

GAPGAN, Opt., AdvSeq, MalGAN

[27]

MalFox: Obfusmal, Stealmal, Hollowmal

[91]

RL-based method

[92]

FGSM, method

[59]

Binary diversification

[46]

gym-malware

[93]

Gym-plus, gym-malware

✓

[48]

RL-based (BFA, DDQN, ACER, AMG-VAC)

✓

[25]

AIMED-RL (DDQN, DuDDQN, ACER, DiDDQN)

[54]

ARMED

[2]

RL-based

✓

✓

[56]

AIMED, ARMED

✓

✓

[55]

Append strings from goodware & Packing

✓

[57]

MDEA

✓

[1]

GAMMA attacks

✓

[62]

binary-search based

✓

[69]

FGM Append, Slack FGM

✓

[94]

AMG-PPO, AMG-random, MAB-Malware

✓2

✓

[63]

Extend, Full DOS, Shift, FGSM padding + slack, GAMMA

✓

✓

[64]

Partial DOS, Full DOS, GAMMA, Gymmalware

✓

✓

[49]

DQEAF

[51]

MalInfo, MalFox

[58]

AMG-PDG

[7]

MalPatch, Random Patch, Benign Patch and Transfer Patch, GAME-UP and BASGAN

[6]

MiniMal, GAMMA and MAB-malware

[95]

GanGenetic, MalGAN, Mab-malware, and GAMMA

✓

[96]

OLIVANDER, AMG and GAMMA

✓

[24]

RL-based

[17]

MAB-Malware, SecML-Malware and GymMalware

[97]

RLAttackNet

[98]

REINFORCE, DQN

[99]

MalRNN

Random,

MalConv

LGBM

✓

✓

VirusTotal

AvastNet

FireEye

Commercial

✓ ✓ ✓

and

Random Forest

✓

Experience-based

✓

✓ ✓ ✓

✓

✓

✓

✓ ✓

✓

✓ ✓

✓

✓ ✓ ✓ ✓

✓

✓

✓

✓✓3 ✓

✓ ✓

✓

✓

✓

✓ ✓

✓

✓

✓

✓

✓

✓ ✓

✓✓

named sections and altering offsets that determine in-memory layouts, imports, and resources. Even small inconsistencies can cause parsers or the Windows loader to fail. By comparison, Gamma-padding appends data at the end of the file without affecting loader-visible metadata, and Gamma-shift makes minor in-file relocations that typically preserve functionality and executability. This structural fragility explains why many Gamma-section binaries fail executability checks, reducing the measured ASR.

F. Attack Effectiveness on the NTS dataset (RQ1) G. Coverage of MAB-malware (RQ2) Table XII presents the coverage of other attacks by samples that successfully evade detection through the MABmalware attack. Although MAB-malware shows high overlap with Gamma-padding (99.57%), MiniMal (99.34%), and Gamma-shift (98.44%), its coverage notably decreases for Gamma-section (93.37%), Game-Up (96.79%), and ARMED (96.93.45%). This suggests that while MAB-malware is highly representative of attacks from the Gamma family, it misses

18

TABLE VI: Problem-space binary transformations and obfuscation techniques. Technique - Abbreviation

Description

Overlay Append - OA

Appends benign contents to the end of a binary (overlay). Typically preserves functionality and alters byte-based features. Writes bytes into unused space within an existing PE section; low risk but can be fragile if offsets are miscomputed. Adds a new named PE section containing benign content; modifies PE headers and section tables and can break binaries if not done correctly. Renames sections to common benign names to evade simple heuristics. Zeroes out or removes the certificate block (affects signing metadata). Strips debug information (reduces static cues about build environment). Strips or invalidates code signatures (alters provenance signals). Adds unused or benign API imports to the import table to change import-based features. Zeroes the optional header checksum, altering PE integrity flags (may affect some loaders). Rewrites instruction sequences to semantically equivalent variants (true code obfuscation). Packs the binary with UPX (packing is a form of obfuscation; requires an unpacking stub at runtime). Removes UPX packing (not an obfuscation, but useful for analysis/unpacking). Modifies or injects bytes in the DOS header padding between the “MZ” signature and the PE header without affecting execution. Appends bytes into section slack space (alignment padding between PE sections) without affecting program execution.

Section Append - SP Section Add - SA Section Rename - SR Remove Certificate - RC Remove Debug - RD Remove Signature - RS Imports Append - IP Break Checksum - BC Code Randomization - CR UPX Compression/Packing - UP UPX Decompression/Unpacking - UPD Modify DOS Header - MD Slack Append - AS

TABLE VII: Summary of datasets used in our experiments. Dataset

Features

Collected Year

Size

Train-Test Split

Availability

SOREL-20m

EMBER-v2 features

Jan 2017 – Apr 2019

9,919,251 malware, 9,470,626 benign

EMBER 2017

EMBER-v1 features

in or before 2017

Extracted features and disarmed malware binaries Extracted features

EMBER 2018

EMBER-v2 features

in or before 2018

Train: 80% and Test: 20%

Extracted features

EMBER 2024

EMBER-v3 features

Sep, 2023 – Sep, 2024

400K malware, 400K benign, 300K unlabeled 400K malware, 400K benign, 200K unlabeled 1,616,000 malware, 1,616,000 benign, and 6,291 challenge

Training: 65.49%, Validation: 12.87%, Test: 21.64% Train: 81.82% and Test: 18.18%

Training: 81.3% and Test: 18.8%

Extracted features

TABLE VIII: Collected vs Original Samples across datasets. Dataset EMBER 2017 EMBER 2017 EMBER 2018 SOREL-20m Commercial Commercial EMBER 2024

Split Train Test Test Test Train Test Test

Malware 51,352 / 300,000 12,500 / 100,000 2,500 / 100,000 12,501 / 1,360,622 60,397 / 60,397 1,670/ 1,670 17,345 / 270,000

H. Performance of Hardened Models on clean samples (RQ4)

Benign 43,983 / 300,000 12,500 / 100,000 2,500 / 100,000 22,282 / 2,834,441 53,084 / 53,084 1,055 / 1,055 12,441 / 270,000

In XIII, we report the performance of hardened models under different defense mechanisms at a fixed false-positive rate (FPR) operating point. Following the same evaluation practice as the original models, most models are calibrated at an FPR of 0.1%, while for the commercial model we use a relaxed operating point of 1.0% due to significantly reduced TPR under stricter constraints. We evaluate recall, precision, F1score, accuracy, and AUC across different defense strategies, including Gamma-padding, Gamma-shift, MAB-malware, and AIMED, to assess their impact on classification performance.

TABLE IX: Comparison of clean malware detectors with respect to threshold-dependent metrics: recall, precision, F1score, accuracy, and AUC. Model

Threshold

FPR

Recall (TPR)

Precision

F1-score

Accuracy

AUC

MalConv MalConv2 EMBER-17 EMBER-18 EMBER-24 FFNN-SOREL LGBM-SOREL Commercial model

0.9940 0.997622 0.8710 0.9996 0.9557 0.9352 0.9492 0.9741

0.1% 1.0% 0.1% 0.1% 0.1% 0.1% 1.0% 1.0%

0.8166 0.70 0.9300 0.8681 0.8850 0.8450 0.7457 0.7214

0.9988 0.9868 0.9989 0.9989 0.9989 0.9867 0.9678 0.9918

0.8985 0.8190 0.9632 0.9289 0.9385 0.9100 0.8414 0.8352

0.9078 0.9322 0.9645 0.9335 0.9420 0.9397 0.9006 0.8256

0.9959 0.9775 0.9991 0.9964 0.9982 0.9884 0.9491 0.9831

I. Performance of Hardened Model on the NTS dataset (RQ4)

many samples evaded by other attack types. These results highlight that MAB-malware does not fully span the evasive behavior space.

19

TABLE X: Mean ASR (%) with and without executability evaluation across models. Bold and underlined values indicate the highest and second-highest ASR, respectively, considering only executability-preserving adversarial examples. Model

MAB-malware

Gamma-shift

Gamma-padding

Gamma-section

MiniMal

w/o exec.

w. exec.

w/o exec.

w. exec.

w/o exec.

w. exec.

w/o exec.

w. exec.

w/o exec.

w. exec.

FFNN-SOREL LGBM-SOREL EMBER-17 EMBER-18 EMBER-24 MalConv MalConv2 Commercial

85.63 95.30 89.03 98.43 83.40 98.93 98.80 100.00

85.63 95.30 89.03 98.33 83.40 98.83 98.70 99.90

15.97 70.50 48.83 91.93 12.73 100.00 88.53 97.80

15.97 70.50 48.83 91.93 12.73 100.00 88.53 97.80

14.17 70.40 48.87 91.10 7.33 98.63 82.07 91.23

14.17 70.40 48.87 91.00 7.33 98.53 81.97 91.13

41.17 26.93 72.23 96.70 0 99.83 90.60 92.70

3.27 3.93 5.43 4.60 0 4.53 2.20 3.20

65.20 81.50 64.80 77.40 30.60 98.30 11.00 97.36

65.20 81.50 64.80 77.40 30.60 98.20 11.00 97.23

Avg

93.69

93.64

65.79

65.79

62.98

62.93

65.02

3.40

65.77

65.74

TABLE XI: ASR% across models on the NTS dataset. For Gamma and MAB attacks, results are shown without (w/o) and with (w.) executability evaluation. Bold indicates highest ASR per row; underline indicates second-highest. MABmalware

Model

Gammashift

Gammapadding

Gammasection

MiniMal

AIMED

AIMEDRL

ARMED

Gameup

w/o

w.

w/o

w.

w/o

w.

w/o

w.

w/o

w.

FFNN-SOREL LGBM-SOREL EMBER-17 EMBER-18 EMBER-24 MalConv MalConv2 Commercial

57.3 99.0 84.8 99.8 9.6 65.2 94.9 99.7

55.3 87.7 84.5 99.7 9.6 65.1 94.8 99.1

15.1 34.7 54.0 97.9 8.1 100.00 82.3 94.6

14.8 33.5 53.7 97.8 8.1 99.6 82.1 94.0

14.2 35.0 52.2 97.9 5.6 60.3 78.2 93.6

13.9 33.4 51.9 97.8 5.6 60.2 78.0 93.0

36.0 25.4 68.6 97.9 0 94.7 92.1 94.6

3.5 3.0 21.8 3.7 0 21.6 4.1 28.6

20.3 54.4 70.5 98.0 4.6 77.6 12.3 98.5

18.0 52.1 70.4 98.0 4.5 77.4 12.3 97.6

6.3 14.2 27.2 30.9 1.4 29.2 58.9 33.2

4.0 12.8 21.9 33.7 0.4 12.8 47.6 25.1

0.8 4.9 20.9 8.8 0.2 6.5 15.5 4.6

1.4 9.1 8.7 5.1 0 23.3 51.2 19.6

Avg

76.29

74.48

60.84

60.45

54.63

54.23

63.66

10.79

54.53

53.79

25.16

19.79

7.78

14.80

TABLE XII: Coverage of other attacks by the samples evaded by MAB-malware across models. Each attack column is divided into two sub-columns: Overlapping / Evasive and Coverage (%). Model

Gammashift

Gammapadding

Gammasection

MiniMal

AIMED

AIMED-RL

ARMED

GameUp

FFNN-SOREL LGBM-SOREL EMBER-2017 EMBER-2018 EMBER-24 MalConv MalConv2 Commercial

146 / 160 692 / 705 476 / 488 919 / 919 118 / 127 989 / 1000 882 / 885 978 / 978

91.25% 98.16% 97.54% 100% 92.91% 98.90% 99.66% 100%

142 / 142 700 / 704 481 / 488 910 / 910 72 / 73 985 / 986 819 / 820 911 / 911

100% 99.43% 98.57% 100% 98.63% 99.90% 99.88% 100%

26 / 30 40 / 40 43 / 54 46 / 46 0/0 44 / 46 21 / 22 31 / 32

86.67% 100% 79.63% 100% — 95.65% 95.45% 96.88%

646 / 652 805 / 815 642 / 648 774 / 774 304 / 306 980 / 982 110 / 110 973 / 973

99.08% 98.77% 99.07% 100% 99.35% 99.80% 100% 100%

32 / 50 162 / 162 8/9 563 / 563 17 / 17 224 / 224 351 / 351 157 / 157

64% 100% 88.89% 100% 100% 100% 100% 100%

5/5 68 / 76 2/2 305 / 305 3/3 140 / 143 220 / 222 77 / 77

100% 89.47% 100% 100% 100% 97.90% 99.10% 100%

6/6 48 / 52 1/2 195 / 195 1/4 131 / 136 190 / 197 83 / 83

100% 92.31% 50% 100% 25% 96.32% 96.45% 100%

0/0 33 / 44 3/3 260 / 260 0/0 117 / 125 254 / 258 95 / 95

— 75% 100% 100% — 93.60% 98.45% 100%

Total

5200 / 5262

98.82%

5020 / 5034

99.72%

251 / 270

92.96%

5234 / 5260

99.51%

1514 / 1533

98.76%

820 / 833

98.44%

655 / 675

97.04%

762 / 785

97.07%

20

TABLE XIII: Comparison of hardened model performance at a fixed FPR operating point, reporting threshold, classification metrics, and AUC Model

EMBER-17

MalConv

Commercial

Hardened Against

Threshold

FPR

Recall (TPR)

Precision

F1-score

Accuracy

AUC

No defense Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

0.8710 0.9451 0.9443 0.9681 0.9559 0.9494

0.1% 0.1% 0.1% 0.1% 0.1% 0.1%

0.9300 0.9879 0.9878 0.9878 0.9870 0.9876

0.9989 0.9990 0.9990 0.9990 0.9990 0.9990

0.9632 0.9879 0.9934 0.9934 0.9930 0.9932

0.9645 0.9934 0.9934 0.9934 0.9930 0.9933

0.9991 0.9998 0.9997 0.9998 0.9997 0.9998

No defense Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

0.9940 0.9883 0.9931 0.9874 0.9921 0.9912

0.1% 0.1% 0.1% 0.1% 0.1% 0.1%

0.8166 0.8209 0.8011 0.8500 0.8258 0.8158

0.9988 0.9988 0.9989 0.9989 0.9988 0.9988

0.8985 0.9012 0.8891 0.9184 0.9041 0.8981

0.9078 0.9100 0.9001 0.9245 0.9124 0.9074

0.9959 0.9951 0.9953 0.9951 0.9957 0.9954

No defense Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

0.9741 0.9636 0.9615 0.9840 0.9848 0.9771

1.0% 1.0% 1.0% 1.0% 1.0% 1.0%

0.7214 0.7310 0.7244 0.6339 0.6201 0.7232

0.9918 0.9919 0.9918 0.9906 0.9904 0.9918

0.8352 0.8417 0.8373 0.7731 0.7627 0.8363

0.8256 0.8315 0.8275 0.7720 0.7636 0.8267

0.9831 0.9826 0.9812 0.9801 0.9792 0.9834

21

TABLE XIV: ∆ASRabs and ∆ASRrel before and after adversarial hardening on the NTS dataset, illustrating differences in robustness and defense transfer across attacks. Model

Hardened Against Test Attack

ASRbefore ASRafter ∆ASRabs ∆ASRrel (%)

Gamma-padding

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

51.9 53.7 84.5 70.4 27.2

0.2 0.3 41.8 47.0 20.1

↓ -51.70 ↓ -53.40 ↓ -42.70 ↓ -23.40 ↓ -7.10

↓ -99.61% ↓ -99.44% ↓ -50.53% ↓ -33.24% ↓ -26.10%

Gamma-shift

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

51.9 53.7 84.5 70.4 27.2

0.6 0.2 49.5 39.7 20.0

↓ -51.30 ↓ -53.50 ↓ -35.00 ↓ -30.70 ↓ -7.20

↓ -98.84% ↓ -99.63% ↓ -41.42% ↓ -43.61% ↓ -26.47%

MAB-malware

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

51.9 53.7 84.5 70.4 27.2

0.6 1.8 26.2 10.8 9.8

↓ -51.30 ↓ -51.90 ↓ -58.30 ↓ -59.60 ↓ -17.40

↓ -98.84% ↓ -96.65% ↓ -68.99% ↓ -84.66% ↓ -63.97%

MiniMal

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

51.9 53.7 84.5 70.4 27.2

1.40 1.8 64.0 0 7.0

↓ -50.50 ↓ -51.9 ↓ -20.50 ↓ -70.40 ↓ 20.2

↓ -97.30% ↓ -96.64% ↓ -24.26% ↓ -100.00% ↓ 74.26%

AIMED

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

51.9 53.7 84.5 70.4 27.2

51.2 53.1 62.5 35.7 0

↓ -0.70 ↓ -0.60 ↓ -22.00 ↓ -34.70 ↓ -27.20

↓ -1.35% ↓ -1.12% ↓ -26.04% ↓ -49.29% ↓ -100.00%

Gamma-padding

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

60.2 99.6 65.1 90.8 29.2

43.3 99.4 66.6 67.9 30.4

↓ -16.90 ↓ -0.20 ↑ 1.50 ↓ -22.90 ↑ 1.20

↓ -28.07% ↓ -0.20% ↑ 2.30% ↓ -25.22% ↑ 4.11%

Gamma-shift

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

60.2 99.6 65.1 90.8 29.2

40.6 95.9 64.0 66.6 42

↓ -19.60 ↓ -3.70 ↓ -1.10 ↓ -24.20 ↑ 12.80

↓ -32.56% ↓ -3.71% ↓ -1.69% ↓ -26.65% ↑ 43.84%

MAB-malware

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

60.2 99.6 65.1 90.8 29.2

40.5 99.7 59.7 56.7 36.6

↓ -19.70 ↑ 0.10 ↓ -5.40 ↓ -34.10 ↑ 7.40

↓ -32.72% ↑ 0.10% ↓ -8.29% ↓ -37.56% ↑ 25.34%

MiniMal

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

60.2 99.6 65.1 90.8 29.2

51.96 99.55 60.33 51.3 9.5

↓ -8.24 ↓ -0.05 ↓ -4.77 ↓ -39.50 ↓ -19.70

↓ -13.69% ↓ -0.05% ↓ -7.33% ↓ -43.50% ↓ -67.47%

AIMED

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

60.2 99.6 65.1 90.8 29.2

59.6 99.7 68.2 64.9 12.7

↓ -0.60 ↑ 0.10 ↑ 3.10 ↓ -25.90 ↓ -16.50

↓ -1.00% ↑ 0.10% ↑ 4.76% ↓ -28.52% ↓ -56.51%

Gamma-padding

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

93 94 99.1 98.5 33.2

7.5 28.9 81.8 81.8 35.9

↓ -85.50 ↓ -65.10 ↓ -17.30 ↓ -16.70 ↑ 2.70

↓ -91.94% ↓ -69.26% ↓ -17.46% ↓ -16.95% ↑ 8.13%

Gamma-shift

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

93 94 99.1 98.5 33.2

21.1 5.9 83.8 83.8 36.5

↓ -71.90 ↓ -88.10 ↓ -15.30 ↓ -14.70 ↑ 3.30

↓ -77.31% ↓ -93.72% ↓ -15.44% ↓ -14.92% ↑ 9.94%

MAB-malware

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

93 94 99.1 98.5 33.2

81.6 90 99.6 72.0 38.5

↓ -11.40 ↓ -4.00 ↑ 0.50 ↓ -26.50 ↑ 5.30

↓ -12.26% ↓ -4.26% ↑ 0.50% ↓ -26.90% ↑ 15.96%

MiniMal

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

93 94 99.1 98.5 33.2

91.83 98.29 100 77.5 16.5

↓ -1.17 ↑ 4.29 ↑ 0.90 ↓ -21.00 ↓ -16.70

↓ -1.26% ↑ 4.56% ↑ 0.91% ↓ -21.32% ↓ -50.30%

AIMED

Gamma-padding Gamma-shift MAB-malware MiniMal AIMED

93 94 99.1 98.5 33.2

97.4 96.8 99.4 86.4 24.3

↑ 4.40 ↑ 2.80 ↑ 0.30 ↓ -12.10 ↓ -8.90

↑ 4.73% ↑ 2.98% ↑ 0.30% ↓ -12.28% ↓ -26.81%

EMBER-17

MalConv

Commercial

22

Record · ID 1108639 · SHA-256 0bcd30745336fdc2
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.