Conceptio › Archive › arXiv CS
arXiv CSopen access

Windows Malware Detector as a Compound AI System: Trade-Offs in Accuracy, Efficiency, and Adversarial Robustness

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

1

Windows Malware Detector as a Compound AI System: Trade-Offs in Accuracy, Efficiency, and Adversarial Robustness

arXiv:2609.08394v1 [cs.CR] 8 Sep 2026

Andrea Ponte, Luca Demetrio, Member, IEEE, Luca Oneto, Senior Member, IEEE, Battista Biggio, Fellow, IEEE, Fabio Roli, Fellow, IEEE

Abstract—Industrial Windows malware detectors are commonly described as Compound AI Systems composed of multiple heterogeneous components, including rule-based mechanisms as well as machine-learning-based static and dynamic analyses. However, due to industrial secrecy and limited public disclosure, the internal architectures of these systems can only be inferred, rendering systematic evaluations of detection accuracy, computational costs, and adversarial robustness largely infeasible. In contrast, academic research provides reproducible and transparent evaluation methodologies, but typically investigates individual detection components in isolation. To bridge the gap between academic research and industrial practice, and inspired by state-of-the-art industrial architectures for Windows malware detection, we propose a novel methodology that (i) explicitly balances the trade-off among detection performance, computational requirements, and robustness, and introduces (ii) system-level threat models that capture how attackers exploit different degrees of knowledge to evade the entire Compound AI System rather than isolated detectors. Experiments conducted on real-world data demonstrate that the Compound AI System training time can be reduced and responsiveness improved while incurring only a marginal loss in detection performance. Leveraging our threat modeling, we show that increasingly knowledgeable attackers craft more effective adversarial examples, revealing the system’s strengths and weaknesses, degrading its responsiveness, and exposing a direct trade-off between efficiency and robustness. Finally, we translate these trade-offs into take-home messages and deployment guidelines, helping practitioners to select the system that best matches their operational constraints. Index Terms—Windows Malware Detection, Compound AI System, Accuracy, Efficiency, Threat Models, Robustness.

I. I NTRODUCTION When browsing antivirus solutions to install or purchase, most products emphasize the adoption of multi-layered detection pipelines. These typically combine traditional detection techniques with advanced Machine Learning (ML) models [1]–[5], often collectively referred to as Compound AI Systems. Due to industrial secrecy and limited public disclosure, the internal structure of these Compound AI Systems can only be partially inferred. They are commonly assumed to operate as sequential pipelines that trade detection effectiveness for computational efficiency [1], [2], [4]. Such pipelines usually begin with a signature-based detector, which is deployed for its reliability in identifying well-known malware samples [6], Andrea Ponte, Luca Demetrio and Luca Oneto are with the Department of Informatics, Bioengineering, Robotics and Systems Engineering, University of Genova, 16126 Genova, Italy (e-mail: [email protected]). Battista Biggio is with the Department of Electrical and Electronic Engineering, University of Cagliari, 09123 Cagliari, Italy. Fabio Roli is with the Department of Informatics, Bioengineering, Robotics and Systems Engineering, University of Genova, 16126 Genova, Italy, and also with the Department of Electrical and Electronic Engineering, University of Cagliari, 09123 Cagliari, Italy.

[7]. This stage is typically followed by ML models trained on static analysis, i.e., inferring maliciousness from code and metadata without executing the program [8]–[11]. Finally, the last line of defense is provided by models trained on dynamic analysis, which infer maliciousness from runtime behavior observed when samples are executed in isolated virtual environments [12]–[14]. These dynamic analysis components may be deployed locally or executed in the cloud. However, the secrecy surrounding these pipelines renders systematic and transparent evaluations infeasible. While competitive benchmarking platforms exist1 and report metrics such as detection rates and false positive counts on closed sets of malicious and benign samples, they do not enable comparisons of Compound AI Systems along other critical dimensions, particularly overlooking computational requirements and robustness. As a result, it remains unclear whether - and to what extent - these systems can be evaded by carefully crafted and minimally modified programs, commonly referred to as adversarial EXEmples [15]. In contrast, academic research offers systematic and transparent evaluations, but solely focused on isolated components rather than fully integrated Compound AI Systems [16], [17]. At design time, most research concentrates on training individual ML models within narrowly defined settings, either exclusively for static [8]– [11], dynamic [12]–[14], or hybrid [18] analysis. Alternatively, some works stack pre-trained models to emulate sequential detection pipelines [19]. At deployment time, due to the lack of system-wide security evaluations and formalized threat models for Compound AI Systems, research continues to focus almost exclusively on isolated models [15], [17], [20]–[23]. In fact, at the time of writing, the only available approach to assess the robustness of Compound AI Systems for Windows malware detection consists of generating attacks against a single ML model and subsequently testing them against the real target system [17], [19], [24], [25]. Consequently, academia has largely overlooked rigorous analyses of integrated systems, contributing to a growing disconnect between academic research and industrial practice. To the best of the authors’ knowledge, there is currently no work systematically investigating the joint optimization of performance, computational requirements, and robustness while accounting for both development and deployment costs. Hence, in this paper, we aim to narrow the gap between industry and academia by proposing a novel methodology for optimizing and evaluating Compound AI Systems for Windows malware detection. Our approach explicitly highlights the inherent trade-offs among detection performance, computational requirements (both at 1 https://www.av-test.org/en/antivirus/business-windows-client/

2

design and deployment stages), and robustness that arise from integrating multiple detection technologies within a single pipeline (Sect. III-A). During the design stage, the primary computational costs stem from training the ML models for static and dynamic analysis. Reducing the amount of training data by exploiting the fact that a subset of samples is filtered out at earlier pipeline stages directly lowers these costs. However, using less training data may negatively impact performance, as reduced data availability often leads to lower model accuracy and may also hinder robustness. In this work, we identify when it is possible to reduce the training data while limiting the impact on performance, and when such reductions become detrimental. At the deployment stage, the pipeline components are naturally ordered by increasing computational cost: deterministic signature-based detectors, which incur negligible overhead (hundredths of a second); probabilistic static ML models, which require slightly higher computation times (fractions of a second); and probabilistic dynamic ML models, which are significantly more expensive, often requiring several seconds per sample. We show that optimizing the probabilistic thresholds used to decide whether to trigger more computationally expensive analyses allows one to jointly optimize performance and computational cost. With respect to robustness evaluation, we first formalize a range of threat models under varying assumptions about the attacker’s knowledge of the system (Sect. III-B). We begin with a zero-knowledge adversary, who has no information about the system, and progressively consider stronger attackers who are aware of its structure but possess detailed knowledge of only a subset of its components. We then restrict our analysis to threat models that are feasible in practice, given the current state of the art in attack libraries targeting individual components or entire Compound AI Systems. Specifically, we focus on manipulations that affect only parts of the pipeline (e.g., the static ML component and/or signature-based detectors), as no effective manipulations are currently available for the remaining components. We conduct a series of experiments on real-world data using OBELISK, our implementation of a Compound AI System for Windows malware detection (Sect. IV) inspired by recent work [19], [25], which is optimized and evaluated according to our proposed methodology (Sect. V). Our results show that the proposed filtering mechanism significantly reduces resource consumption while incurring only a marginal decrease in performance (Sect. V-B). Moreover, our system-level threat models reveal that betterinformed attackers can mount more effective evasion attacks, degrading system responsiveness and exposing the system’s strengths and weaknesses that remain hidden when detectors are evaluated in isolation (Sect. V-C). Also, we show that the same filtering mechanism used to reduce computational cost simultaneously widens the system’s attack surface and, by training on smaller datasets, further weakens robustness, exposing a direct trade-off also between computational savings and robustness (Sect. V-D). Finally, we distill our findings into a set of take-home messages and practical guidelines (Sect. V-E), to help practitioners select the configuration that best matches their operational constraints.

MZ

PE

DOS Header + Stub

PE Header

Optional Header

Section Table

A

B

C

D

.text

.rdata

.data

E

…..

Overlay

Sections

F

Fig. 1: Depiction of the Windows PE file format.

II. BACKGROUND AND R ELATED W ORK A Windows program is represented on disk as a file conforming to the Portable Executable (PE) format2 (Fig. 1). The main components of the PE format are summarized as follows: (A) the DOS header and stub, a valid DOS program retained for backward compatibility; (B) the PE header, which contains metadata describing the program; (C) the Optional Header, which stores information required for loading and executing the program; (D) the Section Table; and (E) the Sections, which contain the actual program content (e.g., the “.text” section contains compiled code). We refer to those Windows malware samples that are crafted to evade ML-based detectors [15] as adversarial EXEmples. These samples are generated through the application of practical manipulations, i.e., modifications to the structure of input programs that preserve their original functionality and do not cause errors at load time. Such manipulations exploit blind spots and redundancies in the PE file format, allowing attackers to insert, remove, or replace content without compromising the correctness of the executable. For example, attackers can manipulate and extend headers (gray area in A [15]), inject new sections [17], fill unused space between sections (gray areas in E [26]), or append arbitrary bytes at the end of the last section (gray area in F [20]). The content of these manipulations is typically determined by an optimization algorithm, which may operate in a black-box setting - assuming no knowledge of the detector’s internal structure - via query-based or transfer attacks, or in a white-box setting - assuming full knowledge of the detector via gradient-based optimization methods if the detector is differentiable [20], [21]. We are currently witnessing a growing gap between academic research and industrial practice. On the one hand, industry increasingly relies on complex detection solutions composed of multiple interacting components [1]– [5]. On the other hand, academic research has largely focused on the design and evaluation of individual models, typically based on static, dynamic, or hybrid analysis [8], [10], [12], [14], [18], [27]. To the best of our knowledge, only a limited number of works have explicitly investigated the properties of Compound AI Systems for Windows malware detection [19], [25]. While relevant, these contributions represent only an initial step toward a systematic study of the trade-offs among accuracy, efficiency, and adversarial robustness in AI-based malware detection. In contrast to this work, prior studies do not (i) explicitly address the trade-off between performance, computational requirements, and robustness or (ii) define and evaluate system-level threat models that capture how attackers exploit different degrees of knowledge to evade a Compound AI System. 2 https://learn.microsoft.com/en-us/windows/win32/debug/pe-format

3

Compound AI System Signatures

Features

Known PE

Training

Static Analysis

𝑷⋚𝜹

Attacks Implementation Available Attacks Implementation Unavailable

Sandboxing/Emulation

Reports

Features

Training

Dynamic Analysis

𝑷⋚𝜹

Fig. 2: Considered Compound AI System for Windows malware detection based on a sequential architecture. Each analysis stage is characterized by its intermediate pre-processing and deployment steps, and the criteria used to forward samples through the pipeline. All components are annotated with the presence or absence of known attacks reported in the literature, as well as with the availability of open-source, stable tools that implement such attacks for assessing the robustness of the ML models.

III. C OMPOUND AI S YSTEMS FOR W INDOWS M ALWARE D ETECTION To bridge the gap between academic research and real-world deployments, it is necessary to move beyond the prevailing assumption of isolated ML components acting as standalone defense mechanisms. Instead, we advocate a Compound AI System paradigm, in which multiple ML components cooperate and are complemented by conventional software responsible for automation and pre-processing, with the explicit goal of trading off detection performance and computational requirements (Sect. III-A). This shift must be accompanied by comprehensive threat modeling that considers varying assumptions about the attacker’s knowledge of the detection pipeline, explicitly accounts for interdependencies among components, and focuses on attacks that are practically feasible given the current state of the art in available tools (Sect. III-B). A. Compound AI System Paradigm Based on the state of the art [1], [2], [4], real-world implementations of Compound AI Systems for Windows malware detection typically adopt a three-layer architecture: (i) signature-based pattern matching; (ii) static analysis of sample metadata; and (iii) behavioral profiling via dynamic analysis, as illustrated in Fig. 2. This ordering reflects a deliberate trade-off between the use of domain knowledge - encoded as fast-to-evaluate, rigid logical constraints (i.e., signatures) and purely data-driven methods that rely on softer decision boundaries learned from computationally expensive features. The layered design enables early filtering of samples that can be classified with high confidence at minimal computational cost, reserving more complex and resource-intensive analyses for cases in which prior knowledge is insufficient. In the absence of computational constraints, such pipelines would naturally collapse into an end-to-end approach. However, a layered architecture introduces dependencies across levels, whereby each component acts as a filter for subsequent ones by progressively reducing and refining the data passed downstream. This mechanism can be advantageous when deploying relatively “shallow” models, which often benefit from smaller but highly representative datasets, especially when features

encode domain expertise. By contrast, aggressive filtering may be detrimental for deep neural networks, which are designed to learn representations directly from data and therefore require large training sets to reduce reliance on handcrafted features. Pattern-matching with Signatures. This level identifies wellknown malicious and benign samples via pattern matching against previously-extracted knowledge, often deployed as a primary (if not the sole) line of defense [4], [6], [28], as shown in Fig. 2. Such mechanisms are explicitly documented in commercial products, including Apple’s native macOS antivirus [6], Bitdefender [29], Microsoft Defender [4], and CrowdStrike Agent [30]. They are typically implemented as blocklists, i.e., collections of byte patterns derived from reverse engineering malware samples; a match to any signature results in an immediate malware verdict, unlike other domains where rules often contribute to a soft anomaly score [31]. Blocklists are commonly complemented by allowlists, i.e., sets of hash values identifying known benign samples, which are essential to mitigate false alarms that may cause financial and reputational damage3 . Signature-based detectors are simple to deploy, as updates only require modifying byte patterns, but they do not natively support tuning the trade-off between detection rate and false alarms. Such tuning can be achieved either by (i) training a classifier that incorporates signature matches as features [31], [32], at the cost of losing negligible inference overhead and effectively shifting signatures to the ML-based level, or (ii) removing signatures that produce excessive false positives. During development, signatures can constrain the data distribution passed to downstream components, reducing their training requirements [25]. After deployment, they immediately block known samples and forward only those lacking prior knowledge, typically including previously unseen samples produced via obfuscation techniques that preserve functionality while altering the byte-level representation4 . Static Analysis on Metadata. This level leverages program representations to assess whether a sample is malicious using a dedicated ML model, which requires an explicit feature 3 https://corelight.com/resources/glossary/false-positives-cybersecurity 4 https://www.darkreading.com/threat-intelligence/ only-half-of-malware-caught-by-signature-av

4

extraction phase (see Fig. 2). Features may be manually engineered using domain knowledge and aggregation techniques [33], or learned end-to-end [8], [10] by exploiting structural dependencies in the data. This enables highly effective detectors that extract knowledge directly from data, reducing reliance on time-consuming - though precise - manual reverse engineering. Once features are defined, developers must (i) train the model on labeled benign and malicious samples and (ii) deploy it by specifying how samples are forwarded to subsequent pipeline stages. Training is computationally expensive, whereas inference is relatively cheap; both can be automated and accelerated via increased computational resources, including GPUs at inference time. Unlike signaturebased detectors, which fundamentally rely on human expertise and manual updates, ML-based approaches scale primarily with available hardware. Training for static analysis can be performed either (i) in isolation, ignoring upstream filtering and leveraging a larger but potentially less representative dataset, or (ii) jointly with the signature-based level, reducing data volume and computational cost while better matching the distribution observed after upstream filtering [25]. At deployment time, the designer must decide how to propagate samples to the dynamic analysis stage, which is typically resource-intensive. Unlike signatures, which produce binary decisions, this propagation is governed by the probabilistic output of the model. Decision thresholds determine whether a prediction is confident enough to terminate the pipeline with a benign or malicious label, or whether the sample should be forwarded for further analysis, trading uncertainty for increased computational cost. The same thresholding mechanism can be used during training by defining two cutoffs: samples below a lower threshold are labeled benign, those above an upper threshold malicious, and those in between are forwarded as training data to the next level. While this strategy introduces a potential attack surface - since adversaries may attempt to keep samples below the lower threshold - it represents an explicit trade-off against system-wide computational constraints that must be considered by designers. Dynamic Analysis of Behavior. This level analyzes the runtime behavior of input samples, which is necessary because static representations cannot fully capture execution dynamics (see Fig. 2). In practice, malware may download additional payloads, inject code into other processes, unpack components in memory, or perform actions observable only at runtime. Behavioral analysis is therefore performed by executing samples in a sandbox, i.e., an instrumented environment that allows programs to manifest their behavior. While sandboxing can, in principle, provide the most comprehensive behavioral view, observing all relevant actions within a limited time budget is often infeasible. This is due to (i) the need for multiple executions to explore different control-flow paths and (ii) deliberate evasion techniques whereby malware detects sandbox environments and suppresses or delays malicious behavior [34]. Additionally, sandboxes enforce strict execution time limits to contain computational costs [35]. Collected runtime traces are serialized into structured reports (e.g., JSON) describing API calls, network activity, and file-system operations. Unlike static features, behavioral reports are difficult to encode with fixed,

handcrafted feature sets and are therefore commonly processed using Natural Language Processing (NLP) techniques [12], [14]. The resulting pipeline typically includes: (i) cleaning, to remove noisy or sample-specific fields (e.g., hashes); (ii) feature extraction, either learned directly from data or obtained via pre-trained models; and (iii) classifier training. Both feature learning and classification are computationally expensive, as they rely on complex neural architectures whose effectiveness scales with dataset size. These costs can be mitigated by using upstream levels as filters, reducing the number of samples processed at this stage and saving both training time and inference latency, with only a limited impact on detection performance (Sect. V). However, such filtering introduces a tradeoff with representation expressiveness. Behavioral modeling is inherently challenging because (i) individual executions generate large volumes of low-level events, often dominated by noise; (ii) benign and malicious programs may exhibit similar behaviors for different purposes; (iii) malware frequently employs obfuscation; and (iv) aggressive pre-processing may map distinct behaviors to similar feature representations. Addressing these issues typically requires very large training datasets, increasing the overall computational burden. After training, the model output is calibrated via decision thresholds that balance detection performance and false alarms. This step is particularly critical, as dynamic analysis constitutes the final stage of the Compound AI System, where decisions are definitive. Although dynamic analysis often underperforms static analysis when considered in isolation [36], it acts as a last line of defense against samples that evade earlier detectors. Finally, independently of automated decisions, a subset of samples is usually analyzed manually by human experts to identify novel threats, variants of known malware, or benign programs not adequately captured by existing models (outside the scope of this work). Error Management. During the development and deployment of Compound AI Systems, practitioners must account for the possibility that any upstream level may fail. For instance, a sandbox may crash on specific API calls, or a static preprocessing module may be unable to extract features due to malformed inputs. Since the system must always return a decision, developers must define how such failures are labeled, explicitly accepting the corresponding trade-offs between increased false alarms and reduced detection rates [19]. B. Evaluating Robustness of Compound AI Systems. Isolated evaluations fail to capture the robustness of a Compound AI System as a whole. They may underestimate security, because upstream components can pre-process inputs and attenuate adversarial perturbations before they reach downstream detectors [19], [25]. Conversely, they may overestimate security, as adversaries can focus on the subset of samples that consistently evade the entire pipeline rather than individual components. Robustness evaluations must hence explicitly account for system-level complexity by considering multiple threat models. A natural starting point is a zero-knowledge (black-box) evaluation [37], in which attackers interact with the system only through its outputs, without access to internal

5

TM1 – Black-box Transfer Signatures

Static Analysis

Dynamic Analysis

TM2 – Knowledge of Signatures Signatures

Static Analysis

Dynamic Analysis

TM3 – Knowledge of Static Model Signatures

Static Analysis

Dynamic Analysis

TM4 – Knowledge of Model and Signatures Signatures

Static Analysis

Dynamic Analysis

Fig. 3: The four threat models we consider, from the black-box one (TM1), i.e., the attacker is unaware of the components of the system, to the acquisition of partial knowledge (TM4), i.e., the attacker knows how the first levels are implemented.

details. This setting includes: (i) transfer-based attacks, which assess robustness using adversarial samples crafted against a surrogate system with similar characteristics [24]; (ii) querybased attacks, which iteratively optimize evasive samples by placing the target system in a feedback loop [17], assuming queries are not severely constrained; and (iii) obfuscation and packing techniques that alter representations while preserving functionality. The latter are widely used in practice but are treated as a separate problem and are not addressed in this work. Due to the high computational cost of querying the full pipeline, where inference on a single sample may take several seconds (as we will show in Sect. V-B), we focus on transfer-based evaluations. The effectiveness of black-box attacks can be further increased by acquiring partial knowledge of the target via reverse engineering or information leakage, leading to a limited-knowledge (gray-box) threat model that has proven effective even against commercial systems5 . At the other extreme, a perfect-knowledge (white-box) attacker has full access to the system, a setting that, while potentially unrealistic, provides insight into worst-case vulnerabilities. Accordingly, we define a spectrum of intermediate threat models - characterized by attacker goals, knowledge, and capabilities - in which the attacker’s knowledge increases, while the goal (evasion) and capabilities (functionality-preserving manipulations at test time) remain consistent across models. Threat Model 1: Black-box Transfer (TM1). As a practical and widely adopted zero-knowledge setting, we consider an attacker with no access to the internals of the target Compound AI System (first row in Fig. 3, unknown components in dark gray). The attacker builds or collects a surrogate model with behavior similar to the target and crafts adversarial samples against it, which are then transferred to the real system, following established evaluation practices [17]. Threat Model 2: Knowledge of Signatures (TM2). Here, the attacker is aware of the presence of signature-based detection and can extract the deployed signatures, but ignores the existence of additional detection layers (second row in Fig. 3). This assumption is realistic, as signatures are often deployed 5 https://www.kb.cert.org/vuls/id/489481/

locally [4]–[7]. The attacker can therefore optimize attacks to avoid artifacts detectable by signatures [19], [25]. Threat Model 3: Knowledge of Static Model (TM3). We assume the attacker can extract the static ML model used in the pipeline, but is unaware of any preceding signature-based filtering (third row in Fig. 3). This scenario is plausible given that vendors deploy ML models on endpoints [4]. The attacker can perform worst-case (white-box) evaluations [15] and test the resulting adversarial EXEmples against the full system. We assume the attacker steals only the model parameters, not the decision thresholds, and therefore selects a reasonable operating point (1% FPR) on an external dataset. Threat Model 4: Knowledge of Signatures and Static Model (TM4). In this setting, the attacker extracts both the static model and the signatures, enabling attacks jointly optimized to evade both components (fourth row in Fig. 3). Again, the attacker selects an operating point corresponding to a 1% FPR on an external dataset, without knowledge of the deployed thresholds. This threat model is particularly relevant when dynamic analysis is performed in the cloud and static analysis is executed locally [4]. By evading the initial stages, the attacker may gain sufficient time to compromise the system before dynamic analysis intervenes. Other Threat Models. In principle, enumerating all possible threat models would require considering every combination of assets an attacker might compromise, leading to an exponential number of cases. In practice, this space is constrained by the availability of techniques and tools for evaluating the security of Compound AI Systems. While a mature literature and stable open-source tools exist for adversarial attacks against static detectors [22], [38], we are not aware of equally stable implementations targeting dynamic analysis (marked with the white symbol in Fig. 2). Proposed approaches either (i) lack public implementations [39], (ii) rely on deprecated and unmaintained platforms such as the Cuckoo sandbox, discontinued in 2021 [23], [40], or (iii) operate only in feature space [41], [42]. Therefore, also considering the observed superiority of static over dynamic analysis [36], we restrict our study to the four threat models in Fig. 3, deferring the analysis of attacks against dynamic analysis components to future work. IV. I MPLEMENTATION D ETAILS In this section, we present OBELISK, our implementation of a Compound AI System for malware detection, along with the attacks that follow the threat models described in Sect. III. The implementation of OBELISK and the attacks are publicly available at https:// github.com/ Andrea-Ponte/ obelisk. We use this framework to quantify the trade-offs among performance, computational requirements, and robustness. A. OBELISK Implementation OBELISK is implemented following the architecture shown in Fig. 2, which is composed of three levels (see Sect. III). Signature Matching with YARA. The first level is implemented using YARA rules6 , a pattern-matching tool for 6 https://github.com/VirusTotal/yara

6

signature-based detection. Our system combines a blocklist, consisting of malicious signatures collected from public sources, with an allowlist, containing hashes of well-known programs distributed with Windows installations. Although some rules achieve high detection rates, they also produce an unacceptable number of false alarms. We hence discard all those that trigger at least one false alarm on the training data. This choice reflects the fact that many rules were developed in isolation and not validated on large and heterogeneous datasets, making false alarms on benign software likely. Static Analysis with EMBER-GBDT. The second level is implemented using a Gradient Boosting Decision Tree (GBDT) [43], [44] from the XGBoost library [45], trained on features extracted with EMBER [33]. Although this feature set was introduced in 2018, it remains one of the most adopted benchmarks in the literature [18], [19], [25], [46]–[48]. Dynamic Analysis with Nebula. The third level is implemented using Nebula [12], a transformer-based model [49] trained on reports generated with Speakeasy7 , a Pythonbased Windows kernel emulator. We select Nebula because it provides an end-to-end detection approach that outperforms previous methods [14], [50] on this type of data. Error Management. In this evaluation, we exclude samples that fail during preprocessing instead of assigning them a default label. This ensures that our assessment of Compound AI System trade-offs is not skewed by inflated false positive rates (from labeling errors as malicious) or degraded detection rates (labeling errors as benign), effect documented in [19]. Training and Inference. We now describe how we implement the progressive reduction of the training and test data introduced in Sect. III-A. Following the methodology outlined in Sect. III, the static model at the second level is trained under two settings: first, on the full dataset, without considering the signature-based filtering level (GBDT-ALL); and second, on the subset of samples that are not identified by either the blocklist or the allowlist (GBDT-YARA). We then use the probability scores produced by the resulting second-level static model to determine which samples should be used to train the computationally more expensive dynamic stage. Similarly, the dynamic model is also trained under two settings: on the full dataset without considering the first and second levels (NEBULA-ALL), and on the subset of samples that remain after applying the previous filtering levels (NEBULA-GBDTYARA). In the latter case, we apply a two-step filtering procedure. First, we retain only the samples that are not identified by either the blocklist or the allowlist. Then, we introduce a parameter δ that defines two decision thresholds on the static model’s estimated probability that a sample is malware. Training samples with probability p ≤ δ are labeled as benign, while samples with probability p ≥ 1 − δ are labeled as malicious. Consequently, only the samples whose probability lies in the interval (δ, 1−δ), which we interpret as the uncertainty region of the static model, are used to train the dynamic model. Such a decision contrasts with the standard approach, where detectors are only using a single threshold to determine the maliciousness of a sample (i.e., a sample is 7 https://github.com/mandiant/speakeasy

considered malicious when the computed probability exceeds a certain threshold). In fact, with only one detection threshold, it is not possible to reduce the computational requirements of a Compound AI System, as all benign programs would be sent to the costly dynamic analysis level [19]. At inference time, the blocklist and allowlist are first used to label known samples as malicious or benign, respectively. Samples that do not match any signature are then passed to the static model, which assigns each of them a probability p of being malware. Those with p ≤ δ are classified as benign, whereas those with p ≥ 1−δ are classified as malicious. The remaining samples, i.e., those falling in the static model’s uncertainty region, are forwarded to the dynamic stage for the final decision. B. Attacks Implementation To evaluate the adversarial robustness of OBELISK, we rely on two established techniques: GAMMA section injection [17] (GAMMA), which inserts non-executable content extracted from benign programs, and padding attacks [20] (PAD), which append bytes to the overlay. Both techniques are used to implement attacks for each threat model introduced in Sect. III-B and illustrated in Fig. 3. Attacks for TM1 (A1). We construct a surrogate for OBELISK following a standard approach: we leverage a pretrained open-source model8 , which in our case is a GBDT classifier trained on the EMBER dataset [33]. We then generate two sets of adversarial examples against this model using the GAMMA and PAD attacks. Attacks for TM2 (A2). In TM2, we have exact knowledge of the first layer of the deployed Compound AI System, namely the signatures. A2 is therefore constructed by exploiting this information from A1. Since YARA rules are triggered by specific artifacts, we enhance GAMMA by explicitly avoiding the injection of artifacts that could activate such signatures, while still inserting carefully crafted content to evade detection; we denote this variant as GAMMA-YARA. We then employ GAMMA-YARA and PAD to attack a surrogate Compound AI System consisting of the signatures followed by the model adopted in A18 . We thus generate two sets of adversarial examples against this model, one using PAD and the other using GAMMA-YARA. Attacks for TM3 (A3). A3 is identical to A1, except that, rather than relying on the surrogate used in A18 , we use the same GBDT model deployed in OBELISK. Note that OBELISK may employ different GBDT models depending on the training set considered (see Sect. IV-A). Accordingly, we generate different sets of adversarial examples against this model, depending on both the attack employed (PAD or GAMMA) and the GBDT variant considered: either the model trained on the full dataset (GBDT-ALL) or the one trained on the data filtered according to the YARA rules (GBDT-YARA). Attacks for TM4 (A4). A4 is identical to A2, except that, as A3 relates to A1, instead of relying on a surrogate model, we use the same GBDT model deployed in OBELISK. Also in this case, OBELISK may employ different GBDT models 8 https://github.com/endgameinc/malware evasion competition/tree/master/ models/ember

7

depending on the training set considered (see Sect. IV-A). We therefore generate different sets of adversarial examples against this model, depending on both the attack employed (PAD or GAMMA-YARA) and the GBDT variant considered (GBDT-ALL and GBDT-YARA). V. E XPERIMENTAL A NALYSIS In this section we first detail our experimental setup (Sect. V-A) and then report our findings in terms of trade-off between performance and computational requirements (Sect. V-B), robustness and computational requirements (Sect. V-C), along with a dedicated analysis of the detection levels of the Compound AI Systems (Sect. V-D). We summarize findings and analysis in take-home messages (THM) and practical deployment guidelines (G). A. Experimental Setup Dataset. We use the Speakeasy dataset [18] as the main data source. It contains PE files divided into training and test splits, collected in January 2022 and April 2022, respectively, and includes malware samples from seven families: Backdoor, Coinminer, Dropper, Keylogger, Ransomware, RAT, and Trojan. After duplicate removal, the training set consists of 71,505 malware samples and 26,059 benign samples, while the test set contains 17,495 malware samples and 10,000 benign samples. To mitigate the class imbalance, we increase the number of benign samples by adding 2,648 Windows system files from the sys32 and syswow64 directories of Windows 8.1, 10, and 11, together with 9,809 programs obtained from Chocolatey9 . To avoid data snooping [51], Windows system files are included only in the training set, whereas Chocolatey samples, collected in 2024, are reserved for testing. The final PE dataset therefore comprises 71,505 malware samples and 28,707 goodware samples for training, and 17,495 malware samples and 19,809 goodware samples for testing. Signatures. We collected the signatures (Sect. IV-A) forming the blocklist in April 2025 from several open-source repositories10,11,12,13,14,15 , excluding those that produced any false positive on the training data, yielding 236 signatures. We also collected 2,648 hashes from Windows 8.1, 10, and 11 system files to build the allowlist. The blocklist and the allowlist are integrated into OBELISK using YARA Python16 . Static Analysis using GBDT. The GBDT model used in the static level (see Sect. IV-A) is trained with 1000 trees, a learning rate of 0.1, and 32 parallel jobs, while keeping the other parameters to their default values [45]. Dynamic Analysis using Nebula. This phase relies on behavioral reports generated with the Speakeasy emulator7 . We use both the reports released with the dataset [18] and additional 9 https://chocolatey.org/ 10 https://github.com/bartblaze/Yara-rules/tree/master/rules 11 https://github.com/elastic/protections-artifacts/tree/main/yara/rules 12 https://github.com/malpedia/signator-rules/tree/main/rules 13 https://github.com/Neo23x0/signature-base/tree/master/yara 14 https://github.com/Yara-Rules/rules/tree/master/malware

reports generated for the newly collected goodware samples from Chocolatey9 . As observed in previous work [18], [19], some files fail during emulation, thereby reducing the number of usable reports. The dataset also contains duplicated reports originating from different samples, i.e., samples identified by different SHA-256 hashes. Upon inspection, we found that some of these correspond to malware variants, whereas others arise from limitations of the emulator. We remove all errors from the training set, counting 78,084 successful reports. Then we retain unique reports for a total of 67,611. Lastly, at test time, we evaluated all systems discarding samples that could not be analyzed as anticipated in Sect. IV-A. The Nebula model used at the dynamic level (see Sect. IV-A) is built using the same hyperparameters and preprocessing steps (filtering, normalization, and tokenization) reported by the authors in the original paper [12] and the associated library17 . Specifically, the tokenizer is implemented using BPE [52] with a vocabulary of 50,000 tokens, and the model takes as input sequences of up to 512 tokens. We train the model for 60 epochs using the AdamW [53] optimizer, with a learning rate of α = 1 × 10−4 , β1 = 0.9, β2 = 0.999, and ϵ = 10−8 and a batch size of 64. The best model is selected based on the validation ROC AUC computed on a 10% split of the training data. Setup of Adversarial Attacks. For A1–A4, we generated the adversarial sample sets described in Sect. IV-B using GAMMA, GAMMA-YARA, and PAD. For GAMMA, we leveraged the injection of 50 .rdata sections from Chocolatey samples (GAMMA-CHOCO) and 10 sections from Windows 11 files (GAMMA-WIN), while setting the regularization parameter to λ = 10−7 . Note that GAMMA-YARA is only needed when injecting sections from Chocolatey, as these may sometimes trigger signatures [19], [25], yielding the GAMMA-YARA-CHOCO variant. Conversely, GAMMAYARA-WIN is not reported, since sections from Windows 11 files never trigger signatures and, therefore, GAMMAWIN coincides with GAMMA-YARA-WIN. For PAD, we configured attacks that inject 0.5 KB, 1 KB, and 1.5 KB of content. Because all the described attacks share the same genetic algorithm as optimizer [17], [54], we set a maximum of 500 queries and a population size of 10, i.e., the number of candidate solutions evaluated at each optimization round. All attacks were instantiated through the SecML Malware library [38]. We generated all attacks using the same pool of 700 malware samples drawn from the Speakeasy testset, comprising 100 samples from each family. Thus, we consider the following sets of adversarial EXEmples for each attack: • A1 comprises three sets of adversarial samples, generated with GAMMA-CHOCO, GAMMA-WIN, and PAD against a surrogate model; • A2 comprises three sets of adversarial samples, generated with GAMMA-YARA-CHOCO, GAMMA-WIN, and PAD. These attacks are in part the same as A1 (i.e., GAMMA-WIN and PAD) but against a different model (YARA plus the surrogate model of A1); • A3 comprises six sets of adversarial samples, generated with GAMMA-CHOCO, GAMMA-WIN, and PAD

15 https://yaraify.abuse.ch/yarahub/ 16 https://github.com/VirusTotal/yara-python

17 https://github.com/dtrizna/nebula

8

•

against GBDT-ALL or GBDT-YARA; A4 comprises six sets of adversarial samples, generated with GAMMA-YARA-CHOCO, GAMMA-WIN, and PAD against GBDT-ALL or GBDT-YARA. These attacks are in part the same as A3 (i.e., GAMMA-WIN and PAD) but against a different model (YARA plus GBDT-ALL or YARA plus GBDT-YARA).

Considered Compound AI Systems. We tested all the adversarial samples previously described against different variants of our Compound AI System OBELISK, named as follows: Standard (STND), where the first level is implemented with YARA rules, the second level with GBDT-ALL with only one detection threshold, and the third level is implemented with Nebula trained on all reports (NEBULAALL); • OBV1, which uses the same structure of STND, but, at inference time, the second level filters data using two thresholds (YARA, GBDT-ALL, NEBULA-ALL); • OBV2, which uses the same structure of OBV1, but the second level deploys GBDT-YARA (hence YARA, GBDT-YARA, NEBULA-ALL); • OBV3, which uses the same structure of OBV2 but, the third level is implemented with Nebula trained on the data filtered by both YARA and GBDT-YARA (NEBULA-GBDT-YARA), resulting in YARA, GBDTYARA, NEBULA-GBDT-YARA. •

Metrics. We quantify the detection performance with three different metrics: the True Positive Rate (TPR), counting the fraction of correctly classified malicious samples; the False Positive Rate (FPR), counting the fraction of raised false alarms, and the F1 Score (F1), which jointly accounts for both TPR and FPR. We also include the transfer rate (TR), which counts the fraction of attacks that evade both the target for which they are computed and the real target system. We quantify the computational requirements with two other metrics: the training time (Tt ) and mean inference time (Ti ), both expressed in seconds. Training time (Tt ) is computed as the sum of the average preprocessing time needed to extract EMBER features, the time needed to train GBDT, the average preprocessing time to produce reports with Speakeasy along with filtering and normalization, the training of the BPE tokenizer, and the time needed to train Nebula. We omit from this analysis the time needed to produce signatures, as it would require estimating the amount of working days needed by reverse engineers (thus being infeasible for the scope of this paper). Inference time Ti is computed as the average time needed to make a prediction using the different versions of OBELISK. It sums the time needed to perform the pattern-matching with signatures (always) and possibly (if the signatures do not label it) the time to extract EMBER features and compute a probability with GBDT and possibly (if the prediction does not reach the desired confidence) the time needed to produce a report with Speakeasy, filtering and normalizing it, and computing a probability with Nebula. Hardware. All experiments have been conducted on a workstation equipped with an Intel® Xeon(R) Gold 5420, two Nvidia L40S GPUs, and 540 GB of RAM.

B. Performance vs Computational Requirements To quantify the trade-off between detection performance and computational requirements, we first conduct an ablation study on the threshold δ that OBV1, OBV2, and OBV3 use to filter samples between levels. We consider twelve increasing values (0.1, 0.27, 0.44, 0.72, 1.18, 1.93, 3.16, 5.18, 8.48, 13.89, 22.76, 37.28) × 10−3 , denoted in ascending order as δ0 , δ1 , . . . , δ11 . This yields twelve configurations for each of OBV1, OBV2, and OBV3 which, together with the STND baseline, amount to the 37 Compound AI Systems considered throughout our evaluation. The amount of data reaching each level differs across systems, and so does the cost of training it. GBDT-ALL, deployed in STND and OBV1, is trained on the full training set of 100,212 samples, whereas GBDT-YARA, deployed in OBV2 and OBV3, is trained only on the 79,319 samples that are not resolved by the signature level. Likewise, NEBULA-ALL, deployed in STND, OBV1, OBV2, and OBV3, is trained on all 67,611 available emulation reports, while NEBULA-GBDT-YARA, deployed in OBV3, is trained only on the reports that fall inside the uncertainty region of the static level; its training set shrinks from 59,461 samples for δ0 down to 3,373 for δ11 . We evaluate all systems on the test set, reporting the resulting ROC curves in Fig. 4, namely OBV1 in Fig. 4a, OBV2 in Fig. 4b, and OBV3 in Fig. 4c, each compared against STND (black dashed line) and each solid curve corresponding to one value of δ. These curves expose a consistent effect of δ on test-time performance: the TPR degrades when a large fraction of samples is forwarded to the dynamic level (i.e., low values of δ), and improves as filtering becomes more aggressive (i.e., high values of δ). This is expected, since delegating more decisions to the dynamic level exposes the system to its lower stand-alone accuracy [36]. To obtain scores that are comparable across systems, we tune the detection threshold of each system (i.e., the threshold of its last level) at ≈1% FPR, selecting the point on the ROC curve that comes closest to this constraint and therefore attaining slightly smaller or slightly higher FPRs (circles and stars in Fig. 4, respectively). On this common operating point we can inspect the trade-off between performance and computational requirements in Fig. 5, which relates TPR and F1 to Tt in Fig. 5a, and to Ti in Fig. 5b, for each of the 37 Compound AI Systems. When considering the trade-off between Tt and performance: THM1 OBV3 is the fastest family of systems to train while retaining high detection performance, the clearest example being OBV3 with δ11 , which scores ≈96% TPR at ≈1% FPR. THM2 STND, OBV1, and OBV2 require an extensive training time (≥ 100 hours), dominated by data preparation rather than by model fitting: feature extraction costs ≈0.26s per sample and, above all, emulation requires ≈8.1s per sample. When observing the trade-off between Ti and performance: THM3 The best system is OBV1 with δ11 , but almost all the others cluster in the bottom-left corner of the plot, with OBV3 at δ11 nearly matching the performance of the best system at a comparable inference cost.

9

0.95 0.9 0.85

10 2

10 1

100

False Positive Rate

1

True Positive Rate

1

True Positive Rate

True Positive Rate

1

0.95 0.9 0.85

10 2

(a) OBV1 STND

0

1

10 1

100

False Positive Rate

0.95 0.9 0.85

10 2

(b) OBV2 2

3

4

5

6

7

8

9

10 1

100

False Positive Rate (c) OBV3

10

Threshold @ 1% FPR

11

First tunable threshold

4 × 100 3 × 100

Mean Inference Time (s) Mean Inference Time (s)

102

Training time (h)

Mean Inference Time (s)

Fig. 4: Receiver Operating Characteristic (ROC) curves for the evaluated Compound AI systems. From (a) to (c), we plot the different configurations of OBV1, OBV2 and OBV3, keeping the STND system always depicted as a black dashed line.

101 1

2 × 100

0.98 0.96 0.94 0.92 F1 Score

0.9

0.88 1

0.95

0.9 TPR

0.85

100

6 × 10 1 1

0.8

(a) Training Time vs Performance

STND

OBV1

OBV2

OBV3

0.98 0.96 0.94 0.92 F1 Score

0.9

0.88 1

0.95

0.9 TPR

0.85

0.8

(b) Mean Inference Time vs Performance 0

1

2

3

4

5

6

7

8

9

10

11

STND

Fig. 5: Pareto fronts representing the trade-off between performance and Tt (a) and between performance and Ti (b). We considered all 37 Compound AI Systems using different markers and colors. THM4 STND requires ≥ 4 seconds per sample on average, being one order of magnitude slower than its competitors. This delay follows directly from its design: malware samples are typically stopped by an early level, whereas goodware samples traverse the entire pipeline and therefore always incur both feature extraction and emulation. Since inference is by far the most frequent operation, such a delay is likely to be unacceptable in a production environment. These results indicate that OBV3 with δ11 is the system that best balances performance and computational cost, making it suitable for production environments as long as only these two aspects are considered. C. Adversarial Robustness vs Computational Requirements We now turn to the trade-off between robustness and computational requirements when the systems are under attack. We take the 37 Compound AI Systems tuned at ≈1% FPR and, against each of them, we evaluate the adversarial EXEmples generated under the four threat models, i.e., A1–A4 (see Sect. IV-B). We quantify robustness through the worst-case TPR, computed by treating the attacks available within each threat model as an ensemble [16], [55]. For every sample, the attack is counted as successful if at least one strategy in the ensemble yields an adversarial EXEmple that evades the target system. The TPR computed on these outcomes therefore describes the worst-case scenario in which the attacker can always select the most effective strategy on a per-sample

basis, and it is the metric we report throughout this section. Fig. 6 shows the resulting trade-off for all 37 systems: the interplay between Ti and the worst-case TPR in Fig. 6a, and the interplay between Ti and TR in Fig. 6b. From these results we report that: THM5 STND is the most robust system across almost all threat models, and it ties with the other top performers under A3. It is followed by OBV1 and OBV2 (at intermediate values of δ), while OBV3 is the least robust family of systems. Neither training the levels on less data nor filtering samples at the static level is therefore beneficial for robustness. THM6 Robustness is not monotone in δ: both OBV1 and OBV2 attain their best robustness at intermediate values (δ6 and δ7 ), and the most robust configuration of OBV3 is again δ6 , at the cost of an average TPR drop of less than 10% w.r.t. STND. THM7 The TR grows as attacker knowledge increases from TM1 to TM4, since knowing more levels of the target lets the attacker craft attacks that are tailored to the deployed components rather than to a surrogate. THM8 Attacks inflate Ti since (i) the manipulations enlarge the input, slowing down feature extraction at the static level, and (ii) the adversarial EXEmples that bypass the first two levels reach the dynamic level, triggering emulation, i.e., the most expensive operation of the pipeline. Thus, robustness and responsiveness are coupled, as attacks degrade the latter even when they

10

A1

A2

A3

A4

Mean Inference Time (s)

8 6 4 2 01

0.8

0.6

0.4

0.2

01

0.8

0.6

0.4

0.2

01

0.8

0.6

0.4

0.2

01

0.8

0.6

0.4

0.2

0

0.8

1

(a) Trade-off between Worst-case TPR and Ti computed on all 37 Compound AI Systems for all attacks A1–A4. A1 A2 A3 A4 Mean Inference Time (s)

8 6 4 2 00

0.2

0.4

0.6

0.8

10

0.2

0.4

0.6

0.8

10

0.2

0.4

0.6

0.8

10

0.2

0.4

0.6

(b) Trade-off between Worst-case TR and Ti computed on all 37 Compound AI Systems for all attacks A1–A4.

STND

OBV1

OBV2

OBV3

0

1

2

3

4

5

6

7

8

9

10

11

STND

Fig. 6: Pareto fronts representing the trade-off between Ti and worst-case TPR (a) and TR (b). We considered all 37 Compound AI Systems using different markers and colors.

do not fully defeat the former. Summarizing, STND is the best choice for the trade-off between robustness and computational requirements under attack: it does not rely on a filtering threshold that widens the attack surface, and its levels are trained on more data, which makes them better at flagging samples that were never observed at training time. D. Level-wise Analysis We now analyze how each level of the Compound AI Systems contributes to both performance and robustness. Rather than repeating this analysis on all 37 systems, we focus on five representative ones: STND, as the most accurate and robust system, at the expense of computational requirements (THM5); OBV1 and OBV2 with δ6 , as they exhibit average behavior across all metrics (THM3 and THM6); OBV3 with δ6 , being among the fastest systems to train while incurring the smallest robustness drop within its family (THM1 and THM6); and OBV3 with δ11 , the fastest system to train and among the fastest at inference time, retaining top-tier performance at the expense of robustness (THM1 and THM3). For each of these systems, Fig. 7 reports the percentage of test samples whose final decision is taken by a specific level, separating correct predictions from misclassifications (solid and hatched bars, respectively). This analysis shows that: THM9 The filtering threshold δ shifts the decision burden onto the static level, which resolves ≈80% of the test set on its own. STND behaves in the opposite way, delegating most benign samples to the dynamic level, which accounts for ≈47% of the correctly classified test set and explains the Ti reported in THM4.

THM10 The signature level correctly classifies ≈13% of the test set while raising a negligible number of false alarms, a direct consequence of discarding every signature that fires false positives on the training set. Turning to robustness, Fig. 8 reports the same level-wise breakdown when the systems are exposed to the adversarial EXEmples of A1–A4, under the worst-case scenario described in Sect. V-C. Here we observe that: THM11 Signatures are effective defenders, but only if the deployed set stays undisclosed: for STND, the share of attacks stopped at this level falls from ≈75% under A1 and A3 to ≈25% under A2 and A4, where the attacker avoids the artifacts that trigger the rules. THM12 The filtering threshold δ should not be set too high, as an overly wide benign region lets the attacker defeat the entire pipeline by evading the static level alone; this is visible as the hatched area on top of the static level for OBV3 with δ11 . THM13 The dynamic level contributes to defense only when trained on the full dataset. The model deployed in OBV3 with δ11 , trained on the smallest number of reports, stops virtually no attack once the first two levels have been bypassed. Taken together, THM11–THM13 show that the robustness of a Compound AI System is not the sum of the robustness of its levels: aggressive filtering both removes decisions from the levels best placed to catch unknown samples and hands the attacker a shortcut through the pipeline.

Percentage of Predictions

11

100% 90% 80% 70% 60% 50% 40% 30% 20% 10% 0%

robust system in our evaluation (THM5). It is viable only where the infrastructure tolerates decisions taken in the order of seconds per sample, as every benign program traverses the full pipeline (THM4, THM9).

STND

Signatures Level

OBV1- 6 Static Level

OBV2- 6

OBV3- 6 OBV3- 11

Dynamic Level

Correctly Classified

Misclassified

Fig. 7: Percentage of test-set predictions per level, with correct (solid) and misclassified (hatched) samples shown separately. 100%

A1

A2

A3

A4

80%

Percentage of Predictions

60% 40% 20% 0% 100% 80% 60% 40%

Signatures Level

Static Level

Dynamic Level

Correctly Classified

OBV3- 11

OBV3- 6

OBV2- 6

OBV1- 6

STND

OBV3- 11

OBV3- 6

OBV2- 6

OBV1- 6

0%

STND

20%

Misclassified

Fig. 8: Percentage of predictions under attacks A1–A4 per level, with correct (solid) and misclassified (hatched) samples shown separately. E. Final Recommendations We conclude by distilling our findings into four deployment guidelines, each tailored to a distinct set of operational constraints, so that practitioners can identify the configuration matching their own setting: G1 No dominant constraint. Deploy OBV3 with δ6 , i.e., a system whose levels are all trained on filtered data and which also filters samples at inference time. It is the balanced compromise across all metrics, excelling at none of them but never collapsing on any (THM1, THM6). G2 Computation constrained at both training and inference time. Deploy OBV3 with δ11 , the fastest system to train and among the fastest to compute predictions, while still reaching ≈96% TPR at ≈1% FPR (THM1, THM3). This speed is paid in robustness: its wide benign region lets attackers bypass the pipeline by evading the static level alone, and its dynamic level, trained on the fewest reports, provides no residual defense (THM12, THM13). G3 Training resources abundant, deployment resources constrained. Deploy OBV1 with δ6 , which combines strong performance, robustness close to STND, and fast inference, at the price of training cost (THM2, THM6). G4 Performance and robustness paramount, latency not a bottleneck. Deploy STND, the most accurate and most

VI. L IMITATIONS Limited Performance of Dynamic Analysis. As previously observed [19], [36], dynamic analysis is less accurate than static analysis, and improving its effectiveness remains a challenging and still insufficiently explored research direction [56]. In fact, Nebula [12] requires large training sets and days of emulation; virtualization would yield richer traces but at considerably higher cost [35]. Nevertheless, to the best of our knowledge, Nebula remains one of the best-performing architectures for dynamic malware detection [12], [19], which makes its inclusion in a Compound AI System representative of what a practitioner could realistically deploy today. The limits we report for the dynamic level should therefore be read as limits of the current state of the art, not as artifacts of our choice of model. Only Static Adversarial EXEmples. Our robustness evaluation relies exclusively on attacks against static models, without targeting the dynamic level within the proposed threat models. As discussed in Sect. III-B, this reflects the current state of available tools: the relevant techniques either lack a public implementation [39], [41] or depend on outdated and unmaintained platforms [23] difficult to set up and execute, in contrast to static attacks, for which established open-source solutions exist [38]. This choice is also consistent with our own findings: adversarial EXEmples that evade the static level are rarely stopped by the dynamic one (THM13), so attacking the static level is where the attacker’s effort is best spent. Extending the threat models to manipulations of runtime behavior would enlarge the attack surface we measure, and we expect our robustness figures to be optimistic in that respect. Missing End-to-End Attacks. We did not compute querybased attacks [17], [22] that optimize directly against the output of the Compound AI System; we instead attacked its first levels and transferred the resulting adversarial EXEmples to the complete pipeline. As anticipated in Sect. III-B, a query-based attack would require waiting for the emulator to terminate at every iteration of the optimization, making the process prohibitively slow for the number of samples and configurations considered here. Transfer attacks provide a lower bound on attack effectiveness, and hence an upper bound on the robustness of the systems under test: an endto-end attacker would be at least as effective as the one we model, which strengthens rather than weakens our comparative conclusions across systems. Robustness to Packing and Obfuscation. We did not explicitly evaluate detection performance under packing and obfuscation, nor did we control for their presence in the data. The Speakeasy dataset [18] was collected in the wild by a security vendor and its authors report that packed samples are included, so such techniques are represented in our evaluation; we rely on this assessment without quantifying their prevalence. Limited Data Sources. While large-scale [33], [57] and more recent [58] datasets are available, they require paid access to

12

VirusTotal18 to retrieve the corresponding programs, which we need to run the attacks and the emulation. Our methodology is nonetheless independent of the specific data source and can be replicated on any dataset providing the raw executables. Temporal Drift. We did not explicitly address temporal drift, i.e., the shift in data distribution over time. Unlike other domains [59], establishing when a sample was first observed in the wild would require either vendor telemetry or VirusTotal access, neither viable here. Our splits preserve a temporal ordering (Speakeasy training/test sets collected in January/April 2022, Chocolatey goodware in 2024) but this is not a substitute for an evaluation over a longer observation window.

a critical defensive role that collapses once the signatures are disclosed; overly aggressive filtering lets attackers bypass the pipeline by evading the static level alone; and dynamic analysis contributes to defense only marginally, and only when trained on the full dataset. Efficiency gains and robustness losses thus stem from the same mechanism, which is precisely the tradeoff a practitioner must resolve. We translate this into four deployment guidelines that map operational constraints to the best configuration. Our work is a first step toward a more realistic evaluation of Compound AI Systems, narrowing the gap between industrial practice and academic research. ACKNOWLEDGEMENTS

VII. F UTURE W ORK AND C ONCLUSIONS Future Work. Our results identify the dynamic level as the weakest component of the pipeline, and improving it is our primary direction for future work: we will investigate richer representations of behavioral data, as well as alternative emulators and virtualization approaches that trade execution speed for fidelity [35]. In parallel, we will extend our treatment of Compound AI Systems to the malware classification task, i.e., attributing a sample to the family it belongs to, so that the same system can reliably both detect and characterize threats. On the attacker side, we plan to broaden the threat models along two axes. First, we will study techniques that evade dynamic detection directly, and the tampering of training data through poisoning attacks [37], [60]. Second, as discussed in Sect. V-C, an attacker who evades both the signature and the static level forces the system to fall back on emulation, inflating inference time. We will investigate whether this behavior can be exploited systematically to mount Denial-ofService (DoS) attacks that overload the dynamic level, turning a robustness weakness into an availability one. Conclusions. In this work, we advocate for a Compound AI System paradigm in which multiple ML components cooperate and are complemented by classical detection techniques, moving beyond the evaluation of standalone models. We pair this paradigm with system-level threat models that consider attackers with increasing knowledge of the deployed components, overcoming the limits of isolated robustness evaluations. Together, these two contributions enable an explicit analysis of the trade-offs among detection performance, computational requirements, and robustness. To this end, we implemented OBELISK, a Compound AI System whose levels reduce and refine the data forwarded from one stage to the next, and we compared several of its configurations, each adopting a different filtering policy at training and/or test time, against a standard version that applies no filtering. Our experiments show that filtering policies substantially reduce training cost and improve responsiveness while sacrificing little detection performance. Our four threat models further show that betterinformed attackers craft adversarial EXEmples that transfer more effectively to the complete system, and that these attacks degrade responsiveness as a side effect. The levelwise analysis explains where the system’s robustness comes from, and where it is lost: signature-based components play 18 https://www.virustotal.com

This project has been partially funded by FISA-2023-00128 “InfoAICert” funded by the MUR program “Fondo Italiano per le Scienze Applicate”, by SERICS (PE00000014) and by PNRR MUR Project (PE0000013) ”Future Artificial Intelligence Research (FAIR)”, funded by the European Union – NextGenerationEU, CUP J33C24000420007. R EFERENCES [1] Crowdstrike, “The Rise of Machine Learning in Cybersecurity,” 2025, accessed in February 2026, https://go.crowdstrike.com/rs/281-OBQ266/images/WhitepaperMachineLearning.pdf. [2] Kaspersky, “Machine Learning for Malware Detection,” 2021, accessed in February 2026, https://media.kaspersky.com/en/enterprisesecurity/Kaspersky-Lab-Whitepaper-Machine-Learning.pdf. [3] Avast, “AI and Machine Learning,” 2025, accessed in February 2026, https://www.avast.com/technology/ai-and-machine-learning. [4] Microsoft, “Advanced technologies at the core of Microsoft Defender Antivirus,” 2025, accessed in February 2026, https://learn.microsoft. com/en-us/defender-endpoint/adv-tech-of-mdav. [5] BitDefender, “An Adaptive and Layered Approach to Endpoint Security,” 2017, accessed in February 2026, https://www.bitdefender.co.th/ wp-content/uploads/gz/ESG-White-Paper-Bitdefender-Jun-2017.pdf. [6] Apple, “Protecting against malware in macOS,” 2025, accessed in February 2026, https://support.apple.com/guide/security/sec469d47bd8. [7] Kaspersky, “About YARA scan in Kaspersky Endpoint Agent,” 2025, accessed in February 2026, https://support.kaspersky.com/kea/3.15/225454. [8] E. Raff, J. Barker, J. Sylvester, R. Brandon, B. Catanzaro, and C. K. Nicholas, “Malware detection by eating a whole exe,” in Workshops at the 32nd AAAI conference on artificial intelligence, 2017. [9] E. Raff, W. Fleshman, R. Zak, H. S. Anderson, B. Filar, and M. McLean, “Classifying sequences of extreme length with constant memory applied to malware detection,” in AAAI Conference on Artificial Intelligence, 2021. [10] S. E. Coull and C. Gardner, “Activation analysis of a byte-based deep neural network for malware classification,” in IEEE Security and Privacy Workshops, 2019. [11] M. Krčál, O. Švec, M. Bálek, and O. Jašek, “Deep convolutional malware classifiers can learn from raw executables and labels only,” in International Conference on Learning Representations Workshop, 2018. [12] D. Trizna, L. Demetrio, B. Biggio, and F. Roli, “Nebula: Self-attention for dynamic malware analysis,” IEEE Transaction on Information Forensics and Security, vol. 19, pp. 6155–6167, 2024. [13] X. Ling, L. Wu, W. Deng, Z. Qu, J. Zhang, S. Zhang, T. Ma, B. Wang, C. Wu, and S. Ji, “Malgraph: Hierarchical graph neural networks for robust windows malware detection,” in IEEE Infocom 2022-IEEE Conference on Computer Communications. IEEE, 2022, pp. 1998– 2007. [14] C. Jindal, C. Salls, H. Aghakhani, K. Long, C. Kruegel, and G. Vigna, “Neurlux: dynamic malware analysis without feature engineering,” in Annual Computer Security Applications Conference, 2019. [15] L. Demetrio, S. E. Coull, B. Biggio, G. Lagorio, A. Armando, and F. Roli, “Adversarial exemples: A survey and experimental evaluation of practical attacks on machine learning for windows malware detection,” ACM Transaction on Privacy and Security, vol. 24, no. 4, pp. 1–31, 2021.

13

[16] F. Croce, M. Andriushchenko, V. Sehwag, E. Debenedetti, N. Flammarion, M. Chiang, P. Mittal, and M. Hein, “Robustbench: a standardized adversarial robustness benchmark,” in Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2021. [17] L. Demetrio, B. Biggio, G. Lagorio, F. Roli, and A. Armando, “Functionality-preserving black-box optimization of adversarial windows malware,” IEEE Transaction on Information Forensics and Security, vol. 16, pp. 3469–3478, 2021. [18] D. Trizna, “Quo vadis: hybrid machine learning meta-model based on contextual and behavioral malware representations,” in ACM Workshop on Artificial Intelligence and Security, 2022. [19] A. Ponte, D. Trizna, L. Demetrio, B. Biggio, I. T. Ogbu, and F. Roli, “Slifer: Investigating performance and robustness of malware detection pipelines,” Computers & Security, vol. 150, p. 104264, 2025. [20] B. Kolosnjaji, A. Demontis, B. Biggio, D. Maiorca, G. Giacinto, C. Eckert, and F. Roli, “Adversarial malware binaries: Evading deep learning for malware detection in executables,” in European signal processing conference, 2018. [21] K. Lucas, M. Sharif, L. Bauer, M. K. Reiter, and S. Shintre, “Malware makeover: Breaking ml-based static analysis by modifying executable bytes,” in ACM Asia Conference on Computer and Communications Security, 2021. [22] W. Song, X. Li, S. Afroz, D. Garg, D. Kuznetsov, and H. Yin, “Mabmalware: A reinforcement learning framework for blackbox generation of adversarial malware,” in ACM on Asia Conference on computer and communications security, 2022. [23] G. Digregorio, S. Maccarrone, M. D’Onghia, L. Gallo, M. Carminati, M. Polino, and S. Zanero, “Tarallo: Evading behavioral malware detectors in the problem space,” in International Conference on Detection of Intrusions and Malware, and Vulnerability Assessment, 2024. [24] A. Demontis, M. Melis, M. Pintor, M. Jagielski, B. Biggio, A. Oprea, C. Nita-Rotaru, and F. Roli, “Why do adversarial attacks transfer? explaining transferability of evasion and poisoning attacks,” in USENIX security symposium, 2019. [25] A. Ponte, L. Demetrio, L. Oneto, I. T. Ogbu, B. Biggio, and F. Roli, “Demystifying the role of rule-based detection in ai systems for windows malware detection,” in IEEE European Symposium on Security and Privacy Workshops, 2025. [26] F. Kreuk, A. Barak, S. Aviv-Reuven, M. Baruch, B. Pinkas, and J. Keshet, “Deceiving end-to-end deep learning malware detectors using adversarial examples,” arXiv preprint arXiv:1802.04528, 2018. [27] R. Chaganti, V. Ravi, and T. D. Pham, “A multi-view feature fusion approach for effective malware classification using deep learning,” Journal of information security and applications, vol. 72, p. 103402, 2023. [28] C. TALOS, “ClamAV,” 2025, accessed in February 2026, https://www.clamav.net/. [29] BitDefender, “BitDefender Antivirus Technology,” 2025, accessed in February 2026, https://www.bitdefender.com/files/Main/file/ BitDefender Antivirus Technology.pdf. [30] R. Horrigan and T. Nguyen, “How CrowdStrike’s malware analysis agent detects malware at machine speed,” 2026, accessed in February 2026. [Online]. Available: https://www.crowdstrike.com/en-us/blog/ how-crowdstrike-detects-malware-at-machine-speed/ [31] C. Scano, G. Floris, B. Montaruli, L. Demetrio, A. Valenza, L. Compagna, D. Ariu, L. Piras, D. Balzarotti, and B. Biggio, “Modseclearn: Boosting modsecurity with machine learning,” in International Symposium on Distributed Computing and Artificial Intelligence, 2024. [32] S. Gupta, F. Lu, A. Barlow, E. Raff, F. Ferraro, C. Matuszek, C. Nicholas, and J. Holt, “Living off the analyst: Harvesting features from yara rules for malware detection,” in IEEE International Conference on Big Data, 2024. [33] H. S. Anderson and P. Roth, “Ember: an open dataset for training static pe malware machine learning models,” arXiv preprint arXiv:1804.04637, 2018. [34] A. Afianian, S. Niksefat, B. Sadeghiyan, and D. Baptiste, “Malware dynamic analysis evasion techniques: A survey,” ACM Computing Surveys, vol. 52, no. 6, pp. 1–28, 2019. [35] A. Küchler, A. Mantovani, Y. Han, L. Bilge, and D. Balzarotti, “Does every second count? time-based evolution of malware behavior in sandboxes,” in Network and Distributed Systems Security Symposium, 2021. [36] S. Dambra, Y. Han, S. Aonzo, P. Kotzias, A. Vitale, J. Caballero, D. Balzarotti, and L. Bilge, “Decoding the secrets of machine learning in malware classification: A deep dive into datasets, feature extraction, and model performance,” in ACM SIGSAC Conf. on Computer and Communications Security, 2023.

[37] B. Biggio and F. Roli, “Wild patterns: Ten years after the rise of adversarial machine learning,” Pattern Recognition, vol. 84, pp. 317– 331, 2018. [38] L. Demetrio and B. Biggio, “Secml-malware: Pentesting windows malware classifiers with adversarial exemples in python,” arXiv preprint arXiv:2104.12848, 2021. [39] I. Rosenberg and E. Gudes, “Bypassing system calls-based intrusion detection systems,” Concurrency and Computation: Practice and Experience, vol. 29, no. 16, p. e4023, 2017. [40] I. Rosenberg, A. Shabtai, L. Rokach, and Y. Elovici, “Generic black-box end-to-end attack against state of the art api call based malware classifiers,” in International Symposium on Research in Attacks, Intrusions, and Defenses, 2018. [41] I. Rosenberg, A. Shabtai, Y. Elovici, and L. Rokach, “Query-efficient black-box attack against sequence-based malware classifiers,” in Annual Computer Security Applications Conference, 2020. [42] X. Chen, L. Cui, H. Wen, Z. Li, H. Zhu, Z. Hao, and L. Sun, “Malader: Decision-based black-box attack against api sequence based malware detectors,” in Annual IEEE/IFIP Interenational Conference on Dependable Systems and Networks, 2023. [43] J. H. Friedman, “Greedy function approximation: a gradient boosting machine,” Annals of statistics, pp. 1189–1232, 2001. [44] ——, “Stochastic gradient boosting,” Computational statistics & data analysis, vol. 38, no. 4, pp. 367–378, 2002. [45] T. Chen and C. Guestrin, “Xgboost: A scalable tree boosting system,” in ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016. [46] M. Kozak, L. Demetrio, D. Trizna, and F. Roli, “Updating windows malware detectors: Balancing robustness and regression against adversarial exemples,” Computers & Security, vol. 155, p. 104466, 2025. [47] L. Yang, W. Guo, Q. Hao, A. Ciptadi, A. Ahmadzadeh, X. Xing, and G. Wang, “{CADE}: Detecting and explaining concept drift samples for security applications,” in USENIX Security Symposium, 2021. [48] G. Severi, J. Meyer, S. Coull, and A. Oprea, “{Explanation-Guided} backdoor poisoning attacks against malware classifiers,” in USENIX security symposium, 2021. [49] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” Neural information processing systems, 2017. [50] Z. Zhang, P. Qi, and W. Wang, “Dynamic malware analysis with feature engineering and feature learning,” in AAAI Conf. on artificial intelligence, 2020. [51] D. Arp, E. Quiring, F. Pendlebury, A. Warnecke, F. Pierazzi, C. Wressnegger, L. Cavallaro, and K. Rieck, “Dos and don’ts of machine learning in computer security,” in USENIX Security Symposium, 2022. [52] P. Gage, “A new algorithm for data compression,” The C Users Journal, vol. 12, no. 2, pp. 23–38, 1994. [53] I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations, 2019. [54] F. A. Fortin, F. M. De Rainville, M. A. G. Gardner, M. Parizeau, and C. Gagné, “Deap: Evolutionary algorithms made easy,” The Journal of Machine Learning Research, vol. 13, no. 1, pp. 2171–2175, 2012. [55] A. E. Cinà, J. Rony, M. Pintor, L. Demetrio, A. Demontis, B. Biggio, I. B. Ayed, and F. Roli, “Attackbench: Evaluating gradient-based attacks for adversarial examples,” in AAAI Conference on Artificial Intelligence, 2025. [56] Y. Kaya, Y. Chen, M. Botacin, S. Saha, F. Pierazzi, L. Cavallaro, D. Wagner, and T. Dumitraş, “Ml-based behavioral malware detection is far from a solved problem,” in IEEE Conference on Secure and Trustworthy Machine Learning, 2025. [57] R. Harang and E. M. Rudd, “Sorel-20m: A large scale benchmark dataset for malicious pe detection,” arXiv preprint arXiv:2012.07634, 2020. [58] R. J. Joyce, G. Miller, P. Roth, R. Zak, E. Zaresky-Williams, H. Anderson, E. Raff, and J. Holt, “Ember2024-a benchmark dataset for holistic evaluation of malware classifiers,” in ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2025. [59] F. Pendlebury, F. Pierazzi, R. Jordaney, J. Kinder, and L. Cavallaro, “{TESSERACT}: Eliminating experimental bias in malware classification across space and time,” in USENIX security symposium, 2019. [60] A. E. Cinà, K. Grosse, A. Demontis, S. Vascon, W. Zellinger, B. A. Moser, A. Oprea, B. Biggio, M. Pelillo, and F. Roli, “Wild patterns reloaded: A survey of machine learning security against training data poisoning,” ACM Computing Surveys, vol. 55, no. 13s, pp. 1–39, 2023.

Record · ID 667899 · SHA-256 841faf3c863f7cfa
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.