Conceptio › Archive › arXiv CS
arXiv CSopen access

Delphi Scanner: efficient and interpretable static malware detection via API sequence modeling

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

Delphi Scanner: efficient and interpretable static malware detection via API sequence modeling Bijied Brahimi1 , Vincent Cohadon1 , Gabriel Glazman1 , Rayan Al Mohaize1 , Omran Berjawi2 , and Rida Khatoun2 1

arXiv:2609.19900v1 [cs.CR] 17 Sep 2026

2

Université Paris Cité, Paris, France Institut Polytechnique de Paris, Télécom Paris, Palaiseau, France

Abstract. Static malware detection for Windows Portable Executable files demands a careful balance between detection effectiveness, computational efficiency, and analytical interpretability. This paper introduces Delphi Scanner, a static malware detection system for Windows PE files that balances efficiency with behavioral interpretation. It uses a convolutional neural network (CNN) to model Windows API sequences to classify PE and a decoupled interpretation layer based on a rule-based layer to categorize APIs into high-level malicious capabilities. Evaluated on over 190,000 Windows PE files, the system achieves 95.35% accuracy with a 1.53 MB model footprint. Robustness experiments on 5,647 out-of-distribution MalwareBazaar samples, paired packed and unpacked executables, and three adversarial manipulation strategies confirm generalization beyond the training distribution and resistance to functionalitypreserving evasion techniques. Overall, these results demonstrate that API sequence-based static analysis offers a practical, interpretable, and efficient foundation for malware triage in local deployment scenarios.

Keywords: Malware detection· Static analysis · Windows API imports · Deep learning Cross-modal signal fusion

1

Introduction

Malicious software targeting the Microsoft Windows ecosystem remains one of the most persistent and impactful threats in contemporary cybersecurity. Windows executables, distributed in the Portable Executable (PE) format, a standard binary structure for Windows programs and libraries, constitute the primary delivery mechanism through which such malware is distributed. Despite decades of research and the deployment of sophisticated commercial antivirus solutions, malware continues to evolve rapidly in volume and evasiveness, exploiting the widespread adoption of Windows-based systems across personal, corporate, and industrial environments [6]. This persistent threat landscape has driven sustained interest in automated malware detection techniques. Traditionally malware detection works in two techniques: Dynamic and static analysis. Dynamic analysis operates by executing

2

Brahimi et al.

suspicious software within an isolated environment and monitoring its runtime behavior, capturing system-level events including API calls, file system and registry modifications, network connections, and memory allocation patterns [13]. This technique can uncover concealed behaviors, but it requires high computational cost and careful environment configuration, and it remains vulnerable to evasion by environment-aware malware. In contrast, static techniques examine executables without execution by inspecting their binary structure, metadata, and embedded artifacts. This approach enables fast, deterministic, and safe analysis, making it attractive for pre-execution screening and large-volume malware triage [30] Early static malware detection relied on manually crafted signatures, which proved brittle against minor code modifications and ineffective against unseen threats. Consequently, research has increasingly adopted machine learning to automatically learn discriminative patterns from static PE file features [4], including PE headers, section statistics, and imported libraries [28]. Among these, Windows API imports are particularly compelling: they offer a compact, semantically meaningful abstraction of program capabilities, revealing dependencies on file system access, network communication, registry manipulation, and process control, and have long served as key indicators for malware analysts [10]. Recent deep learning advances have further enabled modeling API usage as sequential data, allowing classifiers to capture local patterns and co-occurrence relationships among API calls [21]. Despite promising detection performance, many existing approaches prioritize accuracy in isolation and overlook practical deployment constraints, including inference latency, memory footprint, explainability, and user trust [16]. Complex architectures like transformers, while expressive, often introduce substantial computational overhead without commensurate gains for static analysis tasks [3]. A persistent limitation is interpretability: blackbox predictions provide little insight into classification decisions, complicating analyst validation and undermining trust in automated systems [5, 12]. This highlights a critical gap — current systems struggle to simultaneously balance detection effectiveness, computational efficiency, and meaningful interpretability under local, offline deployment constraints. Motivated by this challenge, this paper presents Delphi Scanner, a lightweight static malware detection system for Windows PE files that designed to explicitly balance detection effectiveness, inference efficiency, and analytical interpretability within a locally deployed architecture. The system extracts ordered Windows API call sequences from the Import Address Table (IAT) of PE files—including resolution of ordinal-based imports—and encodes them into fixed-length numerical representations that serve as the primary input to a multi-scale onedimensional convolutional neural network (CNN) classifier optimized for submillisecond inference via ONNX Runtime deployment. In parallel, and independently of the classification pipeline, a rule-based behavioral interpretation layer maps the same extracted API sequences to predefined malicious capability categories each aligned with MITRE ATT&CK technique identifiers [19]. In short, the contributions of this work are summarized as follows:

Title Suppressed Due to Excessive Length

3

– We propose a lightweight static malware detection pipeline that combines ordinal-aware Windows API import extraction with efficient sequence-based deep learning. – We introduce a decoupled post-hoc behavioral interpretation layer grounded in the MITRE ATT&CK framework, which enhances analyst-oriented explainability without influencing classification decisions or degrading inference efficiency. – We developed Delphi Scanner as a fully functional desktop application, and evaluate its detection performance on a real-world malware dataset, demonstrating competitive accuracy with a minimal computational footprint. The remainder of this paper is organized as follows. Section 2 reviews related work, Section 3 describes the proposed system, Section 4 presents the experimental setup, Section 5 reports the results. Section 6 shows a case study of the proposed system, Section 7 discusses limitations and implications, and Section 8 concludes the paper.

2

Related Work

Research on malware detection has progressed from signature-centric pipelines toward learning-based systems that exploit both static and dynamic evidence. This section reviews recent advances most closely related to our work.

2.1

API-centric representations: dynamic call sequences and static imports

Dynamic API call sequence modeling has been extensively explored using recurrent and convolutional architectures, often treating sequences as a language-like signal . For instance, API-MalDetect uses an NLP-inspired encoder together with convolutional and recurrent components to learn from long API call traces [21]. Similarly, hybrid sequence models have been proposed to improve generalization and sample efficiency; Owoh et al. combine GRUs with GAN-based augmentation to enhance detection performance on API call sequences while controlling computational overhead [23]. While dynamic traces require controlled execution environments and are sensitive to sandbox fidelity, many practical systems prefer static analogues of behavioral evidence. Static import-based analysis supports pre-execution screening and complements header/section features. Recent work has also explored API-oriented rule generation and signature assistance: APIARY proposes an API-driven automatic rule generator for YARA, aiming to produce discriminative signatures that can leverage API patterns originating from either static or dynamic sources [11, 7]. Such directions align with the broader trend of using API-level signals to bridge human analyst intuition and automated detection [14, 8].

4

2.2

Brahimi et al.

Deep learning directly over binaries

Beyond engineered feature vectors, end-to-end models have attempted to learn directly from raw executables. MalConv demonstrated the feasibility of bytelevel convolutional modeling over large PE inputs, catalyzing a substantial line of follow-on work [26]. Subsequent improvements addressed scalability and representation capacity. Notably, Raff et al. proposed constant-memory sequence classification to remove input-length constraints and improve training efficiency (often referred to as MalConv2 in the literature and associated tooling) [27]. These models reduce reliance on hand-crafted features but raise challenges in interpretability, susceptibility to spurious correlations, and robustness to distribution shift or deliberate manipulation of input bytes. Transformer-based approaches have also expanded rapidly in malware and binary analysis. I-MAD introduced a transformer architecture (Galaxy Transformer) for representing executable semantics at multiple granularities and paired it with an interpretable classification component [18]. In addition, surveys highlight the breadth of transformer usage across malicious software detection settings, including byte/opcode modeling and hybrid representations, while emphasizing open issues around efficiency, dataset bias, and reproducibility [3]. A recent SoK further systematizes the landscape of transformer-based malware analysis, highlighting design choices (tokenization, context length, pretraining, and multimodality) and stressing evaluation pitfalls under evolving threat conditions [24]. 2.3

Graph-based and relational static representations

A complementary direction models executables as structured objects (graphs) rather than sequences or flat feature vectors. Graph learning can capture relations between heterogeneous static features (e.g., imports co-occurring with section/string attributes) and has been explored to improve robustness under drift. MFGraph, for instance, constructs feature graphs from static PE attributes and applies graph convolution to learn representations that remain comparatively stable under temporal changes [31]. Such approaches suggest that explicitly modeling feature relations may provide resilience when distributions shift due to evolving development toolchains, packing practices, or malware family turnover. 2.4

Adversarial machine learning

Adversarial machine learning has likewise emerged as a practical threat model for both feature-based and raw-byte detectors. GAMBD proposes a gradient-based approach to generate functionality-preserving adversarial malware against MalConv, illustrating that even strong deep detectors can be evaded via constrained PE modifications [17]. Defensive research has responded with mechanisms that explicitly manage or scan the functionality-preserving attack space. Liu et al. propose attack-space management to harden byte-sequence malware detection, aiming to reduce the need for repeated retraining and to improve resilience

Title Suppressed Due to Excessive Length

5

against both white-box and black-box attacks [20]. At a higher level, surveys of PE malware evasion methods emphasize that transformation-, concealment-, and attack-based strategies are routinely combined in practice, challenging detectors that rely on narrow assumptions about observable structure [15]. These findings motivate approaches that (i) prioritize robust static signals, (ii) evaluate under realistic evasion settings, and (iii) maintain operational feasibility. In summary, prior work spans engineered static PE feature modeling [4, 2], dynamic API-call sequence learning [21, 23], end-to-end byte and transformerbased detectors [26, 27, 18, 3, 24], and robustness studies under packing and adversarial manipulation [17, 20, 15]. Our work builds on these insights while focusing on an API-centric static representation tailored for efficient PE analysis and evaluated under realistic constraints, thereby complementing sequenceheavy dynamic approaches and resource-intensive end-to-end models.

3

Delphi Scanner System

This section presents Delphi Scanner, a static malware detection system for local, interpretable, and low-latency analysis of Windows PE files. The system integrates machine learning based classification with rule-based behavioral interpretation to deliver both an automated malware verdict and analyst-oriented explanatory insights. By operating exclusively on static binary features and avoiding code execution, Delphi Scanner enables safe and efficient malware screening suitable for end-user and analyst environments.

Fig. 1. High-level architecture of Delphi Scanner: static feature extraction from the PE Import Address Table, CNN-based malware classification, rule-based behavioral interpretation layer operating in parallel with the CNN, and result presentation in the desktop interface.

6

Brahimi et al.

3.1

System Overview

Delphi Scanner targets four operational requirements: local and safe analysis without executing untrusted binaries, low-latency inference for interactive malware triage, interpretable behavioral output, and lightweight deployment with minimal runtime overhead. Figure 1 presents the high-level architecture. Given an input PE file, the system processes the binary through a modular pipeline comprising four components: static feature extraction, machine learning inference, behavioral interpretation, and user-facing result presentation. Each component operates independently and exchanges data through well-defined representations, facilitating modularity and extensibility. – Static Feature Extraction: The input PE file is statically parsed to extract imported Windows API functions from the Import Address Table (IAT). Ordinal-based imports are resolved where possible, and auxiliary static indicators, including section entropy and dynamic loading patterns, are computed. The resulting API sequence is encoded into a fixed-length numerical representation. – Machine Learning Inference: The encoded feature representation is processed by a CNN that estimates the probability of maliciousness. Inference is optimized for low-latency execution and is performed independently of any handcrafted behavioral rules. – Behavioral Interpretation: In parallel with classification, extracted APIs are mapped to predefined malicious capability categories, such as network communication, registry persistence, and process manipulation. This rule-based layer produces human-readable behavioral indicators without influencing the classifier’s prediction. – User Interface: The outputs of the classification and interpretation stages are aggregated and presented through a desktop graphical interface, including the final verdict, confidence score, and detected behavioral indicators. This architecture enforces a strict separation between statistical detection and behavioral interpretation, enabling accurate predictions while providing transparent explanatory context. 3.2

Static Feature Extraction Module

This module describes the static analysis pipeline used by Delphi Scanner to transform a Windows PE file into a numerical representation suitable for machine learning–based classification. The extraction process targets semantically meaningful behavioral signals while remaining lightweight and resilient to common obfuscation techniques. PE Parsing and Import Extraction Delphi Scanner parses the binary structure of Windows PE files to extract imported Windows API functions from the

Title Suppressed Due to Excessive Length

7

Import Address Table (IAT). The extraction follows a seven-step pipeline, described below. – Load and validate: The binary is read and validated by checking the DOS header magic bytes (MZ) and the PE signature (PE\0\0). Files exceeding 100 MB are rejected prior to parsing to bound memory usage. – Locate the Import Directory: The parser navigates the PE Optional Header to the Data Directories array and retrieves the Import Table entry, which points to the Import Directory Table in the binary’s virtual address space. – Iterate DLLs in import order: Each IMAGE_IMPORT_DESCRIPTOR structure yields the name of a dependent DLL (e.g., KERNEL32.dll). – Extract function names: For each DLL, the associated Import Lookup Table (ILT) is traversed entry by entry. Each entry is either a name reference, containing a function name string (e.g., CreateFileW), or an ordinal entry with a 16-bit ordinal number. – Resolve ordinal imports: Ordinal-based imports are resolved using the OrdinalResolver module, which contains built-in lookup tables for eight major Windows system libraries: kernel32.dll, ntdll.dll, ws2_32.dll, user32.dll, advapi32.dll, _shell32.dll, wininet.dll, and urlmon.dll, covering approximately 300 function mappings. For example, kernel32.dll ordinal 581 resolves to GetProcAddress. Ordinals that cannot be resolved against the built-in tables are mapped to the special <UNK> token. – Construct the ordered sequence: Resolved function names are appended to the sequence in DLL order, then in function order within each DLL. – Vectorize: Each API name is mapped to an integer index via a fixed vocabulary of 10,000 tokens. Sequences are truncated to a maximum length of 512 tokens; shorter sequences are padded with a dedicated <PAD> token to produce uniform-length inputs for the convolutional classifier. API Sequence Construction Following import extraction and ordinal resolution, the system constructs an ordered sequence of API identifiers representing the static capabilities of the executable. Each API name is mapped to a unique integer index using a predefined vocabulary generated during training. To enable batch processing and fixed-size model inputs, sequences are normalized to a uniform length by padding shorter sequences and truncating longer ones. This normalization ensures compatibility with convolutional neural network architectures while preserving the most informative portion of the API sequence. Auxiliary Static Indicators In addition to API sequences, Delphi Scanner computes auxiliary static indicators that provide contextual information about the reliability and limitations of static analysis. These indicators are not used as direct inputs to the classification model but are exposed to the user to support result interpretation: – Entropy: The Shannon entropy of the executable, used as a coarse indicator of packing or compression that may limit the completeness of static import analysis.

8

Brahimi et al.

Algorithm 1 Delphi Scanner Detection Pipeline Require: PE file f , API vocabulary V, max sequence length L, decision threshold τ , rule set R Ensure: Verdict v, confidence p, behavioral indicators B 1: pe ← ParsePE(f ) 2: (A, M ) ← ExtractImports(pe) ▷ A: ordered API list, M : metadata (entropy, ordinals, etc.) 3: A ← ResolveOrdinals(A) 4: s ← BuildSequence(A, V) 5: x ← NormalizeSequence(s, L) ▷ pad/truncate to fixed length L 6: p ← CNNInfer(x) ▷ maliciousness probability 7: B ← InterpretBehaviors(A, R) ▷ map APIs to capability categories + severity 8: if p ≥ τ then 9: v ← "Malware" 10: else 11: v ← "Benign" 12: end if 13: return (v, p, B)

– Dynamic Loading Indicators: The presence of APIs associated with runtime function resolution, which may enable evasion of import-based detection. – API Count: The total number of extracted APIs, reflecting the richness of the available static feature set. The output of the static feature extraction stage consists of a fixed-length numerical representation of the API sequence for machine learning inference, together with auxiliary metadata describing the extraction process, including entropy values, ordinal resolution statistics, and dynamic loading indicators. Feature extraction is fully decoupled from classification, preserving modularity and facilitating future extensions.

3.3

Malware Classification Model

This subsection describes the malware classification model used in Delphi Scanner to estimate the maliciousness of a given PE file, corresponding to the CNNInfer step (line 6) of Algorithm 1. The task is formulated as a binary classification problem in which each input sample is represented as an ordered sequence of Windows API identifiers, as described in Section 3.2. The model produces a probabilistic estimate of maliciousness p ∈ [0, 1], where p denotes the softmax output probability that the input belongs to the malware class (i.e., p = softmax(logits)[1]). This score is compared against a decision threshold τ (default τ = 0.5): samples with p ≥ τ are classified as malware, and those with p < τ as benign, as detailed in Algorithm 1.

Title Suppressed Due to Excessive Length

9

Fig. 2. Architecture of the 1D CNN classifier, combining multi-scale convolutional filters (kernels 3, 5, 7) with global pooling to produce a fixed-length representation for binary malware classification.

Model Architecture Delphi Scanner employs a convolutional neural network (CNN) selected for its favorable balance between detection performance and computational efficiency in static API sequence modeling. The deployed architecture, shown in Figure 2, processes API sequences of fixed length 512. Each API identifier is mapped to a 128-dimensional embedding, yielding an embedded sequence of shape (512, 128). The embedded representation is passed through a multi-scale one-dimensional convolutional module comprising parallel convolutional kernels of sizes 3, 5, and 7, with 128, 128, and 64 filters, respectively, enabling the extraction of local patterns at multiple temporal scales. The resulting feature maps are concatenated and further processed by two additional Conv1D layers with ReLU activation and dropout for regularization. Global max pooling and global average pooling are applied to obtain a fixed 256-dimensional representation, which is subsequently fed into a fully connected layer followed by a final linear layer to produce the logits for binary classification.

Model Deployment For efficient and portable inference, the trained CNN model is exported to the Open Neural Network Exchange (ONNX) format and executed using the ONNX Runtime. This deployment approach decouples model training from inference and enables direct execution within the Rust-based backend without reliance on heavyweight machine learning frameworks. As a result, Delphi Scanner maintains a compact runtime environment while achieving submillisecond inference latency on standard hardware, supporting responsive malware analysis in desktop deployment scenarios.

10

Brahimi et al.

Table 1. Behavioral categories, trigger conditions, severity levels, descriptions, associated API indicators, and MITRE ATT&CK identifiers implemented in the interpretation layer. Category

Trigger Condition

Severity Description

Process Injection ≥2 APIs Critical matched

Credential cess

Ac- ≥1 API Critical matched

Keylogging

≥2 APIs Critical matched

Registry Persis- ≥1 API High tence matched Network Com- ≥2 APIs High munication matched

Service Manipu- ≥2 APIs High lation matched

Anti-Debugging ≥1 API Medium matched

3.4

API Indicators

ATT&CK ID

T1055 APIs for code injection into VirtualAllocEx, other processes WriteProcessMemory, CreateRemoteThread, NtCreateThreadEx T1555 APIs for stealing credentials CredEnumerateA/W, CryptUnprotectData, CryptDecrypt APIs for keyboard monitoring SetWindowsHookExA/W, T1056.001 GetAsyncKeyState, GetKeyboardState APIs for establishing persis- RegSetValueExA/W, T1547.001 tence via registry RegCreateKeyExA/W APIs for network/C2 commu- WSAStartup, socket, T1071 nication connect, send, recv, InternetOpenA/W, URLDownloadToFileA/W APIs for Windows service ma- CreateServiceA/W, T1543.003 nipulation OpenServiceA/W, StartServiceA/W, ControlService APIs to detect or evade de- IsDebuggerPresent, T1622 bugging CheckRemoteDebuggerPresent, NtQueryInformationProcess

Behavioral Interpretation Layer

The behavioral interpretation layer provides explanatory context for malware detection results derived from the MITRE ATT&CK framework [22] and grounded in established malware analysis literature [5]. Each category is mapped to a corresponding technique identifier, as listed in the table above. Its role operates in parallel with the malware classification pipeline; it does not influence the final malware verdict and serves exclusively as an explanatory component. This layer maps observed API functions to predefined behavioral categories, each representing a class of malicious capability observed in real-world malware. Table 1 summarizes the implemented categories, their trigger conditions, and the corresponding API patterns used to identify them. Each category is activated when a minimum number of representative APIs is detected. Per-category thresholds are calibrated to reduce false activations caused by incidental API usage common in legitimate software. A qualitative severity level (e.g., High or Medium) is then assigned based on the sensitivity of the matched capability.

Title Suppressed Due to Excessive Length

11

Table 2. Technology stack used in the implementation of Delphi Scanner.

Layer

Technology Justification React 19 Reactive components, large ecosystem Frontend TypeScript Static typing, improved maintainability Tailwind CSS v4 Utility-first styling, modern UI Rust 1.70+ Memory safety, native performance Tauri v2 Lightweight desktop framework Backend Goblin Pure Rust PE parsing, no C FFI ONNX Runtime 1.22 Optimized inference, GPU support Python 3.10+ Mature ML ecosystem Training PyTorch 2.0 Flexible training and debugging scikit-learn Preprocessing and metrics 3.5

User Interface

The user interface serves as the interaction layer between the analyst and the static analysis backend, aggregating the outputs of malware classification and behavioral interpretation into a unified view. For each analyzed PE file, the interface displays the malware verdict and associated confidence score, along with relevant static analysis metadata such as file characteristics, entropy values, and API counts. Behavioral threat indicators produced by the interpretation layer are grouped by category and visually differentiated by severity to support rapid malware triage. All analysis is performed locally, and the interface operates without reliance on external services, preserving user privacy and enabling lowlatency interaction. 3.6

Implementation Overview

Delphi Scanner is implemented as a lightweight desktop application following a modular client–backend architecture. The backend performs static analysis, feature extraction, machine learning inference, and behavioral interpretation, while the frontend provides an interface for initiating analysis and presenting results. The technology stack is summarized in Table 2. The backend is implemented in Rust to ensure performance and memory safety, and the frontend is developed using React and TypeScript with Tailwind CSS for consistent styling. Frontend and backend components are integrated using the Tauri framework, which enables secure communication between web-based interfaces and native system functionality. Communication between the frontend and backend is performed via a commandbased interface, supporting structured analysis requests and dynamic rendering of results. All analysis is executed locally without dependence on external services, ensuring privacy preservation and low-latency operation. This modular design decouples analysis logic from presentation concerns, improving maintainability and facilitating future extensions, including the integration of additional

12

Brahimi et al.

static features, alternative classification models, or enhanced visualization components. The complete source code of the application, including backend, frontend, and model deployment logic, is publicly available as an open-source repository3 .Figure 3 consolidates the complete Delphi Scanner pipeline across both offline training and deployment-time inference stages described in this section. Shared Preprocessing Raw PE File

Validate & Parse PE

Extract IAT & Resolve Ordinals

Ordered API Sequence

Vectorize & Pad (L=512)

Offline Training Train / Test Split (80/20)

Train CNN (PyTorch)

Export to ONNX

Deployment-Time Inference Verdict + Confidence + Indicators

Behavioral Rule Engine

ONNX Runtime Inference

dashed = secondary data flow

Fig. 3. End-to-end pipeline of Delphi Scanner: shared preprocessing (blue), offline training (orange), and deployment-time inference (green). Dashed arrows indicate secondary data flows to the behavioral rule engine and classifier input.

4

Experiments

4.1

Dataset Description and Data Preprocessing

This study employs a large-scale dataset of Windows Portable Executable (PE) files constructed by combining a widely used public reference dataset with additional complementary sources. The goal of this aggregation is to increase sample diversity, reduce dataset bias, and improve the robustness of the trained models. Dataset Description The primary data source is the PE Malware Machine Learning Dataset released by Practical Security Analytics [25], which contains real-world PE binaries labeled as malicious or benign. The dataset includes executables and dynamic-link libraries collected from multiple sources and time periods, with samples provided as complete binaries rather than pre-extracted features, enabling flexible static analysis. 3

https://github.com/0xf1d0/delphi_scanner

Title Suppressed Due to Excessive Length

13

Table 3. Dataset characteristics. Datasets Sources Source Malware Benign Total PE Malware ML Dataset 114,737 86,812 201,549 VirusShare 20,000 0 20,000 MalBehavD-V1 8,000 2,000 10,000 Windows System32 0 5,000 5,000 Combined Dataset Before and After Deduplication Source Malware Benign Total Before deduplication 142,737 93,812 236,549 After deduplication 120,000 80,000 200,000

To enrich the dataset and improve coverage of recent malware families and clean system binaries, additional data sources were integrated. Malware samples were obtained from publicly available repositories, including VirusShare [29], while benign samples were collected from legitimate Windows system binaries. In addition, a curated dataset containing pre-extracted API sequences was incorporated to increase variability in observed malware behaviors [21]. The composition of all data sources and the resulting dataset statistics before and after deduplication are summarized in Table 3. After aggregation, the combined dataset was deduplicated using SHA-256 cryptographic hashes to remove identical binaries and prevent data leakage between training and evaluation phases. This step reduced the dataset to 200,000 unique samples, as reported in Table 3. Preprocessing Pipeline All PE files were processed using the static feature extraction pipeline described in Section 3.2. Imported Windows API functions were extracted, ordinal-based imports were resolved when possible, and API sequences were transformed into fixed-length numerical representations suitable for machine learning inference. Samples with corrupted PE structures and insufficient API coverage were excluded. During preprocessing, an additional 9,636 samples were removed, resulting in a final dataset of 190,364 samples. The remaining samples were partitioned into training and test subsets using a fixed 80/20 split. Sequence normalization and dataset partitioning were performed prior to model training, with both subsets maintaining a balanced distribution of malicious and benign files. 4.2

Algorithms

To assess the effectiveness of different sequence modeling approaches for static malware detection, multiple deep learning architectures were implemented and evaluated in this study. All models operate on the same fixed-length API sequence representation described in Section 3.2 and are trained to perform binary classification, distinguishing malicious from benign PE files.

14

Brahimi et al. Table 4. Common training configuration used for all evaluated models. Hyperparameter Value Batch size 64 Maximum epochs 50 Initial learning rate 1 × 10−3 Weight decay 1 × 10−4 Optimizer AdamW Scheduler ReduceLROnPlateau Early stopping patience 5 epochs

Evaluated Architectures The evaluated models represent a range of sequence learning paradigms commonly employed in malware detection and sequential data analysis: – Convolutional Neural Networks (CNN): architectures designed to capture local patterns in API sequences using one-dimensional convolutional filters. – Long Short-Term Memory (LSTM) networks: recurrent models capable of modeling long-range dependencies in sequential data. – Gated Recurrent Units (GRU): a computationally efficient alternative to LSTMs with fewer parameters. – Transformer-based models: attention-driven architectures that model global dependencies across the entire API sequence. – Hybrid CNN–LSTM architectures: models that combine convolutional layers for local feature extraction with recurrent layers for sequential dependency modeling. These architectures were selected to provide a comprehensive comparison of convolutional, recurrent, attention-based, and hybrid approaches.

Training Configuration To ensure a fair and reproducible comparison, all models were trained using a unified training configuration. Hyperparameters were selected empirically based on convergence behavior, training stability, and validation performance, rather than exhaustive grid search, reflecting practical constraints commonly encountered in applied security research. The shared training configuration is summarized in Table 4. All models were trained using cross-entropy loss with label smoothing to improve generalization and reduce prediction overconfidence. Training was terminated early when validation performance ceased to improve, ensuring efficient convergence while avoiding unnecessary overfitting. Although the models share identical input representations and training procedures, they differ in architectural complexity, parameter count, and computational characteristics, enabling a meaningful comparative analysis.

Title Suppressed Due to Excessive Length

4.3

15

Evaluation Metrics

All models were trained and evaluated under identical experimental conditions to ensure a fair and reproducible comparison. Specifically, the same dataset splits, preprocessing pipeline, and API sequence representations were used across all evaluated architectures. Model performance was assessed exclusively on a heldout test set that was not used during training. Detection effectiveness was evaluated using standard binary classification metrics commonly adopted in malware detection research: – Accuracy: defined as the proportion of correctly classified samples. – Precision: measuring the fraction of samples predicted as malicious that are truly malicious. – Recall: measuring the fraction of malicious samples correctly identified. – F1-score: computed as the harmonic mean of precision and recall. In addition to classification performance, system-level metrics relevant to practical deployment were also measured. Inference latency was calculated as the average time required to process a single PE file during model execution, while model size was recorded to assess memory footprint. This evaluation protocol ensures that selected models are not only effective in detecting malware, but also feasible for real-world usage within the constraints of a local analysis environment. 4.4

Implementation

All experiments were implemented and executed in a controlled local environment to ensure reproducibility and consistent performance measurements. Model training and evaluation were conducted on a single workstation using a modern multi-core CPU and a dedicated GPU to accelerate deep learning workloads. The experimental setup is summarized in Table 5. The system is equipped with an AMD Ryzen 7 5800H processor and an NVIDIA RTX 3060 Mobile GPU with dedicated VRAM, providing sufficient computational resources for training deep learning models on large-scale PE datasets. The operating system used was Windows 11 Pro, which aligns with the target deployment environment of the proposed malware analysis tool. Model training was implemented using the PyTorch deep learning framework, which was selected for its flexibility, extensive ecosystem, and ease of experimentation. CUDA support was enabled to leverage GPU acceleration during training, significantly reducing training time for recurrent and attention-based models. All preprocessing steps, including API sequence extraction and vectorization, were performed prior to training to minimize runtime overhead during model optimization. Inference experiments were conducted using the trained models exported to the ONNX format, enabling efficient execution through the ONNX Runtime. This separation between training and inference environments reflects the intended system design, where models are trained offline and deployed within a lightweight desktop application for real-time analysis.

16

4.5

Brahimi et al.

Robustness Evaluation

To assess the robustness of the trained CNN model beyond the held-out test set, three complementary evaluation experiments were conducted, targeting generalization to unseen malware, resilience to executable packing, and resistance to adversarial manipulation of PE binary structures. All experiments were performed using the deployed ONNX model without retraining or threshold adjustment. Generalization to Unseen Malware To evaluate the model’s ability to generalize beyond the training distribution, a set of 5,647 malware samples was collected from MalwareBazaar [1], a publicly accessible malware repository aggregating recent and actively distributed threats. All samples were processed through the same static feature extraction pipeline. For this experiment, recall serves as the primary evaluation metric since all collected samples are labeled malicious. UPX Packing Robustness To evaluate the impact of executable packing on detection performance, 106 UPX-packed malware samples were collected and subjected to automated unpacking using the UPX decompressor, yielding a paired set of 106 packed and 106 unpacked binaries derived from identical malware instances. Both variants of each sample were independently analyzed through the static feature extraction pipeline. Adversarial Manipulation Robustness To assess resistance to functionalitypreserving adversarial perturbations of PE files, 300 adversarial samples were generated from malicious samples across three attack techniques, with 100 samples per technique. The evaluated techniques are as follows. API Benign Injection appends legitimate, non-malicious API imports to the malware’s import table to dilute the learned malicious API patterns without altering program functionality. Full DOS Padding (FullDOS) extends the DOS header stub region of the PE file with arbitrary byte content, inflating the binary without modifying executable code or imports. Section Padding appends benign byte sequences to existing PE sections, increasing file size while preserving the executable’s original import structure and behavior.

5

Results and Analysis

This section presents the experimental results obtained using the evaluation protocol described in the previous section. The objective is to assess the detection performance, class-wise behavior, and deployment feasibility of the evaluated models under identical experimental conditions.

Title Suppressed Due to Excessive Length

17

Table 5. Hardware and software configuration used for experiments. Component Specification Processor AMD Ryzen 7 5800H (8 cores) GPU NVIDIA RTX 3060 Mobile (6 GB VRAM) RAM 16 GB DDR4 Operating System Windows 11 Pro CUDA Version 12.6 PyTorch Version 2.0+ Table 6. Comparative benchmark of the evaluated architectures on the test set. Model Accuracy (%) F1-score Inference Time (ms) Model Size (MB) CNN 95.35 0.9536 0.56 1.53 GRU 95.23 0.9524 0.30 1.79 Transformer 95.04 0.9505 1.39 1.60 LSTM 94.95 0.9496 0.33 1.98 CNN–LSTM 94.82 0.9483 0.89 1.72

5.1

Overall Detection Performance

The overall detection performance of the evaluated architectures is summarized in Table 6. All models achieve high accuracy and F1-score on the held-out test set, confirming that static API sequence–based analysis is effective for malware detection. Among the evaluated approaches, the CNN-based model achieves the best trade-off between detection performance and efficiency. It achieves the highest accuracy (95.35%) and F1 Score (0.9536), while maintaining a compact model size and low inference latency. Recurrent and transformer-based models exhibit comparable accuracy but incur higher computational overhead, resulting in increased inference time and larger memory footprints. Figure 4 illustrates the complete training history of the evaluated models in terms of accuracy and loss. The CNN model demonstrates stable and rapid convergence, reaching optimal performance in fewer epochs compared to recurrent and transformer-based architectures. This behavior further supports the suitability of convolutional models for modeling local API patterns in static malware detection. Recurrent and transformer-based models exhibit slower convergence and higher variance during training, indicating increased optimization complexity compared to convolutional architectures. 5.2

Per-Class Performance

To further analyze the behavior of the selected CNN model, class-wise performance metrics were examined. Table 7 reports precision, recall, and F1-score for both benign and malicious classes. The CNN model achieves balanced performance across both classes, with high precision and recall values for malware and benign samples alike. This balance indicates that the model does not exhibit a

18

Brahimi et al.

Fig. 4. Validation accuracy and loss over training epochs for all evaluated architectures. Table 7. Per-class performance metrics of the CNN model on the test set. Class Precision Recall F1-score Support Benign (0) 0.9478 0.9602 0.9540 19,037 Malware (1) 0.9595 0.9468 0.9531 19,036 Macro Avg 0.9537 0.9535 0.9536 38,073 Weighted Avg 0.9537 0.9535 0.9536 38,073

strong bias toward either class, limiting both false alarms and undetected malware, which is particularly important in malware detection scenarios where both false positives and false negatives can have practical consequences. 5.3

Inference Efficiency and Model Size

In addition to detection accuracy, inference latency and model size were evaluated to assess the practicality of deploying the models in a desktop malware analysis tool. As shown in Table 6, the CNN model achieves sub-millisecond inference time while maintaining a small memory footprint. Compared to recurrent and transformer-based architectures, the CNN model offers faster inference and reduced computational overhead, making it well-suited for interactive malware analysis scenarios. These characteristics are particularly important for local, real-time analysis environments where responsiveness and resource efficiency are critical. 5.4

Error Analysis

To further understand the behavior of the evaluated models beyond aggregate performance metrics, an error analysis was conducted focusing on false positive (FP) and false negative (FN) outcomes. In the context of malware detection, false negatives—malicious samples incorrectly classified as benign—are generally considered more critical than false positives, as they represent undetected threats that may execute without warning. Table 8 summarizes the false positive and false negative counts and rates for all evaluated architectures. Across

Title Suppressed Due to Excessive Length

19

Table 8. Error analysis by model showing false positive (FP) and false negative (FN) counts and rates. Model CNN GRU Transformer LSTM CNN–LSTM

FP 758 782 815 836 867

FN FP (%) FN (%) 1,013 2.0 2.7 1,032 2.1 2.7 1,073 2.1 2.8 1,086 2.2 2.9 1,105 2.3 2.9

all models, false negative rates are consistently higher than false positive rates, reflecting the inherent difficulty of identifying stealthy or weakly represented malicious behaviors using static analysis alone. Among the evaluated approaches, the CNN-based model achieves the lowest false positive rate (2.0%) and a comparatively low false negative rate (2.7%), indicating a favorable balance between detection sensitivity and misclassification risk. Recurrent and attention-based models, including LSTM, GRU, and Transformer architectures, exhibit slightly higher false negative rates, suggesting a reduced ability to generalize across diverse malware behaviors in this setting. Hybrid CNN–LSTM models show the highest error rates among the evaluated architectures, likely due to increased model complexity without corresponding gains in discriminative capability. Overall, the error analysis reinforces the suitability of the CNN-based architecture for deployment in security-sensitive environments, where minimizing undetected malware while maintaining acceptable false alarm rates is essential. 5.5

Performance Trade-off Analysis

The results show a clear trade-off between model performance and computational efficiency. The CNN achieves the highest accuracy (95.35%) and the lowest false positive rate (2.0%), while maintaining a compact footprint (1.53 MB) and sub-millisecond inference latency, making it suitable for lightweight malware detection. In contrast, LSTM and GRU models require 1.5–3× more memory and yield higher false negative rates despite their ability to capture sequential dependencies. The Transformer model exhibits the highest inference cost (1.39 ms), approximately 2.5× slower than the CNN, without performance gains, indicating that local API co-occurrence patterns are more informative than global dependencies in static analysis. The behavioral interpretation layer introduces negligible overhead due to its rule-based set-membership lookups and operates independently of the neural pipeline. As a result, explainability is effectively cost-free at inference time compared to post-hoc methods such as SHAP or LIME. The main runtime bottleneck is PE parsing and I/O, which takes approximately 50 ms per sample—several orders of magnitude higher than inference time—and can be mitigated through parallel processing. Training time scales approximately linearly with dataset size,

20

Brahimi et al.

with the full training process completing in about 11 minutes on an RTX 3060, supporting periodic retraining as new malware families emerge.

5.6

Robustness Evaluation Results

The results of the three robustness evaluation experiments are summarized in Table 9. Overall, the CNN model demonstrates consistent resilience across all evaluated conditions, maintaining high recall. Table 9. Robustness evaluation results across generalization, packing, and adversarial manipulation experiments. Experiment Generalization

Condition Recall (%) MalwareBazaar (5,647 samples) 89.13 Packed (106 samples) 92.08 UPX Packing Unpacked (106 samples) 97.16 API Benign Injection 95.00 Adversarial Manipulation Full DOS Padding 97.85 Section Padding 97.01

Generalization to Unseen Malware. The model achieves a recall of 89.13% on the MalwareBazaar sample set, demonstrating meaningful generalization to malware samples outside the training distribution. The observed reduction relative to the test-set malware recall of 94.68% is consistent with the expected effect of distribution shift, as MalwareBazaar aggregates recently submitted and actively distributed threats that may employ API usage patterns underrepresented in the training data. Nevertheless, the result confirms that the model captures generalizable malicious API patterns rather than overfitting to dataset-specific artifacts. UPX Packing Robustness. The paired packing experiment reveals a measurable but moderate impact of UPX packing on detection recall. The model achieves 97.16% recall on unpacked binaries, consistent with performance on the main test set, while recall on the corresponding packed variants decreases to 92.08%, a reduction of 5.08 percentage points. This decline is attributable to the reduced import table visibility in packed executables, where the IAT exposes only the unpacking stub rather than the full malicious API sequence. Despite this degradation, the model retains substantial detection capability on packed samples, suggesting that packing-stub API patterns carry residual discriminative signal.

Title Suppressed Due to Excessive Length

21

Adversarial Manipulation Robustness. The model demonstrates strong resistance to all three evaluated adversarial manipulation techniques. Recall remains high under API Benign Injection (95.00%), Full DOS Padding (97.85%), and Section Padding (97.01%), with no technique reducing recall below 95%. These results indicate that the CNN classifier does not rely on superficial binarylevel features susceptible to padding-based perturbations, and that the injection of benign API imports produces only a limited dilution effect on the learned malicious sequence representations. The relative robustness to API Benign Injection is particularly noteworthy, as this technique directly targets the feature space exploited by the model, yet results in only a modest recall reduction of 0.35 percentage points compared to the original test set performance. 5.7

Summary of Results

Overall, the experimental results demonstrate that convolutional architectures provide an effective and efficient solution for static malware detection based on API sequences. The CNN model achieves strong detection performance, balanced class-wise behavior, and low inference latency, validating its selection as the deployed classifier in the proposed system.

6

Case Study: Delphi Scanner Output

To complement the quantitative evaluation presented in the previous section, we provide a qualitative case study illustrating how Delphi Scanner operates in practice and the type of outputs it produces for analysts. The objective of this case study is to demonstrate the system’s end-to-end behavior, including static feature extraction, machine learning–based classification, and rule-based behavioral interpretation, as exposed through the user interface. Three representative screenshots are used to illustrate (i) the main interface, (ii) the analysis of a benign executable, and (iii) the analysis of a malicious sample. 6.1

User Interface and System Overview

Figure 5 presents the main user interface of Delphi Scanner. The application is designed as a lightweight desktop tool supporting fully offline analysis of Windows PE files. Upon launch, the interface displays system-level information, including the model version, API vocabulary size, maximum sequence length, and model load time. This metadata provides transparency regarding the underlying detection engine and reinforces the system’s emphasis on efficiency and local deployment. The interface supports a simple drag-and-drop workflow for PE files, enabling rapid analysis without requiring external services or sandbox execution. All results are computed locally and displayed within a single consolidated view, combining the malware verdict, confidence score, static file metadata, and behavioral indicators.

22

Brahimi et al.

Fig. 5. Delphi Scanner main user interface.

6.2

Analysis of a Benign File

Figure 6 shows the analysis results for a legitimate executable (ChromeSetup.exe). The classifier assigns a benign verdict with a relatively low confidence score (41.4%), reflecting the model’s calibrated uncertainty rather than an overconfident prediction. This behavior is intentional, as benign software may legitimately exhibit characteristics that overlap with malware, such as dynamic API loading or elevated entropy due to compression. The analysis details reveal several noteworthy static indicators: elevated entropy suggesting possible packing, detection of dynamic loading APIs, and the presence of anti-debugging–related functions. These indicators are surfaced in the Threat Indicators section as a medium-severity warning. Importantly, these behavioral flags do not override the machine learning verdict; instead, they provide contextual information to the analyst. This example illustrates how Delphi Scanner avoids false positives by decoupling behavioral interpretation from classification, while still exposing potentially relevant technical signals. 6.3

Analysis of a Malicious File

Figure 7 illustrates the analysis of a malicious PE file. In this case, the classifier produces a malware verdict with high confidence (94.3%), indicating strong alignment between the observed API sequence and learned malicious patterns. The static analysis details show a higher number of extracted APIs, elevated entropy consistent with packing, unresolved ordinals, and the presence of unknown APIs, all of which are commonly associated with malicious binaries. The behavioral interpretation layer reinforces the classification outcome by highlighting suspicious characteristics, including dynamic loading and packing

Title Suppressed Due to Excessive Length

23

indicators. The combination of a high-confidence verdict and multiple corroborating behavioral signals enables rapid triage and prioritization without requiring execution or dynamic analysis. This example demonstrates how Delphi Scanner supports analyst decision-making by providing both an automated judgment and interpretable evidence. 6.4

Summary of Case Study Findings

This case study demonstrates that Delphi Scanner effectively integrates efficient static malware detection with interpretable, analyst-oriented output. Benign files exhibiting suspicious but legitimate characteristics are handled conservatively, while malicious samples are identified with high confidence and supported by clear behavioral indicators. These examples validate the system’s design goals of transparency, low-latency analysis, and practical usability, complementing the quantitative results reported earlier.

7

Discussion and Limitations

This section interprets the experimental findings in the context of the system’s design objectives, discusses the boundaries of the proposed approach, and outlines directions for future research. 7.1

Robustness and Evasion Considerations

The experimental results demonstrate that the CNN-based classifier achieves the best trade-off among all evaluated architectures, combining the highest accuracy (95.35%) with the smallest model size (1.53 MB) and the lowest inference latency (0.56 ms). This confirms that local API co-occurrence patterns are sufficiently discriminative for static PE malware detection, and that the additional complexity of transformer or recurrent architectures does not translate into measurable gains in this setting. The robustness evaluation reveals that the model generalizes meaningfully beyond its training distribution, retains strong detection capability under packinginduced feature degradation, and resists all three evaluated adversarial manipulation strategies. The most practically significant finding is the recall gap between packed and unpacked samples, which reflects the reduced import table visibility caused by packing rather than a fundamental weakness in the learned representations. Resistance to API benign injection is particularly noteworthy, as this attack directly targets the model’s feature space yet produces only a modest recall reduction, suggesting that the classifier captures discriminative local patterns that remain robust to benign noise injection. 7.2

Limitations of Static Analysis

Static inspection is inherently constrained to metadata present in the binary at analysis time. Packed executables expose only unpacking stub imports, limiting

24

Brahimi et al.

feature coverage, as confirmed empirically by the 5.08 percentage point recall gap observed between packed and unpacked variants. Runtime behaviors such as dynamically resolved APIs, downloaded payloads, and self-modifying code remain unobservable without execution. These constraints are structural to static analysis and represent a deliberate trade-off against the speed, safety, and privacy benefits of local offline analysis. Consequently, malware relying exclusively on dynamic API resolution at runtime remains outside the scope of import-based static analysis and constitutes the most significant open robustness challenge. 7.3

Dataset Bias and Generalization

The 89.13% recall on MalwareBazaar samples confirms meaningful generalization to out-of-distribution threats, while the gap relative to in-distribution performance reflects the expected effect of distributional shift toward newer malware families [9]. Nevertheless, the training dataset remains subject to temporal and family-level biases, and a temporally stratified evaluation training on earlier samples and testing on a later time window, remains an important direction for future work. Addressing longer-term generalization would require continuous dataset updates, broader malware family coverage, and potentially the integration of complementary behavioral signals. 7.4

Practical Deployment Relative to Existing Solutions

Delphi Scanner is not intended to replace full-featured antivirus products or cloud-based malware scanning services. Instead, it occupies a complementary operational niche: fast, privacy-preserving, and interpretable pre-execution screening. Its lightweight ONNX backend makes it architecturally extensible to integration within broader security pipelines, including browser-based download screening, antivirus pre-filtering, and endpoint detection and response (EDR) platforms. In a layered defense architecture, it can serve as an efficient firstpass triage component, feeding high-confidence verdicts to downstream dynamic analysis or YARA rule engines. These integration scenarios are enabled by the system’s sub-millisecond inference latency and its fully offline, modular design. The inherent trade-offs of this approach — namely the absence of real-time protection, signature-based detection, and continuous threat intelligence updates — reflect a deliberate focus on transparency and deployability rather than comprehensive threat coverage. Beyond standalone desktop deployment, Delphi Scanner’s modular architecture supports broader organizational integration scenarios. It can be deployed on internal servers, Network Attached Storage (NAS) systems, or enterprise file-sharing platforms to perform automated pre-execution screening of executable files, limiting the lateral movement of malware within organizational networks. Integration with Cloud Access Security Brokers (CASBs) further extends this capability to monitor executable files shared through cloudbased services, while preserving the system’s lightweight footprint and offline inference design.

Title Suppressed Due to Excessive Length

7.5

25

Future Work

Several directions emerge naturally from the current limitations. First, the integration of dynamic analysis as a complementary stage would address the most significant gap of the current system: the inability to observe runtime behaviors such as dynamically resolved APIs, downloaded payloads, and self-modifying code. In a layered architecture, Delphi Scanner could serve as a fast pre-screening filter, escalating uncertain or high-entropy samples to a dynamic analysis sandbox, thereby combining the efficiency of static analysis with the behavioral depth of dynamic execution. Second, integration of lightweight unpacking heuristics would improve detection coverage for packed executables, directly addressing the recall gap observed between packed and unpacked variants in the robustness evaluation. Third, a temporally stratified evaluation protocol — training on earlier samples and testing on a later time window-would provide rigorous empirical insight into concept drift and real-world generalization to emerging malware families, addressing a limitation acknowledged in the current dataset composition. Fourth, a structured analyst study measuring triage efficiency with and without the behavioral interpretation layer would provide quantitative validation of its practical utility beyond the qualitative case study presented in Section 6. Finally, future research must address the open challenges associated with dynamic malware detection, including fileless malware that resides entirely in memory, AI-powered malware capable of adapting its behavior to evade detection, and adversarial threats such as data poisoning attacks that gradually degrade model accuracy. Privacy and ethical concerns surrounding the execution of untrusted code further complicate dynamic analysis in practice. Hybrid approaches combining static and dynamic analysis represent a promising direction, though their effective integration remains a technically demanding endeavor.

8

Conclusion

This paper presented Delphi Scanner, a static malware detection system for Windows PE files that combines deep learning-based API sequence classification with interpretable behavioral analysis. Experimental results on a large and diverse dataset show that convolutional neural networks provide an effective balance between detection performance and computational efficiency for static API sequence modeling, achieving strong accuracy with a compact model and sub-millisecond inference latency. The decoupled behavioral interpretation layer complements the statistical classifier by translating low-level API evidence into human-readable capability indicators without influencing classification outcomes. Overall, Delphi Scanner demonstrates that static malware detection can be both accurate and interpretable when detection and explanation are explicitly separated, providing a practical foundation for privacy-preserving, analystoriented malware triage.

26

Brahimi et al.

Fig. 6. Analysis output for a benign PE file.

Title Suppressed Due to Excessive Length

Fig. 7. Analysis output for a malicious PE file.

27

Bibliography

[1] Abuse.ch: MalwareBazaar: A Project from Abuse.ch to Share Malware Samples with the Infosec Community. https://bazaar.abuse.ch (2024), accessed: 2024 [2] Baker del Aguila, R., Contreras Pérez, C.D., Silva-Trujillo, A.G., CuevasTello, J.C., Nunez-Varela, J.: Static malware analysis using low-parameter machine learning models. Computers 13(3), 59 (2024). https://doi.org/ 10.3390/computers13030059 [3] Alshomrani, S., Alqurashi, F., Abozinadah, E., Mehmood, R.: Survey of transformer-based malicious software detection systems. Electronics 13(23), 4677 (2024). https://doi.org/10.3390/electronics13234677 [4] Anderson, H.S., Roth, P.: Ember: An open dataset for training static pe malware machine learning models. arXiv preprint arXiv:1804.04637 (2018), https://arxiv.org/abs/1804.04637 [5] Arp, D., Quiring, E., Pendlebury, F., Warnecke, A., Pierazzi, F., Wressnegger, C., Cavallaro, L., Rieck, K.: Dos and don’ts of machine learning in computer security. In: 31st USENIX Security Symposium (USENIX Security 22). pp. 3971–3988 (2022) [6] AV-TEST Institute: Malware statistics & trends report. AV-TEST Research Report (2024), https://www.av-test.org/en/statistics/malware/ [7] Badis, H., Doyen, G., Khatoun, R.: A collaborative approach for a source based detection of botclouds. in 2015 ifip. In: IEEE International Symposium on Integrated Network Management (IM). pp. 906–909 [8] Berjawi, O., El Attar, A., Chbib, F., Khatoun, R., Fahs, W.: Cyberattacks detection through behavior analysis of internet traffic. Procedia Computer Science 224, 52–59 (2023) [9] Berjawi, O., Khatoun, R., Fenza, G.: Evaluating the generalization of machine learning and deep learning models for intrusion detection systems. In: 2025 International Conference on INnovations in Intelligent SysTems and Applications (INISTA). pp. 1–6. IEEE (2025) [10] Cesare, S., Xiang, Y.: Classification of malware using structured control flow. Proceedings of the 8th Australasian Symposium on Parallel and Distributed Computing (2014) [11] Coscia, A., Lorusso, R., Maci, A., Urbano, G.: Apiary: An api-based automatic rule generator for yara to enhance malware detection. Computers & Security 153, 104397 (2025). https://doi.org/10.1016/j.cose.2025. 104397 [12] Doshi-Velez, F., Kim, B.: Towards a rigorous science of interpretable machine learning. arXiv preprint arXiv:1702.08608 (2017) [13] Egele, M., Scholte, T., Kirda, E., Kruegel, C.: A survey on automated dynamic malware-analysis techniques and tools. ACM Computing Surveys 44(2), 1–42 (2012). https://doi.org/10.1145/2089125.2089126

Title Suppressed Due to Excessive Length

29

[14] El Attar, A., Khatoun, R., Lemercier, M.: A gaussian mixture model for dynamic detection of abnormal behavior in smartphone applications. In: 2014 global information infrastructure and networking symposium (GIIS). pp. 1–6. IEEE (2014) [15] Geng, J., Wang, J., Fang, Z., Zhou, Y., Ge, W.: A survey of strategy-driven evasion methods for pe malware: Transformation, concealment, and attack. Computers & Security 137, 103595 (2024). https://doi.org/10.1016/j. cose.2023.103595 [16] Gibert, D., Mateu, C., Planes, J.: The rise of machine learning for detection and classification of malware: Research developments, trends and challenges. Journal of Network and Computer Applications 153, 102526 (2020) [17] Li, K., Guo, W., Zhang, F., Du, J.: Gambd: Generating adversarial malware against malconv. Computers & Security 130, 103279 (2023). https://doi. org/10.1016/j.cose.2023.103279 [18] Li, M.Q., Fung, B.C.M., Charland, P., Ding, S.H.H.: I-mad: Interpretable malware detector using galaxy transformer. Computers & Security 108, 102371 (2021). https://doi.org/10.1016/j.cose.2021.102371 [19] Lipton, Z.C.: The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery. Queue 16(3), 31–57 (2018) [20] Liu, L., Kuang, X., Liu, L., Zhang, L.: Defend against adversarial attacks in malware detection through attack space management. Computers & Security 141, 103841 (2024). https://doi.org/10.1016/j.cose.2024.103841 [21] Maniriho, P., Mahmood, A.N., Chowdhury, M.J.M.: Api-maldetect: Automated malware detection framework for windows based on api calls and deep learning techniques. Journal of Network and Computer Applications 218, 103704 (2023). https://doi.org/10.1016/j.jnca.2023.103704 [22] MITRE Corporation: Mitre att&ck® framework. https://attack.mitre. org/ (2026), accessed: 2026-06-16 [23] Owoh, N., Adejoh, J., Hosseinzadeh, S., Ashawa, M., Osamor, J., Qureshi, A.: Malware detection based on api call sequence analysis: A gated recurrent unit–generative adversarial network model approach. Future Internet 16(10), 369 (2024). https://doi.org/10.3390/fi16100369 [24] Perry, J., et al.: A sok of transformer-based malware analysis. arXiv preprint arXiv:2405.17190 (2024), https://arxiv.org/abs/2405.17190 [25] Practical Security Analytics: Pe malware machine learning dataset. https://practicalsecurityanalytics.com/ pe-malware-machine-learning-dataset/ (2021), accessed: 2024 [26] Raff, E., Barker, J., Sylvester, J., Brandon, R., Catanzaro, B., Nicholas, C.: Malware detection by eating a whole EXE. arXiv preprint arXiv:1710.09435 (2017), https://arxiv.org/abs/1710.09435 [27] Raff, E., Fleshman, W., Zak, R., Anderson, H.S., Filar, B., McLean, M.: Classifying sequences of extreme length with constant memory applied to malware detection. arXiv preprint arXiv:2012.09390 (2020), https: //arxiv.org/abs/2012.09390, also appears in AAAI 2021 (open-access toolchains often refer to this as MalConv2).

30

Brahimi et al.

[28] Saxe, J., Berlin, K.: Deep neural network based malware detection using two dimensional binary program features. In: Proceedings of the 10th International Conference on Malicious and Unwanted Software (MALWARE). pp. 11–20 (2015). https://doi.org/10.1109/MALWARE.2015.7413680 [29] VirusShare: VirusShare Malware Repository. https://virusshare.com/ (2024), accessed: 2024-01 [30] Ye, Y., Li, T., Adjeroh, D., Iyengar, S.S.: A survey on malware detection using data mining techniques. ACM Computing Surveys 50(3), 1–40 (2017). https://doi.org/10.1145/3073559 [31] Zou, B., Cao, C., Wang, L., Cheng, Y., Dang, C., Liu, Y., Sun, J.: Feature graph construction with static features for malware detection. arXiv preprint arXiv:2404.16362 (2024), https://arxiv.org/abs/2404.16362

Record · ID 978361 · SHA-256 e6e66077c2cdf247
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.