Adversarial Malware Generation in Linux ELF Binaries via Semantic-Preserving Transformations Lukáš Hrdonka1 , Martin Jureček1
a
1 Faculty of Information Technology, Czech Technical University in Prague, Thákurova 9, Prague, Czech Republic
arXiv:2604.22639v1 [cs.CR] 24 Apr 2026
{hrdonluk, martin.jurecek}@fit.cvut.cz
Keywords:
adversarial attacks, adversarial malware, machine learning, ELF, Linux.
Abstract:
Malware development and detection have undergone significant changes in recent years as modern concepts, such as machine learning, have been used for both adversarial attacks and defense. Despite intensive research on Windows Portable Executable (PE) files, there is minimal work on Linux Executable and Linkable Format (ELF). In this work, we summarize the academic papers submitted in this field and develop a new adversarial malware generator for the ELF format. Using a variety of metrics, we thoroughly evaluated our generator and achieved an Evasion Rate of 67.74 % while changing the confidence of the malware detector by −0.50 in the mean case for the dataset used. In our approach, we chose MalConv as the target classifier. Using this classifier, we found that the most successful modifications used strings typical of benign files as a data source. We conducted a variety of experiments and concluded that the target classifier appears sensitive to strings at any location within the executable file.
1
INTRODUCTION
IT security specialists, including academic researchers and industry professionals, have been battling against malware for decades. Over this time, the methods of malware deployment have evolved significantly (Cozzi et al., 2018). As a result, antimalware solutions must be continuously updated to keep pace with these advancements and effectively detect emerging threats. Within the scope of this discussion, our primary focus will be on the generation of adversarial malware samples. In this context, attackers modify binaries to create new versions that retain the original functionality but are designed to evade detection by the target model, leading to misclassification. This technique represents a growing challenge in the cybersecurity landscape, as it can undermine the effectiveness of defense mechanisms that rely on machine learning models for threat detection. However, creating perturbations in malware binaries is a complex task because even minor modifications to the binary code can make the executable non-functional or lead to undefined behavior (Kreuk et al., 2018). An adversarial malware generator operates under a key restrictive condition: the program’s functionala
https://orcid.org/0000-0002-6546-8953
ity must remain unchanged. Additionally, there is the optimization to consider: the goal is to maximize the probability that the manipulated malware is misclassified as a benign file by the target detection model. To the best of our knowledge, there is a larger number of academic papers regarding adversarial malware creation for the Windows platform (PE file format), while only a few works focus on the Linux ELF file format. However, with the massive growth of Linux usage in recent years (especially in highperformance computing, cloud services, or IoT), the research needs to investigate the possibilities of adversarial malware for the Linux ELF format. Consequently, this paper will focus mainly on the Linux ELF file format. The main contribution of this paper is the design of an adversarial malware generator targeting the ELF file format. Our approach is based on a genetic algorithm workflow that explores the modification space using 12 modification types and 7 data sources, increasing the diversity and effectiveness of generated samples. The work also addresses interpretability limitations of machine learning based detection through advanced logging that provides insight into the generation process. In addition, we propose evaluation metrics for adversarial generators that extend the basic evasion rate and enable a more comprehensive assessment of performance.
The rest of the paper is organized as follows. Related work done in adversarial attacks is discussed in Section 2. Section 3 then presents the methodology and metrics used in the experiments. The proposed generator is then described in Section 4. Section 5 presents the environment used for the experiments, whose results are then described and discussed in Section 6. At the end of the article, we outline our future work and conclude (Section 7).
2
RELATED WORK
Despite our focus on ELF binaries, we first highlight a few works done in the field of PE files as the research dominates for these. In (Kozák et al., 2024), the authors presented a black-box evasion attack using reinforcement learning algorithms. They defined ten modification types, including appending random benign content to the end of the file or to newly created sections, removing the digital certificates information, or increasing the timestamp. They reported an Evasion Rate of 53.84 % against the GBDT classifier, 11.41 % against MalConv, and an average Evasion Rate of 2.31 % against leading antivirus engines. The work of (Lucas et al., 2021) presented a method of raw byte modification of PE files. They defined two families of transformation types – inplace randomization (e.g., replacing instructions with their equivalent ones of the same length, or reassigning the registers within the same functions) and code displacement (the original code is altered with the jmp instruction that passes control to the displaced code). They achieved evasion success rates up to 85 % against commercial anti-viruses. The authors of (Quertier et al., 2022) presented MERLIN – Malware Evasion with Reinforcement LearnING. They defined 16 modification types of PE files, including adding of benign strings to the end of sections, adding import functions, or packing and unpacking the malicious file. Against MalConv, they achieved an almost 100 % Evasion Rate, mainly because they discovered that the action type called add section strings has a high cumulative score. They presented an Evasion Rate of 80 % against EMBER, and 70 % using a commercial antivirus. In (Song et al., 2022), the authors modeled the action selection problem using the multi-armed bandit (MAB) problem and developed MAB-Malware, which consists of two main modules – Binary Rewriter and Action Minimizer. In Binary Rewriter, they defined eight macro-actions, including appending benign content at the end of the binary, adding random bytes to the unused space at the end of the
section, or zeroing out several fields from PE files. They also implemented micro-actions, which follow some of the principles used in macro-actions but modify only a few bytes. Using Action Minimizer, they then remove unnecessary actions and replace macroactions with micro-actions to produce a sample with only minimal changes to the binary. They presented an Evasion Rate of 74.7 % using EMBER classifier, 97.72 % using MalConv, and 31.99–48.3 % using the top 3 antivirus software. In (Anderson et al., 2018), OpenAI Gym code called gym-malware was released. The authors defined several manipulation types, such as adding a function to the Import Address Table that is never used, changing existing section names, or removing signer information. They also trained their own gradient boosted decision model, and presented an Evasion Rate of 24 % against that model. Moreover, gymmalware was compared with MAB-malware in (Song et al., 2022). The work concluded that gym-malware achieved an Evasion Rate of 32.5 % against MalConv and 15 % against EMBER. While moving towards ELF file format, the authors of (Xue et al., 2024) defined four major types of ELF modification methods – add redundant data to the end of a binary file, modify the sections, add file import information, and encapsulate a file (the file content is converted into data part of Go code). They achieved a detection rate of 25 % for ClamAV and an Evasion Rate of 75.8 % against VirusTotal1 . The authors of (Kosikowski et al., 2023) developed a tool to modify ELF binary files using deep learning algorithms. They used several types of ELF alternations, including Header Alternation, Debug Alternation, Padding Alternation, or Dynamic Extension. They reported an Evasion Rate of 76.6 % against FireEyeNet and 8.4 % against MalConv. In (Ravi et al., 2025), the authors developed the tool called ADVeRL-ELF to generate adversarial ELF malware using reinforcement learning. They used ResNet18, an 18-layer CNN (Convolutional Neural Network), as a target classifier. They mainly target the executable section of ELF files by adding semantic NOPs (no-operation instructions). They evaluated their tool using the IoT malware dataset (binaries compiled for the ARM architecture) and achieved a success rate of 59.5 %. Still, they noted that their framework can be easily extended to support the x8664 architecture as well. To the best of our knowledge, these are the few academic papers that discuss the offensive point of view of adversarial malware for ELF format. However, there are other works, such as (Qiao et al., 2023), 1 https://www.virustotal.com/
or (Ramamoorthy et al., 2025), focusing primarily on the defensive aspect of adversarial malware. Furthermore, authors of (Guesmi et al., 2025) focused on the defense against adversarial malware in lightweight ELF headers, which are used mainly in IoT.
3
for threshold parameter t determining the maximal confidence included in calculation, indicator function 1, and variable αX ′ defined as a confidence in malware label, thus calculated using:
BACKGROUND & EVALUATION METRICS
At this point, we formally define the problems of generating adversarial samples and ML-based malware detection. Our approach, tailored to ELF executable files, also benefits from the work of (Louthánová et al., 2024) on PE files. Further in the section, we use X ∈ D to denote a malware from dataset D , which contains only the files classified with the malware label by the target classifier. In our work, we developed the adversarial malware generator G , which is defined using the formula:
X ′ = G (X ) = X + δ
(1)
where the adversarial perturbation δ is added to the original file X . The output, X ′ , is also ELF file that must remain unchanged in terms of execution. We then define an adversarial sample X ′ so that the malware detector classify X ′ as benign or at least reduce the confidence in malware label in comparison to X . To classify the ELF files, ML-based malware detector is used. For the purpose of this article, we represent this detector as a function: f (X ) = (L , P )
Second, we propose an extension to the ER, the Extended Evasion Rate (EER), defined as: 1 EERt = 1{αX ′ <t} ∗ 100% (4) |D | X ′ |∑ X ∈D
(2)
where X represents an ELF binary file, L ∈ {malware, benign} is the output label, and P ∈ [0.5, 1.0] is the probability (confidence) of that label according to detector f . In our work, we use the following three metrics to evaluate the effectiveness of the proposed generator. We note that the first metric is commonly used in this field, while we developed the other two metrics to uncover additional properties of the generator. First, we use the standard Evasion Rate (ER) defined as: #misclassi f ied ER = ∗ 100% (3) total where #misclassi f ied stands for number of misclassified adversarial files by the target classifier and total is a total number of files submitted to that classifier (after discarding files that were already incorrectly predicted before the actual modification).
( αX ′ =
P 1−P
f (X ′ ) = (malware, P ) f (X ′ ) = (benign, P )
(5)
This metric was designed to better approximate the distribution of confidence in malware labels across the whole set. We propose using this metric for different values of the threshold parameter t and plotting the resulting values to observe the distribution. Third, we propose the Mean Difference in Confidence (MD) metric to observe how the confidence in malware label changes between original and adversarial malware samples in the mean case: MD =
1
|D | (X ,X∑ ′ )|X ∈D
(αX ′ − PX )
(6)
where PX is the confidence in malware label of original malware sample X , and αX ′ is computed using Equation (5).
4
GENERATOR DESCRIPTION
We describe our approach to generating malware samples in this section. We note that the generator is designed to address interpretability issues in ML algorithms by carefully logging all operations it performs. Thus, for each input file, three files are produced – the adversarial file, .log file of all the operations performed (including name of the performed operation, the label of the input data, the size of the data used, and the offset to the .data file where the injected content resides), and the .data file with all the data used.
4.1
Modification Types
We propose 12 modifications to ELF executables that were designed to preserve the original functionality of the binaries. These include: • Add Section. This modification adds a new section near the end of the ELF file. • Modify Padding Between Loadable Segments. We modify the unused space between loadable (p type is PT LOAD) segments.
• Extend Padding Between Loadable Segments. We extend the space between loadable segments to satisfy the alignment constraints (for loadable segments, page-size alignment is required (TIS Committee, 2000)). We use the first suitable nonzero offset for this modification. • Rotate Loadable Segments. We move the first loadable segment after the last loadable segment. Second and forthcoming loadable segments are moved towards the beginning of the file. • Append Benign to Malware. We append benign content to the end of the malware executable without altering its structure. • Append Malware to Benign. We use a benign file and append block of binary zeros to meet the alignment constraints of the malware file, which is pasted afterward. The relevant fields in the ELF header are then redirected to the malware part. • Modify Content of Section. We modify the content of sections that have no impact on the program execution flow. However, these sections can be used by tools. To prevent these tools from producing corrupted outputs we optionally rename the section. • Modify Content of .strtab Section. We modify the static symbols of ELF executable files (symbols in the .symtab section, which are used only by debuggers or other tools). We aim not only to add typical benign-file symbol names to malware files, but also to remove typical malware-file symbol names. • Unregister Section. We propose removing definitions of sections that are unnecessary for the program’s execution flow. Their original content is then used as a space for adversarial perturbations. • Remove Section. We also propose removing the section entirely from the ELF file. This modification is tailored only to a selected set of sections that have no impact on the program’s execution, and are not part of any loadable segments (to avoid issues with program loading). • Change Registers. In this modification, we interchange the registers with equivalent purposes within the executable sections of the binary file. • Change Instructions. We propose interchanging the instructions with the same logical (or mathematical) properties while ensuring the hardware dependencies also remain unchanged.
4.2
Data Sources
For the mentioned modification types, the critical question is which data are suitable for these. We use the following data sources: • Sequence of Binary Zeros. We propose using a sequence of 0x00 bytes. Although applying this sequence is an effective way to remove the original content, the presence of it is common in ELF files (compilers often use it as a padding). • Random. We propose including Random data in the adversarial samples. For this data source, we assume that the content of some parts of executables may differ significantly, so that the distribution of exact byte sequences may appear random. • Static File. Using this data source, the actual path to the file, whose content will serve as a source for adversarial perturbations, is expected. We included two options in our implementation: rodata and text. The former contains extracted manual pages of well-known Linux commands, while the latter contains their .text sections. • Extracted Strings. We extract the strings from the .rodata section of benign and malware files of the specified directories. We then select only the benign strings that differ more than the defined threshold and include them in malware files. • Extracted Symbols – Dynamic. We extract the symbols of benign and malware files from the specified directory. We then select symbols typical of benign files to include in malware files, and also symbols typical of malware files to remove them from these files. This data source is tailored to the .strtab section only. • Extracted Symbols – Static. This data source is also tailored to the .strtab section. However, no symbols are extracted, hence the data source expects a static list of symbols to add and remove. • Benign Executable File. We use this data source exclusively for the Append Malware to Benign modification. We enable using of any valid, benign executable. However, we use /usr/bin/xxd primarily because of its smaller size.
4.3
Generator Workflow
Our approach to generating malware samples is summarized in Algorithm 1 and can be seen as a simplified genetic algorithm, to some extent (as some of its components, such as crossover, which can easily produce corrupted executables, are not included).
Input: generator input = directory of files serving as input to the generator. Output: generator out put = directory of adversarial files. Parameters: max iterations = upper limit to number of iterative modifications to be applied (default: 2). enabled modi f ications iter = list of enabled iterative modifications (default: Add Section, Modify Padding Between Loadable Segments, Extend Padding Between Loadable Segments, Rotate Loadable Segments, Append Benign to Malware, Modify Content of Section, Unregister Section, Remove Section). enabled modi f ications single = list of enabled one-time modifications. (default: Modify Content of .strtab Section, Append Malware to Benign, Change Registers, Change Instructions). data sources = list of enabled data sources (default: Sequence of Binary Zeros, Random, Static File (rodata, text), Extracted Strings, Extracted Symbols – Dynamic, Extracted Symbols – Static, Benign Executable File). modi f ication attempts = number of times the modification is applied (default: 10). top samples = number of files proceeding to the next iteration (default: 3). target con f idence = required confidence in malware label of the output file (default: 0.2). samples out put count = number of adversarial files required as output (default: 1). foreach input f ile ∈ generator input do for iteration ∈ {0, 1, ..., max iterations} do Copy input f ile to generator temp foreach f ile ∈ generator temp do if iteration < max iterations then modi f ication list = enabled modi f ications iter else modi f ication list = enabled modi f ications single end foreach modi f ication ∈ modi f ication list do foreach data source ∈ data sources do Apply the modi f ication using the data source to the f ile (modi f ication attempts)-times. Select modified file with the lowest confidence in malware label. Store selected file to iteration best. Remove other files. end end end Select the top samples with the lowest confidence in malware label from the iteration best. Store them to generator temp. Remove other files from iteration best. if confidence in malware of adversarial file < target con f idence then break end end Select the samples out put count samples with the lowest confidence in malware label from the generator temp. Save these files to generator out put. Remove other files from generator temp. end Algorithm 1: Proposed Process of Adversarial Malware Generation.
Our generator supports two types of modifications: iterative (which can be used multiple times, such as Add Section or Extend Padding Between Loadable Segments) and one-time (which can be used
only once, such as Change Instructions or Append Malware to Benign). Whereas the max iterations parameter limits the number of iterative modifications, the one-time modifications are applied only in the last
additional iteration. The data sources are often used so that two random numbers are chosen from the defined range, serving as the offset and size of the data source used for the adversarial perturbation, to avoid always using the same data. To limit the consequences of that randomization, we decided to apply each modification for the given data source multiple times and select the sample with the lowest confidence in the malware label.
5
EXPERIMENTAL SETUP
Our implementation is written in Python. For implementation of machine learning algorithms, we use PyTorch2 package. We also utilize ELFFile3 to read and parse attributes of the ELF binary file. We use Keystone-engine4 as an assembly framework, and the Capstone5 engine as a disassembly framework. We use the Labeled-Elfs dataset6 containing both malware and benign samples, which are labeled not only by package name but also by target architecture, endianness, ABI, or compiler. However, as the malware files contain only files compiled for 64-bit x8664 architecture, we omit the other architectures from benign set. We also undersampled the benign class to create a balanced dataset. The files were split into training (64 %), validation (16 %), and test (20 %) sets. Specifically, 882 files (441 malware and 441 benign) form the training set, 220 (110 malware and 110 benign) form the validation set, and 160 malware files form the testing set (we do not include the benign files in the test set). To classify ELF executable files, we utilize MalConv (Raff et al., 2017), a widely used CNN-based classifier. We use the MalConv-PyTorch implementation7 , which provides effective parallel execution on a GPU (Graphics Processing Unit) using CUDA (Compute Unified Device Architecture). To prepare MalConv to classify unknown files, we train it for 5 epochs. In each epoch, we use the cross-entropy loss function and the Adam (Adaptive Moment Estimation) optimizer for learning. At the end of the training, we achieved an accuracy of 99.54 % on the validation set. For test set, 96.87 % samples was successfully detected as malware. Additionally, 90.62 % of the test set samples 2 https://docs.pytorch.org/docs/stable/index.html 3 https://pypi.org/project/elffile/ 4 https://pypi.org/project/keystone-engine/ 5 https://pypi.org/project/capstone/ 6 https://github.com/nimrodpar/Labeled-Elfs 7 https://github.com/Alexander-H-Liu/
MalConv-Pytorch
were labeled as malware with confidence higher than 0.80, and 84.38 % of the samples with confidence higher than 0.90.
6
EXPERIMENTAL RESULTS
First, we evaluated the effectiveness of the proposed generator using all the modification types and data sources presented in Section 4. More precisely, we use the default values as described in Algorithm 1. Using this setup, we achieved the ER of 67.74 % and changed the confidence in the malware label by −0.5006 in the mean case. We also calculated the EER metric for threshold values between 0.00 and 1.00 with a step size of 0.01. The output is presented in Figure 1 alongside the distribution of confidence in the malware label of the original samples. We then analyzed the .log files and found that only four modification types account for the majority of operations. Namely, Extend Padding Between Loadable Segments (34.61 %), Modify Content of .strtab Section (28.16 %), Add Section (16.23 %) and Append Benign to Malware (15.99 %), were used. We also found that Extracted Strings is used almost exclusively for modifications that do not require the more tailored data source. For the data in the .strtab section, Extracted Symbols – Dynamic was used in the majority of modifications. Based on these observations, we decided to experiment with modification types and data sources in further experiments. In Experiment 2, we decided to limit the list of enabled modifications while keeping all the data sources enabled. We effectively divide modification types into five groups based on the operations they perform. The precise division is presented in Table 1 and can also be seen as a division based on the riskiness of the proposed modifications. Generator 2.1 is considered the safest because it does not alter any ELF structures. In Generators 2.2, 2.3, and 2.4, the risk of unintentionally affecting the program flow is slightly higher. The Generator 2.5 then uses modifications of executable code, which is considered the highest risk, despite our aim to optimize the modifications so that they do not affect executability. Using this setup, we observe the highest EER for Generator 2.4, but Generators 2.2 and 2.1 provide outputs comparable to the best-performing generator. On the other hand, Generators 2.3 and 2.5 do not achieve any substantive success, as shown in Figure 2. In Experiment 3, we decided to limit the data sources used in the particular generators. In Generators 3.1 and 3.2, we enable all the data sources that extract data from malware and benign files. While
Table 1: Experiment 2 – Enabled Modification Types in Different Generators.
Generator Generator 2.1 Generator 2.2 Generator 2.3 Generator 2.4 Generator 2.5
Enabled modification types Modify Padding Between Loadable Segments, Append Benign to Malware Extend Padding Between Loadable Segments, Rotate Loadable Segments Append Malware to Benign Add Section, Modify Content of Section, Unregister Section, Remove Section, Modify Content of .strtab Section Change Registers, Change Instructions Table 2: Experiment 3 – Enabled Data Sources in Different Generators.
Generator 3.2 Generator 3.3 Generator 3.4 Generator 3.5
Enabled data sources Extracted Strings (included in the training set), Extracted Symbols – Dynamic (included in the training set), Benign Executable File Extracted Strings (not included in the training set), Extracted Symbols – Dynamic (not included in the training set), Benign Executable File Static File (rodata), Extracted Symbols – Static, Benign Executable File Static File (text), Extracted Symbols – Static, Benign Executable File Random, Benign Executable File
Dependency of Extended Evasion Rate on Threshold Parameter
Extended Evasion Rate
1.0 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0
Evasion Rate Original Generator 1
0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0
Threshold Parameter (t)
Figure 1: Experiment 1 – Dependency of Extended Evasion Rate on Threshold Parameter. The blue line (Original) represents the malware samples before modifications are applied, whereas the orange line (Generator 1) represents the Extended Evasion Rate of the adversarial samples.
Dependency of Extended Evasion Rate on Threshold Parameter
Extended Evasion Rate
1.0 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0
Evasion Rate Original Generator 2.1 Generator 2.2 Generator 2.3 Generator 2.4 Generator 2.5
0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0
Threshold Parameter (t)
Figure 2: Experiment 2 – Dependency of Extended Evasion Rate on Threshold Parameter.
Generator 3.1 extracts data from the files used in the training set, Generator 3.2 uses benign files not included in the training set and malware files intended for modification. For Generators 3.3 and 3.4, only data prepared in advance is used. The only differ-
Dependency of Extended Evasion Rate on Threshold Parameter
1.0 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0
Extended Evasion Rate
Generator Generator 3.1
Evasion Rate Original Generator 3.1 Generator 3.2 Generator 3.3 Generator 3.4 Generator 3.5
0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0
Threshold Parameter (t)
Figure 3: Experiment 3 – Dependency of Extended Evasion Rate on Threshold Parameter.
ence between Generators 3.3 and 3.4 is the Static File data source (while Generator 3.3 uses manual pages of well-known Linux commands, Generator 3.4 uses extracted .text sections of the same commands). The Generator 3.5 then uses mainly the Random data source. The precise set of enabled data sources in the particular generators is presented in Table 2. Using this setup, we observe the highest EER for Generator 3.2, but Generator 3.1 remains comparable. Worse properties are then shown for Generator 3.3, which uses the contents of manual pages as primary data source. Moreover, Generators 3.4 and 3.5 achieve no substantive success as shown in Figure 3. We also present Table 3, which compares the calculated Evasion Rate and Mean Difference in Confidence metrics for all the generators examined in this section. We conclude that the target classifier appears sensitive to string-based data sources but does not depend heavily on their position in the executables. However, we based our observations on the dataset that is smaller than the ones usually found in PE files.
Table 3: Comparison of All the Generators Presented.
Generator Generator 1 Generator 2.1 Generator 2.2 Generator 2.3 Generator 2.4 Generator 2.5 Generator 3.1 Generator 3.2 Generator 3.3 Generator 3.4 Generator 3.5
7
ER 67.74 % 49.68 % 54.84 % 1.29 % 64.52 % 0.00 % 67.74 % 67.10 % 14.19 % 3.87 % 3.87 %
MD −0.5006 −0.4004 −0.4182 −0.0092 −0.4621 −0.0014 −0.4926 −0.5216 −0.1826 −0.0549 −0.0571
CONCLUSIONS & FUTURE WORK
This study has examined the current landscape of adversarial malware generation, with focus on ELF file format. In our work, we developed a generator of adversarial malware files for ELF file format. Our workflow is based on the genetic algorithm and accommodates 12 modification types and 7 data sources. Moreover, we aim to address interpretability issues of MLbased algorithm using advanced logging. We conducted a variety of experiments to observe the properties of our generator. We concluded that our generator achieved the Evasion Rate of 67.74 % and Mean Difference in Confidence of −0.50 when all the modification types and data sources were enabled. We also found that only two data sources and four modification types are used in the majority of operations. Based on that, we decided to experiment with enabling only subsets of modification types and data sources to evaluate the generator thoroughly. We divided the modification into five groups and concluded that three of the five groups produce comparable results. In a further experiment, we limited the data sources. We found that the string-based data sources were the most successful, and that the generator performed worse when other data sources were used. We concluded that the target classifier, MalConv, appears extremely sensitive to strings at any position in the executable file. In future work, we will continue developing our generator to achieve a higher Evasion Rate within a reasonable amount of time needed for the actual generation of adversarial malware samples. We will aim to extend our generator to accommodate binaries compiled for the ARM architecture, which is primarily used in IoT. We will dive deeper into execution testing and, consequently, extract more data for dynamic analysis. We will also work on the defensive strategies to strengthen the classifiers.
ACKNOWLEDGEMENTS This work was supported by the 2025 FIT CTU Student Summer Research Program in Prague and by the Grant Agency of the Czech Technical University in Prague, grant No. SGS26/187/OHK3/3T/18 funded by the MEYS of the Czech Republic.
REFERENCES Anderson, H. S., Kharkar, A., Filar, B., Evans, D., and Roth, P. (2018). Learning to Evade Static PE Machine Learning Malware Models via Reinforcement Learning. Cozzi, E., Graziano, M., Fratantonio, Y., and Balzarotti, D. (2018). Understanding Linux Malware. In 2018 IEEE Symposium on Security and Privacy (SP), pages 161– 175. Guesmi, H., Khalfallah, A., and Bouallegue, B. (2025). Lightweight ELF Header Analysis Model for IoT Malwares Detection Based on Machine Learning. Engineering Research Express, 7(2):025213. Kosikowski, A., Cho, D., Ninan, M., Ralescu, A., and Wang, B. (2023). EvilELF: Evasion Attacks on DeepLearning Malware Detection over ELF Files. In 2023 International Conference on Machine Learning and Applications (ICMLA), pages 1702–1709. Kozák, M., Jureček, M., Stamp, M., and Troia, F. D. (2024). Creating Valid Adversarial Examples of Malware. Journal of Computer Virology and Hacking Techniques, 20(4):607–621. Kreuk, F., Barak, A., Aviv-Reuven, S., Baruch, M., Pinkas, B., and Keshet, J. (2018). Adversarial Examples on Discrete Sequences for Beating Whole-Binary Malware Detection. arXiv preprint arXiv:1802.04528, pages 490–510. Louthánová, P., Kozák, M., Jureček, M., Stamp, M., and Di Troia, F. (2024). A Comparison of Adversarial Malware Generators. Journal of Computer Virology and Hacking Techniques, 20(4):623–639. Lucas, K., Sharif, M., Bauer, L., Reiter, M. K., and Shintre, S. (2021). Malware Makeover: Breaking ML-Based Static Analysis by Modifying Executable Bytes. In Proceedings of the 2021 ACM Asia Conference on Computer and Communications Security, ASIA CCS ’21, page 744–758, New York, NY, USA. Association for Computing Machinery. Qiao, Y., Zhang, W., Tian, Z., Yang, L. T., Liu, Y., and Alazab, M. (2023). Adversarial ELF Malware Detection Method Using Model Interpretation. IEEE Transactions on Industrial Informatics, 19(1):605–615. Quertier, T., Marais, B., Morucci, S., and Fournel, B. (2022). MERLIN – Malware Evasion with Reinforcement LearnINg. arXiv preprint arXiv:2203.12980. Raff, E., Barker, J., Sylvester, J., Brandon, R., Catanzaro, B., and Nicholas, C. (2017). Malware Detection by Eating a Whole EXE. arXiv preprint arXiv:1710.09435.
Ramamoorthy, J., Shashidhar, N. K., and Varol, C. (2025). Automated Static Analysis of Linux ELF Malware: Framework and Application. In 2025 13th International Symposium on Digital Forensics and Security (ISDFS), pages 1–5. Ravi, A., Chaturvedi, V., and Shafique, M. (2025). ADVeRL-ELF: ADVersarial ELF Malware Generation using Reinforcement Learning. In 2025 62nd ACM/IEEE Design Automation Conference (DAC), pages 1–7. Song, W., Li, X., Afroz, S., Garg, D., Kuznetsov, D., and Yin, H. (2022). MAB-Malware: A Reinforcement Learning Framework for Blackbox Generation of Adversarial Malware. In Proceedings of the 2022 ACM on Asia Conference on Computer and Communications Security, ASIA CCS ’22, page 990–1003, New York, NY, USA. Association for Computing Machinery. TIS Committee (2000). Tool Interface Standard TIS - Executable and Linkable Format (ELF) Specificatoin — linuxfoundation.org. https://refspecs.linuxfoundation. org/elf/elf.pdf. [Accessed 16-07-2025]. Xue, M., Fu, J., Li, Z., Ni, S., Wu, H., Zhang, L. Y., Zhang, Y., and Liu, W. (2024). A Reinforcement LearningBased ELF Adversarial Malicious Sample Generation Method. IEEE Journal on Emerging and Selected Topics in Circuits and Systems, 14(4):743–757.