SeqShield: A Behavioral Analysis Approach to Uncover Rootkits
Shukla
Paras Ghodeshwar1[0009−0009−2301−007X] , Sandeep K , Anand Handa3[0000−0003−0075−1165] , Nitesh 3[0000−0003−0998−0925] Kumar , and
2[0000−0001−5525−7426]
Department of Computer Science and Enginerring, Indian Institute of Technology, Kanpur, India [email protected] 2 International Institute of Information Technology, Hyderabad, India [email protected] 3 C3iHub Center, Indian Institute of Technology, Kanpur, India {anand, niteshkr}@c3ihub.iitk.ac.in
arXiv:2604.23812v1 [cs.CR] 26 Apr 2026
1
Abstract. Rootkits are among the most elusive types of malware, capable of bypassing traditional static analysis methods due to their metamorphic behavior. Signature-based detection techniques struggle against these threats, necessitating a shift toward dynamic analysis approaches. We propose SeqShield, a behavior-based rootkit detection approach designed specifically for the Windows OS, leveraging API call sequences for dynamic behavior analysis. Instead of relying on static signatures, SeqShield examines the execution patterns of API calls, which inherently reflect malicious intent. Analyzing API sequences, we can effectively identify rootkit-like behavior. We also employed a metamorphic code engine to generate 10× mutated variants of rootkits, demonstrating their obfuscation strategies. SeqShield applies n-gram analysis to extract bigram and trigram features from these API call sequences, enabling effective detection of rootkit-like activity. Among the models tested, Random Forest achieves the highest accuracy of 97.27% (bigram) and 96.17% (trigram). To optimize performance and decrease the dimension, we apply feature importance ranking using the Gini Impurity Index, iteratively selecting the most significant features. The optimized lower-dimensional feature matrix significantly enhances detection efficiency without sacrificing accuracy. Using the optimized feature set, our approach achieves 96.72% accuracy for bigrams and 97.81% accuracy for trigrams. Keywords: Kernel-Level Rootkits· Rootkits· Cybersecurity· NGrams· NLP· Gini Index· MetaMe.
1
Introduction
Malware has become a serious problem in today’s digital world, and rootkits are one of the most dangerous types. They work secretly, hiding deep inside the operating system to avoid being noticed and keeping control over the computer
2
P. Ghodeshwar et al.
for long periods. Their ability to compromise system integrity and security makes detecting rootkits a critical area of cybersecurity research. Adversaries often employ rootkits to stealthily conceal the presence of malicious programs, files, network connections, services, drivers, and other critical system components. Rootkits achieve this by intercepting, hooking, and manipulating operating system API calls to obscure system information from security tools and administrators. Rootkits or rootkit-enabling functionalities can operate at various levels of a system. They may reside in user space, within the kernel, or even deeper in the system architecture, including the hypervisor, Master Boot Record (MBR), or system firmware [2]. Kernel-level rootkits are considered the most dangerous and sophisticated type within the rootkit family, often referred to as next-generation rootkits due to their advanced capabilities, including kernel hooks and data structure manipulation. Unlike first-generation rootkits, which primarily operated in user mode, kernel-level rootkits of the second and third generations transitioned to the kernel, providing attackers with deeper access and greater control over the operating system. These features make kernel-level rootkits exceptionally difficult to detect and remove, posing a significant threat to system security [10]. To conceal malicious activities or modify system behavioral techniques employed to develop kernel-level rootkits, such as IRP (I/O Request Packet) hooks, SSDT (System Service Descriptor Table) hooks, IDT (Interrupt Descriptor Table) hooks, DKOM (Direct Kernel Object Manipulation), and virtual file system hooking. Among these, DKOM is considered a third-generation rootkit technique as it directly manipulates kernel data structure [9]. To tackle these challenges, we propose SeqShield, a novel approach that leverages API call sequences to detect rootkit-like behavior. Our method utilizes NLP techniques, specifically bigram and trigram sequences, to analyze API call patterns for identifying malicious activities. We did our experiment on the Windows operating system because it’s widely used and often targeted by rootkits. Our dataset included a balanced mix of known rootkits and normal (benign) executables, which helped in training the model fairly. Malware writers/developers often use tricks to hide their code and avoid getting caught, we used an obfuscation tool to make our test cases more realistic. To study how the executable behaves, we looked at sequences of API calls made during execution. We created two types of feature sets from this data: bigrams (pairs of API calls) and trigrams (triplets of API calls). These patterns helped us spot suspicious behavior similar to rootkits. Although these feature sets contained both relevant and irrelevant features, we employed a novel approach using the Gini impurity index to rank and extract the most relevant ones. Instead of using a fixed threshold, we divided the sorted features into topranked chunks and identified the subset where the model’s F1-score peaked. This allowed us to significantly reduce the dimensionality of the feature space while retaining high predictive performance. With fewer but more relevant features, our model still performed as well or even better than the full feature set. It also
SeqShield: A Behavioral Analysis Approach to Uncover Rootkits
3
worked well in identifying new rootkits it hadn’t seen before, showing that the model is able to generalize and catch unknown threats. The key contribution of our research are as follows: – Developed a lightweight rootkit detection model using API call sequences, achieving high accuracy with minimal computational overhead. – Used a Metamorphic Code Engine to simulate obfuscation, showing hash changes that evade VirusTotal detection, aligning with the Pyramid of Pain concept [3]. – Applied Gini Impurity Index to rank and select key features, reducing feature size while maintaining model accuracy and efficiency. – Validated that top-ranked features common across Decision Tree and Random Forest classifiers are most effective for rootkit detection. – Relevant datasets, feature matrices, and algorithms will be shared upon request if deemed necessary.
2
Background
The concept of rootkits dates back to the early 1990s when the term rootkit was first introduced in the context of UNIX-like operating systems [19]. A rootkit refers to a toolkit that allows attackers to gain and maintain unauthorized rootlevel access to a system. These tools were initially developed to modify the operating system’s core functionalities, enabling attackers to conceal their activities while maintaining control over the system. The prominence of rootkits grew significantly in 2005 during the infamous Sony BMG copy protection rootkit scandal. The rootkit, called Extended Copy Protection (XCP), was surreptitiously installed on users’ systems through a music CD. Once installed, it restricted user access and functionality, effectively acting as a Digital Rights Management (DRM) mechanism but at the cost of user privacy and system security [8,19]. Since this incident, rootkits have become a focal point for both malware authors and cybersecurity researchers, with their functionality and sophistication evolving substantially over the years. 2.1
Evolution of Rootkits
The evolution of rootkits can be categorized into three distinct generations: First-Generation Rootkits (User-Level Rootkits) Early rootkits operated at the user level and were primarily designed to execute malicious activities in the user mode of the operating system. These included tasks such as stealing passwords, collecting sensitive user information, and altering application-level functionality. While these rootkits were relatively simple, they paved the way for more advanced techniques [10, 11]. Second-Generation Rootkits (Kernel-Level Rootkits) With advancements in malware development, attackers moved beyond user-level operations to target the kernel—the core of the operating system. Kernel-level rootkits operate in kernel mode, granting them deeper access and control over the system.
4
P. Ghodeshwar et al.
These rootkits utilize system-level hooks to manipulate critical OS functionalities. Techniques such as hooking System Call Tables, I/O Request Packets (IRP), and Interrupt Descriptor Tables (IDT) became common, allowing rootkits to hide their presence and perform malicious tasks stealthily [10, 11]. Third-Generation Rootkits (Kernel Data Structure Manipulation) The latest generation of rootkits focuses on manipulating kernel-level data structures to achieve invisibility and persistence. This category includes techniques like Direct Kernel Object Manipulation (DKOM) and Dynamic Kernel Object Hooking. Unlike traditional kernel hooks, DKOM modifies dynamic kernel data structures, making detection significantly more challenging. For instance, attackers can hide processes, kernel device drivers, and active ports by altering kernel-maintained data structures without disrupting normal system functionality [10, 11].
3
Related Work and Motivation
The detection of malicious software, i.e., malware, including rootkits, typically relies on two primary methodologies: static analysis and dynamic analysis. Each methodology has unique approaches, strengths, and limitations that, when combined, provide complementary strategies for identifying and mitigating malicious activity. Rootkit detection also benefits from these methodologies, with each offering distinct advantages in identifying rootkits, including kernel-level rootkits. Malware authors, including those developing rootkits, often employ obfuscation techniques to evade signature-based detection systems. In this research, we utilize MetaMe to mutate rootkit samples, demonstrating how obfuscation can bypass signature-based detection. The recent work has demonstrated the detection of the rootkits via three major factors: Hardware information-based rootkit detection Approach, Memory Forensics based rootkit detection, Behavior-based rootkit detection approaches. Hardware-Based Rootkit Detection Recent studies have leveraged hardwarelevel monitoring to detect rootkits by analyzing control registers and hardware performance metrics. Liwei Zhou et al [21]. introduced a hardware-assisted infrastructure that monitors modifications to Control Register (CR3), identifying unauthorized changes indicative of rootkits. Their approach extracts features at the hardware level but performs off-chip analysis in a trusted software environment using mathematical classification techniques. Similarly, Baljit Singh et al [14]. explored hardware performance counters (HPCs) by creating synthetic rootkits (IRP, SSDT, DKOM) and capturing hardware-level execution data. While they successfully detected SSDT and IRP rootkits, they failed to detect DKOM rootkits due to their indirect manipulation of kernel data structures. However, Boyou Zhou et al [20]. highlighted that HPC-based detection suffers from high false positive rates, making it impractical. Additionally, Joel A. [5] et al. proposed a power-based rootkit detection method, leveraging nonlinear phase-space analysis on power fluctuations to detect anomalies in execution patterns.
SeqShield: A Behavioral Analysis Approach to Uncover Rootkits
5
Memory Forensics-Based Rootkit Detection Memory forensics has proven to be an effective technique for dynamic rootkit analysis by extracting key kernel structures from volatile memory. Mohammad Nadim et al [10]. demonstrated that forensic tools like Volatility can analyze memory artifacts such as EPROCESS, ETHREAD, LDR_DATA_TABLE_ENTRY, and CR0 to identify rootkits at the kernel level. Similarly, Xiao Wang et al [18]. utilized memory dumps from VMs running on the OpenStack platform, ensuring trusted acquisition from the controller node. Their approach combined memory forensic techniques with machine learning, where extracted malicious features were used to train classifiers, achieving automated rootkit detection. By leveraging multiple Volatility plugins [16], their model effectively captured suspicious memory modifications indicative of rootkits.
Behavior-Based Rootkit Detection Behavior-based detection methods focus on identifying anomalies in system activity caused by rootkits. Suresh Kumar et al [15]. proposed a deep learning-based detection framework, where they collected hooked data using McAfee’s Rootkit Detective and trained a reinforcement learning model to recognize abnormal behavior. The trained model was then enhanced using Generative Adversarial Networks (GANs) for anomaly detection. Kruegel et al [6]. introduced a binary analysis approach, statistically analyzing kernel module interfaces for modifications that indicate rootkit presence. Patrick Luckett et al [7]. explored neural network-based rootkit detection, utilizing system call timing analysis to classify infected OS behavior. Their model, implemented in MATLAB’s Neural Network Toolbox, achieved 82.8% accuracy in classifying rootkit-infected system calls, demonstrating the effectiveness of behavioral analysis in rootkit detection. Previous research has explored various approaches, each with its own drawbacks. For instance, in hardware-based detection, the HPC method failed to detect DKOM rootkits and exhibited a high false positive rate, making practical implementation challenging. Memory forensics is one of the most effective techniques available; however, it is primarily suited for analyzing compromised systems and is difficult to implement efficiently in a live system environment. Additionally, detecting hooking-based methods may fail to identify DKOM rootkits, as DKOM does not rely on hooking techniques. This limitation motivated us to explore behavioral analysis as an alternative approach. Unlike traditional methods, behavioral analysis focuses on the actions and interactions of a program rather than its static properties. We were particularly interested in analyzing API call sequences, as they provide critical insights into how rootkits operate within a system. By examining API call patterns, we aim to uncover rootkit-like behavior, even in the presence of obfuscation techniques. We utilized MetaMe to mutate rootkit samples to further validate our approach, demonstrating how obfuscation can successfully evade signature-based detection. However, we can effectively detect rootkit activities regardless of their evasion techniques through behavioral analysis and API call sequence modeling.
6
4
P. Ghodeshwar et al.
SeqShields
SeqShield is a behavior analysis method designed to detect rootkit-like behavior, specifically targeting kernel-level rootkits in the Windows environment. The method is based on analyzing API sequences, which serve as fundamental parameters for a process to execute tasks, whether operating in user mode or kernel mode. API calls act as an essential interface between software and the operating system, enabling applications to request system resources, perform operations, and communicate with hardware. Every program, whether benign or malicious, interacts with the system through a sequence of API calls to execute its intended functionality. These functions reside in ntdll.dll in user mode but internally transition to kernel mode via a system call mechanism. Regardless of whether a rootkit operates at the kernel level or user level, it must rely on specific API call sequences to execute its malicious functionalities. Even kernel-mode rootkits, which operate with elevated privileges, typically require some form of user-mode execution to trigger their behavior, since user-mode processes are responsible for invoking system calls that eventually transition into the kernel. By analyzing these API call sequences, SeqShield provides valuable insights into a program’s behavior, enabling the distinction between legitimate and malicious activities while effectively detecting rootkit-like behavior. To analyze API call sequences, we employed the n-grams method, a widely used technique in natural language processing (NLP) for modeling sequential data. N-grams represent contiguous sequences of n items (in this case, API calls) within a dataset [13]. This approach allows us to capture the immediate relationship between consecutive API calls, enabling a deeper understanding of execution patterns. The ability of n-grams to preserve sequential dependencies makes them particularly effective for behavioral analysis, as they help identify recurring patterns that indicate malicious activity. Unlike other complex models such as Long Short-Term Memory (LSTM) networks may struggle with shorter sequences and require additional mechanisms like attention mechanisms to weigh important sequence segments properly. By leveraging N-grams, we can analyze API call sequences to detect deviations from normal execution flow, making it a robust technique for identifying rootkit-like behavior. In our analysis, we specifically utilized bigram and trigram sequences, which capture the relationships between two and three consecutive API calls, respectively. This approach enables us to better understand the contextual dependencies within API sequences, improving the accuracy of our detection mechanism while maintaining computational efficiency. Architecture This section describes the architecture(Fig 1) of the SeqShield model, which is centered on analyzing API sequence patterns. We utilized multiple tools and modeling techniques to achieve our objectives to derive the desired results.
400 Mutated Rootkits
SeqShield: A Behavioral Analysis Approach to Uncover Rootkits
MetaMe Code Engine
Model Trainig - Selected Importance Features
90 Samples
Windows Guest Machine
310 Samples
Relevent Features Extraction
Behavior Training Data
610 Executables api
300 Benign
Low-dim Model Training
Model Trainig Executables
40 Rootkits
Open Source Databse
7
Ubuntu Host Machine
API Sequence Extraction
Bigram Feature Transformation
Trigram
RF DT XGB LR AdaB SVM KNN
Testing Data
Effective Trained Model
Results
Trained Model
Results
Fig. 1. Architecture of SeqShield
4.1
Experimental Setup
Our experiments utilize multiple tools to collect and process data in a secure environment. The experiments were conducted on an Intel® CoreTM i7-4700 CPU @ 3.40 GHz × 8, running a 64-bit Ubuntu 18.04.6 LTS operating system with 16 GB of RAM and 1 TB of disk space. Ubuntu was used as the host OS, with VirtualBox installed for virtualization. A Windows 7 Pro virtual machine was configured with 2 GB of RAM and 32 GB of disk space. The latest version of Cuckoo [4] was installed on the host machine, while the Cuckoo agent was deployed on the virtual machine to facilitate malware analysis. Malware samples were executed in the Windows virtual machine, and activity reports were transmitted to the host Ubuntu system via the Cuckoo agent. Additionally, we employed MetaMe [12], a metamorphic code engine, to mutate the samples and analyze their obfuscation behavior. To validate obfuscation, VirusTotal’s [17] hash search was utilized for further verification. 4.2
Dataset
To create a dataset of rootkit malware, we downloaded 40 rootkit samples from open sources such as MalwareBazaar [1]. These rootkits exhibited properties such as evading detection, privilege escalation, and file manipulation. Given the limited availability of rootkit samples and to demonstrate obfuscation techniques used by malware authors, we utilized MetaMe [12] to mutate
8
P. Ghodeshwar et al.
each malware sample ten times, resulting in 400 mutated rootkits. These samples were divided into: – 310 samples for analysis and model training – 90 samples as test data (previously unseen by the model) Additionally, we included 300 benign executables, consisting of general-purpose applications like browsers and system processes. In total, our dataset comprised 610 executables for analysis and modeling. 4.3
Obfuscation Engine: MetaMe
MetaMe [12] is a simple metamorphic code engine that generates logically equivalent versions of executable code. Malware authors leverage metamorphic techniques to evade signature-based detection systems. In our experiment, MetaMe was used to mutate rootkit samples, altering their hash values to remain undetected by traditional antivirus solutions [12]. MetaMe Architecture Fig 2
1 Mutated
Rootkit Sample 10 Mutated Samples MetaMe Code Engine 10X
Fig. 2. MetaMe Mutation Architecture
Each malware sample was passed through MetaMe in a recursive process: 1. If A is the original executable, MetaMe generates A1 . 2. A1 is then passed through MetaMe to generate A2 . 3. This process continues iteratively up to A10 , resulting in ten variations of the original sample (A1 , A2 , ..., A10 ). These transformed samples were then analyzed to assess their detectability. When tested against VirusTotal [17], most modified signatures remained undetected, except for a few cases where hash values did not change significantly. 4.4
Experimental Methodology
For our experiment, we prepared 610 executable samples (310 rootkit malware and 300 benignware).
SeqShield: A Behavioral Analysis Approach to Uncover Rootkits
9
Malware and Benignware Execution The sandbox environment was set up by creating a virtual network and configuring Cuckoo [4] to establish an isolated execution space for malware analysis. Once the sandbox was ready, the collected malware samples were executed, generating detailed Cuckoo analysis reports. These reports contained essential behavioral data, including process creation details, API calls made, files created or modified, network connections established, memory dumps, and Indicators of Compromise (IOCs). From these reports, API sequences were extracted alongside their execution timestamps. To enhance the analysis, API calls were grouped into bigrams (pairs of API calls) and trigrams (triplets of API calls) [13], incorporating time as a parameter to capture sequence patterns. Only unique API sequences were retained, and n-gram analysis was performed to identify patterns indicative of rootkit behavior. N-Gram Analysis We employed bi-gram and tri-gram feature engineering techniques. Given a malware sample M containing an API sequence mi = {a1 , a2 , ..., aj }: – Bigram Sequence: {a1 a2 , a2 a3 , ..., aj−1 aj } – Trigram Sequence: {a1 a2 a3 , ..., aj−2 aj−1 aj } Using this method, we obtained, 12,007 unique bi-gram sequences and 68,442 unique tri-gram sequences. One-hot encoding was applied to prepare datasets for machine learning models. Classification Models We created two training and testing datasets: a bigram matrix consisting of 610 rows (executables) and 12,008 columns representing API calls, along with one label column, and a trigram matrix with 610 rows and 68,443 dimensions. To classify rootkit behavior, we employed multiple machine learning classifiers from the scikit-learn package, including Decision Tree (DT), Random Forest (RF), Support Vector Machine (SVM), K-Nearest Neighbors (KNN), Logistic Regression, AdaBoost, XGBoost, and Gradient Boosting. The dataset was split into a 70:30 ratio for training and testing, ensuring a balanced evaluation of model performance. Additionally, we used the joblib library to save trained models, allowing for future classification of previously unseen traces. 4.5
Classification Results
In our experiment, we distributed the training and testing data in a 7:3 ratio. Specifically, 70% of the dataset was used for training the Machine Learning Models, while the remaining 30% was used for testing the models. In this section, we describe the results obtained using the 30% testing dataset. Using this testing dataset, we achieved the best accuracy of 97.2678% for bigram and 96.1749% for trigram API sequence feature matrices using the Random Forest Machine Learning Classifier. Tables 1 and 2 present the results of our evaluation of bigram and trigram models, respectively. These tables showcase key performance metrics, including accuracy, precision, recall, and F1-score, for the various machine learning classifiers used to assess model performance.
10
P. Ghodeshwar et al.
Table 1. Model performance on the test set for Bigram API sequences (values in percentages). Model RF XGBoost LR SVM KNN GB DT AdaBoost
Accuracy Precision Recall F1 Score 97.268 96.175 96.721 96.175 94.536 96.175 96.721 95.628
97.274 97.268 96.180 96.175 96.920 96.721 96.443 96.175 94.617 94.536 96.180 96.175 96.721 96.721 95.974 95.628
97.268 96.174 96.716 96.167 94.531 96.174 96.721 95.616
Table 2. Model performance on the test set for Trigram API sequences (values in percentages). Model RF XGBoost LR SVM KNN GB DT AdaBoost
5
Accuracy Precision Recall F1 Score 96.175 94.536 96.175 96.175 85.246 95.082 93.989 95.082
96.443 96.175 95.066 94.536 96.443 96.175 96.443 96.175 88.566 85.246 95.516 95.082 94.625 93.989 95.516 95.082
96.167 94.514 96.167 96.167 84.879 95.066 93.961 95.066
Key Feature Importance Analysis
In our experiment, we utilized two datasets: one for bigram features and another for trigram features. These datasets contain both relevant and irrelevant features, which can impact the detection results. To identify the most significant contributing features, we employed a feature importance for tree-based classifier, which leverages Machine Learning’s feature importance ranking using Gini Impurity Indexing. This method assigns an importance score between [0, 1] to each feature. Tree-based classifiers, such as Decision Trees and Random Forests, select features for splitting based on their ability to reduce Gini Impurity [22]. Features that contribute to larger reductions in impurity are deemed more important, making them highly relevant for classification. This approach not only aids in feature selection but also helps extract the most informative features that significantly impact detection accuracy. By reducing the dimensionality of the feature matrix, we optimize computational efficiency without compromising detection performance. Our hypothesis is that by selecting only the most relevant features, we can effectively detect rootkit-like behavior with high precision while significantly lowering computational overhead. This ensures a balance between accuracy and efficiency, making our approach fast and scalable.
SeqShield: A Behavioral Analysis Approach to Uncover Rootkits
5.1
11
Feature Selection
To extract the most significant features from bigram and trigram sequences, we utilized Gini impurity-based feature importance indexing. For each feature in the n-gram feature vector, importance scores were computed using tree-based classifiers that rely on Gini impurity to assess the contribution of each feature toward classification performance. These Gini-based importance scores range between 0 and 1, where a higher value indicates a greater role in reducing impurity during the decision-making process. Once computed, we perform the following task based on their Gini importance scores. – For the bigram dataset, we had 12,007 features. These were sorted as {B1 , B2 , ..., B12,007 } based on their importance score, ensuring that Bi > Bi+1 . – Similarly, for the trigram dataset, we had 68,442 features, sorted as {T1 , T2 , ..., T68,442 } where Ti > Ti+1 . Further for classification results, we divided the sorted features into chunks of 100 and trained the classifier incrementally. Rather than selecting features based on a Gini threshold which would lead to variable chunk sizes and potentially skew the analysis, we chose to divide the sorted features into fixed-size chunks of 100. This fixed-size chunking ensures consistency across experimental runs and allows us to evaluate how model performance scales as more features are included. The process is followed in Algorithm 1: Algorithm 1 Incremental Feature Selection and Classification Require: Sorted feature sets for bigram {f b1 , f b2 , ..., f b12,007 } and trigram {f t1 , f t2 , ..., f t68,442 }, Tree-based classifiers {DT, RF, AdaB, XGB, GB} 1: Initialize chunk size k ← 100 2: Initialize maximum F1-score F1max ← 0 3: Initialize best feature count nmax ← 0 4: for i = k to n step k do 5: Select top i features: 6: For bigram: {f b1 , f b2 , ..., f bi } 7: For trigram: {f t1 , f t2 , ..., f ti } 8: for each classifier C in {DT, RF, AdaB, XGB, GB} do 9: Train C with selected features 10: Compute performance metrics (Precision, Recall, F1-score, Accuracy) 11: if F1-score > F1max then 12: Update F1max ← current F1-score 13: Update nmax ← i 14: end if 15: end for 16: end for 17: Output the first nmax features where F1-score is maximized for both bigram and trigram.
By analyzing the recorded F1-scores, we determined that the most contributing features correspond to the first chunk where the F1-score reached its maxi-
12
P. Ghodeshwar et al.
mum. These features were considered the most relevant for detecting malicious behavior. After reaching the maximum F1-score, additional features often resulted in a decline in model performance. This suggests that while the initial n×100 features contributed positively, subsequent features may have introduced noise or irrelevant information, thereby reducing accuracy. Interestingly, we observed fluctuations in results when using Decision Tree and Random Forest classifiers. However, other models such as XGBoost, AdaBoost, and Gradient Boosting showed more stable results across all feature chunks. Therefore, Decision Tree and Random Forest were chosen for further analysis to identify features that exhibit rootkit-like behavior. In Fig.3, we can see the bar graph where the Decision Tree classifier reaches the maximum F1score at Top n × 100 features where n is 14, shorted using Gini Impurity Index. Similarly, we considering the Decision Tree classifier for Trigram with Top n×100 features where n is 69 and 19 in the case of Random Forest, where the F1-score is nearly same both the case.
100
97.81%
7000 6000
97.27% 95.67%
96
5000
96.17%
Top n Features
F1-Score (%)
98
6900
4000 3000
94
1900 1400
92
2000 1000
700
Top n Features
F Tri g
ram Tri g
F1-Score (%)
ram
-D
-R
F -R Big
ram
-D ram Big
T
0
T
90
Fig. 3. Model performance comparison for Top n-feature where F1-score is maximum
5.2
Identifying the Most Contributing Features
To pinpoint the most relevant features, we identified the intersection of the top n×100 features from both Decision Tree and Random Forest classifiers where the F1-score is max. The overlapping features were considered the most informative for rootkit detection. Let us consider the set A, B, C, D for the selected feature sets for bigram and trigram classifiers be defined as follows:
SeqShield: A Behavioral Analysis Approach to Uncover Rootkits
13
DT DT A = {f bDT 1 , f b2 , . . . , f b1,400 } RF RF B = {f bRF 1 , f b2 , . . . , f b700 } DT DT C = {f tDT 1 , f t2 , . . . , f t6,900 } RF RF D = {f tRF 1 , f t2 , . . . , f t1,900 }
To find features that are important across models, we look for overlaps in their top-ranked features. This is done by calculating the intersection of the feature sets from each model. In our case, we compared the most significant features chosen by both the Random Forest and the Decision Tree classifiers, using both bigram and trigram feature vectors. If a feature appears in the lists of both classifiers, it means that feature is consistently recognized as important. These overlapping features are valuable because they show strong and stable influence on the model classification. |A ∩ B| = 76 |C ∩ D| = 483 For bigram we have found 76 common features from the Decision Tree and Random Forest classifier where we have seen the first F1-score max. Similarly, for trigram we have found 483 common features which are really contributing. Our approach demonstrates high reduction of the feature matrix dimension like for bigram, we had 12,007 features contains both relevant and irrelevant features, now we have 76 features which are more relevant and really contributing towards results for rootkit detection, and for the trigram, we had 68,442 features now we have 483 most relevant features. With our approach, in both bigram and trigram we have seen a heavy reduction of dimension from higher dimension to lower dimension. We have further trained the machine learning classifier based on these lower dimension i.e. considering only most relevant features. While have tested this hypothesis using the 30% testing data that was split from feature matrix for model testing for the higher dimension data. We have successfully achieved the accuracy of 96.72% for bigram and 97.81% for trigram using the Random Forest classifier. These results demonstrates our hypothesis for considering the most relevant features meanwhile decrease featured dimension drastically. For bigram and for trigram sequences, a few listed features are listed in Table 3 for Bigram Sequences and Table 4 for Trigram Sequences along with Justification.
6
Model Testing
During our experiment, we utilized MetaMe to mutate malware samples, ensuring that 90 samples remained undetected by the model. These rootkit samples were intentionally excluded from the training phase to evaluate our model’s ability to detect unknown threats. To evaluate the effectiveness of our approach,
14
P. Ghodeshwar et al. Table 3. Top Common Features for Bigram Features.
API Sequence
Justification DeviceIoControl, to interact with device drivers, espe(GetTempPathW, DeviceIo- cially in the temp directory, may indicate attempt exploit Control) vulnerabilities in drivers or install rootkits. Temp directory/file creation in low-level sys call (GetTempPathW, NtCre- (ntoskrnl), could be used for staging malicious payload. ateFile) This rootkit behavior can modify critical system files to facilitate stealth or control. Querying registry keys after system time retrieval may in(GetSystemTimeAsFileTime,dicate malicious configuration checks. Rootkit behavior NtQueryKey) for creating unique payloads or registry based on timestamp. Querying the system directory and accessing the registry (GetSystemDirectoryW, are activities that, when combined, could indicate an atRegOpenKeyExW) tempt to tamper with system configurations. HTTP query combined with registry access suggests (HttpQueryInfoA, Re- network-based configuration tampering. Allows rootkit gOpenKeyExW) to dynamically adapt to instructions from remote servers for long term malware activity. Table 4. Top Common Features for Trigram Features. API Sequence
Justification
Rootkits often use DeviceIoControl to interact with (IsDebuggerPresent, Devi- stealth-enabling kernel components. Malware and rootkceIoControl, NtQuerySys- its often use NtQuerySystemInformation to manipulate temInformation) or hide processes or modules. The sequence has rootkit-like potential if the APIs are be(IsDebuggerPresent, Create- ing used to evade detection, hide malicious files, or comDirectoryW, DeviceIoCon- municate with system-level drivers (IRPs), potentially as trol) part of privilege escalation and maintaining persistence. Checking the debugging environment, and CreateThread (IsDebuggerPresent, Cre- could be used for spawning threads to perform malicious ateThread, NtCreateFile) actions. NtCreateFile could be used to create or manipulate files, including hidden or system files. Both NtAllocateVirtualMemory and NtProtectVir(LdrGetProcedureAddress, tualMemory are associated with memory manipulation, NtAllocateVirtualMemory, commonly used by rootkits to manipulate or hide NtProtectVirtualMemory) malicious code in memory. Debugger detection, system directory access, and mem(IsDebuggerPresent, Get- ory allocation all point toward a potential rootkit atSystemDirectoryW, NtAllo- tempting to hide itself in system memory or files. Rootkit cateVirtualMemory) behavior related to memory manipulation and evasion.
we extracted API sequence features from unknown rootkit samples and tested
SeqShield: A Behavioral Analysis Approach to Uncover Rootkits
15
Trained Model
Rootkit or Benign 90 Samples Not part of Model Training
Fig. 4. Model Prediction for Unseen Samples
them using our trained machine learning classifier. We conducted experiments on both high-dimensional feature matrices (containing both relevant and irrelevant features) and optimized lower-dimensional feature sets (comprising only the most significant features). Using a dataset of 90 previously unseen rootkit samples. For the higher-dimensional feature matrix, where both relevant and irrelevant features were included, the Random Forest classifier correctly identified 77 out of 90 samples as rootkits for bigram sequences and 70 out of 90 samples for trigram sequences. However, when using the lower-dimensional feature matrix, which retained only the most relevant features, the classifier performed better by correctly identifying 80 out of 90 samples for bigrams as a rootkit and achieving a perfect detection rate of 90 out of 90 for trigrams. These findings validate our hypothesis that analyzing API call sequences enhances rootkit detection. The improved detection rates with optimized feature selection confirm that reducing dimensionality enhances classification accuracy and minimizes computational overhead. Additionally, the high detection rate for previously unseen rootkits demonstrates the model’s strong generalization ability, making it a viable approach for detecting sophisticated and evolving malware threats in real-world environments.
7
Limitations and Future Work
Future research can focus on expanding the rootkit database, which is currently limited in scope via publicly available database. Institutions with access to larger and more diverse datasets could facilitate the development of a more robust model by contributing additional samples of known rootkits and rootkit-like behaviors. Further, we can extend our approach to quadgrams or more consecutive sequences. However, incorporating quadgrams would significantly increase the feature matrix, potentially up to 10 times, leading to higher computational overhead. Future work should explore optimized feature selection and dimensionality reduction techniques to balance performance with efficiency when scaling to higher-order N-grams. Meanwhile, our approach has demonstrated potential in detecting rootkits or rootkit-like behavior. It may flag off generic processes as
16
P. Ghodeshwar et al.
rootkits, processes such as privilege escalation, that may resemble rootkit-like activity. But this drawback can be overcome by training on a broader rootkit dataset and incorporating higher-order n-gram features.
8
Conclusion
SeqShield presents a robust approach for extracting API sequences and training machine learning classifiers to detect rootkits. Our method achieved an impressive 97.27% accuracy using a bigram feature matrix and 96.17% accuracy with a trigram feature matrix, both evaluated with the Random Forest Classifier. Additionally, we demonstrated how rootkits leverage obfuscation techniques to evade traditional signature-based IDS. By leveraging the Gini impurity index, we effectively reduced the feature space, retaining only the most impactful features while eliminating irrelevant ones. This reduction significantly decreased computational overhead and reduced the dimensionality of the featured matrix without compromising model accuracy. Furthermore, we validated SeqShield’s detection capabilities by analyzing previously unseen rootkit traces, confirming its effectiveness in identifying modern and evasive kernel-level rootkits.
References 1. Abuse.ch: Malwarebazaar, https://bazaar.abuse.ch/, accessed: March 1, 2025 2. ATT&CK, M.: T1014: Rootkit (2024), https://attack.mitre.org/techniques/ T1014/, accessed: Feb 13, 2025 3. Bianco, D.J.: The pyramid of pain (2013), https://detect-respond.blogspot. com/2013/03/the-pyramid-of-pain.html, accessed: March 1, 2025 4. Cuckoo Sandbox Team: Cuckoo sandbox, https://cuckoosandbox.org/, accessed: Feb 13, 2025 5. Dawson, J.A., McDonald, J.T., Shropshire, J., Andel, T.R., Luckett, P., Hively, L.: Rootkit detection through phase-space analysis of power voltage measurements. In: 2017 12th International Conference on Malicious and Unwanted Software (MALWARE). pp. 19–27 (2017). https://doi.org/10.1109/MALWARE.2017.8323953 6. Kruegel, C., Robertson, W., Vigna, G.: Detecting kernel-level rootkits through binary analysis. In: 20th Annual Computer Security Applications Conference. pp. 91–100 (2004). https://doi.org/10.1109/CSAC.2004.19 7. Luckett, P., McDonald, J.T., Dawson, J.: Neural network analysis of system call timing for rootkit detection. In: 2016 Cybersecurity Symposium (CYBERSEC). pp. 1–6 (2016). https://doi.org/10.1109/CYBERSEC.2016.008 8. Mulligan, D., Perzanowski, A.: The magnificence of the disaster: Reconstructing the sony bmg rootkit incident. Aaron K. Perzanowski (10 2010) 9. Nadim, M., Akopian, D., Lee, W.: A review on learning-based detection approaches of the kernel-level rootkit. In: 2021 International Conference on Engineering and Emerging Technologies (ICEET). pp. 1–6 (2021). https://doi.org/ 10.1109/ICEET53442.2021.9659710 10. Nadim, M., Lee, W., Akopian, D.: Characteristic features of the kernel-level rootkit for learning-based detection model training. Electronic Imaging 33(3), 34– 1–34–1 (2021). https://doi.org/10.2352/ISSN.2470-1173.2021.3.MOBMU-034, https://library.imaging.org/ei/articles/33/3/art00003
SeqShield: A Behavioral Analysis Approach to Uncover Rootkits
17
11. Nadim, M., Lee, W., Akopian, D.: Kernel-level rootkit detection, prevention and behavior profiling: A taxonomy and survey (2023), https://arxiv.org/abs/2304. 00473 12. Ortega, A.: Metame: A metamorphic engine for evasion (2024), https://github. com/a0rtega/metame, accessed: Feb 13, 2025 13. Saha, B., Rani, N., Shukla, S.K.: Malxcap: A method for malware capability extraction. In: Meng, W., Yan, Z., Piuri, V. (eds.) Information Security Practice and Experience. pp. 230–249. Springer Nature Singapore, Singapore (2023) 14. Singh, B., Evtyushkin, D., Elwell, J., Riley, R., Cervesato, I.: On the detection of kernel-level rootkits using hardware performance counters. In: Proceedings of the 2017 ACM on Asia Conference on Computer and Communications Security. p. 483–493. ASIA CCS ’17, Association for Computing Machinery, New York, NY, USA (2017). https://doi.org/10.1145/3052973.3052999, https://doi.org/10. 1145/3052973.3052999 15. Suresh Kumar, S., Stephen, S., Suhainul Rumysia, M.: Rootkit detection using deep learning: A comprehensive survey. In: 2024 10th International Conference on Communication and Signal Processing (ICCSP). pp. 365–370 (2024). https: //doi.org/10.1109/ICCSP60870.2024.10543963 16. The Volatility Foundation: The volatility framework, https:// volatilityfoundation.org, accessed: Feb 13, 2025 17. VirusTotal: Virustotal, https://www.virustotal.com/, accessed: Feb 13, 2025 18. Wang, X., Zhang, J., Zhang, A., Ren, J.: Tkrd: Trusted kernel rootkit detection for cybersecurity of vms based on machine learning and memory forensic analysis. Mathematical Biosciences and Engineering 16, 2650–2667 (03 2019). https://doi. org/10.3934/mbe.2019132 19. Wikipedia contributors: Rootkit (2025), https://en.wikipedia.org/wiki/ Rootkit, accessed: Feb 13, 2025 20. Zhou, B., Gupta, A., Jahanshahi, R., Egele, M., Joshi, A.: Hardware performance counters can detect malware: Myth or fact? In: Proceedings of the 2018 on Asia Conference on Computer and Communications Security. p. 457–468. ASIACCS ’18, Association for Computing Machinery, New York, NY, USA (2018). https://doi. org/10.1145/3196494.3196515, https://doi.org/10.1145/3196494.3196515 21. Zhou, L., Makris, Y.: Hardware-assisted rootkit detection via on-line statistical fingerprinting of process execution. In: 2018 Design, Automation & Test in Europe Conference & Exhibition (DATE). pp. 1580–1585 (2018). https://doi.org/10. 23919/DATE.2018.8342267 22. Zhou, Z., Hooker, G.: Unbiased measurement of feature importance in tree-based methods. ACM Trans. Knowl. Discov. Data 15(2) (Jan 2021). https://doi.org/ 10.1145/3429445, https://doi.org/10.1145/3429445