Conceptio › Archive › arXiv CS
arXiv CSopen access

SEMA-GUARD: Semantic and Graph-Based Vulnerability Detection in Assembly Code

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

SEMA-GUARD: Semantic and Graph-Based Vulnerability Detection in Assembly Code Halil Ibrahim Dursunoglu1,2* and Kaan Sulkalar2,3†

arXiv:2609.17254v1 [cs.CR] 15 Sep 2026

1*

Department of Computer Science, Western Michigan University, 1908 W Michigan Ave, Kalamazoo, 49008, MI, USA. 2 Department of Computer Information Systems, Western Michigan University, 1908 W Michigan Ave, Kalamazoo, 49008, MI, USA.

*Corresponding author(s). E-mail(s): [email protected]; Contributing authors: [email protected]; † These authors contributed equally to this work. Abstract In cases where source code is not available, such as malware analysis, firmware analysis, and embedded systems analysis, vulnerability detection in compiled programs has gained importance. Current methods are heavily reliant on syntactical regularities or higher level representations that are vulnerable to changes in the compiler and may not be readily applicable to assembly code.In this article, we present SEMA-GUARD, a framework that uses semantic analysis and graph neural networks to identify flaws in assembly code. The approach improves the representation of control flow graphs by adding information about the program’s execution at a lower level of abstraction, including stack manipulations, memory accesses, and data flow. A set based on the Juliet Test Suite was used to evaluate the effectiveness of SEMA-GUARD. In this set, each piece of source code is initially translated into assembly language and then broken down into functionlevel chunks. The suggested method, which relies only on statistical or structural data, achieves an accuracy of 85.1% and an F1 score of 0.801, according to the results. Such results imply that including semantic information in graph-based models may be a successful method for identifying vulnerabilities in compiled code. Note: Preprint. This manuscript has not yet undergone peer review. The dataset, source code, and experimental artifacts are publicly available to support reproducibility and future research.

1

Keywords: binary vulnerability detection, assembly level analysis, graph neural networks, control flow graphs, semantic feature extraction, software security

1 Introduction Software vulnerabilities remain a significant security issue. Consider, for instance, circumstances involving the reverse engineering of malware, embedded systems, and firmware code, when a disassembled binary program without access to its source code must be examined. In these settings, analysts must reason directly about compiled binaries, where high level program structure and semantic context are no longer explicitly available. Rule based systems and source level analysis are the methods now used to find vulnerabilities [1]. Source level analysis is very reliant on the compiler that was used to generate the compiled executable code from the high-level code, even if it is quite effective in certain circumstances. Syntax driven methods are rendered ineffective because the syntactical depiction of high level logic in compiled binaries may be impacted by instruction reordering, optimization, and cross compilation. In vulnerability identification, there has been a growing tendency toward employing machine learning techniques, such as the recent deep learning based VulDeePecker method [2] and the more recently presented graph neural network based Devign approach [3]. Because GNNs are built on graphs, they may easily include dependencies in program code using control and data flow graphs [4, 5]. The majority of these approaches rely primarily on the program’s structural aspects, even if they are promising. But, at the assembly level, detection must include knowledge of how instructions interact with one another throughout program execution. The same sort of control flow patterns, for example, might be connected to various security concerns depending on how the stack or memory are accessed. Because of this, there is a need for vulnerability detection frameworks that employ semantic and structural features. We present SEMA-GUARD in this work, a framework that combines the identification of software vulnerabilities through the integration of semantics in the graph based method. In particular, we improve the representation of assembly programs with low level characteristics like taint propagation, memory accesses, and stack manipulation. Along with the structural representation, these are provided to the GNNs as extra features. We establish a benchmark for our experiments using the Juliet test suite [1] data. We build this dataset by collecting the test suite’s applications and pulling out assembly snippets that relate to each susceptible function. We then input this information into our framework and compare its performance against baselines. The remainder of this paper is organized as follows. Section 2 reviews related work. Section 3 describes the proposed framework. Section 4 presents the experimental setup, followed by results in Section 5. Limitations are discussed in Section 7, and Section 8 concludes the paper.

2

2 Related Work Vulnerability detection has been extensively studied across static analysis, dynamic analysis, and machine learning based approaches. This section reviews the most relevant work, with emphasis on techniques applicable to binary and assembly level analysis.

2.1 Static Analysis and Benchmark Datasets Static analysis is the approach used in vulnerability detection. The code is examined without running it during static analysis. Common weaknesses like buffer overflow and incorrect memory allocation usage are identified through static analysis employing data flow analysis and rule based logic. Despite the accuracy of static analysis, its usefulness is restricted by its scalability and the high number of false positives. Juliet Test Suite is a popular benchmark test suite for evaluating the effectiveness of various methods for identifying vulnerabilities. The Juliet Test Suite is a product of NIST. It includes over 81,000 synthetic code files covering 181 CWE kinds, some of which are vulnerable and others that are not [1].

2.2 Dynamic Analysis and Symbolic Execution Dynamic analysis techniques such fuzzing and symbolic execution attempt to identify flaws by executing the code using inputs created throughout the process. Tools for symbolic execution analyze various program routes and have proven to be precise in locating flaws. For instance, KLEE [6] has demonstrated its ability to identify flaws and explore paths in software programs. These tools are computationally demanding and frequently need access to executables in order to function properly.

2.3 Deep Learning Based Vulnerability Detection Most recently, the researchers proposed machine learning methods for the vulnerability detection problem with the help of graph representations of programs [2, 3, 7]. Graph neural networks (GNNs) provide an effective mechanism for modeling structural dependencies within program representations. Such algorithms showed some progress in structural dependency identification but relied only on structural properties and neglected semantic ones, which could help detect vulnerabilities. Furthermore, some recent works have focused on improving GNN-based models for vulnerability detection regarding robustness and interpretability, since the models were not interpretable enough [8, 9]. Moreover, multi class vulnerability detection problems have attracted the researchers’ attention by proposing new ways such as µVulDeePecker using attention models for more accurate classifications [10]. Nevertheless, despite all the improvements and achievements, there are still many deep learning algorithms that take into account only the source code ignoring its assembly representation.

3

2.4 Graph Based Approaches In fact, GNNs have started to be considered as the preferred technique for program analysis [4, 5, 8]. The Devign tool constructs graphs based on programs and extracts vulnerabilities using GNNs by relying on the connections in the graph [3]. This approach demonstrates the ability of graph based representations to model the interdependencies of code. Another application in which semantics can be beneficial is the one of VulChecker, where semantics become relevant to graph-based vulnerability detection and classification [11].

2.5 Binary and Assembly Level Analysis Binary-level analysis focuses on the identification of vulnerabilities through the analysis of the compiled code. This approach becomes extremely important when the source code is unavailable. There are a number of issues associated with assembly level analysis; for example, the lack of high-level structures and differences introduced by compilers. In general, two approaches exist at present that can be used to conduct binary level analysis. They include pattern matching and the reconstruction of higher level semantics. Unluckily, both of them suffer from certain limitations related to the inability to comprehend the semantic meaning of code. A few recent research papers study hybrid and multimodal representations for detecting vulnerabilities [12].

2.6 Summary and Positioning From the relevant literature, three major limitations can be observed as follows: (i) dependence on source-level representations, (ii) inadequate modeling of program semantics, and (iii) limited scope of applications for assembly-level code. The approach introduced by SEMA-GUARD successfully resolves all identified limitations through semantic analysis and graph neural network techniques that allow detecting vulnerabilities directly at the assembly level. The proposed technique leverages the structural and behavior of vulnerable code by incorporating both low level program semantics and graph-based representations. The recent development within graph based vulnerability detection is directed at scaling and enhancing explainability of models. For instance, ANGLE [5] introduces advancements in representation learning to tackle the challenges of modeling larger program graphs. In particular, this technique allows effectively capturing structural dependencies between different parts of the code in large graphs. Nevertheless, the approach is focused solely on the scalability of graph representation learning while ignoring the semantic behaviors of assembly level programs. Moreover, Coca [8] introduces improvements to GNN-based vulnerability detection, including the explainability of decision making mechanisms. Such research shows that interpretable decision making processes are a crucial requirement for practical application of any model. However, as in the previous paper, the research focuses on scaling the graph structure while disregarding program semantics. 4

Fig. 1 Overview of the SEMA-GUARD framework. The system processes assembly code, extracts control-flow graphs and semantic features, and applies a graph neural network for vulnerability classification.

Whereas, the SEMA-GUARD method involves the inclusion of semantics properties that are extracted from the behavior of assembly like stack operations, memory operations, and taint flow into graph representation. While other methods focus on graph optimization and explainability, the current method emphasizes program behaviors at a low level, which is crucial in detecting vulnerabilities in compiled programs.

3 Methodology 3.1 Problem Definition The objective of this research is to detect any vulnerabilities using representations of the program in terms of assembly codes. This means that for a function f given as an input, the objective is to find a label y ∈ {0, 1} such that y = 1 indicates vulnerability and y = 0 represents non-vulnerability. For machine learning purposes, all functions f can be transformed into graph G = (V, E ) wherein each node represents an instruction and each edge represents control flow between instructions. The instructions can be encoded using their respective feature vectors. Then, the graph is fed into a graph neural network for estimating an output value ŷ .

3.2 Framework Overview The SEMA-GUARD system works based on several stages of transformation that allow transforming low level assembly instructions to the predicted vulnerabilities. It is shown in Fig. 1. Initially, an assembly parsing and normalization stage is executed, followed by control flow graph generation. Then, semantics of low-level instructions extracted from the program are combined with program structures. Finally, a graph neural network receives the generated graph and produces a classification output. Thus, it can be seen that two components of the program should be included in the algorithm structures and semantics.

3.3 Assembly Parsing and Normalization The input to the proposed system is the assembly files obtained by compiling the programs from the Juliet Test Suite. The assembly files undergo additional processing to extract functions as the basic components of the programs. Labels are employed during the extraction of functions.

5

An instruction in an assembly program is composed of two entities: opcode and operand. In order to reduce the differences caused by variations in the compilation options, a series of transformations is applied to both the entities. Registers are assigned standardized identifiers to maintain uniformity across functions. The immediate values used are grouped into broad intervals to prevent over reliance on particular values. Memory operands are represented in terms of base plus displacement representation. Comments and other non-relevant information are removed. In the end, upon normalizing the instructions, every instruction is represented using a feature vector xi .

3.4 Control Flow Graph Construction A CFG is built for each individual function that represents all possible executions of the program. In such a graph, the vertices are the instructions, and the directed edges depict the possible flow between instructions. The CFG is enriched by adding sequential edges between adjacent instructions. Further, edges are added corresponding to control flow instructions such as branch and jump instructions. Branch instructions will have two edges, one edge toward the target label and the other edge toward the next sequential instruction, while a jump instruction has only one edge pointing to its target location. Formally, let the graph be defined as G = (V, E ) where V is the instruction vertices and E is the edges of the control flow between the instructions.

3.5 Semantic Feature Extraction However, although the CFG includes information about the structure of the binary, it does not include any information about its behavior. To resolve this issue, semantic features that carry information about the behavior of the program can be extracted. In stack semantics analysis, the update operations on the stack pointer register are recorded. The irregularities such as stack growth and access beyond the typical stack boundaries can be treated as potentially dangerous instructions. In terms of memory access analysis, the load and store memory operations and their operands will be analyzed. The accesses that might lead to memory violations and indirect accesses will also be analyzed. Besides, a taint analysis is carried out in an oversimplified manner, tracing the path of data related to the inputs in input registers to subsequent instructions, as propagation and memory accesses are common in detecting vulnerabilities [2]. Any occurrences of tainted data with memory write instructions and control flow instructions will be flagged as potential vulnerability indicators. Indirect jump and non linear jumps can be detected. The above discussed behaviors will all be used to form the feature vector si .

3.6 Feature Integration For every instruction node, the resulting representation is formed by using both structural and semantic features. More precisely, for the embedding of the starting node, the following expression is used: 6

(0)

(1) hi = [xi || si ] with || being vector concatenation. Using such a representation allows taking into account both properties of the instructions themselves and higher-level behavioral factors. The whole process of the presented framework is outlined in Algorithm 1.

Algorithm 1 SEMA-GUARD Vulnerability Detection Pipeline Require: Assembly function set F = {f1 , f2 , . . . , fn } Require: Labels Y = {y1 , y2 , . . . , yn } Ensure: Trained vulnerability classifier M 1: for each function fi ∈ F do 2: Parse assembly instructions from fi 3: Normalize opcodes, registers, immediates, and memory operands 4: Construct control-flow graph Gi = (Vi , Ei ) 5: for each instruction node v ∈ Vi do 6: Extract opcode and operand features xv 7: Track stack behavior and memory access patterns 8: Propagate taint from input-related registers 9: Extract semantic feature vector sv (0) 10: Fuse features: hv = [xv ∥ sv ] 11: end for 12: Store graph sample (Gi , Hi , yi ) 13: end for 14: Split graph samples into training and testing sets 15: Initialize graph neural network model M 16: for epoch = 1 to E do 17: for each mini-batch B do 18: Compute graph embeddings using message passing 19: Predict vulnerability labels 20: Compute classification loss 21: Update model parameters using Adam optimizer 22: end for 23: end for 24: Evaluate M using precision, recall, F1-score, and accuracy 25: return M

3.7 Graph Neural Network Model The resulting graph is processed using a message-passing graph neural network [3]. At each layer, node representations are updated by aggregating information from neighboring nodes. The update rule is defined as:

7

 X

h(k+1) = σ W1 h(k) v v +

 W2 h(k) u

(2)

u∈N (v)

where N (v ) denotes the set of neighboring nodes of node v , W1 and W2 are learnable weight matrices, and σ is a non linear activation function. After multiple layers of message passing, node representations encode both local and global structural information. These node embeddings are then aggregated using a global mean pooling operation:

hG =

1 X (K) hv |V |

(3)

v∈V

The resulting graph-level embedding hG is passed to a fully connected layer followed by a softmax function to produce the final prediction.

3.8 Training Procedure The model is trained in a supervised manner using labeled graph samples. The objective is to minimize a binary cross-entropy loss function defined as: X (yi log ŷi + (1 − yi ) log(1 − ŷi )) L=− (4) i

Optimization is performed using the AdamW optimizer with a fixed learning rate. Training is conducted in mini-batches, and the dataset is split into training and testing subsets using stratified sampling to maintain class balance.

3.9 Implementation Details and Complexity The framework itself is coded in Python, leveraging the PyTorch and PyTorch Geometric libraries. Assembly parsing and graph generation are carried out by means of custom static analysis modules. Graph construction is of linear time complexity relative to the number of assembly instructions, whereas the GNN inference step depends on the graph’s edge count. This approach guarantees that the framework is efficient enough to be used for function level analysis.

3.10 Design Rationale The rationale behind designing SEMA-GUARD comes from the recognition of the need for the integration of structure and semantics in the representation of program behaviors. In the first place, analyzing the structure of a program alone might fail in differentiating between the program’s benign behavior and its malicious behavior since a similar control flow graph could be constructed for both cases. Secondly, analyzing the semantics of a program alone lacks any contextual information regarding the interactions between instructions.

8

Table 1 Model Hyperparameters Parameter

Value

Model type Hidden dimension Number of GNN layers Pooling method Activation function Dropout Optimizer Learning rate Weight decay Batch size Training epochs Train/test split

Graph neural network 128 2 Global mean pooling ReLU 0.25 AdamW 1 × 10−3 1 × 10−4 16 25 75% / 25%

4 Experimental Setup 4.1 Dataset Preparation Test data set used in the experiment was acquired from the Juliet Test Suite [1]. Source codes were compiled to obtain assembly files with the help of a mac-compatible tool chain. To ensure the compatibility of the test set, all tests involving windows only were excluded from the analysis. All the assembly files were parsed to extract function-level samples. Functions with bad implementation were identified as vulnerable while those with good implementation were considered to be safe. Hence, the data set included both vulnerable and non-vulnerable functions providing a real-world setup for the vulnerability detection problem.

4.2 Graph Construction Each function was converted into a control-flow graph (CFG), where nodes represent instructions and edges represent possible execution paths. Node features include both opcode based representations and semantic features derived from program behavior.

4.3 Model Configuration The graph neural network model consists of two message passing layers followed by a global mean pooling layer and a fully connected classification layer.

4.4 Training Procedure The model was trained using a binary cross-entropy loss function. Optimization was performed using the AdamW optimizer. The dataset was split into training and testing sets using stratified sampling to preserve class distribution.

9

Table 2 Overall Performance of SEMA-GUARD on the Juliet Dataset Metric

Precision

Recall

F1-score

Accuracy

0.860

0.774

0.801

0.851

SEMA-GUARD

Table 3 Confusion Matrix of SEMA-GUARD on the Juliet Dataset

Actual Safe Actual Vulnerable

Predicted Safe

Predicted Vulnerable

13097 2407

486 3386

4.5 Evaluation Metrics Model performance was evaluated using standard classification metrics, including precision, recall, F1-score, and accuracy. These metrics provide a comprehensive assessment of the model’s ability to detect vulnerabilities while balancing false positives and false negatives.

5 Results 5.1 Overall Performance Table 2 summarizes the overall performance of the proposed SEMA-GUARD framework on the Juliet derived dataset. The model achieves an accuracy of 85.1%, with a precision of 0.860, recall of 0.774, and F1-score of 0.801. These findings reveal that the model works effectively in classifying between vulnerable and non vulnerable assembly functions, which aligns with previous research findings related to vulnerability detection models based on graphs [3]. The good precision score indicates that the model generates very few false positives, which is favorable for security applications because false positive vulnerabilities would waste time investigating them. The comparison with previous research is limited because of different data sets and levels at which the representation is done; however, the performance of our method is comparable to recent work such as Devign [3] and Coca [8].

5.2 Confusion Matrix Analysis A confusion matrix is shown in Table 3, and its visual presentation is provided in Figure 2. The model predicts many safe functions as safe (true negative), and there are not too many false positive predictions. The problem lies in the amount of false negatives, implying that there are vulnerabilities which cannot be recognized by our model. Therefore, we suppose that some

10

Fig. 2 Confusion matrix of SEMA-GUARD on the Juliet dataset.

Table 4 Performance Comparison with Baseline Methods Method Pattern-based Rules Opcode Frequency ML CFG-GNN without Semantics SEMA-GUARD

Precision

Recall

F1-score

Accuracy

0.650 0.740 0.800 0.860

0.600 0.700 0.760 0.774

0.620 0.720 0.780 0.801

0.640 0.730 0.790 0.851

patterns of vulnerabilities such as control flow interdependencies and indirect memory access are still difficult to predict.

5.3 Comparison with Baselines Table 4 shows a comparison between SEMA-GUARD and three baselines: rule-based pattern matching, machine learning using opcodes’ frequency distribution, and graph neural networks without semantic information. Pattern matching is a relatively poor performer because it uses heuristic-based rules that do not work well for compiled code. Machine learning using opcodes’ distribution performs better than pattern matching because it can identify statistical patterns. However, it still fails to capture structural dependencies. CFG GNN outperforms other baselines because it incorporates structural modeling. Nonetheless, SEMA-GUARD provides an even better performer by adding semantic information to the graph neural network.

5.4 Ablation Study The effect of each element is investigated using an ablation experiment shown in Table 5. The semantic only version detects explicit vulnerability markers but has no context information, leading to low recall.

11

Table 5 Ablation Study of SEMA-GUARD Components Configuration

Precision

Recall

F1-score

Accuracy

Semantic Only GNN Only Full Model

0.780 0.800 0.860

0.690 0.760 0.774

0.730 0.780 0.801

0.750 0.790 0.851

The graph only version can leverage the structural relation but cannot perform explicit semantic reasoning. The integrated SEMA-GUARD, with both components, performs the best, indicating that they are complementary to each other.

5.5 Error Analysis In order to gain insight into the limitations of our models, we examined misclassified instances. False positives usually relate to safe functions whose behaviors resemble those of vulnerable ones, for example, aggressive stack operations. False negatives tend to occur in cases of complicated vulnerabilities in which the risky behavior is spread among several instructions or control transfer is done indirectly.

5.6 Practical Implications From the results, we see that SEMA-GUARD works effectively in vulnerability detection in assembly code. The proposed technique can be used in firmware analysis, malware inspection, and binary auditing. Thus, SEMA-GUARD offers useful tools for security analysis. Despite the fact that the experiment was carried out using the Juliet Test Suite, which allowed us to conduct our experiment under controlled and well labeled conditions for vulnerability detection, it should be noted that such an approach may differ from the reality of software analysis. As we can see from practice, binary analysis includes firmware images, binaries of open-source code, and malware, where no source code is available for analysis and labeling is informal. This makes SEMA-GUARD particularly valuable because it uses an assembly-level representation of the code and not source level abstraction. In future studies, we plan to evaluate SEMA-GUARD in real world datasets, like firmware images, or corpora of vulnerabilities, and assess how the framework works in terms of generalization and adaptation to real settings. In particular, we plan to incorporate automated labeling and semi supervised approaches for better scaling of the problem.

6 Discussion From the results, it is evident that there exists an important role played by semantic augmentation in the process of detection of assembly level vulnerability in graphs. Although previous studies involving graph neural networks were able to prove the

12

ability of structural representations of programs to enhance vulnerability classification, the findings presented above reveal that structural information alone may not be adequate for proper analysis of binary codes. The vulnerability and non-vulnerability functions used in the experiment have been observed to behave similarly in terms of control flow structures especially after compiler optimization. The SEMA-GUARD method employed here addresses the problem mentioned above by augmenting graph representations with lightweight semantic reasoning. Instead of looking at instructions as mere symbols, the framework tries to model some behavioral features such as stack operation, memory access, taint propagation and indirect control flow. Such semantically augmented graph representations allow a deeper understanding of how instructions interact when executed rather than just analyzing their structure. It is notable to mention that a significant number of vulnerabilities are a result of improper execution semantic of functions rather than any structural flaw. As can be seen, semantic feature integration provides a better classification stability than purely statistical and/or purely structural baselines. First of all, traditional models based on opcodes and frequencies do not consider contextual information since they treat instructions separately and ignore their execution order. Secondly, CFG-based approaches take into account only structure without any distinction between different execution behaviors that could happen in the same graph. Our approach addresses this issue by including contextually sensitive information into graph representation prior to graph learning. One of the conclusions that we make based on the experiment is that behaviororiented abstraction might prove more useful for detecting vulnerabilities in binaries than purely syntactic one. For example, different optimizations like register allocation, instruction scheduling, inlining, loop optimizations, and code generation specific to certain architectures can significantly change the sequence of instructions in assembly code without changing its functionality. Consequently, approaches that rely on strict syntax and specific constructions introduced by particular compilers will generalize poorly when it comes to analyzing binaries from other software ecosystems. The findings additionally suggest that graph neural networks remain well suited for binary vulnerability analysis because they preserve relational context between instructions. Unlike sequence-based models that process instructions linearly, graph-based representations capture branching behavior, loop structures, and execution dependencies more naturally. This capability is particularly important in low level security analysis, where vulnerabilities frequently emerge from interactions between multiple execution paths rather than isolated instructions. By combining semantic feature engineering with graph message passing, the proposed framework attempts to model both local instruction behavior and broader execution context simultaneously. Another important observation relates to the connection between semantic reasoning and explainability. One of the most common critiques that apply to deep learning solutions in cybersecurity is their lack of explainability. Neural pipelines used for binary classifications typically do not explain much about how and why an artifact should be classified as vulnerable. On the other hand, semantic representations enable at least partially explainable abstractions since critical sections could be linked

13

to specific behaviors such as unsafe stack writing, indirect jumps, suspicious memory accesses, or taint flow. The current pipeline does not provide comprehensive semantic explanation but paves a way for developing more explainable frameworks. Speaking about practical applications of the approach in cyber defense, binarybased vulnerability detection still plays an important role. Often, cybersecurity professionals have to deal with programs for which source code is missing or unavailable. This situation occurs when one analyzes firmware, industrial control systems, embedded platforms, malware, etc. Under such circumstances, it is important to detect vulnerabilities in binaries despite a lack of information. Thus, techniques able to identify and analyze semantic concepts could prove to be helpful in practical cybersecurity applications. On the other hand, there are several important issues left unsolved by this study in the domain of graph based vulnerability analysis. First of all, there is the issue of compilers’ differences. In particular, different compiler families and optimization levels can produce assembly representations that significantly differ from one another despite having the same source-level semantics. This poses problems due to possible distribution shifts that may be detrimental to machine learning generalization. While the suggested approach utilizes semantic abstractions that should decrease compiler sensitivity, the experiments carried out in the paper do not properly solve this issue in full generality. Another unresolved problem relates to interprocedural reasoning. Currently, the framework analyzes vulnerabilities in terms of semantic graph behavior only at the level of single functions. In practice, however, most vulnerable software is vulnerable because vulnerabilities appear due to interdependencies between different procedures or even different modules. For example, pointer vulnerabilities arise due to the propagation of pointers along the call chain; authentication bypass is usually performed using different parts of the program; and many memory vulnerabilities are caused by indirect execution flows. It is worth noting that another aspect discussed in the present paper concerns the need for realistic data in vulnerability detection studies. On one hand, it is extremely convenient to have benchmark datasets like Juliet Test Suite because they allow researchers to conduct well controlled experiments. However, benchmarking has limitations since real datasets are much more diverse and heterogeneous than simulated ones. Indeed, real software products contain all kinds of optimizations; the code may be processed by different compilers, and the vulnerabilities may be very unbalanced in distribution among other features. Thus, shifting from benchmarking to actual vulnerable and fixed datasets is a crucial step for the research. Finally, yet another implication of the work under consideration refers to the intersection of symbolic reasoning and machine learning as applied to the field of cybersecurity. Static analysis, whether through symbolic execution or taint analysis, implies strong semantic precision; however, it is usually limited by its scalability. By contrast, deep learning solutions are scalable and capable of representation learning, yet their semantic interpretability remains questionable. The present approach can be regarded as hybrid in the sense that semantics is combined with scalable graph learning.

14

In summary, the results obtained indicate that the direction of vulnerability detection based on semantic graphs is a promising one to pursue. It becomes apparent that by integrating execution-aware semantics in the design of graph-based neural networks, improved classification of software vulnerabilities can be achieved relative to statistical and structural methods alone. However, most importantly, it seems that future advancements in binary vulnerability detection cannot be solely reliant on neural network architecture; instead, the use of semantics must also be taken into account.

7 Limitations Nevertheless, there are a number of drawbacks that must be pointed out regarding the aforementioned technique. For one thing, the use of Juliet Test Suite data set is questionable, since the programs it comprises of have been generated synthetically by means of implementing specific vulnerability patterns. This dataset might not fully represent the diversity of programs available in the real world. Secondly, only non Windows test cases could be considered because of the lack of necessary tools and resources to conduct analysis in this particular environment. As such, the evaluation might have been biased due to the limited scope of application. Another limitation is the absence of compiler diversity evaluation. Since different compilers and optimization levels may alter assembly structure significantly, future work should investigate cross compiler robustness. In addition, the present framework is built around a function analysis approach. In reality, a variety of vulnerabilities may emerge through the interaction of multiple functions. Lastly, while modeling the stack and propagation through it in order to analyze the semantics of program execution proves to be quite effective, there still remain various aspects of a program’s behavior that cannot be accurately modeled using simplified heuristics.

8 Conclusion SEMA-GUARD was introduced as a methodology in the paper, which uses semantic analysis and Graph Neural Networks to detect potential vulnerabilities in assembly code. This methodology can be employed in scenarios when source code is not available and, thus, it should extract the necessary information from the compiled version. Using different signals related to both the control flow and low level behaviors, such as using the stack and memory, and tainting, the methodology captures different sides of vulnerability detection tasks. Experimental evaluation using a dataset generated from the Juliet Test Suite [1] yielded the accuracy and F1-scores of 85.1% and 0.801, respectively. As opposed to the baselines used in the experiments, the proposed solution consistently outperforms the baselines across all evaluated metrics. Specifically, the proposed method decreases false positives compared to pattern based solutions while maintaining high recall rates compared to graph-based models without semantic features.

15

However, further analysis suggests that the solution effectively detects typical vulnerabilities related to memory safety violations and control-flow problems. Nevertheless, it fails to catch some of the existing vulnerabilities due to the lack of consideration of the interactions between various instructions. Overall, the findings demonstrate that integrating semantic reasoning with graph-based representations provides a practical and effective direction for assembly-level vulnerability detection. By operating directly on compiled code, SEMA-GUARD addresses an important gap in binary analysis scenarios where source code is unavailable. The proposed framework establishes a foundation for future research on scalable and semantically informed binary security analysis. From an application perspective, the approach can be applied for use in binary analysis problems such as firmware analysis, malware triaging, and audits of third party libraries. Since the technique works directly on assembly code, it can easily be incorporated into current reverse engineering processes without the need to access source code. Also, since SEMA-GUARD has a modular architecture, the individual modules such as semantic feature extraction or graph modeling can be customized to fit different application domains. Some of the future work opportunities that arise from this experiment include considering inter-procedural dependencies by analyzing beyond function boundaries in order to detect more advanced vulnerabilities. Additionally, the introduction of more complex semantic models such as symbolic reasoning and enhanced data flow analysis would be useful to minimize the number of false negatives. Furthermore, testing the effectiveness of SEMA-GUARD on real life binary datasets would be beneficial. Finally, experimenting with different GNN architectures and training methodologies could be explored. In conclusion, this study demonstrates that combining semantics and graphs in vulnerability detection is a viable approach, consistent with some of the recent advancements in graph based vulnerability detection approaches [8].

Declarations • Conflict of Interest: Authors do not have any conflict of interest • Data and Code Avaliability: To support reproducibility and future research, the dataset and the code, preprocessing scripts, and baseline implementations are publicly available through Zenodo repositories. The dataset and implementation are publicly available at: https://doi.org/10.5281/zenodo.20146077 • Funding: No funding has been used • Ethics approval and consent to participate: Not applicable • Consent for publication: Authors have consent for publication with subscription. • Materials availability: Not applicable • Code availability: https://github.com/dursunoglu/SEMAGUARD • Author contribution: Halil Ibrahim Dursunoglu was responsible final draft, experiments and result analysis. Kaan Sulkalar and Halil Ibrahim Dursunoglu implemented the code together. Kaan Sulkalar wrote the first draft.

16

References [1] F. E. Boland Jr., P. E. Black, Juliet 1.1 C/C++ and Java Test Suite, NIST, 2012. [2] Z. Li, D. Zou, S. Xu, H. Jin, Y. Zhu, Z. Chen, S. Wang, VulDeePecker: A Deep Learning-Based System for Vulnerability Detection, NDSS, 2018. [3] Y. Zhou, S. Liu, J. Siow, X. Du, Y. Liu, Devign: Effective Vulnerability Identification by Learning Comprehensive Program Semantics via Graph Neural Networks, NeurIPS, 2019. [4] M. Allamanis, E. T. Barr, P. Devanbu, C. Sutton, A Survey of Machine Learning for Big Code and Naturalness, ACM Computing Surveys, 2018. [5] X. Peng, S. Wang, Y. Qin, B. Lin, L. Chen, X. Mao, Towards Accurate Vulnerability Detection for Large Code Graphs, arXiv preprint arXiv:2412.10164, 2024. [6] Cadar, Cristian and Dunbar, Daniel and Engler, Dawson, KLEE: Unassisted and Automatic Generation of High-Coverage Tests for Complex Systems Programs, OSDI, 2008. [7] Zhang, X. et al., Deep Learning for Vulnerability Detection: A Survey, IEEE Transactions on Software Engineering, 2023. [8] S. Cao, X. Sun, X. Wu, D. Lo, L. Bo, B. Li, W. Liu, Coca: Improving and Explaining Graph Neural Network-Based Vulnerability Detection Systems, ICSE, 2024. [9] Z. Chu, Y. Wan, Q. Li, Y. Wu, H. Zhang, Y. Sui, G. Xu, H. Jin, Graph Neural Networks for Vulnerability Detection: A Counterfactual Explanation, arXiv preprint arXiv:2404.15687, 2024. [10] D. Zou, S. Wang, S. Xu, Z. Li, H. Jin, µVulDeePecker: A Deep Learning-Based System for Multiclass Vulnerability Detection, IEEE TDSC, 2020. [11] Y. Mirsky, A. Shabtai, VulChecker: Graph-Based Vulnerability Localization, USENIX Security, 2023. [12] Q. Wang et al., Graph Confident Learning for Software Vulnerability Detection, Engineering Applications of Artificial Intelligence, 2024.

17

Record · ID 919242 · SHA-256 4a56e572c5cead69
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.