Deep Reinforcement Learning for Misbehavior Detection Under Partially Observable V2X Data
arXiv:2609.31217v1 [cs.NI] 25 Sep 2026
Roshan Sedar and Charalampos Kalalas Abstract— Misbehavior detection in vehicle-to-everything (V2X) systems is essential for ensuring the semantic correctness of exchanged messages and preventing the dissemination of falsified information. Existing data-centric misbehavior detection approaches largely rely on statistical validation or supervised machine learning models under the implicit assumption of fully observable V2X streams. In practice, however, vehicular environments are inherently partially observable due to hardware failures, intermittent connectivity, and environmental occlusions. Moreover, missingness itself can be strategically exploited by adversaries to evade detection. In this paper, we study misbehavior detection under incomplete V2X observations and propose a deep reinforcement learning (DRL)-based detection framework that learns adaptive policies with incomplete data. We further introduce an adversarial threat model in which attackers exploit or deliberately induce missingness to evade detection, including evasion via natural occlusions and adversarial feature suppression. Extensive experiments conducted on the VeReMi dataset under various missingness patterns demonstrate that DRL significantly outperforms a powerful XGBoost baseline under natural partial observability. However, results also reveal a critical vulnerability: DRL policies can be highly susceptible to evasion attacks that strategically exploit natural missingness. In contrast, DRL exhibits more gradual degradation under direct feature suppression compared to static tree-based models. Index Terms–Misbehavior detection, partial observability, reinforcement learning, adversarial perturbations, feature suppression
I. I NTRODUCTION Misbehavior detection has recently become one of the primary concerns in emerging vehicle-to-everything (V2X) systems characterized by pervasive sensing, computing, and connectivity capabilities. Traditional cryptographic mechanisms, while essential for authentication and integrity, are insufficient to guarantee the semantic correctness of transmitted data. As a result, data-centric misbehavior detection relies on the analysis of spatiotemporal measurement streams, captured at different locations and time instances, to determine the trustworthiness of information by detecting abnormal behaviors with the aid of statistical analysis, rulebased validation, or machine learning (ML) models [1]. Nevertheless, a fundamental yet often overlooked challenge for reliable misbehavior detection resides in the comThis work has been partially supported by the project PID2024-160944OB-I00 (SEASIDE) funded by MICIU/AEI/10.13039/501100011033 and the internal project BARYCENTRE (300960). Roshan Sedar and Charalampos Kalalas are with the Sustainable Artificial Intelligence Research Unit, Centre Tecnològic de Telecomunicacions de Catalunya (CTTC/CERCA), 08860 Castelldefels, Spain
{roshan.sedar, ckalalas}@cttc.es
pleteness of aggregated information. In practice, the emergence of missing/incomplete data in the fused vehicular measurement streams is inevitable [2]. Partial observability of mobility information can be generally attributed to the following factors: i) hardware failures, where the malfunction of vehicle components (e.g., synchronization failures or errors in sensor readings) may result in persistent missing observations for one or multiple state variables of the vehicle; ii) intermittent connectivity, as a result of the high mobility, dynamic network topology, and wireless channel impairments which may result in sporadic outages and packet losses for consecutive time-steps; and iii) environmental factors, such as buildings, large vehicles, adverse weather conditions, and road curvature, which often introduce occlusions, leading to incomplete situational awareness. Collectively, these factors result in a partially observable V2X environment in which detection methods must reason under significant epistemic uncertainty about the true system state. The challenge of partial observability becomes particularly relevant when malicious actors exploit missingness to intentionally obscure inconsistencies (e.g., via stealthy data injections or subtle perturbations) and evade detection. In such scenarios, the adversarial exploitation of missingness aims to degrade the effectiveness of misbehavior detection systems and compromise their trustworthiness. Existing research on data-centric misbehavior detection includes plausibilitybased detectors [3], consistency checks across neighboring vehicles [4], trust and reputation schemes [5], as well as supervised [6], [7] and unsupervised [8] ML models. Despite this growing body of work, the majority of these methods implicitly assume either fully observable V2X streams or the straightforward exclusion of unobserved samples to overcome data incompleteness. Thus, the effectiveness of misbehavior detection under partial observability remains unexplored. In [9], the authors aim to answer whether imputation or missing-tolerant classification yields the best misbehavior detection performance in incomplete V2X streams. Missingness is synthetically induced with different ratios, mechanisms, and distributions using a multi-factor amputation framework that allows a comprehensive benchmark comparison of missing data handling strategies. However, missing data is solely treated as a static classification challenge with merely incomplete feature vectors and without ensuring sequential reasoning over time, since learning relies on supervised classification. As such, adaptive detection decisions (e.g., policy learning instead of fixed classification outputs) cannot be guaranteed. Additionally, to the best of our knowledge,
no prior work considers adversarially induced missingness as a deliberate attack vector against misbehavior detection models. This gap leaves deployed detection systems fundamentally exposed to a class of evasion threats for which they were not designed. All these factors underscore the need for a dedicated study that explicitly accounts for partially observed V2X environments caused by both natural occlusions and adversarial manipulation. To bridge these gaps, this paper proposes a novel deep reinforcement learning (DRL) framework for misbehavior detection specifically designed for partially observable vehicular environments. Unlike prior work that treats missingness as a preprocessing task or static classification problem, we reframe misbehavior detection as a sequential decision-making problem that allows reasoning under incomplete vehicular streams. Detection robustness is empirically evaluated using the VeReMi benchmark dataset [10] by considering three scenarios: i) natural occlusions via standard missingness patterns; ii) adversarial evasion that exploits naturally occurring missing regions to inject stealthy perturbations; and iii) adversarial feature suppression that actively induces missingness as an attack strategy. A comprehensive benchmark assessment against an XGBoost baseline reveals both the robustness advantages and security vulnerabilities of DRL under partial observability. The remainder of this paper is organized as follows. Section II introduces the considered missingness scenarios in our study. Section III presents the proposed DRL-based misbehavior detection mechanism and the adversarial threat model. Section IV details the experimental setup, while Section V provides a comprehensive performance evaluation and benchmark comparison of the proposed framework across all considered scenarios. Finally, Section VI concludes the paper and lists future work. II. M ISSINGNESS S CENARIOS We consider a set of missingness scenarios, as illustrated in Fig. 1, to systematically evaluate the misbehavior detection capabilities of the proposed DRL-based framework under partial observability. The considered scenarios progressively transition from natural missingness to adversarially induced suppression, thereby enabling a comprehensive assessment of detection performance across diverse and challenging observability conditions. Next, we elaborate on the considered scenarios. SC1: Natural occlusions. Due to the dynamic and dense nature of vehicular environments, data missingness in V2X streams is inherent and unavoidable. Complete and synchronized perception of the vehicular system state is thus rarely attainable in practice. As such, the first scenario, shown in Fig. 1a, models partial observability arising from non-malicious hardware, communication, and environmental factors. In real-world vehicular networks, natural occlusions stem from transient sensor failures, multipath fading, physical obstruction between vehicles or buildings, packet collisions, and adverse weather conditions.
V2X Raw Data (BSMs)
Natural Occlusion (Missing Features)
DRL Misbehaviour Detection
(a) Natural occlusions. V2X Raw Data (BSMs)
Natural Occlusion (Missing Features)
Adversary Malicious Perturbations (Exploit Missingness)
DRL Misbehaviour Detection
(b) Natural occlusions and adversarial perturbations. V2X Raw Data (BSMs)
Adversary Feature Suppression (Injected Missingness)
DRL Misbehaviour Detection
(c) Adversarial feature suppression.
Fig. 1: Considered scenarios under data missingness. SC2: Evasion via natural occlusions. Building on SC1, this scenario aims to demonstrate that partial observability not only poses a technical challenge for misbehavior detection systems but simultaneously enlarges the attack surface available to adversarial actors. In particular, SC2 considers an exogenous adversary that exploits the naturally occurring missingness in V2X data streams to launch evasion attacks against the DRL framework. To realize the attack, the adversary leverages existing missing regions induced by natural occlusions, rather than introducing new missing features. By timing subtle false data injections to coincide with periods of natural occlusion or by crafting adversarial perturbations calibrated to blend with the statistical profile of legitimate blind spots, as shown in Fig. 1b, the adversary seeks to minimize its detectability while maximizing the likelihood that the detector misclassifies misbehaving vehicles as benign. SC3: Adversarial feature suppression. This scenario considers an exogenous adversary that actively induces missingness as an attack strategy rather than exploiting existing missing regions. Thus, in contrast to SC1 and SC2, this scenario represents adversarially induced partial observability where missingness itself becomes the attack vector. The original V2X data streams in SC3 are assumed to be complete, and natural occlusions have been addressed through appropriate imputation techniques, prior to adversarial intervention. As shown in Fig. 1c, the adversary executes a feature suppression attack by selectively injecting artificially induced missing values across input features of the V2X data stream. The goal is to reduce the information available to the detector and degrade the classifier’s ability to distinguish between normal and misbehaving vehicles. III. M ISBEHAVIOR D ETECTION A. Model The V2X environment considered in this study follows a Markov decision process (MDP) framework to facilitate the detection of misbehaviors through sequential decisionmaking. Consistent with the MDP formulation, the action of misbehavior detection changes the environment based on
the decision of either legitimate or malicious behavior at time-step t. Subsequently, the decision at time-step t + 1 is influenced by the altered environment from the previous time-step t. In this work, we consider a DRL-based detector deployed at an edge node. i) Agent: The agent takes the V2X data as a time series and prior decisions as its state st , and outputs an action at according to policy π. The agent’s DQN [11] consists of an LSTM layer followed by a fully connected network with linear activation to estimate the Q-values over possible actions. Each interaction is stored as a transition tuple et =< st , at , rt , st+1 >, capturing the full behavioral trace of the misbehavior detector. By replaying this experience, the agent progressively refines its Q(s, a) estimates towards more accurate misbehavior detection. The objective is to maximize the expected sum of future discounted rewards, expressed as XT Rt = γ k−t rk . The Q-values are updated at each step k=t using learning rate α and discount factor γ as Q(st , at ) ← Q(st , at ) + α(rt + γ max Q(st+1 , at+1 ) − Q(st , at )). (1) at+1
ii) States: The state st comprises two components: a sequence of prior actions saction =< at−1 , at , ..., at+n−1 > and a window of current BSM observations stime =< Xt , ..., Xt+n >, where Xt ∈ Rd is a d-dimensional feature vector at time-step t. This design allows the agent to capture temporal dependencies across both past decisions and incoming BSM data, enabling more informed action selection. iii) Actions: The action space is A = {0,1}, where 1 indicates detected misbehavior and 0 represents the genuine behavior. The deterministic policy π : st ∈ S 7−→ at ∈ A maps each state to a single action. In state st , the agent selects the action that maximizes the optimal Q-function, π ∗ = arg max Q∗ (s, a).
(2)
a∈A
iv) Rewards: The reward rt is defined over the four outcomes of the confusion matrix typically used in ML classification problems. Correct decisions yield positive rewards, while errors are penalized, with false negatives (FNs) penalized more heavily than false positives (FPs), since missed misbehavior poses a higher safety risk. Accordingly, the reward of the agent is if at is a TP, a, b, if at is a TN, r(st , at ) = (3) −c, if at is an FP, −d, if at is an FN, where a, b, c, d > 0, with a > b and d > c. B. Adversarial Threat Model We assume the presence of black-box adversaries attempting to evade the DRL-based misbehavior detection system through malicious manipulation of V2X data, by exploiting or introducing missing features. The adversary is assumed to have sufficient knowledge of the data pipeline and the
first-order statistical properties of the input data distribution, such as mean values and standard deviations of the feature vector, which capture the central tendencies and dispersion characteristics of legitimate V2X data [12], [13]. The attacker does not require direct access to the internal parameters of the DRL model. In this context, we consider an exogenous attacker capable of executing two adversarial strategies, each representing a distinct threat vector targeting the integrity of the input data. Next, the two adversarial strategies constituting the adversarial model are described. Evasion via natural occlusions. In this case, the adversary leverages naturally occurring missing regions within the V2X data stream to inject adversarial perturbations that are crafted to blend with the underlying occluded data distribution (shown in Fig. 1b). We assume that the adversary strategically targets high-variance features to maximize perturbation stealthiness, as these regions exhibit high natural variability that masks manipulation [12], [13]. The adversarial input vector ϵadv = C + ρ is derived from the benign mean vector C = µ of the training distribution, where perturbations ρi = sign(Ci ) · k · σi are applied to the topn = ⌊d · r⌋ high-variance features (d: feature dimension, r: missingness rate, k: perturbation scale). The resulting adversarial values ϵadv then fill the missing positions. By doing this, the adversarial examples become indistinguishable from legitimate incomplete observations, penetrating the DRL classifier undetected. This forms an evasion attack where the adversary exploits the inherent issues of the V2X environment as a camouflage for adversarial manipulation. Adversarial feature suppression. In this case, the adversary actively introduces missingness as an attack strategy rather than exploiting naturally occurring occlusions. The original V2X data streams are complete with no missing feature vectors, as in both SC1 and SC2 before amputation. As shown in Fig. 1c, the adversary executes a feature suppression attack by selectively injecting missing values or blocking a subset of input features within the V2X data stream, intending to achieve controlled information suppression. Specifically, missingness is injected at a rate proportional to λ ∈ [0, 1], ranging from 1% to 20%, remaining within plausible natural missingness regions similar to the patterns evaluated in SC1 while maintaining feature suppression stealthiness. The resulting missing entries are then filled with adversarial inputs xadv = xi + ϵi , where i ϵi = (µi − δλ · σi ) − xi drives the feature toward a value between 1.8σi and 3.5σi below µi , a range empirically derived to maximize classifier degradation while maintaining statistical plausibility, with an additive Gaussian noise to reduce detectability [12]. This strategy artificially suppresses the informational content of targeted features, degrading the DRL classifier’s ability to distinguish between normal and misbehaving vehicles. IV. E XPERIMENTS A. VeReMi Dataset In this study, we utilize the open-source VeReMi dataset [10] to assess misbehavior detection performance.
The dataset provides labeled BSMs spanning multiple misbehavior types and traffic densities, making it a well-established benchmark for V2X security research. A representative subset of misbehaviors is selected to provide sufficient coverage of the available misbehavior types in VeReMi, as summarized below. Position falsification attack: A vehicle transmits falsified position coordinates within its communication area while concealing its true position. Among the possible variants, three are considered here: i) constant position, where a fixed position coordinate is repeatedly broadcast; ii) constant position offset, where the true position is transmitted with a fixed offset; and iii) random position, where newly generated random coordinates are broadcast at each transmission. Speed falsification attack: A misbehaving vehicle transmits falsified speed values in its BSMs, following a similar approach to position falsification. Among the possible variants, one is considered here: random speed, where newly generated random speed values are broadcast at each transmission. Delayed messages: A misbehaving vehicle transmits BSMs with correct field values but introduces a deliberate time delay, shifting transmissions away from real-time reporting. B. Missing Patterns in Natural Occlusions In our experimental setup, natural occlusions are introduced in a principled manner by applying various missingness patterns. Specifically, starting from complete BSM records in VeReMi, we employ standard statistical amputation techniques [14], namely, Missing Completely at Random (MCAR), Missing at Random (MAR), and Missing Not at Random (MNAR), to simulate the stochastic and structured nature of naturally occurring occlusions. In MCAR, missing values are injected uniformly at random across all features and samples, independently of both observed and unobserved data, producing missingness with no systematic pattern. To simulate MAR, the missingness probability is determined by the most correlated observed features via a logistic function across multiple data subsets. Under MNAR, features are randomly selected to induce missingness via a logistic function across multiple data subsets, simulating occlusions where a feature’s value determines the likelihood of its own occlusion in the context of the VeReMi misbehavior types. For MAR and MNAR, we adapt the subset-based missingness generation approach of [9] to the specific feature structure of the VeReMi dataset. Moreover, missingness is induced across the complete dataset at missing ratios of 1%, 5%, 10%, and 20% to mimic realistic incomplete V2X environments. V. P ERFORMANCE E VALUATION In this section, we evaluate the performance of the proposed DRL-based misbehavior detection framework across the three scenarios introduced in Section II, using the VeReMi dataset. These scenarios enable a progressive assessment of the framework, ranging from naturally incomplete
observations (SC1), to adversarial exploitation of environmental blind spots (SC2), and finally to deliberate manipulation of observability (SC3). Results are benchmarked against a baseline classifier under identical missing data conditions. A. Detection Performance Detection performance is evaluated in terms of Accuracy (Acc) and F-score (F1): RP TP + TN , F1 = 2 , (4) TP + TN + FP + FN R+P where Acc measures the ratio of correct predictions over all samples, and F1 provides the harmonic mean of Precision (P) and Recall (R), making it the primary indicator of detection performance where both FP and FN rates are critical. Acc =
B. Benchmark Scheme We select XGBoost [15] as our baseline classifier, a highly optimized Gradient Boosted Decision Tree (GBDT) framework that natively handles missing values. To ensure a fair comparison with our DRL-based approach, we employ a sliding-window mechanism similar to that used in the DRL setting. Specifically, each temporal window is flattened into a single high-dimensional feature vector, and labels are assigned based on the final timestep within the window to maintain step-wise classification consistency with DRL. This setup provides XGBoost with identical historical context and input information. We configure XGBoost’s decision boundary hyperparameter to prioritize recall (minimizing missed attacks) while tolerating acceptable FPs, mirroring the DRL reward structure. This window-based strategy evaluates XGBoost’s ability to capture spatio-temporal dependencies within a sequential decision-making framework [16]. Such a design is particularly relevant for edge-based misbehavior detection1 , where mission-critical applications must operate under limited data constraints. By flattening temporal sequences into high-dimensional spatial features, we assess whether the task can be effectively addressed through cross-feature correlations or whether the recurrent dynamics captured by DRL are necessary to effectively reason over stealthy and persistent occlusions. C. Partial Observability under Natural Occlusions To provide a comprehensive assessment under SC1, we evaluate the DRL-based framework’s detection capability across multiple misbehaviors and missing patterns, as summarized in Table I. The results consistently demonstrate the superior robustness of the DRL framework compared to XGBoost across all misbehavior types and missingness mechanisms. Under MCAR, DRL maintains strong performance even at 20% missingness, achieving an accuracy of 0.82 and F1 of 0.58 for Constant Position, an accuracy of 0.89 and F1 of 0.81 for Constant Position Offset, an accuracy of 0.92 and F1 of 0.86 for Random Position, an accuracy of 1 The computational overhead and real-time deployment feasibility of the proposed DRL agent for edge-based misbehavior detection were empirically assessed in our previous work [17], [18].
TABLE I: Detection performance for DRL and XGBoost under natural occlusions in SC1. Misbehavior
Constant Position
Constant Position Offset
Random Position
Random Speed
Delayed Messages
Normal Acc F1
DRL MCAR MNAR Acc F1 Acc F1
MAR Acc F1
Normal Acc F1
XGBoost MCAR MNAR Acc F1 Acc F1
MAR Acc F1
0% 1% 5% 10% 20%
0.99 − − − −
0.98 − − − −
− 0.98 0.92 0.86 0.82
− 0.96 0.85 0.70 0.58
− 0.98 0.97 0.97 0.95
− 0.97 0.96 0.95 0.91
− 0.98 0.98 0.98 0.96
− 0.97 0.97 0.96 0.93
0.92 − − − −
0.88 − − − −
− 0.91 0.83 0.75 0.62
− 0.85 0.76 0.69 0.59
− 0.92 0.90 0.89 0.88
− 0.87 0.84 0.84 0.81
− 0.92 0.91 0.89 0.85
− 0.87 0.85 0.83 0.79
0% 1% 5% 10% 20%
0.99 − − − −
0.98 − − − −
− 0.98 0.95 0.93 0.89
− 0.97 0.91 0.88 0.81
− 0.99 0.98 0.97 0.96
− 0.98 0.96 0.95 0.93
− 0.98 0.98 0.96 0.94
− 0.97 0.96 0.94 0.91
0.78 − − − −
0.72 − − − −
− 0.78 0.77 0.76 0.73
− 0.72 0.70 0.67 0.60
− 0.78 0.78 0.78 0.77
− 0.73 0.72 0.72 0.70
− 0.78 0.78 0.78 0.77
− 0.73 0.73 0.72 0.72
0% 1% 5% 10% 20%
0.98 − − − −
0.97 − − − −
− 0.95 0.94 0.92 0.92
− 0.92 0.88 0.85 0.86
− 0.96 0.95 0.95 0.93
− 0.92 0.92 0.91 0.88
− 0.96 0.96 0.96 0.95
− 0.92 0.92 0.93 0.92
0.83 − − − −
0.78 − − − −
− 0.82 0.77 0.72 0.62
− 0.76 0.72 0.67 0.60
− 0.83 0.82 0.80 0.80
− 0.77 0.76 0.76 0.74
− 0.84 0.83 0.82 0.81
− 0.78 0.77 0.76 0.75
0% 1% 5% 10% 20%
1.0 − − − −
1.0 − − − −
− 0.97 0.95 0.93 0.89
− 0.95 0.91 0.86 0.77
− 0.98 0.96 0.96 0.95
− 0.96 0.93 0.92 0.92
− 0.97 0.96 0.94 0.91
− 0.95 0.93 0.89 0.81
0.99 − − − −
0.98 − − − −
− 0.98 0.95 0.92 0.85
− 0.97 0.93 0.88 0.80
− 0.98 0.98 0.97 0.96
− 0.98 0.96 0.95 0.93
− 0.98 0.98 0.98 0.97
− 0.98 0.98 0.97 0.96
0% 1% 5% 10% 20%
0.91 − − − −
0.82 − − − −
− 0.90 0.90 0.89 0.89
− 0.81 0.80 0.79 0.78
− 0.90 0.90 0.89 0.90
− 0.81 0.80 0.80 0.80
− 0.90 0.90 0.90 0.90
− 0.81 0.81 0.81 0.81
0.80 − − − −
0.55 − − − −
− 0.79 0.78 0.77 0.74
− 0.53 0.51 0.47 0.41
− 0.80 0.80 0.80 0.80
− 0.54 0.54 0.54 0.51
− 0.79 0.79 0.78 0.76
− 0.54 0.53 0.53 0.51
Missingness
0.89 and F1 of 0.77 for Random Speed, and an accuracy of 0.89 and F1 of 0.78 for Delayed Messages. In contrast, XGBoost exhibits significant performance degradation at 20% MCAR, with F1 dropping to the range of 0.41-0.80 under the same conditions. In the presence of MNAR, where missingness depends on the feature values themselves, DRL achieves consistently high accuracy (0.90-0.96) and F1 scores (0.80-0.93) across all misbehavior types at 20% missingness. XGBoost shows markedly lower detection performance, particularly for Constant Position Offset (an F1 of 0.70) and Delayed Messages (an F1 of 0.51). The MAR mechanism reveals similar trends, with DRL maintaining robust detection capabilities with F1 scores in the range of 0.81-0.93 at a 20% missing ratio, while XGBoost drops to F1 scores between 0.51 and 0.79 for most misbehavior types, except for Random Speed, where it retains an F1 of 0.96. Notably, DRL exhibits varying sensitivity to missingness proportions across misbehavior types (Table I), with F1 degradation ranging from 5% to 41% when increasing from 0% to 20% missingness under the most challenging pattern in MCAR. The detection of Delayed Messages shows minimal degradation with F1 decreasing from 0.82 to 0.78 (5% drop), while the detection of Constant Position misbehavior shows the highest performance degradation with F1 decreasing from
0.98 to 0.58 (41% drop). For Random Position and Random Speed, DRL demonstrates moderate robustness with 11% and 23% degradation, respectively. XGBoost suffers comparable or greater degradation under the same conditions. Constant Position detection declines from F1 0.88 to 0.59 (33% drop), Random Position from 0.78 to 0.60 (23% drop), and Delayed Messages from 0.55 to 0.41 (25% drop). DRL demonstrates significantly stronger robustness under structured missingness patterns. Under MAR, F1 scores range from 0.81 to 0.93 at 20% missingness (1.2%-19% drop), while under MNAR, F1 ranges from 0.80 to 0.93 (2.4%-9.3% drop). Notably, MNAR exhibits the lowest worst-case performance degradation for the DRL framework across all misbehavior types shown in Table I. In comparison, XGBoost shows F1 scores ranging from 0.51 to 0.96 under MAR (0%-10% drop) and from 0.51 to 0.93 under MNAR (3%-8% drop). This adaptive robustness stems from DRL’s ability to capture spatio-temporal dependencies within a sequential decision-making framework, which effectively exploits the structure in MAR and MNAR patterns to mitigate the impact of incomplete observations. In contrast, MCAR’s completely random missingness disrupts temporal continuity severely, posing greater challenges even for sequential models.
TABLE II: Comparative performance of DRL and XGBoost under evasion via natural occlusions in SC2. DRL
No Attack
XGBoost
Missingness
Acc
F1
ASR (%)
Acc
F1
ASR (%)
0%
1.00
1.00
0.00
0.99
0.98
0.00
1% 5% 10% 20%
0.90 ± 0.02 0.92 ± 0.01 0.92 ± 0.01 0.93 ± 0.01
0.83 ± 0.03 0.86 ± 0.02 0.87 ± 0.02 0.88 ± 0.01
71.57 ± 1.49 65.19 ± 1.58 62.05 ± 1.57 58.48 ± 0.77
0.98 ± 0.01 0.95 ± 0.01 0.92 ± 0.01 0.86 ± 0.01
0.97 ± 0.02 0.92 ± 0.02 0.86 ± 0.02 0.74 ± 0.02
4.78 ± 0.33 10.62 ± 0.29 18.47 ± 0.23 34.59 ± 0.55
1% 5% 10% 20%
0.89 ± 0.02 0.89 ± 0.02 0.90 ± 0.02 0.90 ± 0.02
0.82 ± 0.04 0.80 ± 0.04 0.82 ± 0.04 0.83 ± 0.04
73.25 ± 1.86 75.44 ± 1.67 72.82 ± 3.93 67.33 ± 7.99
0.99 ± 0.01 0.97 ± 0.01 0.96 ± 0.02 0.96 ± 0.03
0.98 ± 0.02 0.95 ± 0.01 0.93 ± 0.03 0.92 ± 0.05
4.14 ± 5.07 16.43 ± 1.29 15.35 ± 8.81 8.03 ± 12.95
1% 5% 10% 20%
0.89 ± 0.02 0.89 ± 0.02 0.89 ± 0.02 0.88 ± 0.01
0.81 ± 0.04 0.81 ± 0.03 0.81 ± 0.03 0.78 ± 0.03
73.63 ± 1.69 74.29 ± 1.10 75.08 ± 0.99 77.73 ± 1.02
0.97 ± 0.01 0.97 ± 0.01 0.96 ± 0.02 0.95 ± 0.02
0.96 ± 0.02 0.95 ± 0.02 0.94 ± 0.03 0.92 ± 0.03
0.27 ± 0.35 0.37 ± 0.13 0.30 ± 0.05 0.38 ± 0.09
Evasion+MCAR
Evasion+MNAR
Evasion+MAR
D. Partial Observability under Evasion Attacks We evaluate the robustness of DRL under evasion attacks (SC2) considering the Random Speed misbehavior, where both DRL and XGBoost obtained comparable performance across different missingness proportions under natural occlusions in SC1 (Table I). Their adversarial robustness is summarized in Table II. As shown in the table, DRL exhibits significant vulnerability with attack success rates (ASR) consistently exceeding 70% even at minimal missingness rates of 1% under all three evasions via natural missingness patterns. At the highest missingness of 20%, DRL’s ASR ranges from 58.48% (MCAR) to 77.73% (MAR), demonstrating that adversaries successfully evade detection in approximately 75% of attempts by exploiting natural occlusions. Despite maintaining relatively high accuracy between 0.88 and 0.93, this vulnerability arises from DRL’s reliance on sequential temporal dependencies, which makes it challenging to distinguish between legitimate incomplete observations and adversarial manipulations that exploit natural occlusions as a camouflage. In contrast, XGBoost shows notably higher robustness across all evasion attacks via natural missingness (Table II). Under evasion via MAR, XGBoost achieves ASR below 0.4% across all missingness proportions. Against challenging MCAR attacks at 20% missingness, XGBoost limits ASR to 34.59%, an approximately twofold improvement over DRL. In particular, the robustness gap is substantial for evasion via MAR patterns, where XGBoost’s 0.38% of ASR versus DRL’s 77.73% yields a 205× advantage in attack resistance. This robustness arises from XGBoost’s tree ensemble processing temporal windows as flattened high-dimensional feature vectors, treating each window independently rather than enforcing sequential dependencies that make DRL vulnerable. XGBoost’s ensemble architecture dilutes the impact of adversarial perturbations across the larger feature space
TABLE III: Performance comparison under feature suppression attack (SC3). Intensity (λ)
Acc
F1
F1 Drop (%)
DRL No Attack Weak Medium Strong Stronger Strongest
0.0 0.1 0.2 0.3 0.4 0.5
1.00 0.99 ± 0.00 0.98 ± 0.00 0.98 ± 0.00 0.97 ± 0.00 0.96 ± 0.04
1.00 0.98 ± 0.01 0.97 ± 0.01 0.96 ± 0.01 0.95 ± 0.01 0.93 ± 0.07
-2.1 -2.8 -3.9 -5.1 -7.3
XGBoost No Attack Weak Medium Saturated Saturated Saturated
0.0 0.1 0.2 0.3 0.4 0.5
0.99 0.99 ± 0.01 0.88 ± 0.03 0.88 ± 0.03 0.88 ± 0.03 0.88 ± 0.03
0.98 0.98 ± 0.02 0.83 ± 0.04 0.83 ± 0.04 0.83 ± 0.04 0.83 ± 0.04
+0.1 -15.8 -15.7 -15.6 -15.5
and shows greater resistance to distribution-shift attacks compared to the DRL’s sequential approach. E. Partial Observability under Feature Suppression To assess the robustness of DRL under adversarial feature suppression (SC3), we again consider the Random Speed misbehavior, where both classifiers obtained comparable performance in SC1 (Table I), and compare their adversarial robustness. Feature suppression attacks on complete data (i.e., without natural occlusions) reveal contrasting vulnerability patterns between DRL and XGBoost, as shown in Table III and Fig. 2. DRL exhibits gradual performance degradation as suppression intensity (λ) increases from 0.1 to 0.5, with accuracy declining from 0.99 to 0.96 and F1 score from 0.98 to 0.93 (7.3% drop). As illustrated in Fig. 2, this gradual degradation follows a near-linear decline, reflecting how artificial missingness through targeted feature suppression
1.00
0.96 0.94
F1
Accuracy
0.98
0.92 0.90 DRL XGBoost
0.88 0.0
0.1
0.2
0.3
0.4
Suppression Intensity (λ)
0.5
1.00 0.98 0.96 0.94 0.92 0.90 0.88 0.86 0.84 0.82
DRL XGBoost
0.0
0.1
0.2
0.3
0.4
0.5
Suppression Intensity (λ)
Fig. 2: Performance degradation under feature suppression attack (SC3).
icant vulnerability, with high ASRs despite maintaining high accuracy. This shows that adversaries can leverage natural occlusions as a camouflage mechanism against sequential detectors. Third, under adversarial feature suppression, DRL shows gradual degradation in decision-making, whereas XGBoost exhibits abrupt performance collapse once critical features are suppressed. Future work will explore the use of adversarial training to harden the DRL agent’s sequential policy against the camouflage effects of natural occlusions. R EFERENCES
progressively disrupts the spatio-temporal patterns captured by DRL’s recurrent dynamics, which are essential for coherent state representation across sequential timesteps. The DRL agent’s ability to accumulate evidence across sequential timesteps weakens as suppression intensity increases, leading to increasing FNs and a gradual decline in F1 performance. In contrast, XGBoost exhibits a significant performance drop between intensities 0.1 and 0.2, clearly visible in both accuracy and F1 plots (Fig. 2). At intensity 0.1, XGBoost maintains near-baseline performance with accuracy of 0.99 and F1 of 0.98, but at intensity 0.2, performance drops abruptly to accuracy of 0.88 and F1 of 0.83 (15.8% degradation), then plateaus across higher intensities with minimal further degradation. This sudden drop occurs as XGBoost’s tree ensemble loses critical discriminative features once suppression intensity reaches approximately 0.2. Unlike DRL’s gradual degradation, XGBoost’s F1 collapses from 0.98 to 0.83, exhibiting an overly conservative classification where the model flags many benign samples as malicious. These findings reveal that DRL’s graceful degradation to F1 of 0.93 at intensity 0.5 offers more predictable behavior than XGBoost’s abrupt F1 collapse to 0.83, making it preferable for mission-critical V2X deployments. Despite DRL’s more favorable degradation pattern, enhanced defensive mechanisms are required to strengthen its robustness against targeted feature suppression attacks. VI. C ONCLUSIONS This paper investigated the problem of misbehavior detection with partially observable V2X measurements, a setting often overlooked in existing ML-based detection methods. This partial observability, owing to sensor imperfections, intermittent connectivity, and/or environmental occlusions, can be further leveraged by malicious actors via deliberate exploitation of naturally missing regions and adversarial feature suppression. Departing from conventional supervised classification methods, we proposed a DRL-based framework capable of learning robust detection policies under incomplete observations. Experimental results on the VeReMi dataset reveal three key insights. First, under natural occlusions, the proposed DRL approach consistently outperforms a powerful XGBoost baseline across MCAR, MAR, and MNAR patterns, demonstrating the benefit of sequential policy learning in structured partially observable environments. Second, under evasion attacks that exploit natural missingness, DRL exhibits signif-
[1] A. Boualouache and T. Engel, “A Survey on Machine LearningBased Misbehavior Detection Systems for 5G and Beyond Vehicular Networks,” IEEE Comm. Surv. Tutor., vol. 25, no. 2, pp. 1128–1172, 2023. [2] Y. Zhang, X. Kong, W. Zhou, J. Liu, Y. Fu, and G. Shen, “A Comprehensive Survey on Traffic Missing Data Imputation,” Trans. Intell. Transport. Sys., vol. 25, no. 12, p. 252–275, 2024. [3] M. Tsukada, S. Arii, H. Ochiai, and H. Esaki, “Misbehavior Detection Using Collective Perception under Privacy Considerations,” in 2022 IEEE 19th Annual Consumer Communications & Networking Conference (CCNC). IEEE, 2022, pp. 808–814. [4] J. Kamel et al., “Misbehavior Detection in C-ITS: A comparative approach of local detection mechanisms,” in 2019 IEEE Vehicular Networking Conference (VNC), 2019, pp. 1–8. [5] S. Gyawali, Y. Qian, and R. Q. Hu, “Machine Learning and Reputation Based Misbehavior Detection in Vehicular Communication Networks,” IEEE Trans. Veh. Technol., vol. 69, no. 8, pp. 8871–8885, 2020. [6] P. Sharma and H. Liu, “A Machine-Learning-Based Data-Centric Misbehavior Detection Model for Internet of Vehicles,” IEEE Internet of Things Journal, vol. 8, no. 6, pp. 4991–4999, 2021. [7] A. Sharma and A. Jaekel, “Machine Learning Based Misbehaviour Detection in VANET Using Consecutive BSM Approach,” IEEE Open Journal of Vehicular Technology, vol. 3, pp. 1–14, 2022. [8] E. Mármol Campos, A. Gonzalez-Vidal, J. L. Hernández-Ramos, and A. Skarmeta, “Federated learning for misbehaviour detection with variational autoencoders and gaussian mixture models,” International Journal of Information Security, vol. 24, no. 1, p. 95, Mar 2025. [9] R. Razavi-Far, D. Wan, M. Saif, and N. Mozafari, “To Tolerate or To Impute Missing Values in V2X Communications Data?” IEEE Internet of Things Journal, vol. 9, no. 13, pp. 11 442–11 452, 2022. [10] J. Kamel et al., “Veremi extension: A dataset for comparable evaluation of misbehavior detection in vanets,” in IEEE International Conference on Communications (ICC). IEEE, 2020, pp. 1–6. [11] V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al., “Human-level control through deep reinforcement learning,” Nature, vol. 518, no. 7540, pp. 529–533, 2015. [12] N. Papernot et al., “The limitations of deep learning in adversarial settings,” in IEEE European symposium on security and privacy (EuroS&P). IEEE, 2016, pp. 372–387. [13] A. Ilyas, S. Santurkar, D. Tsipras, L. Engstrom, B. Tran, and A. Madry, “Adversarial examples are not bugs, they are features,” Advances in neural information processing systems, vol. 32, 2019. [14] R. M. Schouten et al., “Generating missing values for simulation purposes: A multivariate amputation procedure,” Journal of Statistical Computation and Simulation, vol. 88, no. 15, pp. 2909–2930, 2018. [15] T. Chen and C. Guestrin, “Xgboost: A scalable tree boosting system,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, ser. KDD ’16. New York, NY, USA: Association for Computing Machinery, 2016, p. 785–794. [16] T. G. Dietterich, “Machine learning for sequential data: A review,” in Joint IAPR international workshops on statistical techniques in pattern recognition (SPR) and structural and syntactic pattern recognition (SSPR). Springer, 2002, pp. 15–30. [17] R. Sedar, C. Kalalas, P. Dini, F. Vázquez-Gallego, J. Alonso-Zarate, and L. Alonso, “Knowledge transfer for collaborative misbehavior detection in untrusted vehicular environments,” IEEE Transactions on Vehicular Technology, vol. 74, no. 1, pp. 425–440, 2024. [18] R. Asensio-Garriga et al., “Zsm-based e2e security slice management for ddos attack protection in mec-enabled v2x environments,” IEEE Open Journal of Vehicular Technology, vol. 5, pp. 485–495, 2024.