This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible.
Task-Oriented Semantic Communication for Hazard Warning and Remote Operation in Connected Vehicle Platoons S M Sabit Bananee* , Shahriar Hasan† , Muhammad Mahbub Alam* , Nafiul Rashid‡
arXiv:2609.06242v1 [cs.NI] 5 Sep 2026
*
Department of Computer Science and Engineering, Islamic University of Technology, Gazipur 1704, Bangladesh † Department of Computer Science and Engineering, Mälardalen University, 721 23 Västerås, Sweden ‡ Samsung Research America, Mountain View, CA 94043, USA E-mail: [email protected], [email protected], [email protected], [email protected]
Abstract—Connected and automated vehicle platoons require reliable dissemination of task-relevant hazard information to enable appropriate downstream responses, with unresolved situations additionally requiring remote human intervention. However, the limited channel bandwidth of V2X communication makes high-volume sensor data transmission challenging, motivating the exchange of only task-relevant information rather than raw sensor data. This paper proposes an integrated semantic and task-oriented communication framework for cooperative hazard response and remote operation in automated vehicle platoons. The framework combines onboard instance segmentation and monocular depth estimation for task-oriented hazard decisionmaking with Stable Diffusion Variational Autoencoder (SDVAE)-based semantic scene compression. Task-relevant hazard information and the compact semantic representation are jointly disseminated through a broadcast-based Emergency Semantic Message (ESM) over the existing V2X protocol stack. Evaluation demonstrates that SD-VAE consistently outperforms JPEG and JPEG 2000 on task-relevant reconstruction metrics, exhibiting graceful degradation rather than a sharp cliff effect under lowSINR conditions, demonstrating its feasibility for remote scene reconstruction when human intervention is required. Moreover, an extensive campaign of 7,290 simulation runs shows a low overall inter-vehicle collision rate of 0.99%, with collisions occurring mainly under extreme high-speed and short-gap conditions. Index Terms—Platooning, semantic communication, taskoriented communication, V2X, hazard warning, remote operation
I. I NTRODUCTION The proliferation of Connected and Automated Vehicles (CAVs) is enabling a wide range of applications aimed at making road transport safer, faster, cleaner, and more efficient. However, realizing these benefits requires connected vehicles to support a heterogeneous set of services, each with distinct requirements in terms of bandwidth, latency, and reliability. Such CAV services range from time-critical safety functions, such as cooperative collision avoidance and hazard warning, to bandwidth-intensive services, such as sensor (e.g., camera or LiDAR) data sharing for cooperative perception, highdefinition map distribution, and remote operation, among others. Meeting the requirements of these diverse services over a single shared wireless channel remains a significant challenge for connected vehicle communication. Candidate
vehicular communication technologies, such as IEEE 802.11pbased Dedicated Short-Range Communications (DSRC) and 5G NR-V2X sidelink operating under the UE-autonomous Mode 2 resource allocation scheme, further compound this challenge, as both provide limited channel bandwidth and impose constraints on the maximum allowable payload size that can be transmitted in a single transmission. Semantic communication represents a paradigm shift from conventional Shannon-based digital communication, moving beyond bit-level accuracy toward the transmission of meaning and task-relevant information. This distinction was first introduced by Weaver [1], who described the general communication problem at three levels: Level A (technical), Level B (semantic), and Level C (effectiveness). Conventional vehicular communication technologies, such as DSRC and 5G NR-V2X, are primarily concerned with the Level A problem, i.e., reliably transmitting bits without regard to their semantic content or task relevance. Semantic communication addresses Level B by prioritizing the meaning conveyed by the transmitted information, whereas task- and goal-oriented communication targets Level C by tailoring the transmitted information to the receiver’s intended action. Semantic and task-oriented communication have recently received particular attention in vehicular communication due to stringent bandwidth, latency, and reliability constraints. For task-oriented communication, Eldeeb et al. [2] address trafficsign sharing over vehicle-satellite-vehicle links, Liu et al. [3] study adaptive semantic compression under varying resource constraints, and Yang et al. [4] investigate energy-efficient cooperative relaying. For goal-oriented communication, Yildirim and Arikan [5] investigate action-based communication for autonomous vehicles, whereas Jin et al. [6] study semantic reconstruction under adverse channel conditions. For cooperative perception, Gan et al. [7] focus on bandwidth-efficient LiDARbased 3D object detection, while Sheng et al. [8] investigate importance-guided feature selection under time-varying fading channels. Furthermore, semantic-aware resource management relevant to vehicular communication has also been investigated through spectrum sharing [9] and joint user association with background-knowledge-aware bandwidth allocation [10].
We develop an integrated semantic and task-oriented communication framework for automated vehicle platoons, combining onboard hazard detection, task-oriented decision making, semantic scene compression, and broadcast-based ESM dissemination within the existing V2X protocol stack. • We employ a Stable Diffusion Variational Autoencoder (SD-VAE)-based semantic encoder together with hazardaware spatial masking and resolution-adaptive encoding to reduce the amount of scene information transmitted over the V2X channel. The proposed approach allows the semantic encoding rate to be adapted while preserving perceptual and task-relevant information under varying channel conditions. • We integrate the proposed framework into the PLEXE [14] simulation environment and evaluate its semantic reconstruction, communication, and task-oriented safety performance under realistic vehicular communication scenarios. •
The remainder of this paper is organized as follows. Section II presents the overall system architecture, while Section III describes the end-to-end encoder-decoder pipeline, including the onboard hazard-decision process and SD-VAE-based semantic scene compression. Section IV describes the simulation scenario, parameter space, and performance metrics, and Section V presents the performance evaluation in terms of semantic and task-relevant reconstruction quality and taskoriented safety performance. Finally, Section VI concludes the paper.
Inter- and intra- platoon semantic and task- oriented V2X communication Remote Station
Backhaul link
Semantic and task- oriented feature extraction
Task- oriented response Emergency Semantic Message
Semantic meaning extraction compact latent representation
task- oriented description + semantic meaning
Classif ication and masking hazard class + instance mask
Zone 1: Immediate maneuver ESM broadcast
Hazard distance estimation
Zone 2: Precautionary adjustment Zone 3: Situational awareness Executed through a CACC or emergency braking strategy
monocular hazard distance
Roadside Unit
Remote Operation Station
Forwards the ESM to the respective remote station
Semantic reconstruction: Remote driving, assistance, monitoring
Fig. 1: System architecture for semantic and task-oriented hazard warning and remote operation.
II. S YSTEM A RCHITECTURE Onboard sensors
Despite these advances, several important gaps remain. Most existing frameworks are designed and evaluated under point-to-point communication links, disregarding the broadcast nature that is fundamental to vehicular communication [11]. Moreover, there is still limited understanding of how semantic and task-oriented communication can be realized within existing protocol stacks, particularly as semantic communication is unlikely to fully replace conventional communication in the near term [12]. In addition, the integration of semantic and task-oriented communication into multi-agent system control loops, e.g., automated vehicle platoons, remains largely unexplored. Finally, the performance of semantic and task-oriented communication under high mobility, dense traffic, and realistic vehicular propagation conditions remains insufficiently investigated [13]. To address these gaps, this paper proposes a semantic and task-oriented communication framework for cooperative hazard response in automated vehicle platoons. When a platoon encounters a hazard, its Lead vehicle (LV) extracts the taskrelevant hazard information together with a compact semantic representation of the surrounding scene and combines them into an Emergency Semantic Message (ESM) for broadcast over the existing V2X communication stack. Upon receiving the ESM, downstream platoons independently determine the required response based on the communicated hazard information and their own operating conditions, while the semantic representation can be reconstructed when additional scene information or remote assistance is required. The main contributions of this work are summarized as follows:
A schematic representation of the overall system architecture of the proposed semantic and task-oriented hazard warning and remote operation framework is depicted in Fig. 1. As an example, Fig. 1 shows three platoons, each comprising three CAVs, driving on a highway. During cruising, the Cooperative Adaptive Cruise Control (CACC) algorithm is selected locally; as a result, the intra-platoon communication topology is dictated by the CACC algorithm adopted by each platoon. Although the focus of this work is not to study the string stability of automated vehicle platoons, the maintained gap at the time of braking, which is dictated by the control algorithm, is important for avoiding inter-vehicle collisions. We assume that a platoon encounters a hazard of common interest, e.g., debris, unannounced road work, accidents, potholes, requiring emergency braking and dissemination of hazard messages. In such a scenario, other platoons within the immediate vicinity may also need to perform emergency braking, whereas platoons further away from the hazard may choose to perform other maneuvers, e.g., soft deceleration, lane changing, and rerouting. Moreover, if the hazard type is unknown or a CAV in a platoon cannot determine the next course of action, the vehicle can seek human intervention from a remote station according to the SAE J3016 202104 standard [15]. The semantic and task-oriented framework in Fig. 1 supports platoon cruising, hazard detection, emergency braking, as well as transmission of the camera scene from the hazard site to a remote station for remote driving, assistance, or monitoring; the framework also remains generalizable to both DSRC and 5G NR-V2X communication technologies. The constituent elements of the system architecture in Fig. 1 are outlined below. Semantic and Task-Oriented Feature Extraction: This is typically performed by the LV of the first platoon that encounters the hazard. Using its onboard camera, the LV performs semantic feature extraction, hazard classification and masking, and hazard distance estimation, as depicted in the left-hand inset of Fig. 1. If the hazard cannot be resolved due to an unknown class or low confidence, the semantic scene is communicated to the remote station.
Emergency Semantic Message (ESM): The compact latent representation of the camera frame, hazard class and instance mask, and monocular hazard distance are combined to form the ESM within the payload constraints of the DSRC and 5G NR-V2X protocol stacks. The ESM contains both the taskoriented hazard information and the semantic representation of the scene and is broadcast by the affected platoon’s LV. Task-Oriented and Zone-Based Hazard Response: Upon receiving an ESM, each platoon LV locally determines the appropriate maneuver based on its distance from the hazard, hazard type, and current state. For instance, a nearby platoon may perform emergency braking, while a platoon further away may decelerate using its CACC or seek remote assistance. Semantic Reconstruction and Remote Operation: A receiving platoon or remote station semantically decodes and reconstructs the transmitted scene. The reconstructed scene can support local situational awareness or, at the remote station, enable remote assistance or remote driving. III. S EMANTIC AND TASK -O RIENTED E NCODER -D ECODER D ESIGN Figure 2 shows the end-to-end architecture of the proposed semantic and task-oriented communication framework, from onboard perception and hazard decision at the affected platoon to ESM transmission, zone-based adaptive response, and semantic scene reconstruction at the receiver. A. Task-Oriented Information Extraction The onboard perception pipeline of the platoon LV transforms the captured camera frame into a compact and taskrelevant representation of the surrounding scene. The LV employs YOLOv8n-seg [16] and Depth Anything V2 Small [17] in parallel for instance segmentation and monocular depth estimation, respectively. These model outputs are subsequently fused by a lightweight, rule-based decision engine to construct an actionable state for the detected hazard. Instance Segmentation: The YOLOv8n-seg model is trained on the BDD100K-10K [18] dataset for eight road-relevant object classes, i.e., pedestrian, rider, car, truck, bus, train, motorcycle, and bicycle. For each detection with a confidence score of p ≥ 0.35, the model produces a bounding box, class label, confidence score, and pixel-level instance mask. From these outputs, we compute the mask-area fraction to estimate the spatial extent of the detected object and the normalized horizontal-centre fraction to determine its relevance to the ego or adjacent lane. The detected hazard region is also projected onto the semantic latent grid according to lat h lat w , cy = y · , (1) cx = x · orig w orig h where (x, y) denotes the pixel coordinate of the hazard bounding box and (cx , cy ) the corresponding latent-cell index. This projection enables the Region of Interest (ROI)-based sparse semantic encoding described in the following subsection. Monocular Depth Estimation: In parallel with instance segmentation, the Depth Anything V2 model is employed
to estimate the relative depth of the objects detected within the same camera frame. Unlike stereo- or LiDAR-based approaches, the model produces a dense monocular depth map directly from the camera image. The resulting depth map is inverted and normalized such that pixels corresponding to objects closer to the LV approach a normalized depth score of 1.0. For each detected object, the depth values within its ROI are averaged to obtain a representative depth score. Since the subsequent decision thresholds operate on the normalized depth score rather than directly on metric depth, no explicit camera intrinsic calibration is required for this stage. Fusion and Task-Oriented Hazard Decision: The outputs of the instance segmentation and depth estimation models are subsequently fused to form a three-element task-oriented feature vector for every valid detection, comprising the normalized depth score, mask-area fraction, and horizontal-center fraction. A rule-based decision table with class-dependent thresholds maps this vector to one of five hazard states: Do Nothing, Decelerate, Lane Change, Hard Brake, or Cannot Resolve. The Cannot Resolve state is used when the LV cannot reliably determine the appropriate response, for instance, because the detected object belongs to an unknown class or the detection confidence is insufficient. When multiple objects are detected within the same frame, each detection is independently assigned a hazard state. The overall decision of the LV is subsequently selected according to the highestpriority state among all detections, s∗ = max priority(si ), i
(2)
where si represents the state assigned to detection i and priority(·) ranks the five states from 0 for Do Nothing to 4 for Cannot Resolve. The resulting state s∗ is encoded into the ESM controlAction field together with the estimated hazard distance. As illustrated in Fig. 2, these task-relevant data are subsequently packed into the ESM together with the routing information and semantic payload before transmission over the V2X channel. The hazard distance is subsequently used by receiving platoons for zone classification, while the transmitted action represents the response selected by the affected platoon. B. Semantic Scene Compression using SD-VAE In parallel with the hazard decision process, the same camera frame is independently encoded using the pretrained VAE from the Stable Diffusion Architecture [19]. This allows a receiving platoon or the remote operation station to reconstruct the surrounding scene when additional semantic information is required. The encoder maps an input frame to a continuous W H latent representation z = E(x) ∈ R4× 8 × 8 , corresponding to an eightfold spatial downscaling. The latent representation is subsequently standardized to approximately follow a N (0, 1) distribution, which improves float16 representation and facilitates the compression applied at the subsequent stage. The latent representation is stored using float16 rather than int8 precision. Although int8 further reduces the
32 16
Fig. 2: End-to-end encoder-decoder architecture of the proposed semantic and task-oriented framework. The platoon LV extracts hazard features and sparse latent semantics, and the receiver reconstructs the scene and performs zone-based response.
payload size, its coarser quantization introduces smoothing artifacts that degrade the structural details required for downstream object detection and depth estimation. Therefore, float16 is adopted as a suitable trade-off between reconstruction quality and transmission overhead. In addition, the encoder supports any input resolution divisible by 8. Consequently, the dimensions of the latent grid, and hence the resulting Bits Per Pixel (BPP), can be adjusted according to the input resolution, ranging from approximately 0.01 BPP at 64 px to 0.28 BPP at 320 px. This allows the transmitter to adapt the semantic encoding rate according to the available channel capacity and ESM payload budget without modifying the underlying architecture. A more substantial reduction in payload size is achieved through spatial masking. At an input resolution of 320 px, the full latent representation contains up to 40 × 40 = 1600 spatial cells. For time-critical transmission, only the latent cells corresponding to the detected hazard region are retained, using the same pixel-to-latent projection defined in Eq. (1). The retained latent cells are then losslessly compressed using zlib, which exploits the correlation among neighboring latent values to provide further reduction without additional reconstruction loss. Together, spatial masking and lossless compression reduce the payload associated with a typical hazard region by approximately 5–8 times compared with full-frame latent encoding, while preserving the scene information relevant to the detected hazard. The exact payload sizes for the considered encoding configurations, together with their corresponding reconstruction performance, are reported in detail in Section V. Furthermore, a 12-byte header is appended to each payload to specify the latent dimensions and original image resolution, allowing the receiver to decode the representation without requiring any additional signaling. At the receiver, the ESM is first parsed to separate the taskrelevant information from the semantic payload, as shown in Fig. 2. The task-relevant information is directly used for zone classification and the corresponding adaptive response, whereas the semantic payload is decompressed and decoded when scene reconstruction is required. The reconstructed
frame is then upscaled to the original image dimensions specified in the header and subjected to lightweight sharpening to restore edge details attenuated by the spatial bottleneck. For sparse hazard-ROI encoding, latent cells that are not included in the received payload are initialized to zero before decoding. The pretrained decoder can then reconstruct the surrounding background context while preserving the transmitted hazard region, thereby providing a coherent representation of the scene without requiring the full latent representation to be transmitted. IV. S IMULATION S CENARIO AND P ERFORMANCE M ETRICS We integrate the proposed semantic and task-oriented framework in Fig. 2 into the PLEXE simulation framework [14]. PLEXE is built on top of the VANET simulator Veins [20], which bidirectionally couples OMNeT++ and SUMO and implements the IEEE 802.11p protocol stack. We consider a multi-platoon highway scenario comprising four platoons cruising across two lanes, while a third lane is occupied by non-platooning CAVs that generate periodic beacons at 10 Hz. These additional CAVs are introduced to create background communication traffic and, consequently, a more realistic level of channel contention and interference. When the LV of the first platoon encounters an imaginary hazard at a predefined time, it starts its emergency braking maneuver, constructs the ESM as per Fig. 2, and broadcasts the ESM. Each platoon LV in the downstream direction relays this ESM and starts the emergency braking maneuver following their respective control algorithms. To account for multipath propagation and path-loss effects, we employ the Nakagamim fading model with α = 3 together with the free-space path-loss model. Moreover, to represent the remote operation scenario, an RSU is introduced into the simulation, which can receive the transmitted ESM and reconstruct the corresponding scene from the received semantic representation. Similarly, the following platoons can also reconstruct the scene from the received ESM. An extensive simulation campaign comprising 7,290 runs is performed by varying the parameters listed in Table I. For
Platoon size non-platooning vehicles Leader speed CACC algorithm Inter-platoon gap Deceleration rate Repetitions Hazard action
4, 6, 8 vehicles 50, 100, 200 80, 100, 120 km/h Rajamani et al. [21], Ploeg et al. [22] 50, 100, 150 m −5, −6, −8 m/s2 3 per combination Do Nothing, Lane Change, Decelerate, Hard Brake, Cannot Resolve
Total
Count 3 3 3 2 3 3 3 5 7,290
320 288 256 192 160 144 128 112 96 80 64 0.00
V. P ERFORMANCE E VALUATION This section evaluates the reconstruction quality and taskoriented safety performance of the proposed framework under the simulation scenario described above. A. Semantic and Task-Relevant Reconstruction Quality The resolution-adaptive encoding of SD-VAE allows the transmission rate to be controlled through the selected input resolution. As shown in Fig. 3, the required Bits Per Pixel (BPP) gradually increases from approximately 0.01 at 64 px to 0.28 at 320 px. This provides flexibility for selecting the semantic encoding rate according to the available communication resources and the required reconstruction quality.
0.10 0.15 0.20 BPP (bits per pixel)
0.25
0.30
SSIM ↑
Fig. 3: SD-VAE BPP as a function of input resolution, illustrating continuous rate control via resolution-adaptive encoding. 0.8
0.8
0.6
0.6
0.4 SD-VAE JPEG 2000 JPEG
0.2 0.0 0.0
0.6
0.4
0.0 0.0
0.1 0.2 BPP (bits per pixel)
0.1 0.2 BPP (bits per pixel)
0.8 SD-VAE JPEG 2000 JPEG
0.4 0.2 0.0 0.0
SD-VAE JPEG 2000 JPEG
0.2
0.8
mIoU ↑
platoon cruising, we employ the CACC controllers proposed by Rajamani et al. [21] and Ploeg et al. [22]. For the Rajamani CACC, an inter-vehicle gap of 5 m is considered, whereas the Ploeg CACC operates with a time headway of 0.5 s. Other PHY-layer parameters, e.g., transmission power, bit rate, and noise floor, follow the IEEE 802.11p configuration. For the hazard action parameter in Table I, emergency braking for both the decelerate and hard-brake actions is performed at deceleration rates of −5, −6, or −8 m/s2 ; milder rates were found not to affect collision occurrence and are therefore omitted for conciseness. Nevertheless, deceleration at lower rates can have an impact on string stability, the assessment of which remains outside the scope of this paper. Further, for the cannot resolve action, we consider a scenario where the platoon brakes at −5, −6, or −8 m/s2 to reduce its speed to 20 km/h from its initial speed. To evaluate the reconstructed scene, we employ four complementary metrics. Structural Similarity Index Measure (SSIM) and Learned Perceptual Image Patch Similarity (LPIPS) assess pixel-level structure and perceptual similarity, respectively. To measure task-relevant reconstruction quality, we additionally use mean Intersection over Union (mIoU) and mean Average Precision at an IoU threshold of 0.5 ([email protected]), which evaluate the preservation of semantic information for downstream segmentation and object detection. To assess platoon safety under the proposed task-oriented hazard response, we record inter-vehicle collision occurrences across the 7,290 simulation runs and perform a decisiontree-based root-cause analysis, detailed in Subsection V-B, to identify the parameter combinations associated with collisions.
0.05
LPIPS ↓
Values
0.1 0.2 BPP (bits per pixel)
Parameter
Input Resolution (pixels)
TABLE I: Simulation parameter space (7,290 runs).
0.6 0.4 0.2 0.0 0.0
SD-VAE JPEG 2000 JPEG
0.1 0.2 BPP (bits per pixel)
Fig. 4: Rate-distortion performance of SD-VAE compared against JPEG and JPEG 2000 across pixel-level (SSIM, top left), perceptual (LPIPS, top right), and task-relevant (mIoU, bottom left; [email protected], bottom right) reconstruction quality as a function of BPP.
Fig. 4 compares SD-VAE against the JPEG and JPEG 2000 baselines for the same BPP range considered in Fig. 3 using four metrics. In Fig. 4, notice that JPEG does not have a comparable operating point below 0.167 BPP, as its blockbased Discrete Cosine Transform (DCT) coding lacks the adaptability needed for scalable low-rate encoding. Among the four considered metrics, the advantage of SD-VAE is most evident for the two task-relevant metrics. SD-VAE reaches an [email protected] of approximately 0.66 at only 0.06 BPP, which already exceeds the best value achieved by JPEG 2000 over the entire considered range, i.e., approximately 0.60 at 0.20 BPP. A similar trend can be observed for mIoU, where SD-VAE reaches approximately 0.30 at 0.09 BPP, while JPEG 2000 does not reach this value even at its highest tested rate of approximately 0.28 BPP. For LPIPS, SD-VAE reaches its best perceptual score of approximately 0.15 at 0.15 BPP, which is significantly lower than JPEG 2000 within the considered
1.0 SD-VAE (160 px) JPEG 2000 JPEG (Q=50)
LPIPS ↓
SSIM ↑
0.8 0.6 0.4 0.2 0.0
−5
0
5 10 SINR (dB)
15
0.8
Class
0.4
SD-VAE (160 px) JPEG 2000 JPEG (Q=50)
−5
0
5 10 SINR (dB)
15
20
0.8
0.4 0.2
−5
0
5 10 SINR (dB)
15
20
Box Precision
Box Recall
Box F1 Score
Box mAP50
Mask mAP50
Pedestrian Car Truck Bus Motorcycle Bicycle Rider Train
0.756 0.784 0.603 0.702 0.638 0.397 0.359 1.000
0.419 0.623 0.343 0.311 0.255 0.062 0.133 0.000
0.539 0.694 0.437 0.431 0.364 0.107 0.194 0.000
0.500 0.673 0.409 0.363 0.325 0.094 0.145 0.006
0.464 0.611 0.384 0.360 0.324 0.079 0.102 0.006
All
0.655
0.268
0.380
0.314
0.291
0.6
0.0
20
SD-VAE (160 px) JPEG 2000 JPEG (Q=50)
0.6 mIoU ↑
TABLE II: YOLOv8n-seg validation results on BDD100K-10K.
0.2
0.8
0.0
1.0
SD-VAE (160 px) JPEG 2000 JPEG (Q=50)
0.6 0.4 0.2 0.0
Total simulation runs N = 7,290
−5
0
5 10 SINR (dB)
15
20
Fig. 5: Decoded scene quality of SD-VAE, evaluated across all tested input resolutions (thin lines, 64–320 px, with 160 px highlighted in bold), compared to JPEG and JPEG 2000 for varying SINR (dB).
B. Task-Oriented Safety Performance Table II reports the validation performance of YOLOv8nseg on BDD100K-10K after fine-tuning for 100 epochs on an RTX 3090 GPU. The detection performance varies considerably across classes; for instance, car, the most frequently represented class, achieves the highest Box mAP50 of 0.673, whereas less represented classes such as bicycle and rider
Inter- platoon gap = 10 0 / 150 m N = 4860
No collisions
Speed = 80 or 10 0 km/ h N = 1620
No collisions
Controller = Rajamoni CACC N = 405
No collisions
Repetition seed = 1 or 2 N = 270
Remaining
Inter- platoon gap = 50 m N = 2,430 Remaining
Speed = 120 km/ h N = 810 Remaining
Controller = Ploeg CACC N = 405
range. In terms of SSIM, SD-VAE performs better at low-tomoderate rates, whereas at the highest tested BPP, JPEG and JPEG 2000 slightly outperform SD-VAE. This is mainly due to the pixel-level nature of SSIM, whereas the higher mIoU and [email protected] values with SD-VAE indicate better preservation of the information required for downstream segmentation and object detection, which is the main goal of semantic and taskoriented communication. Fig. 5 presents the same four metrics as a function of SINR. JPEG and JPEG 2000 are evaluated at fixed operating points, while the thin purple curves represent the different SD-VAE input resolutions from 64 to 320 px, and the bold curve highlights the 160 px operating point used for direct comparison. An important observation from Fig. 5 is that the JPEG and JPEG 2000 baselines exhibit the cliff effect, i.e., their performance remains close to the minimum level over most of the low-SINR region and then improves sharply between approximately 10 and 12 dB. In contrast, SD-VAE begins to recover from around 4–6 dB and improves more gradually before reaching its plateau at approximately 10–12 dB. This behavior can also be observed consistently across the different SD-VAE input resolutions. The ability of SD-VAE to preserve useful perceptual and task-relevant information under low-to-moderate SINR conditions is particularly important for vehicular communication, where the channel condition may deteriorate rapidly due to high-speed mobility, dynamic topology changes, and variations in traffic conditions.
No collisions
Remaining
Repetition seed = 0 N = 135
No collisions
No braking N = 54
Collisions
Braking N = 54
Remaining
Cannot resolve reduce speed to 20 km/h N = 27 Collisions
No collisions
N=9
N = 18
Fig. 6: Decision path for collision occurrence across all 7,290 simulation runs, identified by a decision tree classifier. Every leaf reaches exactly 0% or 100% collision rate.
achieve much lower values of 0.094 and 0.145, respectively. Recall is also consistently lower than precision across almost all classes, with overall values of 0.268 and 0.655, indicating that the model produces fewer false detections but misses a larger fraction of actual instances, particularly for the less represented classes. The train class reports a precision of 1.0 with a recall of 0, which is attributable to its very limited representation in the validation set. A total of 72 collisions occurred across the 7,290 simulation runs, corresponding to an overall collision rate of 0.99%. In order to identify the parameter combinations in Table I that dictate the collision occurrence, we trained a decision tree classifier (scikit-learn, Gini criterion, unrestricted depth) on the results of the 7,290 simulation runs, using the eight parameters in Table I as input features and collision occurrence as the binary target. The classifier achieves 100% accuracy, and Fig. 6 shows the resulting decision path. The inter-vehicle collisions mainly occur when the interplatoon gap is 50 m, the leader speed is 120 km/h, and the Ploeg CACC is used. Fig. 6 further shows that when
the platoons reduce the speed to 20 km/h due to a Cannot resolve action at a deceleration rate −6 or −8 m/s2 , there are 18 collision cases. In addition, the largest number of collisions (i.e., 54) is due to the emergency braking maneuver. Importantly, the observed collisions are not attributable to communication outages. All 72 collision-positive runs occurred at an Emergency Semantic Message-Delivery Ratio (ESM-DR) of 1.0, confirming that the collisions arise from the vehicle-following response to a successfully delivered taskrelevant command rather than from packet loss. Therefore, the collisions can likely be attributed to an inter-platoon collision occurring at the 50 m gap and 120 km/h speed with the Ploeg CACC. The leading vehicle of the following platoon begins braking upon receiving the ESM broadcast from the leading vehicle of the platoon ahead. At high speed, however, the braking distance required by the leading platoon increases substantially, leaving insufficient room within the 50 m interplatoon gap for the following platoon to respond to the ESM in time. VI. C ONCLUSION This paper proposed an integrated semantic and taskoriented communication framework for cooperative hazard response and remote operation in automated vehicle platoons. The framework combines onboard hazard detection and task-oriented decision-making with compact semantic scene transmission through broadcast-based ESM dissemination over the existing V2X protocol stack. Simulation results show that the SD-VAE encoder supports flexible selection of the semantic encoding rate to balance available communication resources and the desired reconstruction quality. It consistently outperforms the conventional JPEG and JPEG 2000 baselines on task-relevant reconstruction metrics, reaching an [email protected] of 0.66 at only 0.06 bits per pixel, while exhibiting graceful degradation under low-SINR conditions instead of the sharp cliff effect observed for the conventional baselines. These results demonstrate its potential for supporting remote operation of connected vehicles. Furthermore, across 7,290 simulation runs, the proposed hazard-aware, task-oriented ESM broadcast results in a low overall inter-vehicle collision rate of 0.99%, with collisions occurring only under extreme conditions, i.e., at 120 km/h with a short inter-platoon gap. Future work will investigate lightweight and computationally efficient models for real-time deployment on vehicular platforms. R EFERENCES [1] W. Weaver, “Recent contributions to the mathematical theory of communication,” ETC: A Review of General Semantics, vol. 10, no. 4, pp. 261–281, 1953. [2] E. Eldeeb, M. Shehab, and H. Alves, “A multi-task oriented semantic communication framework for autonomous vehicles,” IEEE Wireless Commun. Lett., vol. 13, no. 12, pp. 3469–3473, 2024. [3] C. Liu, C. Guo, Y. Yang, and N. Jiang, “Adaptable semantic compression and resource allocation for task-oriented communications,” IEEE Trans. Cogn. Commun. Netw., vol. 10, no. 3, pp. 769–782, 2024. [4] W. Yang, X. Chi, L. Zhao, Z. Xiong, and W. Jiang, “Task-driven semantic-aware green cooperative transmission strategy for vehicular networks,” IEEE Trans. Commun., vol. 71, no. 10, pp. 5783–5798, 2023.
[5] Y. Yildirim and O. Arikan, “A goal-oriented effective communication scheme via action innovations,” in Proc. IEEE PIMRC, 2025, pp. 1–6. [6] Z. Jin, T. Song, X. Song, and J. Hu, “Goal-oriented communication with semantic reconstruction in vehicular networks,” in Proc. IEEE VTC2025Spring, 2025, pp. 1–6. [7] J. Gan, Y. Sheng, H. Zhang, L. Liang, H. Ye, C. Guo, and S. Jin, “Scomcp: Task-oriented semantic communication for collaborative perception,” IEEE Trans. Veh. Technol., vol. 75, no. 7, pp. 13 080–13 094, 2026. [8] Y. Sheng, H. Ye, L. Liang, S. Jin, and G. Y. Li, “Semantic communication for cooperative perception based on importance map,” J. Franklin Inst., vol. 361, no. 6, p. 106739, 2024. [9] Z. Shao, Q. Wu, P. Fan, N. Cheng, W. Chen, J. Wang, and K. Ben Letaief, “Semantic-aware spectrum sharing in internet of vehicles based on deep reinforcement learning,” IEEE Internet Things J., vol. 11, no. 23, pp. 38 521–38 536, 2024. [10] L. Xia, Y. Sun, D. Niyato, X. Li, and M. A. Imran, “Joint user association and bandwidth allocation in semantic communication networks,” IEEE Trans. Veh. Technol., vol. 73, no. 2, pp. 2699–2711, 2024. [11] S. Ma, Z. Zhang, Y. Wu, H. Li, G. Shi, D. Gao, Y. Shi, S. Li, and N. Al-Dhahir, “Features disentangled semantic broadcast communication networks,” IEEE Trans. Wireless Commun., vol. 23, no. 6, pp. 6580– 6594, 2024. [12] C. Chaccour, W. Saad, M. Debbah, Z. Han, and H. Vincent Poor, “Less data, more knowledge: Building next-generation semantic communication networks,” IEEE Commun. Surveys Tuts., vol. 27, no. 1, pp. 37–76, 2025. [13] S. Hasan, R. H. Ratul, M. K. Mahadi, A.-H. Fahim, and N. Mehereen, “Semantic and task-oriented v2x communication for connected and automated vehicles: A comprehensive review,” IEEE Open J. Commun. Soc., vol. 7, pp. 10 540–10 583, 2026. [14] M. Segata, R. L. Cigno, T. Hardes, J. Heinovski, M. Schettler, B. Bloessl, C. Sommer, and F. Dressler, “Multi-technology cooperative driving: An analysis based on plexe,” IEEE Trans. Mob. Comput., vol. 22, no. 8, pp. 4792–4806, 2023. [15] SAE International Recommended Practice, “Taxonomy and definitions for terms related to driving automation systems for on-road motor vehicles,” SAE Standard J3016 202104, 2021, revised April 2021, Issued January 2014. [16] G. Jocher, A. Chaurasia, and J. Qiu, “Ultralytics yolov8,” 2023. [Online]. Available: https://github.com/ultralytics/ultralytics [17] L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao, “Depth anything v2,” in Advances in Neural Information Processing Systems, vol. 37. Curran Associates, Inc., 2024. [18] F. Yu, H. Chen, X. Wang, W. Xian, Y. Chen, F. Liu, V. Madhavan, and T. Darrell, “Bdd100k: A diverse driving dataset for heterogeneous multitask learning,” in Proc. 2020 IEEE/CVF CVPR, 2020, pp. 2633– 2642. [19] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “Highresolution image synthesis with latent diffusion models,” in Proc. 2022 IEEE/CVF CVPR, June 2022, pp. 10 684–10 695. [20] C. Sommer, R. German, and F. Dressler, “Bidirectionally coupled network and road traffic simulation for improved ivc analysis,” IEEE Trans. Mob. Comput., vol. 10, no. 1, pp. 3–15, 2011. [21] R. Rajamani, H.-S. Tan, B. K. Law, and W.-B. Zhang, “Demonstration of integrated longitudinal and lateral control for the operation of automated vehicles in platoons,” IEEE T. Contr. Syst. T., vol. 8, no. 4, pp. 695–708, 2000. [22] J. Ploeg, B. T. M. Scheepers, E. van Nunen, N. van de Wouw, and H. Nijmeijer, “Design and experimental evaluation of cooperative adaptive cruise control,” in Proc. 2011 ITSC, 2011, pp. 260–265.