ConceptioArchivearXiv CS
arXiv CSopen access

Lightweight Multi-Vehicle Collaborative Perception Acceleration with Fusion Position Adjustment

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
clouddistributedcomputingparallelcomputing
distributed computing, parallel computing, cloud

Lightweight Multi-Vehicle Collaborative Perception Acceleration with Fusion Position Adjustment Wenzhao Zhang∗‡ , Shujun Han†‡ , Haixiao Gao∗‡ , Mengying Sun∗‡ , Bizhu Wang∗‡ , Xiaodong Xu∗‡§ ∗ State Key Laboratory of Networking and Switching Technology † National Engineering Research Center for Mobile Network Technologies ‡ Beijing University of Posts and Telecommunications, Beijing, China § Department of Broadband Communication, Peng Cheng Laboratory, Shenzhen, Guangdong, China

arXiv:2606.27750v1 [cs.DC] 26 Jun 2026

Email: {zhangwenzhao, hanshujun, haixiao, smy bupt, wangbizhu 7, xuxiaodong}@bupt.edu.cn

Abstract—Multi-vehicle collaborative perception (MvCP) is considered as a key technology to facilitate automated driving (AD), where real-time MvCP under limited resources is significant for reliable AD. In this paper, we formulate a lightweight acceleration scheme for intermediate-fusion (IF) MvCP, which can adapt to both situations of limited computation and communication resources. We provide a relaxed definition conditional additivity and analyze the conditional additivity for various DNN linear layers. On this basis, we focus on the IF-MvCP based on additive feature fusion, and derive the MvCP precision consistency of the forward and backward feature fusion position (FP) adjustments among linear layers. Through experiments, we further validate the precision consistency of the FP adjustment method. Moreover, we propose an FP adjustment among linear layers (FALL) scheme for MvCP acceleration without precision loss theoretically. Simulation results show that the proposed FALL can reduce MvCP latency by up to 74.8% under limited communication resources and by up to 30.3% under limited computation resources. Index Terms—Multi-vehicle collaborative perception, edge intelligence, conditional additivity, inference acceleration, fusion position adjustment.

I. I NTRODUCTION With the breakthroughs in artificial intelligence (AI), automated driving (AD) technology becomes a critical component of intelligent transportation systems (ITS), where accurate positioning and environment perception are significant for reliable intelligent driving [1] [2]. Despite recent advances in single-vehicle perception as the development of multi-modal sensors and computer vision techniques, the challenge for accurate perception remains due to the occlusions and sparse sensor observations [3]. As a solution, researchers investigated multi-vehicle collaborative perception (MvCP) technology leveraging vehicle-to-vehicle (V2V) communication [4]. This cooperative approach of MvCP can share sensing information among connected automated vehicles (CAVs), thereby compensating for the degradation in perception precision caused by their individual single-view limitations [5]. However, there are still challenges in deploying MvCP services in real-time due to the significant computational resource demands and dynamic wireless channel qualities of ITS [6]. The work presented in this paper is funded by the National Natural Science Foundation of China No. 62201079, the Beijing Natural Science Foundation No. L232051.

The traditional cloud-based processing paradigm can not satisfy the real-time requirement of MvCP, which may suffer congestion with massive data transmission [7]. Besides, the edge devices (i.e., automated vehicles and road side units) are now equipped with more powerful computational capabilities, enabling them to process computation-intensive intelligent services [8]. In this context, edge intelligence (EI) emerges as a promising solution to decrease transmission latency by processing intelligent services at the edge devices [9]. To facilitate the deployment of intelligent perception services on the edge devices, several studies have been proposed to reduce the processing overhead of perception tasks, including communication overhead [10] [11] and computation overhead [12] [13]. Specifically, the authors of [10] and [11] introduced the sensor data compression method to reduce the communication overhead. Lu et al. [10] proposed a joint optimization problem of cooperative vehicles selection and compression ratio selection to reduce the size of sensing data, while guaranteeing the perception precision requirement. Wang et al. [11] adapted the variational image compression algorithm to compress the intermediate representations, and then quantized and encoded the latent representation with few bits for transmission. Additionally, the authors of [12] and [13] designed pruning and quantization methods to reduce the computation overhead. Lu et al. [12] proposed a modal cooperative pruning framework designed for camera-LiDAR fused perception in autonomous driving, which attained superior pruning ratios while minimizing precision loss. S et al. [13] focused on optimizing model performance through integration of pruning and quantization techniques, achieving faster inference speed with minimal impact on precision. However, the aforementioned literature [10]–[13] achieved perception acceleration by sacrificing precision. Moreover, these approaches for perception acceleration considered the overhead decrease of computation or communication independently, and they are difficult to adapt simultaneously to both situations with limited computation resources and limited communication resources. To tackle the above challenges, we propose a lightweight fusion position (FP) adjustment among linear layers (FALL) scheme to accelerate MvCP in edge intelligence empowered ITS. The proposed FALL scheme can achieve MvCP acceler-

ation under both limited computation and limited communication resources situations, which can obtain consistent perception precision with MvCP under the original FP theoretically. The main contributions are summarized as follows: • We provide a relaxed definition conditional additivity based on the concept of additivity. Furthermore, we analyze the conditional additivity of DNN linear layers and the DNN model consisting of multiple linear layers. • We derive the MvCP precision consistency of the forward and backward FP adjustments among linear layers. Additionally, we analyze the FP adjustment range of MvCP based on the PIXOR model, and validate the precision consistency under different FPs via experiments. • We propose the FALL scheme, which can achieve MvCP acceleration under both limited computation and limited communication resources situations without precision loss theoretically. Besides, we examine the acceleration performance of FALL under different transmission rates, showing a latency reduction of up to 74.8%. II. S YSTEM M ODEL In this paper, we focus on the intermediate fusion (IF) based MvCP service [5], which requires less transmission bandwidth than early fusion and provides more comprehensive information than late fusion [14]. The inference process of IFMvCP can be concluded as three stages: 1) feature extraction (FE); 2) feature fusion (FF); 3) object detection (OD). As shown in Fig. 1, we consider a scenario of IF-MvCP based on additive feature fusion, containing one ego-vehicle and a set N = {1, ..., N } of collaborative vehicles (co-vehicles). The vehicles can communicate with each other via PC5based V2V sidelink [15]. Each vehicle is equipped with a computing unit for task processing, where the model before FP is deployed and processed at co-vehicles and the ego-vehicle in parallel (FE), and the model behind FP is deployed and processed at ego-vehicle centrally (OD) after additive FF. Considering that there are more affine transformations rather than linear transformations within DNN inference, we define the layers based on the affine transformation (e.g., fully connected, convolution, and batch normalization) as linear layers. Note Pn that linear Pntransformations F satisfy additivity, i.e., F( i=1 xi ) = i=1 F(xi ), while affine transmissions do not satisfy additivity because they consist of both linear transformations and translations. To extend the discussion of 1st stage: feature extraction (FE) … input

Co-1

FP

input

FP

V2V sidelink

Co-N

2nd stage: additive feature fusion (FF)

W

C

C

H

H

H W

detection (OD)

C H

W

FP

3rd stage: object Ego

C

input

V2V sidelink

Output W

Fused feature maps

FP

Fig. 1: IF-MvCP with additive feature fusion system model.

additivity to DNN inference, we provide a relaxed definition conditional additivity drawn inspiration from the separable concept introduced in [16], based on which we will construct the MvCP acceleration scheme. Specifically, the conditional additivity is defined as follows. Definition 1. (Conditional Additivity) The function F is conditional P additive if exits Pnfunctions Fi , for any xi , i ∈ [1, n], n satisfies F( i=1 xi ) = i=1 Fi (xi ). A. Conditional Additivity Analysis of DNN Linear Layers In this subsection, we provide the analysis of conditional additivity for various DNN linear layers, including fullyconnected, convolution, deconvolution, batch normalization, and average pooling. Proposition 1. (Conditional Additivity of Fully-Connected Layer) The computation of fully-connected layer is an affine transformation, which is represented as Ff c (x) = x · W + b. W is P the weight matrix and b is the bias vector. Given n as: Ff c (x) = x =P i=1 xi , FfP c can be transformed P n n n F ( x ) = x · W + b = i i=1 P i=1 i i=1 xi · W + Pfnc Pn n b = (x · W + b ) = F i i i=1 i i=1 i=1 f c,i (xi ), where P n Ff c,i (xi ) = (xi · W + bi ), b i=1 i = b. Therefore, the fully-connected layer satisfies conditional additivity. Proposition 2. (Conditional Additivity of Convolution Layer) The computation of each feature patch for a convolution layer is an affine which is represented as P transformation, P Fconv (x) = x (j + w, k + h) · Kc (w, h) + b. c c w,h xc is the input feature map of channel c, Kc is the kernel of channel c. (j, k) represents the starting pixel of convolution, (w, h) represents the Pnelement pixel of kernel, b is the bias. Given x = i , Fconv can i=1 xP n be transformed as: F (x) = F ( = conv conv i=1xi ) Pn P P x (j + w, k + h) · K (w, h) + b = c i c w,h c,i Pi=1 P P n i=1 Fconv,i (xi ), where Fconv,i c w,h xc,i (j + Pn(xi ) = w, k + h) · Kc (w, h) + bi , b = b. Therefore, the i i=1 convolution layer satisfies conditional additivity. Notably, if the hyperparameter paddingP p of F is not zero, the padding n pi of Fconv,i should satisfy i=1 pi = p. Proposition 3. (Conditional Additivity of Deconvolution Layer) The computation of deconvolution is the same as convolution, which can be represented as Fdeconv (x) = Fconv (x). The main difference between deconvolution and convolution is the size of output, which has no affect on the computation process. Therefore, the deconvolution layer satisfies conditional additivity as well, where Fdeconv,i (xi ) = Fconv,i (xi ). Proposition 4. (Conditional Additivity of Batch Normalization Layer) The computation of batch normalization during inference simplifies to anaffine transformation, which is rep 2 +β. µ and δ are the mean resented as Fbn (x) = γ √x−µ δ 2 +ε and variance counted according to the training data, γ and β are trainable parameters, and ε is a small constant avoiding division by zero. The above parameters are fixed during

Pn inference. Given x = xi , Fbn P can be transformed  P n Pn i=1 xi − n i=1 √ i=1 µi as: Fbn (x) = Fbn ( i=1 xi ) = γ + 2 δ +ε  P Pn Pn  γ(xi −µi ) n √ + β = i=1 Fbn,i (xi ), where i=1 βi = δ 2 +ε  i i=1 Pn Pn −µi ) √i + βi , i=1 βi = β, i=1 µi = µ. Fbn,i (xi ) = γ(x δ 2 +ε Therefore, the batch normalization layer satisfies conditional additivity. Proposition 5. (Conditional Additivity of Average Pooling) The computation of average pooling is a linear transformation. To simplify the illustration, we only analyze the calculation within a single P pooling window, which is represented as Fap (x) = W1H w=1:W,h=1:H x (j + w, k + h). (j, k) is the starting pixel coordinate of the average pooling,PW × H is the size of pooling window. Given n x =P i=1 xi , Fap can P be transformed Pn as: Fap (x) = n Fap ( i=1 xi )  = W1H w=1:W, h=1:H i=1 xi (j + w, k + Pn P 1 h) = = i=1 W H w=1:W, h=1:H xi (j + w, k + h) Pn i=1 Fap,i (xi ), where Fap,i (xi ) = Fap (xi ). Therefore, the average pooling layer satisfies conditional additivity (more precisely, it satisfies additivity). B. Conditional Additivity Analysis of DNN Model Consisting of Multiple Linear Layers On the one hand, the DNN model has a multi-layer structure, where the output of the former layer is the input of the latter one. Thus, the DNN inference can be considered as a composition function. On the other hand, there are many shortcut and skip connection structures in the DNN model [17], which can be regarded as a linear combination function. In this section, we will present the conditional additivity analysis for the composition and linear combination of conditional additive functions (CAFs). Theorem 1. (Composition of CAFs) If functions F1 , F2 satisfy conditional additivity, their composition F1 (F2 (·)) satisfies conditional additivity: n n X X F1,i (F2,i (xi )). (1) xi )) = F1 (F2 ( i=1

i=1

Pn Pn Proof. F1 (F2 ( i=1 xi )) = P F1 ( i=1 F2,i (xi )) n = i=1 F1,i (F2,i (xi )). Theorem 2. (Linear combination of CAFs) If functions F1 , F2 satisfy conditional additivity, their linear combination αF1 (·) + βF2 (·) satisfies conditional additivity: n X

αF1 (

xi ) + βF2 (

i=1

n X i=1

Pn

xi ) =

n X

A. Precision Consistency of FP Adjusted Forward Given a trained MvCP model, we denote the original FP as Loriginal . Besides, we denote the adjusted forward FP as Lf orward , which is ahead of Loriginal . Fig. 2 illustrates the inference processes under FP at Loriginal and Lf orward . If the intermediate features are fused at Loriginal , the output o at Loriginal can be calculated as o=

N X

Fi (fi ),

(3)

i=0

where f0 represents the intermediate feature at Lf orward of ego-vehicle, and fi (1 ≤ i ≤ N ) represents that of co-vehicle i. F0 represents the model between Lf orward and Loriginal deployed at the ego-vehicle, and Fi (1 ≤ i ≤ N ) represents that deployed at co-vehicle i. If the perception features are fused at Lf orward , the output ′ o at Loriginal can be calculated as ′

o = F(

N X

(fi )),

(4)

i=0

where F represents the model between Lf orward and Loriginal deployed at the ego-vehicle centrally. According to the definition of conditional additivity, if F is conditional additive, the outputs at Loriginal under both cases ′ of FP at Loriginal and Lf orward are consistent (i.e., o = o ). Therefore, we can derive the precision consistency for the case of forward FP adjustment. Theorem 3. For an MvCP service with the original FP Loriginal , and Lf orward ahead of Loriginal , assume that the model F between Lf orward and Loriginal satisfies conditional additivity. Then, the perception precision under F P = Lorignal and F P = Lf orward is consistent. Proof. The inference results under F P = Lorignal and F P = ′ Lf orward are denoted as r and r , respectively. The model FE: Ego

Co-1

Co-N

FE: Ego

Co-1

(αF1,i (xi ) + βF2,i (xi )).

i=1

Pn

be adjusted among the linear layers with the same inference precision. In this section, we first derive the precision consistency of the forward and backward FP adjustments, respectively. Afterward, we validate the precision consistency of MvCP service based on the PIXOR model under different FP adjustments.

(2) Pn

Proof. αF1 ( i=1 i=1 (αF1,i ) + P Pxni ) + βF2 ( i=1 xi ) = n i=1 (βF2,i ) = i=1 (αF1,i (xi ) + βF2,i (xi )). III. P RECISION C ONSISTENCY OF F USION P OSITION A DJUSTMENT A MONG L INEAR L AYERS Based on the analysis for conditional additivity of DNN linear layers, it can be derived that the FP of IF-MvCP can

f0 1

0

FF:

fN

f1

o

o0

FF:

o1

OD:

oN Ego

(a) FP= Loriginal

f

L forward

f1

f0

fN

OD:

N

Loriginal

Co-N

o'

Ego

(b) FP= L forward

Fig. 2: The inference processes under FP at Loriginal and Lf orward .

PN ′ after Loriginal is denoted as F1 . Since o = F( i=0 (fi )) = PN ′ ′ i=0 Fi (fi ) = o, r = F1 (o) = F1 (o ) = r . Therefore, Theorem 3 is proved. B. Precision Consistency of FP Adjusted Backward We denote the adjusted backward FP as Lbackward , which is behind Loriginal . Fig. 3 illustrates the inference processes under feature fusion at Loriginal and Lbackward . If the perception features are fused at Loriginal , the output b at Lbackward can be calculated as ′

b=F (

N X

oi ),

(5)

i=0

where o0 represents the output at Loriginal of ego-vehicle, and ′ oi (1 ≤ i ≤ N ) represents that of co-vehicle i. F represents the model between Loriginal and Lbackward deployed at the ego-vehicle centrally. If the perception features are fused at Lbackward , the output ′ b at Lbackward can be calculated as ′

b =

N X

Fi (oi ),

(6)

i=0 ′

where F0 represents the model between Loriginal and ′ Lbackward deployed at the ego-vehicle, and Fi (1 ≤ i ≤ N ) represents that deployed at co-vehicle i. ′ According to the definition of conditional additivity, if F is conditional additive, the outputs at Lbackward under both cases of FP at Loriginal and Lbackward are consistent (i.e., ′ b = b ). Therefore, we can derive the precision consistency for the case of backward FP adjustment. Theorem 4. For an MvCP service with the original FP Loriginal , and Lbackward behind Loriginal , assume that the ′ model F between Loriginal and Lbackward satisfies conditional additivity. Then, the perception precision under F P = Loriginal and F P = Lbackward is consistent. Proof. The inference results under F P = Loriginal and ′ F P = Lbackward are denoted as r and r , respectively. ThePmodel afterPLbackward is denoted as F2 . Since b = ′ ′ ′ ′ N N F ( i=0 oi ) = i=0 Fi (oi ) = b , r = F2 (b) = F2 (b ) = ′ r . Therefore, Theorem 4 is proved.

C. Precision Consistency Validation for FP Adjustment Specifically, we utilize the MvCP service based on the state-of-the-art model PIXOR [18] to validate the precision consistency of FP adjustment. Fig. 4 illustrates the architecture of PIXOR, which consists of a backbone model for feature extraction and a header model for object detection. It is shown that the UpSample layer of PIXOR is composed of the addition of a convolution and a deconvolution, which is represented as Fupsample (x) = Fconv (x) + Fdeconv (x). Because Fconv and Fdeconv satisfy conditional additivity (refer to Proposition 2 and 3), it can be derived that Fupsample satisfy conditional additivity referring to Theorem 2. In addition, since the ResBlocks contain nonlinear structure (i.e., ReLU), they do not satisfy conditional additivity. According to Theorem 1, the composition of CAFs satisfies conditional additivity. Therefore, the FP of PIXOR can be adjusted among the linear layers from Lf orward,min to Lbackward,max (as shown in Fig. 4) with consistent precision, which can be drawn from Theorem 3 and 4. Fig. 5 shows the detection results of MvCP with three CAVs (including one ego-vehicle and two co-vehicles) under the adjusted forward FP, the original FP, and the adjusted backward FP. The average precisions (AP) at different intersection-overunion (IoU) threshold under different adjusted FPs within the linear layers are illustrated in Table I. We denote FP adjusted forward i layers as fi , and FP adjusted backward i layers as bi . The results indicate that the MvCP precision under different adjusted FPs within linear layers is approximately consistent with that under the original FP (with the maximum error not exceeding 0.05), which can further validate the theoretical derivation for the precision consistency of FP adjustment.

FP adjusted range Backbone

1/16

1/8

1/4

ResBlock 3

Co-N

FE: Ego

Co-1

OD:

Loriginal

o1

o0

Conv Conv Conv Conv

ResBlock 2 Conv

Conv

(a) FP= Loriginal

Conv Output

Lbackward ,max

Co-N

… o0

oN

Ego

o1

oN

'

' N

' 0

FF:

'

b

+

Loriginal

Fig. 4: The architecture of PIXOR.

… FF: o

Conv

UpSample 7

Header

Input

Co-1

Deconv

UpSample 6

ResBlock 4

Conv

FE: Ego

L forward ,min

Conv

ResBlock 5

1

b'

b0

b1

Lbackward

TABLE I: Average precisions of different fusion positions. bN Ego

OD:

(b) FP= Lbackward

Fig. 3: The inference processes under FP at Loriginal and Lbackward .

IoU original f3 f2 f1 b1 b2 b3 b4 b5

0.3 0.88 0.87 0.84 0.84 0.84 0.83 0.83 0.88 0.88

0.5 0.85 0.84 0.82 0.82 0.81 0.81 0.81 0.86 0.86

0.7 0.64 0.63 0.60 0.60 0.61 0.60 0.60 0.59 0.59

Then, the comparison of computation latency under FP=L1 and FP=L2 is discussed as follows: PL1 Tc (L1 ) − Tc (L2 ) = (a) Adjusted forward FP (b) Original FP (c) Adjusted backward FP

Fig. 5: The detection results of IF-MvCP with three CAVs under different FP.

A. Computation Latency Considering the randomness of co-vehicle selection, the computation resources of co-vehicles are variable and potentially lower than ego-vehicle. In general, we assume that the computation resource of each co-vehicle fco,i (1 ≤ i ≤ N ) is less than that of the ego-vehicle fego (i.e., fco,i < fego ). Thus, the former FP corresponds to less computation latency, which can be derived as follows. Firstly, we denote the computation latency of FP as Tc (F P ), which is calculated by PF P PF P PLmax j=1 Cj j=1 Cj j=F P +1 Cj Tc (F P ) = max( , )+ fco,i fego fego PF P PLmax j=1 Cj j=F P +1 Cj = + , (7) fco,min fego where Cj represents the computation workload of the j-th layer in PIXOR, Lmax represents the last layer of PIXOR. We consider two different FPs L1 and L2 , which satisfy L1 < L2 .

L forward ,min

Loriginal

Lbackward ,max

=

PLmax +

fco,min PL2 j=1 Cj fco,min PL2

j=L1 +1 Cj

fego PLmax

+

j=L1 +1 Cj

j=L2 +1 Cj

!

fego PL2

j=L1 +1 Cj + fco,min fego   L 2 X fco,min − fego Cj < 0. (8) fco,min · fego

=−

IV. M V CP ACCELERATION S CHEME BASED ON F USION P OSITION A DJUSTMENT W ITHOUT P RECISION L OSS Based on the above analysis for the precision consistency of FP adjustment, we propose the lightweight MvCP acceleration scheme based on FP adjustment among linear layers (FALL) without precision loss. Specifically, the FP can be dynamically adjusted according to the system resource situation to achieve MvCP acceleration. Subsequently, we analyze the acceleration capability of the FALL scheme using MvCP based on PIXOR as an example. Fig. 6 shows the computation workload and intermediate feature size of each layer in PIXOR.

j=1 Cj

j=L1 +1

Therefore, Tc (L1 ) < Tc (L2 ), which means the computation latency of a former FP is less than that of a latter FP. B. Transmission Latency The transmission latency of FP Tt (F P ) is calculated by IF P , (9) R where IF P represents the intermediate feature size at FP, and R represents the transmission rate of the V2V sidelink. It can be seen from Fig. 6 that the intermediate feature size for PIXOR of the latter layer is mostly less than (or equal to, such as L10 ∼ L13 ) that of the former layer. This feature is also generally applicable in other models [19]. Thus, we represent the intermediate feature size of the former FP L1 as I1 and that of the latter FP L2 as I2 , which satisfy I1 ≥ I2 . Thus, Tt (L1 ) = IR1 ≥ Tt (L2 ) = IR2 , which means the transmission latency of a former FP is larger than or equal to that of a latter FP. Tt (F P ) =

C. Total Latency The total latency of MvCP Ttotal is composed of the computation and transmission latency, which is calculated as Ttotal = Tc + Tt . From the above discussion about the effect of FP adjustment on Tc and Tt , it can be observed that the proposed FALL can reduce Tc and Tt , while there is a trade-off between them. Therefore, the optimal FP with minimum latency differs depending on whether the limitation is on computation resources or communication resources. For example, if the computation resources become the performance bottleneck, the optimal FP tends to favor the former layer, while if the communication resources become the performance bottleneck, the optimal FP tends to favor the latter layer. The specific evaluation of FALL for MvCP acceleration under different resource limitations is provided in Section V. V. P ERFORMANCE E VALUATION

Fig. 6: The computation workload and intermediate feature size of each layer in PIXOR.

In this section, we present the simulation results to compare the acceleration performance of our proposed FALL under different transmission rates. Specifically, the simulation is carried out based on PIXOR, where each parameter size is set as 4 Bytes. The computation resource of each co-vehicle

23.6%

74.8%

(a) R = 5 × 108 (bps)

30.3%

(b) R = 5 × 109 (bps)

(c) R = 5 × 1010 (bps)

Fig. 7: The total perception latency of different FPs under transmission rates R ranging from 5 × 108 bps to 5 × 1010 bps.

is set to 0.5 TOPS, and the computation resource of the egovehicle is set to 30 TOPS. Fig. 7 shows the total perception latency of different FP under transmission rates R ranging from 5 × 108 bps to 5 × 1010 bps. When R = 5 × 108 bps, the communication latency becomes the performance bottleneck. In this case, the optimal FP is b5 with the minimum transferred feature size, which can reduce the total latency by 74.8% compared to the maximum value at f3 . When R = 5 × 1010 bps, the computation latency becomes the performance bottleneck. In this case, the optimal FP is f3 , where the most linear layers are processed at the ego-vehicle with more computation resources than co-vehicles. The total latency can be reduced by 30.3% compared to the maximum value at b4 . When R = 5 × 109 bps, the communication latency and computation latency are relatively close. Thus, the optimal FP depends on the trade-off between the communication latency and computation latency, as shown in Fig. 7b at the original FP o. In this case, the total latency can be reduced by 23.6% compared to the maximum value at b4 . VI. C ONCLUSION In this paper, we investigated a lightweight acceleration scheme for IF-MvCP based on additive feature fusion. Firstly, the analysis of the conditional additivity for various DNN linear layers and the DNN model consisting of multiple linear layers was presented. Besides, the precision consistency of the FP adjustment among linear layers was derived. Furthermore, the FALL scheme was proposed to accelerate MvCP while maintaining the perception precision, which can adapt to both situations of limited computation and communication resources. Simulation results validated the effectiveness of the proposed FALL under different limited resource situations. R EFERENCES [1] X. Gao, X. Zhang, Y. Lu, Y. Huang, L. Yang, Y. Xiong, and P. Liu, “A survey of collaborative perception in intelligent vehicles at intersections,” IEEE Trans. Intell. Veh., pp. 1–20, May. 2024, early access. [2] Y. Yang, M. Chen, Y. Blankenship, J. Lee, Z. Ghassemlooy, J. Cheng, and S. Mao, “Positioning using wireless networks: Applications, recent progress and future challenges,” IEEE J. Sel. Areas Commun., vol. 42, no. 9, pp. 2149–2178, Sep. 2024. [3] R. Wang and G. Cao, “Occlusion-aware camera selection in vehicular networks,” IEEE Trans. Veh. Technol., Apr. 2025, early access.

[4] M.-Q. Dao, J. S. Berrio, V. Frémont, M. Shan, E. Héry, and S. Worrall, “Practical collaborative perception: A framework for asynchronous and multi-agent 3d object detection,” IEEE Trans. Intell. Transp. Syst., vol. 25, no. 9, pp. 12 163–12 175, Sep. 2024. [5] R. Xu, H. Xiang, X. Xia, X. Han, J. Li, and J. Ma, “OPV2V: An open benchmark dataset and fusion pipeline for perception with vehicleto-vehicle communication,” in 2022 IEEE Int. Conf. Robot. Autom.n (ICRA), Jul. 2022, pp. 2583–2589. [6] S. Li, H. Chen, F. Tan, N. Zhang, S. Lin, and T. Q. Quek, “Computation offloading in air-ground integrated vehicular edge computing networks,” in IEEE Globecom Workshops, (GC Wkshps), Mar. 2024, pp. 497–502. [7] W. Zhang, S. Han, X. Xu, and P. Zhang, “Joint service placement and model partitioning for accelerating DNN inference in edge intelligence empowered vehicle networks,” IEEE Trans. Veh. Technol., Apr. 2025, early access. [8] X. Liu, J. Liu, and W. Li, “Truthful mechanism for resource allocation and pricing in vehicle-assisted mobile edge computing,” IEEE Trans. Veh. Technol., vol. 74, no. 5, pp. 8171–8186, May. 2025. [9] C. Chen, C. Wang, B. Liu, C. He, L. Cong, and S. Wan, “Edge intelligence empowered vehicle detection and image segmentation for autonomous vehicles,” IEEE Trans. Intell. Transp. Syst., vol. 24, no. 11, pp. 13 023–13 034, Nov. 2023. [10] B. Lu, X. Huang, Y. Wu, L. Qian, S. Zhou, and D. Niyato, “Joint optimization of compression, transmission and computation for cooperative perception aided intelligent vehicular networks,” IEEE Trans. Veh. Technol., vol. 74, no. 5, pp. 8201–8214, May. 2025. [11] T.-H. Wang, S. Manivasagam, M. Liang, B. Yang, W. Zeng, and R. Urtasun, “V2VNet: Vehicle-to-vehicle communication for joint perception and prediction,” in Eur. Conf. Comput. Vis., Aug. 2020, pp. 605–621. [12] Y. Lu, B. Jiang, N. Liu, Y. Li, J. Chen, Y. Zhang, and Z. Wan, “CrossPrune: Cooperative pruning for camera–LiDAR fused perception models of autonomous driving,” Knowl Based Syst, vol. 289, p. 111522, Apr. 2024. [13] R. Jafarpourmarzouni, Y. Luo, S. Lu, Z. Dong et al., “Towards real-time and efficient perception workflows in software-defined vehicles,” IEEE Internet Things J., vol. 12, no. 6, pp. 7240–7258, Nov. 2024. [14] T. Tang, C. Zhang, G. Chen et al., “RoCooper: Robust cooperative perception under vehicle-to-vehicle communication impairments,” in Proc IEEE INFOCOM, May. 2025, pp. 1–10. [15] 3GPP, “Study on Vehicle-to-Everything,” Sohpia Antipolis, France, TR 38.885 V2.0.0, Mar. 2019. [16] L. Sun, H. Li, Y. Peng, and J. Cui, “Serpens: Privacy-preserving inference through conditional separable of convolutional neural networks,” in Proc. ACM Int. Conf. Inf. Knowl. Manage., Oct. 2022, pp. 1837–1847. [17] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conf. Comput. Vis. Pattern Recognit., Jun. 2016, pp. 770–778. [18] B. Yang, W. Luo, and R. Urtasun, “PIXOR: Real-time 3D object detection from point clouds,” in IEEE Conf. Comput. Vis. Pattern Recognit., Jun. 2018, pp. 7652–7660. [19] Z. Liu, H. Du, J. Lin, Z. Gao, L. Huang, S. Hosseinalipour, and D. Niyato, “DNN partitioning, task offloading, and resource allocation in dynamic vehicular networks: A Lyapunov-guided diffusion-based reinforcement learning approach,” IEEE Trans. Mob. Comput., vol. 24, no. 3, pp. 1945–1962, Mar. 2025.

Record · ID 319647 · SHA-256 5a1c0de6f47ec1c7
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.