1
Task-Oriented Communication with Hybrid-Precision Models Digital Comm
Binary NN
Binary Feature
Full-Precision NN
arXiv:2607.16766v1 [eess.SP] 18 Jul 2026
Songjie Xie, Graduate Student Member, IEEE, Wei Guo, Member, IEEE, Shenghui Song, Senior Member, IEEE, Edgeand DeviceKhaled B. Letaief, Fellow, Edge Server Jun Zhang, Fellow, IEEE, Ying-Jun Angela Zhang, Fellow, IEEE, IEEE
Abstract—Edge inference has emerged as a promising solution for the proliferation of artificial intelligence (AI) services by deploying models at the network edge to circumvent cloud-routing latency. Existing edge inference approaches mainly focused on either cooperative inference to reduce latency or lightweight model design to fit resource-constrained devices. These solutions often address the communication and computation challenges separately, and thus struggle to achieve a balanced trade-off among transmission efficiency, on-device processing cost, and inference accuracy. To bridge this gap, this paper proposes a hybrid-precision task-oriented communication framework for edge inference to holistically balance communication, on-device computation, and utility. In this framework, a binarized frontend is deployed on the edge device to extract and transmit binary features via orthogonal frequency-division multiplexing (OFDM) signals, while a full-precision back-end on the edge server performs the final inference. To ensure model consistency, we introduce an on-device binarization method tailored for split inference and develop an integrated channel-aware transmission scheme featuring subcarrier-based feature calibration. Furthermore, a knowledge distillation (KD)-based training strategy, supported by specialized gradient estimators, is developed to optimize the end-to-end system and inherit semantic knowledge from a full-precision teacher model. Extensive experiments on the large-scale ImageNet dataset demonstrate the superiority of the proposed hybrid system. Our analysis confirms that this design achieves an optimal trade-off among communication efficiency, on-device computational cost, and inference accuracy, outperforming existing edge inference solutions. Index Terms—Edge Inference, Binary Neural Networks, TaskOriented Communication, OFDM
I. I NTRODUCTION The rapid evolution of artificial intelligence (AI) is fundamentally transforming wireless networks, driving the vision of the future wireless towards supporting ubiquitous intelligent services [1]–[3]. Across a wide spectrum of emerging applications, ranging from virtual/augmented reality (VR/AR) to autonomous driving, the inference based on deep neural networks (DNNs) has emerged as a pivotal enabler for realtime perception and intelligent decision-making [4], [5]. However, the traditional paradigm of centralized cloud inference, where massive volumes of high-dimensional data are offloaded to remote data centers, suffers from prohibitive communication overhead and high routing latency. To circumvent The authors are with the Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology, Hong Kong (e-mail: [email protected], {eeweiguo, eeshsong, eejzhang, eekhaled}@ust.hk). Ying-Jun Angela Zhang is with the Department of Information Engineering, The Chinese University of Hong Kong, Hong Kong (email: [email protected]).
Binary Feature Digital Comm
Binary NN Edge Device
Full-Precision NN Edge Server
Fig. 1. The hybrid-precision framework for edge inference, where the BNN and FPNN are deployed on the edge device and the edge server, respectively, and the binary features are transmitted via digital communication systems.
these systemic bottlenecks and support latency-sensitive tasks, edge inference has emerged as a promising paradigm, where intensive data processing and model execution are strategically shifted from the distant cloud to the network edge [6]–[8]. Despite its great potential, edge deployment is hampered by the tension between prohibitive data transmission overhead and the resource constraints of edge devices. A straightforward strategy is server-based inference, where edge devices offload raw data to an edge server for full DNN execution [9]. Nevertheless, offloading heavy computation, transmitting massive raw data results in high communication overhead and raises data privacy concerns [10]. To mitigate these issues, edgedevice cooperative inference was proposed by splitting the DNN between the resource-constrained edge device and the edge server, allowing the device to transmit intermediate features instead of raw data [11], [12]. This cooperative paradigm is further refined by the emerging task-oriented communication, which focuses on extracting and transmitting only taskessential features to reduce the communication overhead and latency [13]–[15]. This efficiency is further augmented by feature compression and quantization techniques, which map high-dimensional intermediate representations into compact and low-bit discrete values for efficient transmission [16], [17]. Nevertheless, from a hardware perspective, the existing split-inference and task-oriented schemes still need to run a full-precision neural network (FPNN) on edge devices, and the required floating-point operations (FLOPs) can remain prohibitive under the stringent compute and power budgets of typical edge devices. To address the on-device hardware bottleneck, a distinct line of research aims to replace computationally expensive FLOPs with highly efficient bitwise operations, named binary neural networks (BNNs) [23], [26]. Over the years, numerous BNN methods have been developed to binarize AI models endto-end to enable fully on-device execution [24], [25], [27]– [31]. However, while avoiding high FLOPs demand, BNNs still incur significant binary operations (BOPs) on the edge
2
TABLE I E DGE I NFERENCE M ETHOD C OMPARISON . Method
Local Data Exposure
Communication Overhead
Feature Value Type
On-Device FLOPs
On-Device BOPs
Accuracy
Server-Based FPNN [18]
Yes (Local data exposed to the edge server) No
High
Discrete
Minimal
Minimal
Very High
Continuous
High (Partial FPNN)
Low
High
Continuous
Low
High
Discrete (Quantized features) Not Applicable
Very High (Partial FPNN + compression) Very High (Partial FPNN + VQ)
Low
High
Low
High (Whole BNN)
Low
Low
Moderate (Reduced due to E2E binarization) Very High
FPNN-Split [19], [20] FPNN-SplitCompression [12], [13], [16] FPNN-SplitVQ [17], [21], [22]
No
High (Highdimensional features) Low
No
Low
Full BNN [23]–[25]
No (No data transmission)
Minimal (No data transmission)
Proposed Hybrid-Precision
No
Low
Discrete (Binary features)
device. Moreover, the end-to-end binarization severely restricts model expressivity and introduces substantial quantization errors, resulting in performance degradation that limits their suitability for high-accuracy, mission-critical applications. Synthesizing these developments, we observe that existing schemes optimize one dimension at the expense of others. Server-based inference minimizes on-device workload but burdens the communication; model-splitting reduces transmission but strains device FLOPs; and end-to-end BNNs reduce FLOPs but sacrifice utility. These limitations underscore that a viable edge inference framework must move beyond fragmented considerations to holistically balance three critical dimensions: 1) Communication: The system must achieve low communication overhead and ensure compatibility with existing digital communication systems. 2) On-Device Computation: The inference process needs to adhere to the hardware constraints of edge devices, specifically by minimizing both FLOPs and BOPs to ensure execution efficiency. 3) Utility: The system must guarantee sufficient inference accuracy, overcoming the performance degradation often associated with quantization or compression. To combine the best of all the worlds, we propose a hybridprecision task-oriented communication framework for edge inference. As illustrated in Fig. 1, a binarized front-end network is implemented on the edge device, performing the initial processing to extract task-relevant information into binary features. Then, the encoded binary features are transmitted via the digital communication system to the edge server, where the large full-precision back-end network is deployed to perform the final inference. Table I presents the comparison with the existing methods and highlights the strategic advantages of the proposed method on multiple aspects across communication, on-device computation, and utility. Particularly, the communication overhead is measured by the dimensionality of the transmitted data or features, and the feature value type indicates the compatibility of digital communication systems. The negative for this part is highlighted in red, and otherwise in green.
Despite the expected advantages, developing such a hybridprecision task-oriented framework navigates several unique technical challenges. The first fundamental challenge is the feature consistency and utility loss of integrating on-device architectural binarization with model splitting. Training a partial BNN on the edge device to generate intermediate binary features risks a severe degradation in overall model accuracy, as these binarized features must retain the essential information required by the full-precision server-side model for the task. Moreover, most existing edge inference studies assume simple channel conditions, such as AWGN or flat fading, neglecting the frequency-selective fading characteristic of realistic wireless environments. For highly compact binary features, the bit corruption caused by the inter-symbol interference can lead to catastrophic utility degradation. To cope with frequency-selective fading, orthogonal frequency division multiplexing (OFDM) needs to be integrated into the task-oriented communication system. However, the nondifferentiable nature of OFDM systems together with ondevice binarization pose significant challenges to the end-toend optimization for task-oriented communication systems. In this work, we investigate the design of a hybrid-precision edge inference system that seamlessly integrates efficient architectural binarization with co-inference frameworks. To the best of our knowledge, this work introduces the first hybridprecision solution for edge inference with OFDM systems. To address the presented challenges, including utility loss from feature binarization, reliable feature transmission, and the need for holistic optimization, we develop an efficient and fully integrated framework with holistic optimization. Our core contributions are summarized as follows: •
We develop a binarization method for task-oriented wireless edge inference, where the edge device employs a binarized front-end network as a lightweight semantic feature encoder. Unlike conventional BNNs designed for standalone on-device inference, the proposed method is optimized to generate compact binary intermediate features that are both communication-efficient for wireless transmission and semantically compatible with the full-
3
precision back-end deployed at the edge server. We design a channel-aware binary feature transmission and recovery mechanism over multipath fading channels. Specifically, we propose a subcarrier-based calibration method that maps modulated binary feature segments to OFDM subcarriers according to channel reliability, enabling task-relevant feature dimensions to be preferentially delivered through high-quality wireless resources. At the edge server, a reliability-aware feature recovery mask suppresses corrupted binary features and reconstructs real-valued representations for subsequent fullprecision inference. • We develop a holistic knowledge distillation (KD)-based end-to-end optimization framework for the proposed hybrid-precision task-oriented communication system. The KD-based teacher-student strategy maximizes taskrelevant information by transferring knowledge from a full-precision teacher model, while the binary feature representation, subcarrier-based calibration, and wirelessinduced feature recovery naturally impose an information bottleneck (IB) on the transmitted features. In this way, the proposed training framework encourages the edgedevice encoder to extract compact task-relevant representations while suppressing task-irrelevant redundancy under architectural and wireless transmission constraints. STE-based gradient propagation is further employed to handle the non-differentiability introduced by neural binarization, digital modulation and demodulation, and wireless transmission. •
The rest of the paper is organized as follows. Section II presents the related works in edge inference and Section III introduces the system model of the hybrid-precision edge inference system with OFDM. Section IV presents the hybridprecision edge inference framework, including the design of on-device BNNs, feature calibration, and server-based FPNNs. The KD-based optimization and gradient propagation strategy of the proposed frameworks are presented in Section V. In Section VI, we provide extensive simulation results to evaluate the performance and effectiveness of the proposed hybridprecision design. Finally, Section VII concludes the paper. II. R ELATED W ORK A. Edge Inference Edge inference has shifted from server-based execution to collaborative device-edge intelligence. Early server-centric solutions suffered from high latency and privacy risks when offloading raw data [18]. While edge co-inference mitigates this by partitioning DNNs between devices and servers, it faces the data amplification issue, where intermediate features exceed the original input size [12]. Thus, specialized compression is vital to reduce dimensionality and achieve low-latency inference [16]. Concurrently, the advancement of learning-based communication has redefined feature transmission strategies. Joint source-channel coding (JSCC) utilizes DNN-based encoders and decoders to transmit information over wireless channels,
showing significant success in image and text transmission under various channel models [19], [20], [32]. Building on this, the task-oriented communication paradigm was developed to shift the objective from data reconstruction to successful task completion [33]. By leveraging the information bottleneck (IB) principle, these frameworks discard task-irrelevant information, significantly enhancing communication efficiency [13]. Recent studies have further extended this principle to address multi-device cooperation [34], [35], out-of-distribution (OOD) robustness through invariant risk minimization [36], and crossmodel alignment for heterogeneous edge environments [37]. Despite the efficiency of task-oriented schemes, the reliance on continuous floating-point feature values poses significant integration challenges for modern digital radio frequency (RF) systems [22]. This has motivated the incorporation of discrete representation learning techniques, such as vector quantization (VQ) [38]. Frameworks utilizing VQ-based IB or vector quantized-variational autoencoders (VQ-VAEs) enable digital transmission by mapping features into discrete codebooks, often coupled with adaptive modulation to enhance robustness against channel noise [17]. However, existing model-splitting and discrete representation solutions frequently overlook the hardware-level constraints of edge devices, particularly regarding FLOPs capacity. The deployment of standard DNN layers, alongside additional compression and VQ modules, often exceeds the computational budget of low-power edge hardware. This computational bottleneck remains a critical barrier to the practical deployment of sophisticated edge AI models, necessitating the exploration of ultra-low-precision architectures like the one proposed in this work. B. Binary Neural Network The fundamental motivation for BNNs in edge intelligence is to replace resource-intensive floating-point operations with bitwise XNOR and popcount operations. However, the nondifferentiable nature of binarization and the inherent loss of information represent significant hurdles. Existing research to address these challenges can be categorized into two main trajectories: precision-centric optimization and architecture and deployment efficiency. Early research focused on minimizing the quantization error between full-precision and binary representations. XNORNet [23] introduced analytic scaling factors to close the accuracy gap, while ABC-Net [27] utilized multiple binary bases to approximate full-precision weights. To stabilize the training process and mitigate gradient mismatch caused by the sign function, IR-Net [28] and ProxyBNN [29] proposed information-retention and proxy-function methods. Recent advances like ReActNet [25] and ReCU [30] further refined this by introducing generalized activation functions and rectified clamping to preserve distribution entropy. Despite these innovations, a persistent performance gap remains compared to full-precision counterparts. Parallel to training optimizations, architectural innovations like Bi-Real-Net [24] introduced shortcut connections to preserve signal flow. This evolved into automated discovery via Neural Architecture Search (NAS), with frameworks
4
Binary Features Source
QAM + OFDM Modulation
BNN
Multipath Channel
Corrupted Features Target
FPNN
QAM Mapping
QAM + OFDM Demodulation
Add CP
IFFT
QAM + OFDM Modulation QAM Demapping
Channel Equal.
c
FFT
Remove CP
QAM + OFDM Demodulation
Fig. 2. The system model of the hybrid edge inference system with OFDM.
like BNAS [39] finding that wider layers and specific skipconnections are more resilient to binarization noise. To further reduce the overhead of deployment, DIR-Net [40] and AdaBin [31] focused on discriminative features and adaptive quantization levels to ensure the network remains lightweight yet task-effective. Despite these advances, BNNs are typically treated as monolithic entities for on-device deployment. They still encounter the resource-accuracy paradox: end-to-end BNNs often lack the precision for complex tasks, yet scaling them to improve accuracy significantly increases on-device BOPs. Furthermore, existing BNNs are not optimized for the splitinference paradigm, where binary features must be both computationally efficient and robust enough to serve as descriptors for a high-precision server-side back-end. III. S YSTEM M ODEL As shown in Fig. 2, we consider a point-to-point OFDMbased task-oriented communication system for edge inference. Specifically, the edge device utilizes a BNN for feature encoding and outputs the binary features. The binary features are represented by the quadrature amplitude modulation (QAM) and fed to the OFDM extension to generate OFDM symbols. After transmission over the multipath channels, the edge server decodes the received signals using the QAM and OFDM demodulator and performs the inference based on the FPNN. The details are stated as follows. A. The BNN-Based Transmitter with OFDM At the edge device, there is a data source x ∈ X associated with its target y ∈ Y (e.g., the label for the image) with an underlying joint distribution p(x, y). The BNN fθ̄ served as a feature encoder deployed at the edge device with the binarized parameters θ̄. The data x is encoded by fθ̄ into the bitstream b = fθ̄ (x),
(1)
such as Km = Kb /4 for 16-QAM. Then, we need to reshape m into Ns OFDM symbols, S = (s1 , s2 , . . . , sNs ) ∈ CNs ×NFFT ,
(3)
where S is the collection of Ns OFDM symbols and NFFT is the number of subcarriers. Specifically, the reshaping operation ensures Km = Ns × NFFT . Subsequently, the inverse fast Fourier transform (IFFT) is applied to each symbol, and the cyclic prefix (CP) with the length of NCP is inserted into the OFDM signals to ensure subcarrier orthogonality. After that, the transmitter transmits the time domain signals z ∈ CNs ×(NFFT +NCP ) propagates through the multipath fading channel. B. Signal Propagation over Wireless Multipath Channel We consider a Ψ-tap discrete-time multipath fading channel. The received time-domain signal is modeled as the convolution of the transmitted OFDM signal with the channel impulse response (CIR) plus additive noise, i.e., z̃ = η ∗ z + ϵ,
(4) Ψ
where ∗ denotes linear convolution, η ∈ C denotes the CIR with Ψ propagation paths, each experiences independent Rayleigh fading satisfying ηψ ∼ CN (0, σψ2 ) for ψ = 1, 2, . . . , Ψ, and ϵ ∼ CN (0, σ 2 I) is the additive complex Gaussian noise. The power of each path follows an exponential decay profile σψ2 = αψ exp(− ψγ ), where αψ is a normalization PΨ coefficient to satisfy ψ=1 σψ2 = 1 and γ is the delay spread constant. For the transmitted OFDM signals, the channel response in the frequency domain for the subcarriers is denoted as h = (h1 , h2 , . . . , hNFFT ) ∈ CNFFT . Then, for the i-th OFDM symbol, the received frequency-domain symbols s̃i = (s̃i,1 , s̃i,2 , . . . , s̃i,NFFT ) satisfy s̃i,k = hk si,k + ni,k ,
k ∈ {1, 2, . . . , NFFT },
(5)
where b ∈ {0, 1}Kb denote the encoded binary features and Kb is the length of the encoded bitstream. Then, the transmitter modulates encoded bitstream b with fixed-size constellations (e.g., 16-QAM) by
where ni,k denotes the noise on the k-th subcarrier of the i-th received OFDM symbol.
m = M(b),
At the edge server, the receiver first removes the CP from the received time-domain signal s̃, and then transforms it into frequency-domain signals by applying fast Fourier transform (FFT), which denoted as S̃ = (s̃1 , s̃2 , . . . , s̃Ns ) ∈ CNs ×NFFT . We assume that each sub-carrier has a bandwidth that is much smaller than the coherence bandwidth of the channel. The instantaneous channel estimations on all the subcarriers can be
(2)
where m = (m1 , m2 , . . . , mKm ) ∈ CKm denotes the modulated complex signals, M(·) stands for the fixed-size modulation, and Km is the length of modulated signals. The relation between the dimensions of m and b is determined by the number of bits represented by the constellations in M,
C. The FPNN-based Receiver with OFDM extension
0.2, -1.6, 1.4, … 0.5, 3.4, -0.9, … 3.0, -4.1, 0.1, ...
≈ 0.1 0.5
⊙ 1, 1, -1, …
0.2, -1.6 -0.5, 3.4
1, -1, 1, ...
-1.3
-"!
-"# = sign(-"! )
"
.
(a) On-Device Binary Neural Network
(b) Subcarrier-Based Feature Calibration
%# = sign(%! )
Weight Binarization
4# $
/%,% /%,' /%,( /%,,
%!
%"!
/+!,+""#
1, -1 -1, 1
≈ α⊙
0.2, -1.6 -0.5, 3.4
%# = sign(%! )
1 0 1 1
Binary Activation (Sign) 1, -1, 1, …
≈ 1, 1, -1, …
/- % /- ' /- ( /- , IFFT
0.2, -1.6, 1.4, … 0.5, 3.4, -0.9, … 3.0, -4.1, 0.1, ...
0 − 23% 0 − 23' 0 − 23( 0 − 23,
4"
1, -1, 1, ...
-"# = sign(-"! )
1 1 0
On-Device BNN
5
(c) Server-Based Full-Precision Network
LBias
PReLU
LBias
BN
1-bit Conv
Sign
LBias
4!
6
1, -1 -1, 1
/',% /',' /(,'
Basic Block
-"!
%!
%"!
≈ α⊙
/- +$+""# .' /- +$+""# .% /- +$+""#
0 − 23/.' 0 − 23/.% 0 − 23/
0.0 0.2 0.3
⊙ 1.5
-0.2
6
7
=
-3.3
1, -1, 1, … 1, 1, -1, … 1, -1, 1, ...
0.0 0.3 -1.0
0.0
⊙ 0.3
-1.0
1 0 1 1 1 0 1
Binary Feature Recovery
Server-Based FPNN
! "
Fig. 3. The proposed hybrid-precision task-oriented communication system for edge inference.
estimated at the receiver, denoted as ĥ = (ĥ1 , ĥ2 , . . . , ĥNFFT ). Consequently, conventional equalizers such as the zero-forcing (ZF) or minimum mean square error (MMSE) equalizers can be adopted to efficiently mitigate frequency-selective fading. For instance, the OFDM signal obtained through simple frequency-domain equalization is shown as, ŝi,k =
s̃i,k ĥk
hk
=
ĥk
si,k +
ni,k ĥk
,
(6)
where ŝi,k is the k-th dimension of the equalized symbol ŝi = (ŝi,1 , ŝi,2 , . . . , ŝi,NFFT ). The equalized signals Ŝ = (ŝ1 , ŝ2 , . . . , ŝNs ) will be reshaped into the sequence of complex signals m̃ ∈ CKm and then m̃ is demodulated into the bitstream b̃ with fixed-size constellations by −1
b̃ = M
(m̃).
(7)
The obtained bitstream, treated as binary features, is input into the FPNN fϕ deployed at the edge server with full-precision neural network parameters ϕ. Then, the FPNN outputs the prediction of the target ẑ by ẑ = fϕ (b̃). IV. H YBRID P RECISION E DGE I NFERENCE F RAMEWORK In this section, we present the hybrid-precision task-oriented communication system by developing the on-device binarization and server-based FPNNs. Furthermore, the subcarrierbased feature calibration and binary feature recovery are developed to enhance the binary feature transmission via OFDM. The overall design of the proposed edge inference framework is illustrated in Fig. 3. A. On-Device BNN To facilitate low-latency execution on resource-constrained hardware, we employ an L-layer binary convolutional neural network (CNN) on the edge device. For the l-th layer, let Il ∈ RCl ×Hl ×Wl denote the input activation tensor and
Wl ∈ RCl+1 ×Cl ×K×K represent the weight filters. To maintain clarity in our formulation, we distinguish between fullprecision and binary representations using subscripts: reall , while their corvalued tensors are denoted as IlR and WR l , responding binarized counterparts are denoted as IlB and WB 1 respectively . 1) Weight Binarization: To preserve model capacity while transitioning to discrete weights, we approximate each reall l using a binary filter WB ∈ WB valued filter WR ∈ WR + scaled by a positive factor α ∈ R . The objective is to minimize the reconstruction error such that WR ≈ αWB . ∗ Formally, we seek the optimal WB and α∗ that minimize the following least-squares objective: ∗ WB , α∗ = argmin ∥WR − αWB ∥2 .
(8)
WB ,α
Following the pioneer work of XNOR-Net [23], by expanding the least squared objective of (8), we obtain, T T T α 2 WB WB − 2αWR WB + W R WR .
(9)
T Since WB ∈ {+1, −1}n , we have WB WB = n, where n T is the number of elements in the filter. Given that WR WR is constant for a fixed real-valued weight, the optimization T objective simplifies to α2 n − 2αWR WB + const. Therefore, ∗ the optimal WB is achieved by taking the element-wise sign of the real-valued tensor WR : +1 WR,i ≥ 0 ∗ WB,i = sign(WR )i = . (10) −1 WR,i < 0 ∗ By substituting WB = sign(WR ) back into the objective function, the optimal scaling factor α∗ is derived as the mean of the absolute values of the real-valued filter: T ∗ WR WB ∥WR ∥1 = , (11) α∗ = n n 1 The real-valued tensors Il and Wl are latent parameters used exclusively R R during the training phase for gradient updates. Upon deployment, the model l ), thereby executes inference using only the binary parameters (IlB , WB incurring no additional storage or computational overhead.
6
where ∥ · ∥1 denotes the ℓ1 -norm. While weight binarization significantly reduces computational complexity, achieving the full computational benefits of BNNs requires substituting expensive floating-point convolutions with efficient bitwise operations. This necessitates the binarization of input activations at each layer through a binary activation function. 2) Binary Activation: The binary activation IB is obtained by applying the element-wise sign function to the real-valued features IR , such that IB = sign(IR ). Note that unlike weight binarization, the magnitude of the activation tensor IlR is intentionally discarded to further reduce FLOPs. By combining l l the binary weights WB = sign(WB ) and binary activations l IB , the computationally intensive floating-point convolution in the l-th layer can be approximated using highly efficient bitwise operations: l l WR ∗ IlR ≈ α∗ (sign(WR ) ⊛ IlB ),
IL+1 [:, i, j]∗ , Υ[i, j]∗ = B argmin IL+1 [:,i,j],Υ[i,j] B
2 ∥Υ[i, j]IL+1 [:, i, j] − IL+1 B R [:, i, j]∥ .
IL+1 [:, i, j]∗ = sign(IL+1 B R [:, i, j]) Υ[i, j] =
(13)
To further mitigate the fidelity loss induced by binarization, we incorporate a residual shortcut for real-valued features that bypasses the binary convolution. Additionally, learnable biases β1 and β2 are introduced following the binary convolution and the shortcut summation, respectively. These parameters serve to shift and scale the feature distribution, optimizing the activation range of the PReLU function. Consequently, the output features of the l-th layer are expressed as: ∗ l l l Il+1 R = PReLU(α (sign(WR ) ⊛ IB ) + β1 ) + β2 + IR . (14)
4) Feature Binarization before Transmission: Following the L-layer binarized front-end, the input x is transformed into the output feature map IL+1 ∈ RCL ×HL ×WL . Crucially, R while the internal weights and activations of the front-end are binarized, the inclusion of residual shortcuts results in these final features remaining real-valued. To facilitate transmission over a digital communication, these real-valued features IL+1 R must be decomposed into a binary bit sequence according to our proposed encoding mechanism. To preserve the semantic integrity of the features and keep model consistency during this binarization, we adopt scaling factors Υ ∈ R+,WL ×HL and
(15)
Following a derivation similar to the weight binarization in (8), the optimal binary features and scaling factors are analytically determined by:
(12)
where ∗ denotes the standard real-value convolution operation and ⊛ represents the binary convolution, which can be implemented via efficient bitwise logic, such as XNOR and popcount [23]. This substitution of floating-point multiplyaccumulate (MAC) operations significantly reduces the power consumption and memory footprint of the edge device. 3) Basic Architectural Block: However, static scaling factors often fail to accommodate the varying distributions of binarized input tensors, leading to significant information loss. To dynamically adapt to these distributions, we introduce a trainable bias to the input tensors prior to binarization. This allows the network to learn an optimal activation threshold during training. Inspired by the learnable thresholds in ReActNet [25], we propose an elastic binarization function characterized by the learnable parameter β0 ∈ R: IlB = sign(IlR + β0 ).
perform spatial-wise scaling for the binarization. Specifically, for each spatial coordinate (i, j), where 0 < i ≤ HL and 0 < j ≤ WL , we approximate the feature vector IL+1 R [:, i, j] using a binary vector IL+1 [:, i, j] ∈ {−1, 1}CL and a local B scaling factor Υ[i, j] ∈ R+ . We formulate this as a reconstruction optimization problem to minimize the binarization error:
∥IL+1 R [:, i, j]∥1 CL
.
(16) (17)
To finalize the transmission payload, the scaling factors Υ are quantized and concatenated with the binary feature maps IL+1 . This combined representation is reshaped into a unified B bit sequence b ∈ {0, 1}Kb . Finally, this sequence is mapped onto complex signal constellations using M -QAM modulation, enabling efficient transmission over the physical layer. B. Subcarrier-Based Feature Calibration To ensure robustness in frequency-selective fading environments, we introduce a subcarrier-based feature calibration scheme that aligns feature transmission with subcarrier reliability. In each transmission block, there are NFFT subcarriers for each OFDM symbol, and each data point x requires Ns symbols to fully transmit the binary features b. Each subcarrier is identified by a pair (t, k), where t = 1, 2, . . . , Ns is the OFDM symbol index, and k = 1, 2, . . . , NFFT is the subcarrier index within the symbol. Furthermore, the subcarrier (t, k) corresponds to the OFDM signal in the OFDM symbols st,k ∈ S. The quality of subcarrier (t, k) is represented as its channel gain |ht,k |2 , where |·|2 denotes the squared magnitude of the complex channel coefficient. We sort these subcarriers in descending order of their channel gains to obtain a reliability permutation π, such that |hπ(1) |2 ≥ |hπ(2) |2 ≥ · · · ≥ |hπ(Ns NFFT ) |2 ,
(18)
where π(m) = (tm , km ) maps the m-th best subcarrier to its corresponding OFDM symbol tm and the subcarrier index km . Rather than utilizing a random mapping, which would subject critical semantic bits to unpredictable channel fades, our scheme establishes a deterministic hierarchy of transmission slots. We consecutively map the modulated feature segments M = (m1 , . . . , mKm ) to the ordered subcarriers: mi → sπ(i) ,
for i = 1, . . . , Ns NFFT .
(19)
By fixing these safe positions according to channel reliability, the framework enables the model to implicitly learn during training which feature dimensions are most critical and should be encoded into these prioritized slots. As illustrated in Fig. 3,
7
this ordered mapping ensures that the most essential taskrelevant information is consistently protected by the highestquality channel resources, effectively maximizing inference utility in challenging multipath conditions2 .
Subcarriers
Correct Bits
Corrupted Bits
Mask
FPNN
C. Server-Based FPNN Following OFDM transmission over the multipath fading channel, the received bitstream b̃ is reshaped into binary feature maps ĨL+1 ∈ {−1, +1}CL ×HL ×WL and provided to the B server-side FPNN alongside the scaling factors Υ. However, directly performing inference on features corrupted by wireless noise leads to a catastrophic degradation in task utility. To mitigate this, we adopt a real-valued tensor Ω ∈ RCL ×HL ×WL as a mask and apply it on ĨL+1 to filter bit errors while B mapping the received bits back into the continuous feature domain. First, to characterize the impairments introduced during transmission, each element of the transmitted feature tensor IL+1 is modeled as passing through an equivalent binary B symmetric channel (BSC). Specifically, for a given spatial coordinate (i, j), where 0 ≤ i < HL and 0 ≤ j < WL , the bit errors are abstracted as a vector of independent Bernoulli variables N[:, i, j] = (N1 , N2 , . . . , NCL )T . Each variable Nk follows the distribution: pk a = −1 Pr[Nk = a] = , (20) 1 − pk a = +1 where 0.5 ≥ pk ≥ 0 is the error probability of the BSC, which can be estimated as the bit error rate of QAM modulation under the corresponding subcarrier channel gain and noise level. Therefore, the corrupted binary feature vector ĨL+1 [:, i, j] can B L+1 be expressed as ĨL+1 [:, i, j] = N[:, i, j] ⊙ I [:, i, j], where B B ⊙ denotes the element-wise multiplication. For notational clarity and owing to the spatial symmetry of the feature maps, we henceforth omit the spatial indices and utilize IB , IR , n, ω and υ to represent IL+1 [:, i, j], IL+1 B R [:, i, j], N[:, i, j], Ω[:, i, j] and Υ[i, j], respectively. Similar to the binarization process, we aim to minimize the information loss between υ · ω ⊙ n ⊙ IB and IR , which can be formulated as the following stochastic optimization problem ω ∗ = argmin E[∥υ · ω ⊙ n ⊙ IB − IR ∥2 ].
(21)
ω
E[υ 2 (ω ⊙ n ⊙ IB )T (ω ⊙ n ⊙ IB )] − 2υE[(ω ⊙ n ⊙ IB )T IR ] + ITR IR
=υ 2
CL X k=1 CL X
E[ωk2 n2k I2B,k ] − 2υ
CL X
E[ωk nk IB,k IR,k ] +
k=1
ωk2 I2B,k − 2υ
k=1
CL X k=1
(1 − 2pk )ωk IB,k IR,k +
CL X
I2R,k
k=1 CL X
I2R,k ,
k=1
(22) 2 Furthermore,
FPNN
Bottleneck
Fig. 4. The feature bottleneck induced by the proposed feature calibration and binary feature recovery. When the channel condition is getting worse and more features are corrupted, more dimensionality of transmitted features is masked to reduce the impact of channel noise and reduce the information complexity of received features.
where the last equivalence is obtained by taking the first and second order moments of the Bernoulli variable Nk E[Nk ] = 1 − 2pk
(23)
E[Nk2 ] = (−1)2 pk + 12 (1 − pk ) = 1.
(24)
By taking the derivative of (22), the optimal scaling factor can also be obtained as ωk∗ =
(1 − 2pk )IB,k IR,k . υ
(25)
As shown in (10) and (17), we have IB,k = sign(IR,k ) and υ = ∥IR ∥1 /CL , and we can obtain (1 − 2pk )sign(IR,k )IR,k ∥IR ∥1 /CL CL |IR,k | = (1 − 2pk ) PCL . k=1 |IR,k |
ωk∗ =
(26) (27)
As derived, the optimal mask ωk∗ comprises two components: the channel reliability factor (1 − 2pk ) and the feature magniC |I | tude ratio PCLL R,k . However, as transmitting the individual j=1 |IR,j |
By expanding the objective function in (21), we have
=υ 2
FPNN
the quantized bits of the spatial scaling factors Υ are prioritized for transmission over the highest-quality subcarriers. This ensures that these critical values remain error-free, providing a reliable foundation for the subsequent feature recovery process.
feature magnitudes |IR,k | would incur prohibitive communication overhead, this is strictly unavailable at the receiver. To address this, we approximate the optimal mask by assuming a uniform magnitude distribution for each spatial position, which reduces the magnitude ratio to 1. Therefore, in our framework, we adopt ωk∗ = 1 − 2pk and the mask Ω[:, i, j] = ω ∗ for each spatial coordinate (i, j). The recovered real-valued features ÎL+1 are reconstructed as: R ÎL+1 = Υ ⊙ Ω ⊙ ĨL+1 , R B
(28)
where the spatial scaling factors Υ ∈ RHL ×WL is broadcast L+1 across the channel dimension CL such that IˆR [k, i, j] = L+1 ˜ Υ[i, j]Ω[k, i, j]IB [k, i, j]. Then, the server-based FPNN leverage ÎL+1 to perform the inference. R
8
Remark 1. The binary feature recovery (28) can be interpreted as a noise filter on Ĩl+1 based on the quality of B subcarriers. To be specific, if the quality of the subcarrier is really bad, the error probability pk → 0.5 then Ωk → 0. If the quality of the subcarrier is really good, the pk → 0 then Ωk → 1. Because of the calibration and the mask Ω of binary feature recovery, it will induce a bottleneck, as shown in Fig. 4. The induced bottleneck not only contributes to controlling the uncertainty of noisy feature processing but also enhances the training of the feature encoding that will be discussed in the next section. V. T HE E ND -T O -E ND T RAINING FOR H YBRID S YSTEMS In this section, we develop the KD-based objective function and the model training strategy for the hybrid-precision taskoriented communication system. A. KD-Based Optimization for Hybrid Frameworks Training the hybrid-precision architecture presents significant challenges due to the restricted expressivity of the on-device binarized neural parameters. Such binarized architectures often lack the flexibility required to capture the complex data distributions only implied by sparse categorical labels, leading to suboptimal convergence when trained solely with standard cross-entropy (CE) loss. Therefore, we employ knowledge distillation (KD) to bridge this expressivity gap and provide richer supervisory signals. Specifically, a pretrained FPNN is adopted as a teacher model, pT (y|x), to guide the hybrid model by minimizing the Kullback-Leibler (KL) divergence between their respective outputs. Consequently, the holistic training objective L is formulated as a combination of the CE loss with ground-truth labels and the KD loss: L = Ep(b̃|x)p(x,y) [λCE(y, qϕ (y|b̃))+
sign sign
: 0. 0.<:< appsgn appsgn
..:=: = : 0. 0.<:<
1 appsgn(0) appsgn(0) 1 −1 −1 −1 −1
11
00
1appsgn 1appsgn00 11 10 10
00 (a) The approximation for sign function and its derivative −1 −1
11
non-differentiable non-differentiable channel channelmodel model
01 01
XNOR XNOR
11
22 1 1 22 001 1
binary binarynoise noise (b) STE for non-differentiable channel models Fig. 5. Gradient approximation of sign function and non-differentiable channel models.
is implicitly enforced by the architectural and physical design of our framework. Unlike continuous real-valued representations, the entropy of the discrete binary features b̃ ∈ {−1, +1}Kb establishes a hard upper bound on the mutual information, such that I(X; b̃) ≤ H(b̃) ≤ Kb . Furthermore, the stochastic bit corruption introduced by the frequency-selective multipath channel, coupled with the selective nature of our subcarrierbased feature calibration, acts as an additional severe physical bottleneck. As illustrated in Fig. 4, these compounding structural limitations naturally restrict I(X; b̃). Therefore, training our hybrid-precision framework under these stringent architectural and wireless constraints inherently fulfills the IB optimization, extracting a maximally compact and highly taskrelevant representation for edge inference. B. Gradient Propagation via STE
(1 − λ)KL(pT (y|x)∥qϕ (y|b̃))]. (29) where λ balances the contributions of the ground-truth labels and the teacher’s soft predictions, and qϕ (y|b̃) denotes the predictive distribution of the server-side FPNN given the received binary features b̃. From an information-theoretic perspective, minimizing the cross-entropy objective fundamentally drives the system to extract features that are maximally informative about the downstream inference task. Specifically, utilizing the variational distribution qϕ (y|b̃) to approximate the intractable true posterior mathematically maximizes the variational lower bound of the mutual information I(Y ; b̃): I(Y ; b̃) ≥ Ep(y,x)p(b̃|x) [log qϕ (y|b̃)].
..:<: <
(30)
Crucially, this optimization framework intrinsically aligns with the information bottleneck (IB) principle, which is the essential core of task-oriented communication systems. The standard IB principle dictates maximizing task-relevant information I(Y ; b̃) while strictly constraining the total extracted information I(X; b̃) to avoid transmitting task-irrelevant redundancy. Although our loss function L does not require an explicit regularization term to penalize I(X; b̃), this information constraint
The end-to-end training of the proposed hybrid precision framework under the KD-based loss function L involves the backpropagation with non-differentiable parts, including weight binarization, binary activation, and the digital communication modules (e.g., QAM modulation). To address this issue, we adopt the straight-through estimator (STE) to ensure reasonable gradients for effective backpropagation. For the weight binarization, the derivative of the sign function using STE can be expressed as ∂L ∂L ∂ sign(WR ) STE ∂L . (31) = ≈ ∂WR ∂ sign(WR ) ∂WR ∂ sign(WR ) For the binary activation, using STE fails to learn weights near the borders of −1 and +1, which greatly harms the updating ability of back propagation and restricts the expressiveness of the forward features. To address this, a polynomial step function is used to approximate the forward sign function +1, x>1 2x − x2 , 1≥x>0 appsgn(x) = . (32) 2x + x2 , 0 ≥ x > −1 −1, −1 ≥ x
,,
9
TABLE II T HE N EURAL N ETWORK A RCHITECTURE Model Type
Name
BackBone
Layers
BOPs (×109 )
FLOPs (×108 )
#Params (×106 )
Storage (MB)
On-Device BNN
B-9 B-15
ResNet-18 ResNet-34
Conv + [1-bit conv]×8 Conv + [1-bit conv]×14
0.88 1.57
1.20 1.22
0.68 1.34
0.08 0.16
Server-Based FPNN
FP-9 FP-19 FP-28 FP-79
ResNet-18 ResNet-34 ResNet-50 ResNet-101
ResBlock×4 + Dense ResBlock×9 + Dense Bottleneck×9 + Dense Bottleneck×26 + Dense
0.00 0.00 0.00 0.00
8.25 19.85 22.74 59.80
11.01 20.43 24.11 43.11
42.01 77.97 91.99 164.45
Therefore, the derivative of the input features can be expressed as, ∂L ∂L ∂ sign(IR ) = ∂ IR ∂ sign(IR ) ∂IR STE ∂ appsgn(IR ) ∂L , ≈ ∂ sign(IR ) ∂IR
(33) (34)
Similarly, the derivative of the learnable bias β can be expressed as, ∂ sign(IR + β) STE ≈ 1. ∂β
(35)
For the non-differentiable channel models and the digital communication modules, STE can be directly applied as Fig. 5. The proposed STE estimators approximate gradients for the non-differentiable sign and transmission functions during backpropagation. This allows the model to be trained end-toend by integrating IB-based theoretical foundation with KD. Ultimately, this unified strategy ensures robust convergence and maximizes the preservation of task-relevant information in the hybrid-precision model under varying channel conditions. VI. E XPERIMENTS A. Experimental Setup 1) Dataset: The experiments are conducted on the ILSVRC12 ImageNet dataset. ImageNet [41] is a large-scale dataset with 1000 classes and 1.2 million training images and 50k validation images. Compared to other small-scale datasets commonly used in edge inference studies, it is more challenging and suitable for recent edge inference application scenarios due to its large scale and great diversity. 2) Model Architectures: We adopt ResNet [42] architectures as the backbones of the experiments and employ a group of ResNet architectures from small to large sizes, including ResNet-18, ResNet-34, ResNet-50, and ResNet-101. Then, we split the standard ResNet at the middle point into two parts, where the front half is binarized into the on-device BNN using the proposed method, and the latter half is adopted as the server-based FPNN. We name the on-device BNN and the server-based FPNN as B-n and FP-n, respectively, where n is the number of layers. As the size of the on-device BNN is expected to be small to fit a wide spectrum of devices, we only consider B-9 and B-15 (i.e., binarized first half of the ResNet-18 and ResNet-34) in our experiments. The summary of layers, BOPs, and FLOPs of all the considered on-device BNNs and server-based FPNNs is presented in Table II.
3) Baselines: We compare the proposed hybrid framework with other categories of edge inference. As compared in Table I, the server-based FPNN, FPNN-Split-Compression, FPNN-Split-VQ, and Full BNN are adopted in the experiments. For the FPNN-Split category, as it exhibits prohibitive communication overheads due to the data amplification effect, we only consider its enhanced versions with additional compression modules and VQ modules, FPNN-Split-Compression and FPNN-Split-VQ. • Server-Based FPNN: This baseline is to transmit the local data to the edge server, and the FPNN (e.g., ResNet50) performs inference based on the received local data. • FPNN-Split-Compression: Following the previous works [13], [16], the bottleneck modules are introduced in these benchmarks to compress the dimensionality of the intermediate features by reducing the channel number from 256 to 4, 5, and 6. For ease of fair comparison, the floating-point value is formatted with double-precision using 64 bits. • FPNN-Split-VQ: Following the previous works [17], vector quantization modules are introduced to make the floating-point features into the discrete representations. By adopting the codebook with 256 codewords, each 7dimensional feature vector is quantized. • Full BNN: To comprehensively compare the performance of the proposed hybrid framework and BNNs, we select representative BNNs with different binarization techniques, as summarized in Section II-B. They include XNOR-Net [23], ABC-Net [27], Bi-Real-Net [24], IRNet [28], BNAS [39], NASB [43], Si-BNN [44], ProxyBNN [29], RBNN [45], ReActNet [25], ReCU [30], Bihalf [46], SiMaN [47], AdaBin [31], and DIR-Net [40]. 4) Implementation Details: The considered on-device BNNs encode the image from the ImageNet dataset and output the binary features IB with the size IB ∈ {−1, 1}256×14×14 with the scaling factors Υ ∈ R14×14 . Each floating-point value in Υ is quantized into int8 format and represented by 8 bits. Then, the bit sequence will be modulated by QAM-16 modulation into complex signals. We simulate the multipath fading channel with Ψ = 8 signal paths and the delay spread γ = 4. For the OFDM transmission, each OFDM symbol has NFFT = 256 subcarriers and CP length is set to NCP = 16, ensuring it exceeds the delay spread of the multipath channel. For each image, the binary features IB take 49 OFDM symbols and the binarized Υ take 2 OFDM symbols. Therefore, 51 OFDM symbols are transmitted for the proposed hybrid framework. For the KD-based training of the proposed framework, a full-precision ResNet-101 serves as the teacher for the ResNet-
10
ResNet-101
78 ResNet-50
Accuracy (%)
76 ResNet-34
74 72 ResNet-18
70 Original FPNN Hybrid (B-9) Hybrid (B-15)
68 66 FP-9
FP-19
FP-28
FP-79
Server-Based FPNN Fig. 6. The accuracy of the hybrid framework with different pairs of on-device BNNs and server-based FPNNs, and the corresponding original FPNNs. TABLE III P ERFORMANCE COMPARISON OF FULL BNN S AND THE PROPOSED HYBRID FRAMEWORK BASED ON R ES N ET-18 BACKBONE Methods
Acc (%)
BOPs (×109 )
XNOR-Net ABC-Net Bi-Real-Net IR-Net BNAS NASB Si-BNN ProxyBNN RBNN ReActNet ReCU Bi-half SiMaN AdaBin DIR-Net Hybrid (B-9)
51.20 61.00 56.40 58.10 58.76 60.50 58.90 63.70 59.90 65.50 61.00 60.40 60.10 63.10 60.40 65.90
1.70 5.10 1.68 1.68 1.68 1.68 1.70 1.68 1.68 1.68 1.68 1.68 1.68 1.68 1.68 0.88
FLOPs (×108 ) 1.41 4.89 1.63 1.63 1.63 1.63 1.63 1.63 1.63 1.63 1.63 1.63 1.63 1.63 1.63 1.20
101-based hybrid model, while a full-precision ResNet-50 is employed as the teacher for all other configurations. The inference performance is evaluated by the top-1 accuracy. The binary and floating-point computation overhead are evaluated by BOP and FLOP. Note that OP = BOP = FLOP/64 if we comprehensively consider the computational overhead of BOP and FLOPs into a single metric OPs. B. Model Capacity of Hybrid Edge Inference Before comparing the proposed hybrid framework with coinference frameworks and on-device BNNs, we first investigate the model capacity and expressivity of the hybrid-precision design by comparing it with the original FPNN. We train hybrid frameworks with on-device BNNs, B-9 and B-15, with the server-based FPNNs including FP-9, FP-19, FP-28, and FP-79. Classification accuracy of all the evaluated methods is presented in Fig. 6. The original FPNN is expected to be the performance upper bound, showing the performance gap between the hybrid model and the original FPNN model. This performance gap is most evident at the FP-9 level, where around 4% accuracy gap can be observed. And the accuracy of all methods converges as the server-side model scales toward FP-79 (ResNet-101). The initial 4% performance gap between Hybrid (B-9) and the original FPNN at the lower tier narrows
TABLE IV P ERFORMANCE COMPARISON OF FULL BNN S AND THE PROPOSED HYBRID FRAMEWORK BASED ON R ES N ET-34 BACKBONE Methods
Acc (%)
BOPs (×109 )
FLOPs (×108 )
XNOR-Net ABC-Net Bi-Real-Net IR-Net BNAS NASB Si-BNN ProxyBNN RBNN ReCU Bi-half SiMaN AdaBin DIR-Net Hybrid (B-9) Hybrid (B-15)
56.49 66.70 62.20 62.90 59.81 64.00 63.30 66.30 63.10 65.10 64.17 63.90 66.40 64.10 69.15 71.33
3.53 10.59 3.53 3.53 3.53 3.53 − 3.53 3.53 3.53 3.53 3.53 3.53 3.53 0.88 1.57
2.50 2.80 2.50 2.50 2.50 2.50 − 2.50 2.50 2.50 2.50 2.50 2.50 2.50 1.20 1.22
to a negligible margin at the highest tier. This suggests that a powerful server-side model can refine coarse and binarized features to a level nearly identical to full-precision performance. Furthermore, since the on-device BNNs only utilize the front halves of ResNet-18 and ResNet-34, configurations such as B-9/FP-79 and B-15/FP-79 have fewer total layers than their corresponding full-precision counterparts, ResNet-50 and ResNet-101. This highlights the superior balance of efficiency and accuracy achieved by the proposed hybrid solution. C. Performance Comparison: Hybrid Framework vs. Full BNNs We compare the proposed hybrid frameworks against stateof-the-art full BNNs deployed solely on the edge device. For the hybrid framework, binary features are transmitted over multipath fading channels with an SNR = 20dB to the server-based FPNN. In contrast, full BNNs execute the entire inference process locally on the edge device. However, the existing BNNs research is mainly based on lightweight backbones such as ResNet-18 and ResNet-34. To ensure the fair comparison, we only compare under the backbones (i.e., ResNet-18 and ResNet-34) of the existing BNNs considered. In Table III and IV, the experimental results demonstrate that the proposed hybrid-precision framework achieves superior accuracy while maintaining significantly lower on-device computational costs. For the ResNet-18 configuration, Hybrid (B-9) reaches 65.90% accuracy, surpassing leading binarized models such as ReActNet and ProxyBNN. Notably, the hybrid approach achieves this while requiring only 0.88 × 109 BOPs and 1.20 × 108 FLOPs, which is nearly half the computational workload of standard BNNs. This performance advantage extends to the ResNet-34 backbone, where Hybrid (B-15) achieves 71.33% accuracy, outperforming the best-performing traditional BNN, ABC-Net, by a substantial margin of 4.63%. Furthermore, while traditional BNNs are often incompatible with large-scale backbones like ResNet-101 due to the prohibitive local memory and computational requirements, the proposed hybrid framework overcomes these limitations. Distributing the workload, it enables resource-constrained edge devices to leverage the power of deep architectures.
11
FPNN-Split-Comp. Error Rate FPNN-Split-VQ Hybrid
FPNN-Split-Comp. Error Rate FPNN-Split-VQ Hybrid
FPNN-Split-Comp. Error Rate FPNN-Split-VQ Hybrid
0.40
0.36
0.36
0.36
0.30
0.30
0.32
0.24
0.24
49.00
2.70
55.00
49.00 6.00
61.00
4.10
55.00 9.30
61.00
OPs
# Symbols
49.00 9.50 14.90
10.25
61.00
OPs
# Symbols
(a) ResNet-18 Backbone
4.40
55.00
16.10
OPs
# Symbols
(b) ResNet-34 Backbone
(c) ResNet-101 Backbone
Fig. 7. Performance comparison of the proposed hybrid, FPNN-Split-Compression, and FPNN-Split-VQ under SNR=20dB. FPNN-Split-Comp. Error Rate FPNN-Split-VQ Hybrid
FPNN-Split-Comp. Error Rate FPNN-Split-VQ Hybrid
0.40
0.38
0.36
0.36
0.34
0.30
0.32 49.00 55.00
FPNN-Split-Comp. Error Rate FPNN-Split-VQ Hybrid
0.31 2.70
49.00 6.00
61.00
55.00 9.30
(a) ResNet-18 Backbone
49.00 9.50
61.00
OPs
# Symbols
0.24 4.10 55.00 14.90
10.25
61.00
OPs
# Symbols
4.40
(b) ResNet-34 Backbone
16.10
OPs
# Symbols
(c) ResNet-101 Backbone
Fig. 8. Performance comparison of the proposed hybrid, FPNN-Split-Compression, and FPNN-Split-VQ under SNR= 10dB.
D. Performance Comparison: Hybrid Framework vs. Edge Co-Inference Frameworks Then, we compare the proposed hybrid design against two established edge co-inference baselines: FPNN-Split-Comp. and FPNN-Split-VQ. The comparison is conducted under multipath fading channels at two SNR levels: 20dB and 10dB. We analyze the frameworks across three critical performance dimensions: communication overhead, computational complexity, and classification error rate. In Fig. 7, the experimental results at SNR = 20dB demonstrate that the hybrid framework significantly outperforms both splitting-based solutions across all backbone architectures. The hybrid model consistently occupies the innermost region of the radar charts, showing the lowest error rate, the fewest required OFDM symbols, and drastically reduced computational OPs. For example, in the ResNet-101 configuration, the hybrid framework achieves an error rate of even less than 0.24, whereas the FPNN-Split-VQ and FPNN-Split-Comp. suffer from higher error rates and nearly triple the computational cost. This highlights the hybrid framework’s superior ability to extract and transmit compact, noise-resilient features. When the channel conditions degrade to an SNR of 10dB, as shown in Fig. 8, all models experience a decline in classification accuracy due to increased signal interference. However, the per-
formance hierarchy remains unchanged: the hybrid framework continues to provide the best balance of resource efficiency and accuracy, proving its robustness in more challenging wireless environments despite the lower absolute accuracy. E. Robustness To evaluate the robustness of the proposed hybrid framework, we analyze the classification accuracy under varying testing SNR conditions (SNRtest ) for models trained at noise levels SNRtrain = 5, 10, 15, and 20 dB. As illustrated in Fig. 9, there is a clear trade-off between accuracy and noise resilience based on the training environment. Models trained at high SNRs (e.g., 20 dB) achieve the highest peak accuracy when testing conditions are ideal but suffer from a cliff effect, where performance drops to near zero as SNRtest decreases below 5 dB. In contrast, models trained under noisier conditions (e.g., 5 dB) exhibit significantly higher robustness; for the Hybrid (B-9, FP-9) model, training at 5 dB maintains approximately 60% accuracy even at 0 dB SNRtest , whereas the 20 dB trained model fails completely. This suggests that low-SNR training effectively acts as a noise-regularization technique, allowing the server-based FPNN to learn to decode heavily distorted binary features. The impact of the serverbased backbone complexity is also evident when comparing
12
64 60
Accuracy (%)
50
Accuracy (%)
62
66.0 65.5
40
65.0 64.5
30
64
58
62
56
60 5.0
54
64.0
0.0
63.0 5.0
7.5
10.0
12.5
15.0
17.5
2.5
5.0
SNRtrain = 5dB
SNRtrain = 15dB
SNRtrain = 10dB
SNRtrain = 20dB
12.5
15.0
17.5
7.5
10.0
12.5
20.0
Hybrid
15.0
17.5
20.0
(a) SNRtrain = 5dB 65
0.0
2.5
5.0
7.5
10.0
12.5
15.0
17.5
20.0
Accuracy (%)
SNRtest (dB)
(a) Hybrid Precision (B-9, FP-9) 80 70
60
66 64
55
62
50
60 7.5
45
60
78.0 77.5
50
0.0
77.0
40
10.0
12.5
15.0
17.5
Hybrid w/o Calibration 2.5
5.0
7.5
10.0
12.5
20.0
Hybrid
15.0
17.5
20.0
SNRtest (dB)
76.5
(b) SNRtrain = 10dB
76.0
30
75.5
60
75.0
20
5.0
10 0 0.0
2.5
5.0
7.5
7.5
10.0
12.5
15.0
17.5
20.0
SNRtrain = 5dB
SNRtrain = 15dB
SNRtrain = 10dB
SNRtrain = 20dB
10.0
12.5
15.0
17.5
20.0
SNRtest (dB)
Accuracy (%)
Accuracy (%)
10.0
SNRtest (dB)
20.0
10
0
7.5
Hybrid w/o Calibration
52
63.5
20
60
(b) Hybrid Precision (B-9, FP-79)
64
40
62
30
60
20
8
10
Fig. 9. Performance of the hybrid precision framework under varying testing SNR.
10
12
14
16
18
Hybrid w/o Calibration 0.0
2.5
5.0
7.5
10.0
12.5
20
Hybrid
15.0
17.5
16
18
20.0
SNRtest (dB)
(c) SNRtrain = 15dB 60
Accuracy (%)
the FP-9 and FP-79 configurations. While both frameworks follow identical robustness trends, the hybrid (B-9, FP-79) system provides a substantially higher performance ceiling, reaching a peak accuracy of approximately 78% compared to the 66% achieved by the FP-9 variant. Notably, the larger server model also demonstrates superior baseline robustness; when trained at 5 dB, it maintains over 72% accuracy at 0 dB SNRtest . This indicates that a more powerful server-side backbone not only improves absolute accuracy but also possesses a stronger capacity to reconstruct information from noisy edgetransmitted features. Consequently, for deployment in highly volatile wireless environments, a hybrid configuration utilizing a low SNRtrain and a high-capacity server backbone offers the most reliable balance of efficiency and all-terrain performance. To evaluate the subcarrier-based feature calibration module, we conducted an ablation study comparing the standard hybrid precision framework against a version without calibration named Hybrid w/o Calibration. As illustrated in Fig. 10, the calibration module consistently enhances classification accuracy and robustness within reasonable operational regions by effectively mitigating binary feature distortion caused by multipath fading. This specialization introduces an unavoidable trade-off resulting in increased sensitivity and sharper performance declines when testing SNRs fall significantly below training thresholds. By enabling the server-side FPNN to recover high-fidelity information from noise-corrupted streams, the calibration mechanism ensures superior accuracy across the most critical operating environments. To achieve full adaptivity in various channel conditions, recent techniques
66
50
50
66
40
65
30
64
20
63 10
10
12
14
Hybrid w/o Calibration
0 0.0
2.5
5.0
7.5
10.0
12.5
15.0
20
Hybrid 17.5
20.0
SNRtest (dB)
(d) SNRtrain = 20dB Fig. 10. Performance comparison on subcarrier-based feature calibration under varying testing SNR
such as attention modules [48] and Hypernetworks [49] can be integrated into the proposed hybrid-precision framework. VII. C ONCLUSION This paper investigated the design of a hybrid-precision task-oriented communication that integrates efficient architectural binarization with edge co-inference. By deploying a binarized front-end on resource-constrained edge devices and a full-precision back-end on the edge server, the proposed framework successfully addresses the limitations of traditional model splitting and full BNN implementations. The inclusion of subcarrier-based calibration and binary feature recovery ensures robust performance over challenging multipath fading wireless environments. Combined with a KD-based optimization strategy, this hybrid approach effectively bridges the utility gap caused by binarization while maintaining minimal
13
on-device computational overhead. Experimental results on the large-scale ImageNet dataset validate the effectiveness of the proposed system. Future research will extend this hybrid-precision paradigm to multi-user collaborative scenarios, exploring distributed feature extraction and subcarrier allocation for large-scale edge networks. Furthermore, we aim to adapt this framework to multi-modal large language models (LLMs) for maintaining cross-modal semantic alignment while significantly reducing the memory and communication overhead of generative AI tasks at the wireless edge. R EFERENCES [1] K. B. Letaief, W. Chen, Y. Shi, J. Zhang, and Y.-J. A. Zhang, “The roadmap to 6G: AI empowered wireless networks,” IEEE Commun. Mag., vol. 57, no. 8, pp. 84–90, 2019. [2] C.-X. Wang, X. You, X. Gao, X. Zhu, Z. Li, C. Zhang, H. Wang, Y. Huang, Y. Chen, H. Haas et al., “On the road to 6G: Visions, requirements, key technologies, and testbeds,” IEEE Commun. Surveys Tuts., vol. 25, no. 2, pp. 905–974, 2023. [3] S. Xie, H. Li, Z. Wang, S. Song, J. Zhang, and K. B. Letaief, “Towards reasoning-empowered task-oriented communication for agent networks,” npj Wireless Technol., vol. 2, no. 1, p. 25, 2026. [4] E. Li, L. Zeng, Z. Zhou, and X. Chen, “Edge AI: On-demand accelerating deep neural network inference via edge computing,” IEEE Trans. Wireless Commun., vol. 19, no. 1, pp. 447–457, Jan. 2019. [5] G. Zhu, D. Liu, Y. Du, C. You, J. Zhang, and K. Huang, “Toward an intelligent edge: Wireless communication meets machine learning,” IEEE Commun. Mag., vol. 58, no. 1, pp. 19–25, 2020. [6] Y. Mao, X. Yu, K. Huang, Y.-J. A. Zhang, and J. Zhang, “Green edge ai: A contemporary survey,” Proc. IEEE, 2024. [7] W. Xu, Z. Yang, D. W. K. Ng, M. Levorato, Y. C. Eldar, and M. Debbah, “Edge learning for b5g networks with distributed signal processing: Semantic communication, edge computing, and wireless sensing,” IEEE J. Sel. Topics Signal Process., vol. 17, no. 1, pp. 9–39, 2023. [8] K. B. Letaief, Y. Shi, J. Lu, and J. Lu, “Edge artificial intelligence for 6G: Vision, enabling technologies, and applications,” IEEE J. Sel. Areas Commun., vol. 40, no. 1, pp. 5–36, Jan. 2021. [9] Z. Zhou, X. Chen, E. Li, L. Zeng, K. Luo, and J. Zhang, “Edge intelligence: Paving the last mile of artificial intelligence with edge computing,” Proc. IEEE, vol. 107, no. 8, pp. 1738–1762, Aug. 2019. [10] J. Chen and X. Ran, “Deep learning with edge computing: A review,” Proc. IEEE, vol. 107, no. 8, pp. 1655–1674, 2019. [11] A. E. Eshratifar, M. S. Abrishami, and M. Pedram, “Jointdnn: An efficient training and inference engine for intelligent mobile cloud computing services,” IEEE Trans. Mobile Comput., vol. 20, no. 2, pp. 565–576, 2019. [12] J. Shao and J. Zhang, “Communication-computation trade-off in resource-constrained edge inference,” IEEE Trans. Commun., vol. 58, no. 12, pp. 20–26, 2020. [13] J. Shao, Y. Mao, and J. Zhang, “Learning task-oriented communication for edge inference: An information bottleneck approach,” IEEE J. Sel. Areas Commun., vol. 40, no. 1, pp. 197–211, Jan. 2021. [14] D. Gündüz, Z. Qin, I. E. Aguerri, H. S. Dhillon, Z. Yang, A. Yener, K. K. Wong, and C.-B. Chae, “Beyond transmitting bits: Context, semantics, and task-oriented communications,” IEEE J. Sel. Areas Commun., vol. 41, no. 1, pp. 5–41, Nov. 2022. [15] H. Li, S. Xie, J. Shao, Z. Wang, S. Song, H. He, J. Zhang, and K. B. Letaief, “Mutual information-empowered task-oriented communication: Principles, applications and challenges,” IEEE Commun. Mag., Jan. 2026. [16] J. Shao and J. Zhang, “Bottlenet++: An end-to-end approach for feature compression in device-edge co-inference systems,” in Proc. IEEE Int. Conf. Commun. Workshops, Jun. 2020, pp. 1–6. [17] S. Xie, S. Ma, M. Ding, Y. Shi, M. Tang, and Y. Wu, “Robust information bottleneck for task-oriented communication with digital modulation,” IEEE J. Sel. Areas Commun., vol. 41, no. 8, pp. 2577–2591, Aug. 2023. [18] Y. Cheng, D. Wang, P. Zhou, and T. Zhang, “Model compression and acceleration for deep neural networks: The principles, progress, and challenges,” IEEE Signal Process. Mag., vol. 35, no. 1, pp. 126–136, 2018.
[19] E. Bourtsoulatze, D. B. Kurka, and D. Gündüz, “Deep joint sourcechannel coding for wireless image transmission,” IEEE Trans. Cogn. Commun. Netw., vol. 5, no. 3, pp. 567–579, 2019. [20] D. Gündüz, M. A. Wigger, T.-Y. Tung, P. Zhang, and Y. Xiao, “Joint source–channel coding: Fundamentals and recent progress in practical designs,” Proc. IEEE, 2024. [21] Q. Hu, G. Zhang, Z. Qin, Y. Cai, G. Yu, and G. Y. Li, “Robust semantic communications with masked vq-vae enabled codebook,” IEEE Trans. Wireless Commun., vol. 22, no. 12, pp. 8707–8722, 2023. [22] J. Park, Y. Oh, S. Kim, and Y.-S. Jeon, “Joint source-channel coding for channel-adaptive digital semantic communications,” IEEE Trans. Cogn. Commun. Netw., vol. 11, no. 1, pp. 75–89, 2024. [23] M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi, “Xnor-net: Imagenet classification using binary convolutional neural networks,” in Proc. Eur. Conf. Comput. Vis. Springer, 2016, pp. 525–542. [24] Z. Liu, B. Wu, W. Luo, X. Yang, W. Liu, and K.-T. Cheng, “Bi-real net: Enhancing the performance of 1-bit cnns with improved representational capability and advanced training algorithm,” in Proc. Eur. Conf. Comput. Vis. Springer, 2018, pp. 722–737. [25] Z. Liu, Z. Shen, M. Savvides, and K.-T. Cheng, “Reactnet: Towards precise binary neural network with generalized activation functions,” in Proc. Eur. Conf. Comput. Vis. Springer, 2020, pp. 143–159. [26] I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio, “Binarized neural networks,” Proc. Adv. Neural Inf. Process. Syst., vol. 29, 2016. [27] X. Lin, C. Zhao, and W. Pan, “Towards accurate binary convolutional neural network,” Proc. Adv. Neural Inf. Process. Syst., vol. 30, 2017. [28] H. Qin, R. Gong, X. Liu, M. Shen, Z. Wei, F. Yu, and J. Song, “Forward and backward information retention for accurate binary neural networks,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2020, pp. 2250–2259. [29] X. He, Z. Mo, K. Cheng, W. Xu, Q. Hu, P. Wang, Q. Liu, and J. Cheng, “Proxybnn: Learning binarized neural networks via proxy matrices,” in Proc. Eur. Conf. Comput. Vis. Springer, 2020, pp. 223–241. [30] Z. Xu, M. Lin, J. Liu, J. Chen, L. Shao, Y. Gao, Y. Tian, and R. Ji, “Recu: Reviving the dead weights in binary neural networks,” in Proc. IEEE Int. Conf. Comput. Vis., 2021, pp. 5198–5208. [31] Z. Tu, X. Chen, P. Ren, and Y. Wang, “Adabin: Improving binary neural networks with adaptive binary sets,” in Proc. Eur. Conf. Comput. Vis. Springer, 2022, pp. 379–395. [32] M. Yang, C. Bian, and H.-S. Kim, “Ofdm-guided deep joint source channel coding for wireless multipath fading channels,” IEEE Trans. Cogn. Commun. Netw., vol. 8, no. 2, pp. 584–599, 2022. [33] Y. Shi, Y. Zhou, D. Wen, Y. Wu, C. Jiang, and K. B. Letaief, “Taskoriented communications for 6G: Vision, principles, and technologies,” IEEE Wireless Commun., vol. 30, no. 3, pp. 78–85, 2023. [34] J. Shao, Y. Mao, and J. Zhang, “Task-oriented communication for multidevice cooperative edge inference,” IEEE Trans. Wireless Commun., vol. 22, no. 1, pp. 73–87, 2022. [35] C. Cai, X. Yuan, and Y.-J. A. Zhang, “Multi-device task-oriented communication via maximal coding rate reduction,” IEEE Trans. Wireless Commun., vol. 23, no. 12, pp. 18 096–18 110, 2024. [36] H. Li, J. Shao, H. He, S. Song, J. Zhang, and K. B. Letaief, “Tackling distribution shifts in task-oriented communication with information bottleneck,” arXiv preprint arXiv:2405.09514, 2024. [37] S. Xie, H. He, S. Song, J. Zhang, and K. B. Letaief, “Toward realtime edge ai: Model-agnostic task-oriented communication with visual feature alignment,” IEEE J. Sel. Areas Commun., vol. 43, no. 12, pp. 4262–4276, 2025. [38] H. Xie, Z. Qin, X. Tao, and K. B. Letaief, “Task-oriented multi-user semantic communications,” IEEE J. Sel. Areas Commun., vol. 40, no. 9, pp. 2584–2597, 2022. [39] Z. Ding, Y. Chen, N. Li, D. Zhao, Z. Sun, and C. P. Chen, “Bnas: Efficient neural architecture search using broad scalable architecture,” IEEE Trans. Neural Netw. Learn. Syst., vol. 33, no. 9, pp. 5004–5018, 2021. [40] H. Qin, X. Zhang, R. Gong, Y. Ding, Y. Xu, and X. Liu, “Distributionsensitive information retention for accurate binary neural network,” Int. J. Comput. Vis., vol. 131, no. 1, pp. 26–47, 2023. [41] O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al., “Imagenet large scale visual recognition challenge,” Int. J. Comput. Vis., vol. 115, no. 3, pp. 211–252, Dec. 2015. [42] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2016, pp. 770–778.
14
[43] B. Zhu, Z. Al-Ars, and H. P. Hofstee, “Nasb: Neural architecture search for binary convolutional neural networks,” in Proc. Int. Joint Conf. Neural Netw. IEEE, 2020, pp. 1–8. [44] P. Wang, X. He, G. Li, T. Zhao, and J. Cheng, “Sparsity-inducing binarized neural networks,” in Proc. AAAI Conf. Artif. Intell., vol. 34, no. 07, 2020, pp. 12 192–12 199. [45] H. Qiu, H. Ma, Z. Zhang, Y. Gao, Y. Zheng, A. Fu, P. Zhou, D. Abbott, and S. F. Al-Sarawi, “Rbnn: Memory-efficient reconfigurable deep binary neural network with ip protection for internet of things,” IEEE Trans. Comput.-Aided Des. Integr. Circuits Syst., vol. 42, no. 4, pp. 1185–1198, 2022. [46] Y. Li, S.-L. Pintea, and J. C. Van Gemert, “Equal bits: Enforcing equally distributed binary network weights,” in Proc. AAAI Conf. Artif. Intell., vol. 36, no. 2, 2022, pp. 1491–1499. [47] M. Lin, R. Ji, Z. Xu, B. Zhang, F. Chao, C.-W. Lin, and L. Shao, “Siman: Sign-to-magnitude network binarization,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 5, pp. 6277–6288, 2022. [48] J. Xu, B. Ai, W. Chen, A. Yang, P. Sun, and M. Rodrigues, “Wireless image transmission using deep source channel coding with attention modules,” IEEE Trans. Circuits Syst. Video Technol., vol. 32, no. 4, pp. 2315–2328, 2021. [49] S. Xie, H. He, H. Li, S. Song, J. Zhang, Y.-J. A. Zhang, and K. B. Letaief, “Deep learning-based adaptive joint source-channel coding using hypernetworks,” in Proc. IEEE Int. Mediterranean Conf. Commun. Netw., 2024, pp. 191–196.