1
Proportional-Fair Resource Allocation and Dual-Threshold Early-Exit Inference for Secure Cooperative Multi-Layer Edge Intelligence
arXiv:2609.15847v1 [cs.NI] 14 Sep 2026
Thai T. Vu, John Le, Tu N. Nguyen, Jun Shen, Quang Vinh Duong, Ha Nguyen
Abstract—This paper proposes FREDI (Fair Resource Allocation for Edge Dual-Threshold Inference), a secure wireless edgeintelligence framework for event-triggered inference in a cooperative user equipment (UE)–edge server (ES)–cloud system. Each UE performs early-exit convolutional neural network (CNN) screening using dual confidence thresholds, while critical events are securely offloaded to an edge server for detailed classification. We formulate a proportionally-fair utility maximization problem that jointly optimizes UE–ES association, wireless and processing resources, and confidence thresholds. FREDI decomposes the problem into proportional-fair resource allocation and dualthreshold inference optimization. We prove that the detectedcritical event set is set-monotone non-increasing in both thresholds, and exploit the finite empirical confidence domain for exact threshold optimization. An empirical resource–utility response envelope yields a computable global suboptimality bound and a sufficient condition for global optimality. By pre-eliminating infeasible UE–ES pairs and exactly projecting out bandwidth and transmit-power variables, the resource-allocation subproblem is reduced to a mixed-integer exponential-cone program solvable to the certified global optimality within a prescribed gap. Numerical results with early-exit MobileNetV2 and ShuffleNetV2 demonstrate near-perfect UE fairness with aggregate utility close to a Sum-Utility benchmark, reveal security-induced resource fragmentation, and demonstrate the Stage-A scalability from 6 to 144 UEs with median solving time below 0.1 s in the tested configurations. Index Terms—Wireless edge intelligence, fair resource allocation, dual-threshold inference, early-exit CNN, secure offloading, mixed-integer conic optimization.
I. I NTRODUCTION The growing demand for latency-sensitive intelligent services has accelerated the implementation of wireless edge intelligence, in which communication and nearby computing resources are jointly orchestrated for distributed inferences [1], [2]. In a three-layer deployment, user equipment (UEs) observes events close to their sources, edge servers (ESs) provide more capable inference, and a cloud server (CS) periodically updates the models. This hierarchy reduces reliance on remotecloud inference without requiring resource-constrained UEs to execute the complete learning pipeline. Thai T. Vu, John Le, Jun Shen, and Quang Vinh Duong are with the School of Computing and Information Technology, University of Wollongong, Wollongong, NSW, Australia (e-mail: [email protected]; [email protected]; [email protected]; [email protected]). Tu N. Nguyen is with the Department of Computer Science, Kennesaw State University, Georgia, USA (e-mail: [email protected]). Ha Nguyen is a scientist at CodeZX Software Company Limited, Vietnam (e-mail: [email protected]). This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice.
Early-exit neural networks (EENNs) are well-established for adaptive inference, where intermediate classifiers allow sufficiently confident samples to terminate before the final layer, thereby reducing computation [3]–[5]. In particular, confidence-threshold exit policies are common in literature. More generally, two- or multiple-threshold decision rules that partition a confidence domain into decision and rejection/undecided regions have been widely used well before the emergence of edge AI [6]–[8]. In this paper, we study how early-exit convolutional neural networks (CNNs) and dualthreshold decision rules would interact with shared wireless and processing resources in a secure and cooperative multilayer edge system. Recent studies have increasingly coupled deep neural network (DNN) inference with edge-resource management. Liu et al. [9] optimize multi-user communication and computation with batching and early exits in a single-ES system, while Kim and Lee [10] jointly optimize edge resources and DNN splitting in an edge–cloud setting. Zheng et al. [11] further consider multi-terminal, multi-base station (BS) semantic transmission with communication/computation allocation. Liu et al. [12] jointly optimize model partitioning, exit selection, association, bandwidth, and computation in multi-server mobile edge computing (MEC). Zhang et al. [13] and Yuan et al. [14] further couple service placement and model splitting with resource allocation. These studies demonstrate the value of co-designing DNN inference and edge-resource allocation, but none jointly considers proportional fairness, securityconstrained association, and per-UE dual-threshold screening. Most closely related, Zhou et al. [15] use dual-threshold multi-exit inference for event-triggered offloading between a single device and an edge server. Our FREDI (Fair Resource Alloca- tion for Edge Dual-Threshold Inference) framework differs primarily at the network level: multiple UEs compete for communication and processing resources across multiple ESs, coupling threshold selection with UE–ES association, bandwidth, transmit power, ES capacity, security eligibility, and proportional fairness. This turns device-level inference and offloading control into network-wide resource orchestration with combinatorial association. Fairness and security have largely been studied separately. Xu et al. [16] use proportional fairness for distributed DNNinference assignment and load balancing, while [17] studies energy-based proportional fairness for task offloading and resource allocation in cooperative edge computing, without considering early-exit inference. Security-aware offloading
2
couples execution decisions with protection requirements and resource use [18], [19]. Hybrid quantum–classical models have also been explored for classification and distributed learning [20]–[23]. In FREDI, lightweight binary early-exit screening is performed at UEs, whereas heavier hybrid CNNQNN (quantum neural network) inference [23] is performed at ESs. Table I compares FREDI with representative studies. A checkmark denotes explicit treatment, ◦ for partial or related treatment, and “–” indicating an absence as a principal component. The comparison highlights differences in system scope and analytical treatment rather than a comparison of the novelty in standard early-exit or threshold mechanisms. To address these gaps, we develop FREDI (Fair Resource Allocation for Edge Dual-Threshold Inference). Its novelty lies not in early-exit or dual-threshold inference alone, but in their network-scale formulation, structural analysis, and provably controlled optimization within a secure cooperative multi-layer edge-AI system. The main contributions are: • Network-scale cooperative formulation: We formulate a proportional-fair multi-UE, multi-ES problem that jointly optimizes per-UE dual thresholds, security-constrained UE–ES association, wireless resources, and shared ES processing capacity, thereby capturing competition for heterogeneous communication and computing resources. • Structural analysis and optimization guarantees: We prove that the detected-critical event set is set-monotone non-increasing in both thresholds, enabling exact finitedomain threshold optimization with directional pruning. We further eliminate infeasible UE–ES pairs and exactly project out bandwidth and power variables, yielding a mixed-integer exponential-cone resource-allocation problem with a certified optimality gap. A computable resource–utility response envelope bounds the end-to-end decomposition loss and gives a sufficient condition for global optimality. • Fairness, security, and scalability evaluation: Experiments with early-exit MobileNetV2 and ShuffleNetV2 quantify the dual- versus coupled single-parameter threshold trade-off, fairness–utility trade-off, and securityinduced resource fragmentation, while testing scalability from 𝑁 = 6 to 𝑁 = 144 UEs. FREDI achieves nearuniform UE utility with practical computation over the tested network sizes. The rest of this paper is organized as follows. Section II presents the system model, performance metrics, and joint optimization problem. Section III details FREDI, including its two-stage decomposition and optimality analysis. Section IV reports experimental results on thresholding, fairness, security constraints, and scalability. Section V concludes the paper. II. S YSTEM M ODEL AND P ROBLEM F ORMULATION A. System Model We consider a cooperative three-layer Edge AI system comprising 𝑁 local user devices (i.e., surveillance cameras), referred to as UEs and denoted by N = {1, 2, . . . , 𝑁 }; 𝑀 edge servers (ESs), denoted by M = {1, 2, . . . , 𝑀 }; and one cloud
Fig. 1: Three-layer Edge-AI system entities (left) and cooperative inference and periodic model-update workflow (right).
server (CS). The three-layer system entities are shown on the left of Fig. 1, while the right side summarizes the cooperative inference and periodic model-update workflow. During each scheduling period, UE 𝑛 captures a sequence of independent events (or image), denoted by Φ𝑛 = {𝐼𝑛1 , 𝐼 𝑛2 , . . . , 𝐼𝑛|Φ𝑛 | }, where 𝐼𝑛𝑘 denotes the 𝑘th event and |Φ𝑛 | is the number of events. Each UE 𝑛 ∈ N employs a lightweight CNN with dual-threshold early exits to classify locally observed events as normal or critical. Events detected as normal remain at the UE. For an event detected as critical, the UE securely offloads its extracted feature representation to one associated ES. The selected ES executes a hybrid CNN–QNN model for detailed multi-class classification and returns the resulting label or response action. The CS collects training information or model updates from the ESs, updates the classical CNN, featureadapter, and QNN parameters, and distributes the updated hybrid models back to the ESs. The model parameters are treated as fixed during each scheduling and inference period considered by the optimization. Let S = {1, . . . , 𝑆} denote the set of security levels assigned to UEs, ESs, and the CS, where a larger index represents a stronger security level. Each UE 𝑛 serves one application with a required security level 𝑠UE ∈ S; consequently, all events 𝑛 generated by that UE share the same security requirement. Their critical-event features can therefore be offloaded only to UE an ES 𝑗 satisfying 𝑠ES 𝑗 ≥ 𝑠 𝑛 . We assume that the CS satisfies the required security and privacy protections. 1) Early-Exit Inference at UEs: The UE-side CNN consists of convolutional feature extractors, pooling or subsampling layers, nonlinear activations, and lightweight classification heads [24]. Because events vary in classification difficulty, we adopt early-exit inference [3]–[5] with a two-threshold confidence rule [6], [7]. At each exit, an event is classified as critical if its confidence reaches the upper threshold, as normal if it reaches the lower threshold, and otherwise proceeds to a deeper exit. The UE-specific thresholds are denoted by (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ), with 0 ≤ 𝛼𝑛𝑙 ≤ 𝛼𝑛𝑢 ≤ 1. Fig. 2 illustrates this policy. Our focus is the coupling of threshold-controlled inference loads with secure multi-UE/multi-ES resource orchestration, rather than the threshold rule itself. Confidence Score. Let 𝐿 be the number of layers of the lightweight CNN at UE 𝑛. Given the extracted feature of event 𝐼𝑛𝑘 ∈ Φ𝑛 as input, the CNN generates a confidence score at each layer to assess whether the event is critical. The (𝑘 ) confidence score 𝐶𝑛𝑞 at layer 𝑞 for event 𝐼𝑛𝑘 is defined as
3
TABLE I: Comparison with representative edge-inference and resource-allocation studies. Columns denote: confidence-based early exiting; two independently chosen thresholds; more than one device; more than one server; joint optimization of communication and server-side compute; an explicit proportional-fair objective; association restricted by a security or trust level; a proved structural property of the threshold rule; and a reported runtime study over increasing network size. Entries record the aspects each work treats as a principal component; a dash does not imply the aspect is unattainable within that framework. Early Security Joint network Dual MultiMultiThreshold Network exit threshold UE/device ES/server resource alloc. PF constraints structure/proof scalability
Study
✓ – ✓ ✓ ✓ ✓ ✓
Liu et al., JSAC’23 [9] Xu et al., IoTJ’23 [16] Kim and Lee, IoTJ’24 [10] Zheng et al., TWC’25 [11] Zhou et al., TCOM’26 [15] Liu et al., CJE’26 [12] FREDI (this work)
✓ ◦ ◦ ✓ – ✓ ✓
– – – – ✓ – ✓
– ✓ ◦ ✓ – ✓ ✓
(𝑘) ,critical
𝑒 𝑓𝑛𝑞 (𝑘) ,critical
𝑒 𝑓𝑛𝑞
(𝑘) ,normal
+ 𝑒 𝑓𝑛𝑞
,
(1)
(𝑘 ),critical (𝑘 ),normal where 𝑓𝑛𝑞 and 𝑓𝑛𝑞 are the critical- and normalclass logits, respectively, at layer 𝑞 for event 𝐼𝑛𝑘 . Early-Exit Event Detection. The CNN classifies each event as critical (1) or normal (0) as early as possible. For 𝛼𝑛𝑙 < 𝛼𝑛𝑢 , processing stops at the first exit 𝑞 whose confidence leaves the undecided interval I𝑛 ≜ (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ): event 𝐼𝑛𝑘 is declared (𝑘 ) (𝑘 ) normal if 𝐶𝑛𝑞 ≤ 𝛼𝑛𝑙 and critical if 𝐶𝑛𝑞 ≥ 𝛼𝑛𝑢 . If all exits remain within the interval, the event is conservatively declared normal at the final exit. When 𝛼𝑛𝑙 = 𝛼𝑛𝑢 , the policy reduces to a single decision at the first exit, with equality assigned to the normal class. Accordingly, the predicted label 𝑦ˆ 𝑛𝑘 ∈ {1, 0} is (𝑘 ) (𝑘 ) 0, ∃ 𝑞 ≤ 𝐿 : 𝐶𝑛𝑞 ≤ 𝛼𝑛𝑙 , 𝐶𝑛𝑡 ∈ I𝑛 , ∀𝑡 < 𝑞, (𝑘 ) (𝑘 ) 𝑢 𝑙 1, ∃ 𝑞 ≤ 𝐿 : 𝐶𝑛𝑞 ≥ 𝛼𝑛 > 𝛼𝑛 , 𝐶𝑛𝑡 ∈ I𝑛 , ∀𝑡 < 𝑞, 𝑦ˆ 𝑛𝑘 = (𝑘 ) 0, n𝐶𝑛𝑡 ∈ I𝑛 ,o ∀𝑡 ≤ 𝐿, 1 𝐶 (𝑘 ) > 𝛼𝑙 , 𝛼𝑙 = 𝛼𝑢 . 𝑛 𝑛 𝑛 𝑛1
– ✓ – – – – ✓
– – – – – – ✓
– – – – ◦ – ✓
– – – – – – ✓
and transmit power allocated from UE 𝑛 to ES 𝑗, with perlink limits 𝑏 max and 𝑝 max . Each UE is associated with one ES, represented by x = {𝑥 𝑛 𝑗 } ∈ {0, 1} 𝑁 × 𝑀 , where 𝑥 𝑛 𝑗 = 1 denotes association between UE 𝑛 and ES 𝑗. Let 𝐷 𝑛 denote the representative size, in bits, of the feature data offloaded by UE 𝑛 for one event. The uplink rate from UE 𝑛 to ES 𝑗 is ! 𝑔𝑛 𝑗 𝑝 𝑛 𝑗 , (3) 𝑟 𝑛 𝑗 (𝑏 𝑛 𝑗 , 𝑝 𝑛 𝑗 ) = 𝑏 𝑛 𝑗 log2 1 + 2 𝜎𝑛 𝑗 𝑏 𝑛 𝑗
Fig. 2: CNN with dual-threshold early exits.
(𝑘 ) 𝐶𝑛𝑞 =
✓ ◦ ✓ ✓ ◦ ✓ ✓
(2) where 1{·} is the indicator function. The four cases are mutually exclusive: when 𝛼𝑛𝑙 < 𝛼𝑛𝑢 the last case is inactive, and when 𝛼𝑛𝑙 = 𝛼𝑛𝑢 the interval I𝑛 is empty, so only the first exit is examined and confidence exactly equal to the common threshold is assigned to the normal class. 2) Secure Communication Between UEs and ESs: We employ FDMA for UE–ES uplink communication. Let 𝒃 = {𝑏 𝑛 𝑗 } ∈ R 𝑁 × 𝑀 and 𝒑 = {𝑝 𝑛 𝑗 } ∈ R 𝑁 × 𝑀 denote the bandwidth
where 𝑔𝑛 𝑗 and 𝜎𝑛2 𝑗 are the channel gain and noise power spectral density of the legitimate link, respectively. To model physical-layer confidentiality, we consider an eavesdropper that attempts to intercept transmissions from UE 𝑛. Its achievable rate on the bandwidth allocated to link (𝑛, 𝑗) is ! 𝑔𝑛EV 𝑝 𝑛 𝑗 EV 𝑟 𝑛 𝑗 (𝑏 𝑛 𝑗 , 𝑝 𝑛 𝑗 ) = 𝑏 𝑛 𝑗 log2 1 + EV 2 , (4) (𝜎𝑛 ) 𝑏 𝑛 𝑗 where 𝑔𝑛EV and (𝜎𝑛EV ) 2 denote the corresponding eavesdropper-channel gain and noise power spectral density. The secure transmission h rate [25] is i 𝑟 𝑛se𝑗 (𝑏 𝑛 𝑗 , 𝑝 𝑛 𝑗 ) = 𝑟 𝑛 𝑗 (𝑏 𝑛 𝑗 , 𝑝 𝑛 𝑗 ) − 𝑟 𝑛EV𝑗 (𝑏 𝑛 𝑗 , 𝑝 𝑛 𝑗 )
+
,
(5)
where [𝑧] + = max{0, 𝑧}. Consistent with the FDMA model, transmissions are orthogonal in frequency, so (3) and (4) contain no inter-user or inter-cell interference term; the eavesdropper channel statistics 𝑔𝑛EV and (𝜎𝑛EV ) 2 are assumed known at the scheduler, which is the standard worst-case assumption in secrecy-rate resource allocation and yields a conservative feasible set. The selected secure rate of UE 𝑛 is 𝑀 Õ 𝑟 𝑛se (𝒃 𝑛 , 𝒑 𝑛 , 𝒙 𝑛 ) = 𝑥 𝑛 𝑗 𝑟 𝑛se𝑗 (𝑏 𝑛 𝑗 , 𝑝 𝑛 𝑗 ). (6) 𝑗=1
For a positive selected secure rate, the secure uplink transmission delay is 𝐷𝑛 𝑡 𝑛tx (𝒃 𝑛 , 𝒑 𝑛 , 𝒙 𝑛 ) = se . (7) 𝑟 𝑛 (𝒃 𝑛 , 𝒑 𝑛 , 𝒙 𝑛 ) Let 𝑡 𝑛r denote the maximum allowable uplink transmission delay of UE 𝑛. The real-time latency model considers only the secure uplink transmission delay from UEs to ESs. The execution and queueing delays associated with hybrid-model inference
4
at the ESs, as well as the comparatively slower CS modelupdate delay, are beyond the scope of this work. 3) Hybrid CNN–QNN Multi-Class Inference at ESs: ES 𝑗 is characterized by the tuple (𝐵 𝑗 , 𝑊 𝑗 , 𝑠ES 𝑗 ), where 𝐵 𝑗 is its available UE-uplink bandwidth, 𝑊 𝑗 is its aggregate eventprocessing capacity over the considered scheduling period, and 𝑠ES 𝑗 ∈ S is its security level. Each ES hosts a hybrid CNN– QNN classifier that processes the received representation and produces the final multi-class prediction. We measure 𝑊 𝑗 in event-processing units, with one unit corresponding to the workload required to process one offloaded event through the complete hybrid inference pipeline, assuming approximately homogeneous per-event workloads. If UE 𝑛 is associated with ES 𝑗, the system allocates uplink bandwidth 𝑏 𝑛 𝑗 and reserves event-processing capacity 𝑤 𝑛 𝑗 at ES 𝑗. Let 𝑤 max denote the maximum processing capacity that an ES can reserve for one UE during a scheduling period. When an event is detected as critical at UE 𝑛, its feature representation is offloaded to the selected ES for multi-class classification and the corresponding response. Rather than explicitly modeling ES execution time, processing feasibility is captured by the aggregate capacity 𝑊 𝑗 and per-UE allocation 𝑤 𝑛 𝑗 ; the optimization is therefore agnostic to the ES classifier. We adopt the hybrid CNN–QNN instantiation as a forwardlooking design point for when variational quantum classifiers become practical, and it is not evaluated in this paper. 4) Periodic Model Aggregation and Updating at the CS: Periodically, the ESs transmit approved training information, model parameters, or feature summaries to the CS over secure backhaul links, as illustrated by the upper model-update path on the right of Fig. 1. The CS aggregates this information, updates the hybrid CNN–QNN model, and redistributes the updated model to the ESs for subsequent deployment. Cloudtraining delay and backhaul-resource allocation are outside the optimization scope.
B. Performance Metrics 1) Categories of Input Events and Output Results: For each event 𝐼𝑛𝑘 ∈ Φ𝑛 , UE 𝑛 uses its lightweight CNN to classify the event as critical or normal. Let 𝑦 𝑛𝑘 and 𝑦ˆ 𝑛𝑘 denote the ground-truth and predicted labels of event 𝐼𝑛𝑘 , respectively. The event is correctly classified if 𝑦ˆ 𝑛𝑘 = 𝑦 𝑛𝑘 ; the four possible outcomes are the usual true/false positive/negative combinations, summarized as event sets in Table II. Over a given time period, UE 𝑛 captures the event set P Φ𝑛 = {𝐼𝑛1 , . . . , 𝐼𝑛|Φ𝑛 | }. Let ΦN 𝑛 and Φ𝑛 denote the sets of normal and critical events at UE 𝑛, respectively. Let Φ̂TP 𝑛 , TN , and Φ̂FN denote the sets of true-positive (TP), Φ̂FP , Φ̂ 𝑛 𝑛 𝑛 false-positive (FP), true-negative (TN), and false-negative (FN) events, respectively, as summarized in Table II. During the considered scheduling period, every event predicted as critical is offloaded to an ES. We therefore define the offloaded load of UE 𝑛 as FP 𝐿 𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ) ≜ | Φ̂TP (8) 𝑛 | + | Φ̂𝑛 |.
TABLE II: Categories of events at UE 𝑛. Set
Definition
Type
Φ𝑛 ΦN 𝑛 ΦP𝑛 Φ̂TP 𝑛 Φ̂FP 𝑛 Φ̂TN 𝑛 FN Φ̂𝑛
All events at UE 𝑛 Normal events at UE 𝑛 Critical events at UE 𝑛 Correctly detected critical events Normal events detected as critical Correctly detected normal events Critical events detected as normal
Input Input Input Output Output Output Output
Using the association and processing-allocation variables, define the event-processing capacity allocated to UE 𝑛 as 𝑀 Õ 𝑅𝑛 ≜ 𝑥𝑛 𝑗 𝑤 𝑛 𝑗 . (9) 𝑗=1
Because 𝑤 𝑛 𝑗 is measured in event-processing units over the same period, the allocated capacity must cover the offloaded load but need not exceed the number of events generated by the UE. These processing QoS requirements are specified in Problem P0 . We then define the performance metrics that are influenced by the use of dual confidence thresholds. 2) User Utility Functions: Because correctly detecting critical events is the primary objective of the system, we define the utility of UE 𝑛, denoted by U𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ), as the proportion of critical events correctly classified as critical. This quantity is the true-positive rate (TPR). From Table II we have | Φ̂TP | U𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ) = 𝑛P . (10) |Φ𝑛 | The communication and edge-processing resources required by UE 𝑛 increase with its offloaded load 𝐿 𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ). We therefore optimize the UE utilities subject to the wireless and processing resources available at the ESs. C. Problem Formulation Events arrive sequentially at the UEs, and only those predicted as critical are offloaded to ESs for hybrid CNN– QNN multi-class classification. Over one scheduling period, the joint design contains two coupled decision groups that motivate FREDI: (i) fair UE–ES association and wireless/edge resource allocation through (x, 𝒃, 𝒑, 𝒘), and (ii) early-exit CNN inference control through the lower and upper confidence thresholds (𝜶𝑙 , 𝜶𝑢 ). The CNN and QNN model parameters distributed by the CS are fixed during this period; periodic cloud training is therefore outside P0 . Maximizing aggregate utility alone may allocate few communication and processing resources to UEs whose utility gain per unit resource is comparatively small. FREDI therefore associates fairness with the resource-allocation layer: we adopt weighted proportional fairness across UEs, while the dual thresholds remain inference-control variables. We assume U𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ) > 0 for every UE 𝑛. Weighted proportional fairness is induced by maximizing the standard network utility Í𝑁 𝜌 ln(U𝑛 ), where 𝜌 𝑛 > 0 is the fairness weight of 𝑛 𝑛=1 UE 𝑛 [26]. The end-to-end problem nevertheless optimizes both groups jointly: the dual confidence thresholds 𝜶𝑙 and 𝜶𝑢 determine
5
early-exit CNN decisions and offloaded load, whereas x, 𝒃, 𝒑, and 𝒘 determine fair secure access to wireless and ES processing resources. The resulting problem is P0 :
maximize
𝜶𝑙 ,𝜶 𝑢 ,x,𝒃,𝒑,𝒘
𝑁 Õ 𝑛=1
𝜌 𝑛 ln U𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 )
(11)
subject to 0 ≤ 𝑏 𝑛 𝑗 ≤ 𝑏 max ,
∀(𝑛, 𝑗) ∈ N × M,
(11a)
0 ≤ 𝑝 𝑛 𝑗 ≤ 𝑝 max ,
∀(𝑛, 𝑗) ∈ N × M,
(11b)
0 ≤ 𝑤 𝑛 𝑗 ≤ 𝑤 max , 𝑁 Õ 𝑥𝑛 𝑗 𝑏𝑛 𝑗 ≤ 𝐵 𝑗 ,
∀(𝑛, 𝑗) ∈ N × M,
(11c)
∀ 𝑗 ∈ M,
(11d)
𝑥𝑛 𝑗 𝑤 𝑛 𝑗 ≤ 𝑊 𝑗 ,
∀ 𝑗 ∈ M,
(11e)
UE 𝑥 𝑛 𝑗 𝑠ES 𝑗 ≥ 𝑠𝑛 ,
∀𝑛 ∈ N,
(11f)
∀𝑛 ∈ N,
(11g)
∀𝑛 ∈ N,
(11h)
∀𝑛 ∈ N,
(11i)
𝑥 𝑛 𝑗 ∈ {0, 1},
∀(𝑛, 𝑗) ∈ N × M,
(11j)
𝐿 𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ) ≤ 𝑅𝑛 , 0 ≤ 𝛼𝑛𝑙 ≤ 𝛼𝑛𝑢 ≤ 1,
∀𝑛 ∈ N,
(11k)
𝑛=1 𝑁 Õ 𝑛=1 𝑀 Õ
𝑗=1 𝑡 𝑛tx 𝒃 𝑛 , 𝒑 𝑛 , 𝒙 𝑛 ≤ 𝑡 𝑛r ,
𝑅𝑛 ≤ |Φ𝑛 |, 𝑀 Õ 𝑥 𝑛 𝑗 = 1, 𝑗=1
∀𝑛 ∈ N. (11l) Constraints (11a)–(11c) specify the per-link bandwidth, transmission-power, and event-processing-reservation limits, respectively. Constraint (11i) associates each UE with exactly one ES. Constraints (11d) and (11e) enforce the aggregate bandwidth and event-processing capacities of each ES. Constraint (11f) ensures that a UE is associated only with an ES whose security level is sufficiently strong, while (11g) bounds its secure uplink delay. Finally, constraints (11k) and (11h) enforce the processing QoS chain 𝐿 𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ) ≤ 𝑅𝑛 ≤ |Φ𝑛 |. The coupling between the resource-allocation variables and the dual-threshold load constraint (11k) is the central structure exploited by FREDI in Section III. III. T HE P ROPOSED FREDI F RAMEWORK In this section, we introduce the proposed method, Fair Resource Allocation for Edge Dual-Threshold Inference (FREDI). Its two stages deliberately assign proportional fairness to multi-UE wireless/edge resource allocation and dualthreshold optimization to early-exit CNN inference. A. Problem Characterization Problem P0 is a nonlinear mixed-integer problem: it contains binary association variables 𝑥 𝑛 𝑗 , empirical and generally non-differentiable utilities U𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ), and products coupling the association decisions to the continuous allocations (𝑏 𝑛 𝑗 , 𝑝 𝑛 𝑗 , 𝑤 𝑛 𝑗 ). FREDI exploits this structure through
two complementary stages. Stage A performs proportionalfair UE–ES association and wireless/edge resource allocation and is transformed exactly into a mixed-integer exponentialcone problem. Stage B then performs dual-threshold inference optimization for each UE over the finite set of observed confidence scores under the processing capacity allocated by Stage A. Hence, fairness is governed by Stage A, whereas inference and offloading decisions are controlled by Stage B. The following analysis distinguishes the exactness of each stage from the end-to-end loss induced by their sequential composition. Lemma 1 (Set-monotonicity of dual-threshold early exiting). For fixed CNN confidence traces, let Φ̂critical (𝛼𝑙 , 𝛼𝑢 ) denote the events classified as critical under thresholds 0 ≤ 𝛼𝑙 ≤ 𝛼𝑢 ≤ 1. For any two admissible threshold pairs satisfying 𝛼𝑙′ ≥ 𝛼𝑙 and 𝛼𝑢′ ≥ 𝛼𝑢 , Φ̂critical (𝛼𝑙′ , 𝛼𝑢′ ) ⊆ Φ̂critical (𝛼𝑙 , 𝛼𝑢 ). (12) critical Consequently, | Φ̂ | is monotonically non-increasing in either threshold over the admissible domain. Proof: The proof is shown in Appendix A. B. Fair Resource Allocation for Edge Dual-Threshold Inference The empirical utility depends directly on the inference thresholds (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ), whereas the wireless and processing allocations determine which threshold pairs are feasible through the Í 𝑀offloaded-load constraint (11k). Recall from (9) that 𝑅𝑛 = 𝑗=1 𝑥 𝑛 𝑗 𝑤 𝑛 𝑗 is the processing capacity allocated to UE 𝑛. To expose the resource–inference coupling without assuming that the reserved capacity is fully utilized, let 𝜉 > 0 denote a generic processing-capacity level and define the optimal Stage-B response of UE 𝑛 as 𝑉𝑛 (𝜉) ≜ max U𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ). (13) 0≤ 𝛼𝑛𝑙 ≤ 𝛼𝑛𝑢 ≤1 𝐿𝑛 ( 𝛼𝑛𝑙 , 𝛼𝑛𝑢 ) ≤ 𝜉
Only capacity levels for which the feasible set contains a positive-utility threshold pair are considered, consistently with the logarithmic domain assumed in P0 . Define the corresponding empirical response factor as 𝑉𝑛 (𝜉) 𝜅 𝑛 (𝜉) ≜ . (14) 𝜉/|Φ𝑛 | Consequently, the optimized end-to-end objective associated with a feasible resource allocation admits the exact decomposition Õ 𝑁 𝑁 𝑁 Õ Õ 𝑅𝑛 𝜌 𝑛 ln 𝜌 𝑛 ln 𝑉𝑛 (𝑅𝑛 ) = 𝜌 𝑛 ln 𝜅 𝑛 (𝑅𝑛 ) . (15) + |Φ𝑛 | 𝑛=1 𝑛=1 𝑛=1 {z } | {z } | 𝐴(R)
𝐾 (R)
FREDI optimizes the processing-coverage term 𝐴(R) in Stage A and then realizes the best empirical inference response 𝑉𝑛 (𝑅𝑛 ) in Stage B. Hence, the only end-to-end approximation introduced by the decomposition is that Stage A does not optimize the response term 𝐾 (R); this term characterizes the departure of the empirical inference response from exact proportionality to the normalized reserved capacity. FREDI Stage A: Fair UE–ES association and resource allocation. The variables (x, 𝒃, 𝒑, 𝒘) are optimized to allocate
6
secure communication resources and event-processing capacity proportionally across UEs. The Stage-A problem is ! Í𝑀 𝑁 Õ 𝑗=1 𝑥 𝑛 𝑗 𝑤 𝑛 𝑗 P 𝐴 : maximize (16) 𝜌 𝑛 ln |Φ𝑛 | x,𝒃,𝒑,𝒘 𝑛=1 subject to constraints (11a)–(11g) and (11h)–(11j). Let (x∗ , 𝒃 ∗ , 𝒑 ∗ , Í 𝒘 ∗ ) denote a globally optimal Stage-A allo∗ ∗ ∗ cation and 𝑅𝑛 = 𝑀 fixed within an 𝑗=1 𝑥 Í𝑛 𝑗 𝑤 𝑛 𝑗 . Since |Φ𝑛 | is Í optimization instance, 𝑛 𝜌 𝑛 ln(𝑅𝑛 /|Φ𝑛 |) = 𝑛 𝜌 𝑛 ln(𝑅𝑛 ) − Í 𝑛 𝜌 𝑛 ln |Φ𝑛 |; therefore, the normalized and unnormalized Stage-A objectives have identical maximizers. The normalization is retained to interpret 𝑅𝑛 /|Φ𝑛 | as the fraction of the UE workload for which processing capacity is reserved. FREDI Stage B: Dual-threshold inference optimization. With the Stage-A allocation fixed, the thresholds are selected by 𝑁 Õ P 𝐵 : maximize 𝜌 𝑛 ln U𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ) (17) 𝜶𝑙 ,𝜶 𝑢
𝑛=1
subject to the threshold-domain constraint (11l) and the perUE capacity constraints ∀𝑛 ∈ N. (18) 𝐿 𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ) ≤ 𝑅𝑛∗ , Problem P 𝐵 is separable across UEs and attains 𝑉𝑛 (𝑅𝑛∗ ) for every UE when each per-UE empirical search is solved exactly. Let 𝐹0★ denote the globally optimal objective value of the Í empirical problem P0 , and define 𝐹0FREDI ≜ 𝑛 𝜌 𝑛 ln 𝑉𝑛 (𝑅𝑛∗ ) as the objective obtained by globally solving P 𝐴 followed by the exact Stage-B searches. T HEOREM 1 (End-to-end global performance bound). Suppose that the Stage-A solution admits a positive-utility Stage-B solution for every UE. Let Q𝑛+ denote the finite set of positiveutility threshold candidates for UE 𝑛, with 𝐿 𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ) > 0, and define the computable response envelope |Φ𝑛 |U𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ) 𝜅𝑛 ≜ max . (19) 𝐿 𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ) ( 𝛼𝑛𝑙 , 𝛼𝑛𝑢 ) ∈ Q 𝑛+ Thus, 𝜅 𝑛 is obtained during the same finite candidate enumeration used to solve Stage B and requires no solution of P0 . Then 𝑁 Õ 𝜅𝑛 0 ≤ 𝐹0★ − 𝐹0FREDI ≤ 𝜌 𝑛 ln . (20) 𝜅 𝑛 (𝑅𝑛∗ ) 𝑛=1 Proof: The proof is shown in Appendix C. Corollary 1.1 (Conditional global optimality). If 𝜅 𝑛 (𝜉) is constant over the positive-utility capacity domain of every UE 𝑛, then FREDI globally solves the empirical problem P0 . Proof. Under the stated condition, 𝜅 𝑛 = 𝜅 𝑛 (𝑅𝑛∗ ) for every UE, so the upper bound in (20) is zero. ■ The factor 𝜅 𝑛 (𝜉) = |Φ𝑛 |𝑉𝑛 (𝜉)/𝜉 measures the optimized inference utility obtained per unit of normalized reserved capacity. It therefore reflects both the predictive yield of the selected threshold pair and any unused capacity caused by the discrete empirical operating points. If the Stage-B load is smaller than 𝑅𝑛∗ , the unused reservation can be released without changing the thresholds or the objective of P0 ; this release does not retroactively change the Stage-A allocation decision.
Fig. 3: FREDI algorithm workflow.
Remark 1 (Solver-certified bound). If the mixed-integer solver b terminates before proving the exact Stage-A optimum, let R b its denote its incumbent processing allocation, 𝐴inc = 𝐴( R) incumbent objective, and 𝐴ub its certified upper bound. Then 𝑁 Õ 𝐹0★ ≤ 𝐴ub + 𝜌 𝑛 ln 𝜅 𝑛 . (21) 𝑛=1
b the resulting feasiAfter exact Stage-B optimization under R, FREDI b ble objective 𝐹0 ( R) satisfies ! 𝑁 Õ 𝜅 𝑛 b ≤ 𝐴ub − 𝐴inc + . (22) 𝐹0★ − 𝐹0FREDI ( R) 𝜌 𝑛 ln | {z } b𝑛 ) 𝜅𝑛 ( 𝑅 𝑛=1 Stage-A solver gap | {z } decomposition bound
The second term is evaluated from the Stage-B candidate sets and the realized FREDI operating points; no iterative exchange between the two stages is required. The FREDI workflow is illustrated in Fig. 3. Stage A allocates secure wireless and ES processing resources proportionally across UEs, whereas Stage B optimizes the dual inference thresholds under the resulting capacities. The 𝑁 threshold-selection problems in Stage B are separable and can be solved in parallel using Algorithm 1. Theorem 1 bounds the end-to-end loss relative to the empirical global optimum, and Corollary 1.1 identifies the proportional-response condition under which the bound vanishes. The detailed solutions of the two stages are presented in Sections III-C and III-D, respectively. C. FREDI Stage A: Fair Resource Allocation FREDI Stage A implements proportional fairness through the UE–ES association and wireless/edge resource-allocation problem P 𝐴. Problem P 𝐴 contains products between binary association variables and continuous resource variables. Moreover, when there is no connection (𝑥 𝑛 𝑗 = 0) between UE 𝑛 and ES 𝑗, the resource allocation variables can be set to 0. Thus, constraints (11a), (11b), and (11c) equivalently become 0 ≤ 𝑏 𝑛 𝑗 ≤ 𝑏 max 𝑥 𝑛 𝑗 , ∀(𝑛, 𝑗) ∈ N × M (23) 0 ≤ 𝑝 𝑛 𝑗 ≤ 𝑝 max 𝑥 𝑛 𝑗 ,
∀(𝑛, 𝑗) ∈ N × M
(24)
0 ≤ 𝑤𝑛 𝑗 ≤ 𝑤
∀(𝑛, 𝑗) ∈ N × M
(25)
max
𝑥𝑛 𝑗 ,
7
Once these constraints are imposed, the products 𝑥 𝑛 𝑗 𝑏 𝑛 𝑗 and 𝑥 𝑛 𝑗 𝑤 𝑛 𝑗 are unnecessary in the aggregate capacity constraints. Thus, constraints (11d), (11e), and (11h) reduce, respectively, to 𝑁 Õ 𝑏𝑛 𝑗 ≤ 𝐵 𝑗 , ∀ 𝑗 ∈ M, (26) 𝑛=1
𝑁 Õ 𝑛=1 𝑀 Õ
𝑤𝑛 𝑗 ≤ 𝑊 𝑗 ,
∀ 𝑗 ∈ M,
(27)
𝑤 𝑛 𝑗 ≤ |Φ𝑛 |,
∀𝑛 ∈ N.
(28)
𝑗=1
For the discrete ( security-level requirement, define UE 1, 𝑠ES 𝑗 ≥ 𝑠𝑛 , 𝑎𝑛 𝑗 = ∀(𝑛, 𝑗) ∈ N × M. (29) 0, otherwise, The original security constraint (11f) can then be replaced by 𝑥𝑛 𝑗 ≤ 𝑎𝑛 𝑗 , ∀(𝑛, 𝑗) ∈ N × M. (30) 1) Secure-Rate and Delay Reformulation: Define the normalized legitimate and eavesdropper channel gains as 𝑔𝑛 𝑗 𝑔𝑛EV 𝛾𝑛 𝑗 ≜ 2 , . (31) 𝛾𝑛EV ≜ 𝜎𝑛 𝑗 (𝜎𝑛EV ) 2 For a link that can provide positive secrecy, we require 𝛾𝑛 𝑗 > 𝛾𝑛EV . The secure rate can then be written without the positivepart operator as 𝛾𝑛 𝑗 𝑝 𝑛 𝑗 se 𝑟 𝑛 𝑗 (𝑏 𝑛 𝑗 , 𝑝 𝑛 𝑗 ) = 𝑏 𝑛 𝑗 log2 1 + 𝑏𝑛 𝑗 ! (32) 𝛾𝑛EV 𝑝 𝑛 𝑗 − 𝑏 𝑛 𝑗 log2 1 + , 𝑏𝑛 𝑗 se with the continuous extension 𝑟 𝑛 𝑗 (0, 𝑝) = 0 for 𝑝 ≥ 0. If 𝑥 𝑛 𝑗 = 1, the communication-delay requirement is 𝐷𝑛 ≤ 𝑡 𝑛r . (33) 𝑟 𝑛se𝑗 (𝑏 𝑛 𝑗 , 𝑝 𝑛 𝑗 ) Using the binary association variable, this condition is equivalently represented for every link by 𝐷𝑛 𝑟 𝑛se𝑗 (𝑏 𝑛 𝑗 , 𝑝 𝑛 𝑗 ) ≥ 𝑟¯𝑛 𝑥 𝑛 𝑗 , (34) 𝑟¯𝑛 ≜ r , 𝑡𝑛 for all (𝑛, 𝑗) ∈ N × M. When 𝑥 𝑛 𝑗 = 0, the linking constraints set 𝑏 𝑛 𝑗 = 𝑝 𝑛 𝑗 = 0, and (34) reduces to 0 ≥ 0. When 𝑥 𝑛 𝑗 = 1, it is equivalent to (33). To characterize the curvature of the secure-rate constraint, define 𝑞 𝑛 𝑗 (𝑧) = log2 (1 + 𝛾𝑛 𝑗 𝑧) − log2 (1 + 𝛾𝑛EV 𝑧). (35) For 𝛾𝑛 𝑗 > 𝛾𝑛EV , 𝑞 𝑛 𝑗 (𝑧) is concave for 𝑧 ≥ 0. This property is used below to establish the mixed-integer convex structure of P 𝐴1 . Because |Φ𝑛 | is constant for every UE within the scheduling Í period, maximizing 𝜌 ln(𝑅 𝑛 𝑛 /|Φ𝑛 |) is equivalent to max𝑛 Í imizing 𝜌 ln(𝑅 ). After applying the linking constraints, 𝑛 𝑛 Í 𝑛 𝑅𝑛 = 𝑗 𝑤 𝑛 𝑗 , so we use the latter form to obtain a simpler conic representation. The resulting resource-allocation problem is 𝑁 𝑀 Õ ©Õ ª P 𝐴1 : maximize 𝜌 𝑛 ln 𝑤𝑛 𝑗 ® (36) x,𝒃,𝒑,𝒘 𝑛=1 𝑗=1 « ¬ subject to (23)–(25), (26)–(28), (30), (34), (11i), and (11j).
2) Feasible-Pair Preprocessing: A UE–ES pair is individually feasible only if it satisfies the security-level requirement and can support the required secure transmission rate under the per-link resource limits. Define n o
Mf𝑛 ≜ 𝑗 ∈ M 𝑎 𝑛 𝑗 = 1, 𝛾𝑛 𝑗 > 𝛾𝑛EV , 𝑟 𝑛se𝑗 (𝑏 max , 𝑝 max ) ≥ 𝑟¯𝑛 . (37) All infeasible association variables are fixed according to 𝑥 𝑛 𝑗 = 0, ∀𝑛 ∈ N, 𝑗 ∉ Mf𝑛 . (38) f If M𝑛 = ∅ for any UE 𝑛, then P 𝐴1 is infeasible.
After this fixing, P 𝐴1 is a mixed-integer convex problem whose only source of nonconvexity is the binary association. Indeed, every remaining pair satisfies 𝛾𝑛 𝑗 > 𝛾𝑛EV , so 𝑞 ′′ 𝑛 𝑗 (𝑧) = 1 EV ) 2 (1 + 𝛾 EV 𝑧) −2 − 𝛾 2 (1 + 𝛾 𝑧) −2 ] < 0 for 𝑧 ≥ 0 [(𝛾 𝑛 𝑗 𝑛 𝑛 𝑛𝑗 ln 2 and 𝑞 𝑛 𝑗 is concave; its closed perspective 𝑏 𝑛 𝑗 𝑞 𝑛 𝑗 ( 𝑝 𝑛 𝑗 /𝑏 𝑛 𝑗 ) = 𝑟 𝑛se𝑗 (𝑏 𝑛 𝑗 , 𝑝 𝑛 𝑗 ) is therefore jointly concave in (𝑏 𝑛 𝑗 , 𝑝 𝑛 𝑗 ) [27, Sec. 3.2.6], making 𝑟¯𝑛 𝑥 𝑛 𝑗 − 𝑟 𝑛se𝑗 (𝑏 𝑛 𝑗 , 𝑝 𝑛 𝑗 ) ≤ 0 convex. All other constraints of the continuous relaxation are affine, and the objective is concave because 𝜌 𝑛 > 0 and ln(·) is concave and increasing. For every feasible pair, define the minimum bandwidth required at the maximum transmission power: o n 𝑏𝑛 𝑗 ≜
min
0≤𝑏≤𝑏max
𝑏 𝑟 𝑛se𝑗 (𝑏, 𝑝 max ) ≥ 𝑟¯𝑛 .
(39)
The one-dimensional value 𝑏 𝑛 𝑗 can be computed efficiently by bisection because 𝑟 𝑛se𝑗 (𝑏, 𝑝 max ) is continuous and nondecreasing in 𝑏 over the feasible interval. It is also bounded: 𝑏 𝑞 𝑛 𝑗 ( 𝑝/𝑏) → 𝑝(𝛾𝑛 𝑗 − 𝛾𝑛EV )/ln 2 as 𝑏 → ∞, so a link admits a feasible allocation only if 𝑟¯𝑛 < 𝑝 max (𝛾𝑛 𝑗 −𝛾𝑛EV )/ln 2, regardless of the available bandwidth. Security-induced infeasibility is therefore a property of the secrecy advantage rather than of spectrum scarcity alone. 3) Projection of Communication Variables: The power variables do not appear in the objective, and no aggregate power constraint is imposed. This structure permits the communication variables to be eliminated exactly rather than handled by a custom branch-and-bound procedure. Lemma 2. Consider a binary association matrix x satisfying Í𝑀 f 𝑥 𝑗=1 𝑛 𝑗 = 1 and 𝑥 𝑛 𝑗 = 0 for all 𝑗 ∉ M𝑛 . There exist bandwidth and power allocations (𝒃, 𝒑) satisfying (23), (24), (26), and (34) if and only if 𝑁 Õ 𝑏𝑛 𝑗 𝑥𝑛 𝑗 ≤ 𝐵 𝑗 , ∀ 𝑗 ∈ M. (40) 𝑛=1
Whenever (40) holds, one feasible recovery is 𝑏𝑛 𝑗 = 𝑏𝑛 𝑗 𝑥𝑛 𝑗 , 𝑝 𝑛 𝑗 = 𝑝 max 𝑥 𝑛 𝑗 .
(41)
Proof: See Appendix B. The recovery in (41) selects one representative allocation among possibly many that share the same projected point; it excludes no feasible association or processing decision. Remark 2. The exact projection relies on the absence of a transmission-energy objective or an aggregate power budget. If either feature is introduced, setting every selected link to 𝑝 max may no longer be optimal, and the variables (𝒃, 𝒑) must be retained in the optimization.
8
Algorithm 1: Dual-Threshold Selection for UE 𝑛 Input: Observed confidence-score set C𝑛 ; event labels; capacity 𝑅𝑛∗ Output: Optimal feasible thresholds (𝛼𝑛𝑙,∗ , 𝛼𝑛𝑢,∗ ) and response envelope 𝜅 𝑛 1 Sort the unique values in C𝑛 ∪ {0, 1} into 𝑐1 < · · · < 𝑐 𝐾 ; best ← −∞; 2 Set 𝑈 3 Set 𝜅 𝑛 ← 0; 4 for 𝑖 ← 1 to 𝐾 do 5 for 𝑗 ← 𝑖 to 𝐾 do 6 Set (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ) ← (𝑐 𝑖 , 𝑐 𝑗 ); 7 Evaluate 𝐿 𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ) and U𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ); 8 if 𝐿 𝑛 > 0 and U𝑛 > 0 then 9 𝜅 𝑛 ← max{𝜅 𝑛 , |Φ𝑛 |U𝑛 /𝐿 𝑛 }; 10 end 𝑁 Õ 11 if 𝐿 𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ) ≤ 𝑅𝑛∗ and U𝑛 > 𝑈 best then proj P 𝐴1 : maximize 𝜌𝑛 𝜂𝑛 (45) 12 Store the threshold pair and update 𝑈 best ; x,𝒘,𝜼 𝑛=1 13 end subject to (44), (11i), (40), (27), (25), (28), (38), and (11j). 14 end proj Problem P 𝐴1 is accordingly a mixed-integer exponential- 15 end cone problem: all constraints are affine except for the 16 return (𝛼𝑙,∗ , 𝛼𝑢,∗ ) and 𝜅 ; 𝑛 𝑛 𝑛 exponential-cone memberships, and the only discrete variables are the binary association variables.
4) Mixed-Integer Exponential-Cone Reformulation: Introduce an auxiliary variable 𝜂 𝑛 for each UE and impose 𝑀 ©Õ ª 𝜂 𝑛 ≤ ln 𝑤𝑛 𝑗 ® . (42) « 𝑗=1 ¬ Using the primal exponential cone Kexp ≜ cl (𝑢, 𝑣, 𝑧) ∈ R3 𝑣 > 0, 𝑢 ≥ 𝑣 exp(𝑧/𝑣) , (43) constraint (42) is equivalently represented as 𝑀 ©Õ ª 𝑤 𝑛 𝑗 , 1, 𝜂 𝑛 ® ∈ Kexp . (44) « 𝑗=1 ¬ Í Because 𝜌 𝑛 > 0 and the objective maximizes 𝑛 𝜌 𝑛 𝜂 𝑛 , (44) is active at the optimum, so introducing 𝜂 𝑛 linearizes the objective without changing it. After applying Lemma 2, Subproblem A is reformulated as
proj
Proposition 1. Problems P 𝐴1 and P 𝐴1 have the same optimal objective value. Furthermore, any global optimizer (x∗ , 𝒘 ∗ , 𝜼∗ ) proj of P 𝐴1 , together with the communication recovery in (41), yields a global optimizer of P 𝐴1 . Conversely, every feasible proj solution of P 𝐴1 maps to a feasible solution of P 𝐴1 with the same objective value by setting 𝑀 ©Õ ª 𝜂 𝑛 = ln 𝑤𝑛 𝑗 ® . (46) 𝑗=1 « ¬ Í Proof. Lemma 2 and 𝜂 𝑛 = ln( 𝑗 𝑤 𝑛 𝑗 ) map every original feasible solution to an equally valued projected one, so proj opt(P 𝐴1 ) ≥ opt(P 𝐴1 ). Conversely, the lemma recovers (𝒃, 𝒑), while 𝜌 𝑛 > 0 makes (44) tight at optimum, yielding the reverse inequality and equivalence. ■ proj
We solve P 𝐴1 with the MOSEK mixed-integer conic solver [28] and recover (𝒃 ∗ , 𝒑 ∗ ) via (41); the solver’s certified gap therefore applies to P 𝐴1 itself. D. FREDI Stage B: Dual-Threshold Inference Optimization FREDI Stage B acts on the inference layer: with (x∗ , 𝒃 ∗ , 𝒑 ∗ , 𝒘 ∗ ) fixed, Subproblem P 𝐵 separates across UEs. Because 𝜌 𝑛 > 0 and ln(·) is strictly increasing, the threshold pair of UE 𝑛 can be obtained from (47) SP𝑛 : maximize U𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ) 𝛼𝑛𝑙 , 𝛼𝑛𝑢
subject to
0 ≤ 𝛼𝑛𝑙 ≤ 𝛼𝑛𝑢 ≤ 1,
𝐿 𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ) ≤ 𝑅𝑛∗ ,
(48)
Stage B therefore controls the detection–offloading behavior of the CNN through the label rule (2), rather than introducing fairness into the inference rule. Role of monotonicity. Lemma 1 implies that the empirical utility and offloaded load are non-increasing in both thresholds. Hence, an optimal feasible pair lies on the lower boundary of the feasible threshold region, which permits directional pruning during an exact finite search. Exact empirical search. The empirical utility and load are piecewise constant because they are obtained by counting samples. Their values can change only when a threshold crosses an observed confidence score. It is therefore sufficient to search threshold pairs drawn from the finite set of observed scores. The exhaustive search is globally optimal over the empirical threshold domain. The additional envelope update evaluates (19) without changing the selected thresholds. Monotonicity can be used to stop exploring a row or column once all remaining pairs are dominated or infeasible, provided that any pruned candidates cannot increase either the feasible utility or the response envelope. IV. N UMERICAL R ESULTS We evaluate FREDI from three complementary perspectives. Scenario S1 isolates the dual-threshold (DT) inference component against a single-threshold (ST) policy; Scenario S2 evaluates the end-to-end fairness–utility trade-off and validates the two-stage decomposition; and Scenario S3 examines securityconstrained association and Stage-A scalability.
(49)
Í ∗ ∗ where 𝑅𝑛∗ = 𝑀 𝑗=1 𝑥 𝑛 𝑗 𝑤 𝑛 𝑗 . By construction of Subproblem A, ∗ 𝑅𝑛 ≤ |Φ𝑛 |; hence every feasible threshold solution satisfies the QoS chain 𝐿 𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ) ≤ 𝑅𝑛∗ ≤ |Φ𝑛 |.
A. Experimental Setup 1) Dataset, Models, and Implementation: We use a binary cats–dogs corpus with 4,001 cat and 4,006 dog images
9
for training and 1,011 cat and 1,012 dog images for testing [29]. Images are resized to 224×224 and normalized using ImageNet statistics [30]. UE-side screening uses early-exit MobileNetV2 [31] and ShuffleNetV2 [32], with intermediate heads formed by global average pooling and a two-class linear projection. MobileNetV2 and ShuffleNetV2 contain 18 and 9 exits, respectively. These lightweight backbones are representative of confidence-based early-exit inference [3], [4]. Both models are trained on Google Cloud Platform using four vCPUs, 15 GB RAM, and one NVIDIA T4 GPU for 30 epochs, with batch size 64, learning rate 10−3 , weight decay 10−4 , and equally weighted cross-entropy losses across exits. Final-exit validation performance selects the checkpoint. For S1, exit-wise logits on the complete 2,023-image heldout test set are converted to positive-class softmax confidence traces; all threshold policies therefore operate on identical CNN outputs without retraining. For S2 and S3, each UE generates |Φ𝑛 | = 100 events per scheduling window and MobileNetV2 is used throughout. Unless otherwise stated, 𝑀 = 3, 𝐵 𝑗 = 20 MHz, 𝐷 𝑛 = 128 KiB, and 𝑡 𝑛r = 100 ms. Random geometry generates physically consistent legitimate and eavesdropper channel gains but is proj not itself swept. The projected problem P 𝐴1 is implemented using the MOSEK Fusion API [28]; minimum-bandwidth coefficients 𝑏 𝑛 𝑗 are computed by bisection and infeasible UE– ES pairs are removed before model construction. In S3-B, the relative MIP gap is 10−4 and one warm-up solve is discarded before timing. Per-link caps are 𝑏 max = 20 MHz, 𝑝 max = 0.2 W (23 dBm) and 𝑤 max = 100 events/window, the noise PSD is −174 dBm/Hz. UEs are uniformly distributed in a 70m triangular serving region under distance based path loss (exponent 4) and Rayleigh fading. The threshold grid uses Δ 𝛼 = 0.0025. 2) Scenario Design and Compared Methods: The following defines the compared policies, normalized loads, and statistical protocol. Scenario S1: Threshold policies for early-exit screening. Scenario S1 isolates threshold selection from wireless and multi-UE resource-allocation effects. The allocated edgeprocessing budget is parameterized by the normalized ratio 𝑅★ 𝜂𝑅 = , (50) |Φ| with 𝜂 𝑅 ∈ {0.1, 0.2, . . . , 1.0}. The proposed dual-threshold (DT) policy optimizes (𝛼𝑙 , 𝛼𝑢 ) over 0 ≤ 𝛼𝑙 ≤ 𝛼𝑢 ≤ 1, whereas the coupled single-parameter (ST) policy uses 𝜏 ∈ [0.5, 1] with 𝛼𝑙 = 1 − 𝜏 and 𝛼𝑢 = 𝜏, the symmetric confidence rule used in conventional early-exit inference. Because 1 − 𝜏 ≤ 𝜏 for every admissible 𝜏, the ST family is a one-dimensional subset of the DT family; S1 is therefore an ablation that isolates the value of decoupling the two thresholds, and DT is optimal for ST whenever the two are compared under the same processing budget. The comparison quantifies how much this additional degree of freedom is worth, and where the coupling becomes binding. Both use the same threshold-grid resolution Δ 𝛼 = 0.0025. Refining to Δ 𝛼 = 0.00125 at the most sensitive MobileNetV2 point, 𝜂 𝑅 = 0.3, leaves the selected optimum and reported metrics unchanged. TPR-optimal ties are resolved by lower offloaded load, lower normalized MACs, higher F1
score, and then threshold order. Scenario S2: Utility and decomposition validation. Scenario S2 is the main system-level experiment. Each UE stream contains 50 positive and 50 negative held-out events sampled without replacement within a realization, with identical streams and network realization shared across methods. Sweeping 𝑊 𝑗 ∈ {200, 160, 120, 100, 80, 60} gives the normalized processing load Í𝑁 |Φ𝑛 | 200 𝜆𝑊 = Í𝑛=1 = ∈ {1, 1.25, 1.67, 2, 2.5, 3.33}. (51) 𝑀 𝑊𝑗 𝑗=1 𝑊 𝑗 FREDI is compared with a Sum-Utility benchmark over the Í same feasible associations and resources, maximizing 𝑛 U𝑛 . Since multiple optima may yield different fairness, we report the minimum and maximum Jain’s indices. We also validate the decomposition by exactly solving a finite Í empirical form of P0 at 𝜆𝑊 ∈ {1.25, 2, 3.33}, maximizing 𝑛 𝜌 𝑛 ln(U𝑛 ) over the same feasible candidates. Scenario S3: security constraints and scalability. Scenario S3 stress-tests FREDI Stage A through two complementary sub-experiments. In S3-A, FREDI is compared with an Unrestricted-PF reference under identical channels and resources. Unrestricted-PF removes only the UE–ES securityeligibility constraint while retaining the same proportionalfair resource allocation. In S3-B, FREDI is evaluated as the number of UEs increases. In S3-A, seven UE security-demand profiles produce 𝜅elig ∈ {1, 0.89, . . . , 0.33}, each with 50 paired channel realizations sharing seeds across the two methods. Security-induced infeasibility is retained as an experimental outcome rather than resampled away. In S3-B, 𝑁 ∈ {6, 12, 24, 36, 60, 96, 144} with approximately balanced UE security tiers. To separate runtime growth from increasing resource scarcity, the per-ES capacities scale as 100𝑁 𝑁 𝑊 𝑗 (𝑁) = , 𝐵 𝑗 (𝑁) = 20 MHz. (52) 3 9 Fifty random realizations are used for each system size. S1 is deterministic on the fixed test set. S2 and S3 use 50 realizations per operating point. Continuous quantities are reported as sample means with 95% confidence intervals unless otherwise stated; S3-B additionally reports median solve time, whereas the binary S3-A infeasibility rate uses 95% Wilson confidence intervals. For S2, the confidence intervals describe variability over the paired event-stream and network realizations drawn from the fixed held-out pool. 3) Performance Metrics: For S1, the primary detection metrics are true-positive rate (TPR), false-positive rate (FPR), and F1 score. The computation saving is 1 Õ MAC(𝑞 𝑖 ) 𝐺 comp = 1 − , (53) |Φ| 𝑖 ∈Φ MAC(𝑄) where 𝑞 𝑖 is the exit used for event 𝑖, MAC(𝑞 𝑖 ) is the cumulative MAC count up to that exit, and 𝑄 is the final exit. For a nonnegative vector z = (𝑧 1 , . . . , 𝑧 𝑁 ), Jain’s index [33] is Í 2 𝑁 𝑛=1 𝑧 𝑛 𝐽 (z) = (54) Í𝑁 2 . 𝑁 𝑛=1 𝑧𝑛 For S2, let 𝑈𝑛 ≡ U𝑛 (𝛼𝑛𝑙 , 𝛼𝑛𝑢 ) denote the final empirical
𝑛=1
𝑛=1
(57) The direct formulation excludes zero-utility empirical operating points because the logarithmic objective is undefined there. The aggregate service ratio used in S3-A is Í𝑁 𝑅𝑛 . (58) 𝑆 = Í 𝑁𝑛=1 𝑛=1 |Φ𝑛 | For S3-A, define 𝑒 𝑛 𝑗 = 1 when UE 𝑛 is security-eligible for ES 𝑗, and 𝑒 𝑛 𝑗 = 0 otherwise. The eligible-association ratio is 𝑁 𝑀 1 ÕÕ 𝜅elig = 𝑒𝑛 𝑗 , (59) 𝑁 𝑀 𝑛=1 𝑗=1 and the infeasible-realization rate is 𝐾infeas 𝑃infeas = . (60) 𝐾 We additionally report the actual feasible-pair ratio after preprocessing and ES processing and bandwidth utilization. For S3-B, we report MOSEK solve time and solver optimality status. B. Dual- versus Single-Threshold Screening Performance Fig. 4 and Table III compare DT and ST under the same allocated edge-processing budget. Both policies use identical backbones and identical exit-wise confidence traces, so the differences isolate the effect of decoupling the two thresholds. The coupling is most damaging under tight budgets. Because ST ties 𝛼𝑙 = 1 − 𝜏 to 𝛼𝑢 = 𝜏, reducing the offloaded load forces 𝜏 → 1, which raises the critical-exit threshold and widens the undecided interval I at the same time. For MobileNetV2 at 𝜂 𝑅 ∈ {0.1, 0.2} the load cap admits exactly one ST setting, 𝜏 = 1: then (𝛼𝑙 , 𝛼𝑢 ) = (0, 1), I spans the whole confidence range, no sample crosses either threshold, and all are declared normal at the final exit by (2). Table III records the consequence — a true-positive rate of exactly zero, and no computation saving either, since nothing exits early. DT decouples the two decisions and attains TPRs of 0.197 and 0.396 with computation savings of 91.1% and 67.5% at the same budgets. ShuffleNetV2 does not collapse (𝜏 = 0.9975 stays feasible), so the effect depends on how sharply a backbone concentrates its confidence traces; DT still leads it on every metric there, most visibly in computation saving (55.0% against 5.6% at 𝜂 𝑅 = 0.1). At the intermediate budget 𝜂 𝑅 = 0.3, DT improves MobileNetV2 TPR from 0.576 to 0.593 and F1 from 0.729 to 0.742 while raising computation saving from 35.2% to 58.8%, at essentially the same offloaded load (29.0% to 29.9%). For ShuffleNetV2 the two policies select operating points with identical TPR, FPR, F1 and offloaded load, and DT gains only
Computation saving
utility. We report the Jain’s indices 𝑁 𝑁 𝐽𝑅 = 𝐽 ({𝑅𝑛∗ /|Φ𝑛 |} 𝑛=1 ), 𝐽𝑈 = 𝐽 ({𝑈𝑛 } 𝑛=1 ), (55) Í and the aggregate utility 𝑛 𝑈𝑛 . To characterize the empirical response factor in (14) at the realized FREDI allocation, define std({𝜅 ∗𝑛 }) 𝑈𝑛 𝜅 ∗𝑛 ≡ 𝜅 𝑛 (𝑅𝑛∗ ) = ∗ , CV 𝜅 = . (56) 𝑅𝑛 /|Φ𝑛 | mean({𝜅 ∗𝑛 }) Finally, the decomposition is compared with direct empirical optimization through 𝑁 𝑁 Õ Õ 𝐹0FREDI = 𝜌 𝑛 ln 𝑈𝑛FREDI , 𝐹0direct = 𝜌 𝑛 ln 𝑈𝑛direct .
True positive rate
10
100% 80% 60% 40%
MobileNetV2 - DT ShuffleNetV2 - DT MobileNetV2 - ST ShuffleNetV2 - ST
20% 0% 100% 80% 60% 40% 20% 0% 0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
0.9
1.0
Allocated edge-processing budget, ηR
Fig. 4: True positive rate (top) and computation saving under the dual-threshold (DT) and coupled single-parameter (ST) policies as the allocated edge-processing budget 𝜂 𝑅 increases. in computation saving (30.9% to 35.9%). At 𝜂 𝑅 = 0.5 both backbones show small TPR and F1 gains, and the computation benefit becomes model-dependent. Above 𝜂 𝑅 = 0.5 the comparison changes character, because ST cannot spend the budget it is given. Its offloaded set is {events whose confidence reaches 𝜏 at some exit}, which on the balanced test corpus saturates near half the events as 𝜏 → 0.5: Table III shows the ST load rising only from 49.9% at 𝜂 𝑅 = 0.5 to 50.5% at 𝜂 𝑅 = 0.8, while DT reaches 79.7%. The apparently unfavourable DT entries at 𝜂 𝑅 = 0.8 — FPR 0.599 and F1 0.767 against 0.094 and 0.912 for ST — therefore compare a policy operating at its budget against one operating far below it. They also reflect the objective rather than the policy class: SP𝑛 maximizes TPR subject to the load cap and never penalizes false positives, so DT spends whatever capacity Stage A reserves. Where precision matters, the same machinery accepts an F1- or FPR-constrained utility without any change to Stage A. The value of the second threshold is therefore not a uniform gain on every classification metric, but the removal of a structural constraint: it keeps detection alive where the coupled policy collapses, and it makes the full range of processing budgets reachable. C. Utility Performance and Decomposition Validation The top panel of Fig. 5 compares the final-utility fairness of FREDI with the fairness envelope of Sum-Utility as processing contention increases. With sufficient processing capacity (𝜆𝑊 = 1), both methods attain 𝐽𝑈 = 1. FREDI remains essentially perfectly fair throughout the sweep, with mean 𝐽𝑈 ≥ 0.9993. In contrast, the Sum-Utility envelope widens under heavier scarcity: at 𝜆𝑊 = 2.5, its mean lower and upper boundaries are approximately 0.9824 and 0.9953, and at 𝜆𝑊 = 3.33 they become approximately 0.8561 and 0.9851. Hence, maximizing aggregate utility alone can admit substantially less fair optima even though an arbitrary solver tie-break may conceal this behavior.
11
TABLE III: Representative S1 operating points on the 2,023image held-out test set.
TABLE IV: Empirical validation of the FREDI decomposition and direct finite empirical optimization of P0 .
CNN
Policy
𝜂𝑅
TPR
FPR
F1
Offloaded load
Comp. saving
𝜆𝑊
𝐽𝑅
MobileNetV2 MobileNetV2 MobileNetV2 MobileNetV2 MobileNetV2 MobileNetV2 MobileNetV2 MobileNetV2
DT ST DT ST DT ST DT ST
0.1 0.1 0.2 0.2 0.3 0.3 0.8 0.8
0.197 0.000 0.396 0.000 0.593 0.576 0.994 0.916
0.003 0.000 0.003 0.000 0.005 0.004 0.599 0.094
0.328 0.000 0.566 0.000 0.742 0.729 0.767 0.912
10.0% 0.0% 20.0% 0.0% 29.9% 29.0% 79.7% 50.5%
91.1% 0.0% 67.5% 0.0% 58.8% 35.2% 83.1% 67.8%
1.00 1.25 1.67 2.00 2.50 3.33
1.0 1.0000 1.0 0.9999 1.0 0.9995 1.0 0.9993 1.0 0.9996 1.0 0.9998
ShuffleNetV2 ShuffleNetV2 ShuffleNetV2 ShuffleNetV2
DT ST DT ST
0.3 0.3 0.8 0.8
0.579 0.020 0.724 0.579 0.020 0.724 0.975 0.615 0.753 0.869 0.204 0.838
30.0% 30.0% 79.5% 53.6%
35.9% 30.9% 56.4% 56.6%
DT and ST are evaluated using identical exit-wise confidence traces. Bold entries highlight selected stronger detection or computation results at the same processing budget.
Jain's index, JU
0.95 0.90
n
Aggregate utility, ∑Un
FREDI Sum-Utility (Jmax) Sum-Utility (Jmin)
6.0 5.5 5.0 4.5 4.0 3.5 1.00 1.25
1.67
CV 𝜅 (%)
𝐹0FREDI
𝐹0direct
0.00 1.06 2.27 2.71 2.06 1.34
– −0.0409 ± 0.0065 – −0.4969 ± 0.0138 – −3.1138 ± 0.0109
– −0.0040 ± 0.0023 – −0.4228 ± 0.0128 – −3.1008 ± 0.0084
Values are averaged over 50 paired realizations. 𝐽𝑅 = 𝐽 ( { 𝑅𝑛∗ /|Φ𝑛 | } ) and 𝐽𝑈 = 𝐽 ( {𝑈𝑛 } ). The 𝐹0 entries report mean ± 95% confidenceinterval half-width. Direct-P0 validation is evaluated at three representative processing loads. Bold entries mark the most stringent observed FREDI fairness-consistency point, where the minimum 𝐽𝑈 and maximum CV 𝜅 occur.
Second, we compare FREDI with direct finite empirical optimization of the original logarithmic objective. All 150 direct-P0 validation instances (50 paired realizations at each of three representative loads) are solved to optimal integer status. As required, 𝐹0direct ≥ 𝐹0FREDI in every case within numerical tolerance. The mean objectives remain close, particularly under severe scarcity, where they are −3.1138 and −3.1008 for FREDI and direct optimization, respectively. This direct comparison gives a realized empirical assessment of the end-to-end loss, complementing the general upper bound in Theorem 1. Together, the fairness consistency, small CV 𝜅 , and direct-objective comparison support the FREDI two-stage decomposition over the evaluated operating regime without presuming exact end-to-end equivalence.
1.00
0.85
𝐽𝑈
2.00
2.50
3.33
Normalized processing load, λW
Fig. 5: Jain’s index 𝐽𝑈 (top) and aggregate utility (bottom) for FREDI and the Sum-Utility benchmark as the normalized processing load 𝜆𝑊 increases. The fairness improvement is obtained with a small aggregate-utility cost. The bottom panel of Fig. Í5 shows that Sum-Utility, by construction, gives the largest 𝑛 𝑈𝑛 , while FREDI remains close over the full load range. At 𝜆 𝑊 = 1, both attain an aggregate utility of 6. As contention increases, the largest mean difference over the six operating points is 0.0692 at 𝜆𝑊 = 2, where FREDI and Sum-Utility obtain 5.5252 and 5.5944, respectively. Under severe scarcity (𝜆𝑊 = 3.33), the corresponding values are 3.5712 and 3.5936. Thus, FREDI maintains near-uniform user utility while sacrificing little aggregate detection utility. Table IV assesses the decomposition from two complementary viewpoints. First, the normalized processing allocation is perfectly balanced at every tested load, 𝐽𝑅 = 1, and the final utility remains extremely close to this allocation-level fairness, with 𝐽𝑈 ≥ 0.9993. At the realized FREDI operating points, the mean coefficient of variation of 𝜅 ∗𝑛 = 𝑈𝑛 /(𝑅𝑛∗ /|Φ𝑛 |) never exceeds 2.71%. The repeated small cross-UE dispersion values show that the realized response residual remains balanced across UEs over the load sweep; they do not by themselves establish the per-UE constancy required by Corollary 1.1 or evaluate the response envelope 𝜅 𝑛 in Theorem 1.
D. Impact of Security Constraints and System Scalability 1) Security-Constrained Association: Fig. 6 jointly shows the service and feasibility effects of security-induced restrictions on UE–ES association. As shown in the top panel, when 𝜅elig ≥ 0.778, FREDI serves essentially all offered workload and remains close to the eligibility-unconstrained PF resourceallocation upper bound (Unrestricted-PF). Below this region, the service ratio deteriorates rapidly, falling to about 0.661 at 𝜅elig = 0.444 and to 0.333 at 𝜅elig = 0.333. The bottom panel of Fig. 6 shows the corresponding infeasible-realization rate of the security-constrained FREDI formulation. All realizations remain feasible for 𝜅 elig ≥ 0.778, whereas 𝑃infeas increases to 0.14 at 𝜅 elig = 0.444 and 0.22 at 𝜅 elig = 0.333. Hence, moderate security restrictions can be absorbed without loss of feasibility, whereas severe restrictions simultaneously reduce service capability and increase the probability of an infeasible association/resource-allocation instance. The mechanism is resource fragmentation rather than a lack of aggregate capacity. The measured feasible-pair ratio closely tracks 𝜅 elig , confirming that security eligibility dominates the sweep, while the maximum ES bandwidth utilization rises from about 0.30 to 0.64. Under the most restrictive profile, mean processing utilization is only about 0.333 although one eligible ES is fully utilized. Capacity therefore remains stranded at ESs that some UEs are not permitted to access. Moreover, Jain fairness can still be high when service is poor
12
Service ratio, S
100% 80% 60% 40% 20%
FREDI Unrestricted-PF
Infeasibility rate, Pinfeas
0%
30% 20% 10%
A PPENDIX A P ROOF OF L EMMA 1
0% 0.33
0.44
0.56
0.67
0.78
0.89
1.00
Eligible association ratio, κelig
Fig. 6: Service ratio (top) and infeasible-realization rate (bottom) of FREDI as the eligible association ratio 𝜅 elig increases. Optimization time (ms)
empirical confidence-score domain for each UE. A computable response envelope bounds the loss relative to the empirical global optimum and yields a proportional-response condition for global optimality. Numerically, FREDI maintains nearperfect utility fairness (𝐽𝑈 ≥ 0.9993 on average) while remaining close to the Sum-Utility benchmark; direct optimization of representative P0 instances further supports the twostage decomposition. The results also quantify the effects of security-constrained association and confirm practical Stage-A scalability over the tested system sizes.
60
FREDI Stage A
40
20
0 6 12
24
36
60
96
144
Number of UEs N
Fig. 7: Optimization time of FREDI Stage A as the number of UEs 𝑁 increases. (e.g., 𝐽 = 1 at 𝜅elig = 0.333), showing that fairness and service capability must be interpreted jointly. 2) Solver Scalability: Fig. 7 reports the computational scalability of FREDI Stage A, with communication and processing resources scaled with 𝑁 to maintain a comparable operating load. The median MOSEK solve time increases from about 2 ms at 𝑁 = 6 to 59 ms at 𝑁 = 144 and remains below 0.1 s even for the largest tested system. The wider interquartile range at large 𝑁 reflects realization-dependent mixed-integer search induced by different feasible UE–ES association sets. All 350 measured instances reach an optimal integer status within the prescribed relative-gap tolerance of 10−4 . Together with preprocessing that removes infeasible UE–ES pairs before model construction, these results show that the FREDI security-constrained resource-allocation formulation remains computationally practical over the investigated system sizes without relaxing solution quality.
Proof. Consider two feasible threshold pairs (𝛼𝑙 , 𝛼𝑢 ) and 𝛼𝑢 ) such that e 𝛼𝑙 ≥ 𝛼𝑙 and e 𝛼𝑢 ≥ 𝛼𝑢 . Since the CNN (e 𝛼𝑙 , e (𝑘 ) 𝐿 is fixed, the confidence trace 𝐶𝑞 𝑞=1 of each event 𝐼 𝑘 is independent of the thresholds. Suppose that 𝐼 𝑘 is classified as critical under (e 𝛼𝑙 , e 𝛼𝑢 ), first ★ ★ at layer 𝑞 . For every 𝑞 < 𝑞 , the event has not terminated as normal, so 𝐶𝑞(𝑘 ) > e 𝛼𝑙 ≥ 𝛼𝑙 . If 𝐶𝑞(𝑘 ) ≥ 𝛼𝑢 at any such layer, then 𝐼 𝑘 is already classified as critical under (𝛼𝑙 , 𝛼𝑢 ); otherwise it remains undecided. At layer 𝑞★, critical classification under the larger thresholds gives 𝐶𝑞(𝑘★) ≥ e 𝛼𝑢 ≥ 𝛼𝑢 , so the event is critical 𝑙 𝑢 under (𝛼 , 𝛼 ) as well. The same argument applies at the final layer, while the equal-threshold convention in (2) preserves this implication when the undecided interval vanishes. b critical (e b critical (𝛼𝑙 , 𝛼𝑢 ), Hence, 𝐼 𝑘 ∈ Φ 𝛼𝑙 , e 𝛼𝑢 ) =⇒ 𝐼 𝑘 ∈ Φ critical 𝑙 𝑢 critical b b and therefore Φ (e 𝛼 ,e 𝛼 ) ⊆ Φ (𝛼𝑙 , 𝛼𝑢 ), which critical b proves (12). Consequently, | Φ | is non-increasing in either threshold. ■ A PPENDIX B P ROOF OF L EMMA 2
V. C ONCLUSION
Proof. For a feasible UE–ES pair satisfying 𝛾𝑛 𝑗 > 𝛾𝑛EV and fixed 𝑏 > 0, 𝜕𝑟 𝑛se𝑗 𝛾𝑛 𝑗 𝛾𝑛EV 𝑏 = − > 0, (61) 𝜕𝑝 ln 2 𝑏 + 𝛾𝑛 𝑗 𝑝 𝑏 + 𝛾𝑛EV 𝑝 so the rate increases in 𝑝. For fixed 𝑝 > 0, write 𝑟 𝑛se𝑗 (𝑏, 𝑝) = 𝑏𝑞 𝑛 𝑗 ( 𝑝/𝑏), where 𝑞 𝑛 𝑗 is concave and 𝑞 𝑛 𝑗 (0) = 0. Since 𝑞 𝑛 𝑗 (𝑧)/𝑧 is non-increasing, 𝑟 𝑛se𝑗 (𝑏, 𝑝) = 𝑝𝑞 𝑛 𝑗 ( 𝑝/𝑏)/( 𝑝/𝑏) is non-decreasing in 𝑏. For necessity, any feasible selected link satisfies 𝑟 𝑛se𝑗 (𝑏 𝑛 𝑗 , 𝑝 max ) ≥ 𝑟 𝑛se𝑗 (𝑏 𝑛 𝑗 , 𝑝 𝑛 𝑗 ) ≥ 𝑟¯𝑛 and hence 𝑏 𝑛 𝑗 ≥ 𝑏 𝑛 𝑗 . For an unselected link, 𝑏 𝑛 𝑗 ≥ 𝑏 𝑛 𝑗 𝑥 𝑛 𝑗 is immediate. Summing over UEs gives Õ Õ 𝑏𝑛 𝑗 𝑥𝑛 𝑗 ≤ 𝑏𝑛 𝑗 ≤ 𝐵 𝑗 , (62)
This paper presented FREDI, a secure cooperative wireless edge-intelligence framework that jointly coordinates dualthreshold inference, UE–ES association, uplink bandwidth and power, and shared ES processing. Stage A uses proportional fairness, feasible-pair preprocessing, and an exact communication-variable projection to produce a mixed-integer exponential-cone formulation; Stage B exactly searches the
which proves (40). Conversely, if this constraint holds, set 𝑏 𝑛 𝑗 = 𝑏 𝑛 𝑗 𝑥 𝑛 𝑗 and 𝑝 𝑛 𝑗 = 𝑝 max 𝑥 𝑛 𝑗 . Each selected feasible pair meets the secure-rate requirement by the definition of 𝑏 𝑛 𝑗 ; each unselected pair has zero rate and demand. The per-link bounds follow from 𝑏 𝑛 𝑗 ≤ 𝑏 max on Mf𝑛 , and the aggregate constraint holds by construction. Thus, the recovered communication allocation is feasible. ■
𝑛
𝑛
13
A PPENDIX C P ROOF OF T HEOREM 1 𝑙, 𝜉 𝑢, 𝜉 Proof. For any admissible capacity 𝜉, let (𝛼𝑛 , 𝛼𝑛 ) attain 𝜉 𝜉 𝑉𝑛 (𝜉) and let 𝐿 𝑛 be its load. Since 0 < 𝐿 𝑛 ≤ 𝜉, 𝑙, 𝜉 𝑢, 𝜉 |Φ𝑛 |U𝑛 (𝛼𝑛 , 𝛼𝑛 ) |Φ𝑛 |𝑉𝑛 (𝜉) 𝜅 𝑛 (𝜉) = ≤ ≤ 𝜅𝑛. 𝜉 𝜉 𝐿𝑛 Conversely, evaluating 𝑉𝑛 (𝜉) at the load of a candidate attaining (19) gives the reverse inequality for the supremum; hence, (19) is exactly the upper envelope of 𝜅 𝑛 (𝜉) over the positive-utility capacity domain. For the resource allocation of a global optimizer of P0 , replacing any threshold pair by a maximizer in (13) cannot decrease the objective. Its resource variables also satisfy the Stage-A constraints; hence, its processing-coverage term is no larger than the globally optimal Stage-A value 𝐴(R∗ ). Applying (15) and the preceding response bound therefore gives 𝑁 Õ 𝐹0★ ≤ 𝐴(R∗ ) + 𝜌 𝑛 ln 𝜅 𝑛 . 𝑛=1
The exact Stage-B searches yield 𝐹0FREDI = 𝐴(R∗ ) + Í ∗ 𝑛 𝜌 𝑛 ln 𝜅 𝑛 (𝑅 𝑛 ). Subtracting the latter identity from the upper bound proves (20); the lower inequality follows because the FREDI solution is feasible for P0 . ■ R EFERENCES [1] K. B. Letaief, Y. Shi, J. Lu, and J. Lu, “Edge artificial intelligence for 6g: Vision, enabling technologies, and applications,” IEEE journal on selected areas in communications, vol. 40, no. 1, pp. 5–36, 2021. [2] J. Mendez, K. Bierzynski, M. P. Cuéllar, and D. P. Morales, “Edge intelligence: concepts, architectures, applications, and future directions,” ACM Transactions on Embedded Computing Systems (TECS), vol. 21, no. 5, pp. 1–41, 2022. [3] S. Teerapittayanon, B. McDanel, and H.-T. Kung, “Branchynet: Fast inference via early exiting from deep neural networks,” in 2016 23rd international conference on pattern recognition (ICPR). IEEE, 2016, pp. 2464–2469. [4] G. Huang, D. Chen, T. Li, F. Wu, L. van der Maaten, and K. Weinberger, “Multi-scale dense networks for resource efficient image classification,” in International Conference on Learning Representations, 2018. [5] H. Rahmath P, V. Srivastava, K. Chaurasia, R. G. Pacheco, and R. S. Couto, “Early-exit deep neural network-a comprehensive survey,” ACM Computing Surveys, vol. 57, no. 3, pp. 75:1–75:37, 2025. [6] C. Chow, “On optimum recognition error and reject tradeoff,” IEEE Transactions on information theory, vol. 16, no. 1, pp. 41–46, 1970. [7] G. Fumera, F. Roli, and G. Giacinto, “Reject option with multiple thresholds,” Pattern recognition, vol. 33, no. 12, pp. 2099–2101, 2000. [8] P. L. Bartlett and M. H. Wegkamp, “Classification with a reject option using a hinge loss,” Journal of Machine Learning Research, vol. 9, no. 8, 2008. [9] Z. Liu, Q. Lan, and K. Huang, “Resource allocation for multiuser edge inference with batching and early exiting,” IEEE Journal on Selected Areas in Communications, vol. 41, no. 4, pp. 1186–1200, 2023. [10] J.-W. Kim and H.-S. Lee, “Early exiting-aware joint resource allocation and dnn splitting for multisensor digital twin in edge–cloud collaborative system,” IEEE Internet of Things Journal, vol. 11, no. 22, pp. 36 933– 36 949, 2024. [11] Y. Zheng, T. Zhang, X. Mu, Y. Liu, and R. Huang, “Joint semantic transmission and resource allocation for intelligent computation task offloading in mec systems,” IEEE Transactions on Wireless Communications, vol. 24, no. 10, pp. 8756–8770, 2025. [12] X. Liu, H. Zhai, X. Zhou, H. Zhang, and Q. Liu, “Joint resource allocation and computation offloading for dnn inference with model partition and early exit in mec networks,” Chinese Journal of Electronics, vol. 35, no. 1, pp. 215–232, 2026. [13] Z. Zhang, T. Zhang, T. Shi, and Y. Wang, “Joint service placement and resource allocation for long-term dnn inference accuracy in dynamic mec networks,” IEEE Transactions on Vehicular Technology, 2025, early access.
[14] X. Yuan, N. Li, Q. Chen, W. Xu, and S. Guo, “Era: A qoe-aware collaborative inference algorithm for noma-based edge intelligence,” IEEE Transactions on Mobile Computing, 2025, early access. [15] Y. Zhou, C. You, and K. Huang, “Communication efficient cooperative edge ai via event-triggered computation offloading,” IEEE Transactions on Communications, vol. 74, pp. 3190–3205, 2026. [16] Y. Xu, T. Mohammed, M. Di Francesco, and C. Fischione, “Distributed assignment with load balancing for dnn inference at the edge,” IEEE Internet of Things Journal, vol. 10, no. 2, pp. 1053–1065, 2023. [17] T. T. Vu, N. H. Chu, K. T. Phan, D. T. Hoang, D. N. Nguyen, and E. Dutkiewicz, “Energy-based proportional fairness in cooperative edge computing,” IEEE Transactions on Mobile Computing, vol. 23, no. 12, pp. 12 229–12 246, 2024. [18] H. Xiao, Q. Pei, X. Song, and W. Shi, “Authentication security level and resource optimization of computation offloading in edge computing systems,” IEEE Internet of Things Journal, vol. 9, no. 15, pp. 13 010– 13 023, 2022. [19] K. Peng, P. Xiao, S. Wang, and V. C. Leung, “Scof: Security-aware computation offloading using federated reinforcement learning in industrial internet of things with edge computing,” IEEE Transactions on Services Computing, vol. 17, no. 4, pp. 1780–1792, 2024. [20] A. Abbas, D. Sutter, C. Zoufal, A. Lucchi, A. Figalli, and S. Woerner, “The power of quantum neural networks,” Nature computational science, vol. 1, no. 6, pp. 403–409, 2021. [21] C. Ren, R. Yan, H. Zhu, H. Yu, M. Xu, Y. Shen, Y. Xu, M. Xiao, Z. Y. Dong, M. Skoglund et al., “Toward quantum federated learning,” IEEE Transactions on Neural Networks and Learning Systems, vol. 36, no. 9, pp. 15 580–15 600, 2025. [22] C. Long, M. Huang, X. Ye, Y. Futamura, and T. Sakurai, “Hybrid quantum-classical-quantum convolutional neural networks,” Scientific Reports, vol. 15, no. 1, p. 31780, 2025. [23] F. Fan, Y. Shi, T. Guggemos, and X. X. Zhu, “Hybrid quantum-classical convolutional neural network model for image classification,” IEEE Transactions on Neural Networks and Learning Systems, vol. 35, no. 12, pp. 18 145–18 159, 2024. [24] A. Patil and M. Rane, “Convolutional neural networks: an overview and its applications in pattern recognition,” in International Conference on Information and Communication Technology for Intelligent Systems. Springer, 2020, pp. 21–30. [25] A. D. Wyner, “The wire-tap channel,” Bell system technical journal, vol. 54, no. 8, pp. 1355–1387, 1975. [26] F. P. Kelly, A. K. Maulloo, and D. K. H. Tan, “Rate control for communication networks: shadow prices, proportional fairness and stability,” Journal of the Operational Research society, vol. 49, no. 3, pp. 237–252, 1998. [27] S. Boyd and L. Vandenberghe, Convex Optimization. Cambridge, U.K.: Cambridge University Press, 2004. [28] MOSEK ApS, MOSEK Fusion API for Python 11.2.2, 2026, [Online]. Available: https://docs.mosek.com. [29] S. Schubert, “Cats and dogs dataset to train a dl model,” 2018, kaggle. [Online]. Available: https://www.kaggle.com/datasets/tongpython/catand-dog. [30] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition. Ieee, 2009, pp. 248–255. [31] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,” in 2018 IEEE/CVF conference on computer vision and pattern recognition. Ieee, 2018, pp. 4510–4520. [32] N. Ma, X. Zhang, H.-T. Zheng, and J. Sun, “Shufflenet v2: Practical guidelines for efficient cnn architecture design,” in European conference on computer vision. Springer, 2018, pp. 122–138. [33] R. Jain, D.-M. Chiu, and W. Hawe, “A quantitative measure of fairness and discrimination,” Digital Equipment Corporation, Tech. Rep. TR-301, 1984.