SCALE: Sensitivity-Aware Federated Unlearning with Information Freshness Optimization for Mobile Edge Computing Zihao Ding, Beining Wu, and Jun Huang
arXiv:2605.22589v1 [cs.NI] 21 May 2026
Department of Electrical Engineering and Computer Science, South Dakota State University, Brookings, SD 57007, USA Email: {zihao.ding, wu.beining}@jacks.sdstate.edu, [email protected] Abstract—Federated Unlearning (FU) is emerging as a powerful tool that enables the selective removal of client data to effectively address data contamination and meet strict privacy regulations in mobile edge computing (MEC) systems. Although FU has recently drawn attention in the AI community, existing approaches suffer from low unlearning precision and lack temporal information reflection, which results in suboptimal forgetting performance. To address these issues, we propose SCALE, a dual-level unlearning framework combining historical contribution analysis with information freshness-aware adaptive sparsification. Our framework first employs a historical contribution-based layer sensitivity analysis to identify layers most influenced by target clients, then performs fine-grained unlearning through adaptive sparsification at the weight subgroup level to balance information freshness with forgetting effectiveness. Through theoretical analysis, the proposed framework demonstrates the convergence properties and acceleration advantages. Our experiments and testbed results demonstrate superior unlearning effectiveness compared to state-of-the-art baselines, with significantly improved forgetting performance. Index Terms—Federated Unlearning, Age-of-Information, Reinforcement Learning, Mobile Edge Computing
I. I NTRODUCTION While the integration of Mobile Edge Computing (MEC) and Federated Learning (FL) lays a foundation for privacypreserving and resource-efficient edge intelligence, it raises new issues when data needs to be removed from trained models [1]. Real-world applications often require selective deletion of client data due to privacy regulations, user requests, or data contamination [2]. In smart traffic management, for example, traffic monitoring devices must send sensitive location and movement data to central servers for model training, which can expose private user information and yield massive communication overhead [3]–[5]. FL frameworks, however, lack built-in mechanisms for eliminating the influence of such data once it has been incorporated into the global model [6]– [8]. Federated Unlearning (FU) has emerged as a powerful tool for MEC systems through the selective removal of contributions from specific devices while retaining knowledge from trusted clients. FU is of particular importance when devices introduce corrupted data, experience security breaches, or when regulatory requirements mandate data deletion [9]–[13]. This work was supported by the National Science Foundation under grant CNS-2348422.
A key challenge in applying FU to MEC arises from the need to maintain both system responsiveness and information freshness within strict resource constraints. Age of Information (AoI) plays a critical role in this regard, as it prioritizes timely updates and prevents outdated parameters from degrading the effectiveness of forgetting and overall system performance [12], [14]–[16]. To see this, again, in smart traffic management, cameras monitoring construction zones may capture abnormal patterns that become irrelevant once the construction concludes, while compromised devices may inject malicious data that disrupts citywide traffic optimization. Evolving seasons and infrastructure modifications frequently diminish the relevance of historical data, which calls for its careful removal from existing models [17], [18]. As such, the goal of this work is to achieve federated unlearning in MEC systems to ensure model quality and unlearning efficiency. Recent advances in federated unlearning have explored different strategies for selectively removing data in distributed systems, falling into two main camps: retrainingbased [6], [10], [11], [19]–[22] and model manipulation-based approaches [23]–[26]. Retraining-based methods guarantee perfect unlearning by removing target data and starting training from scratch [27], [28], but they demand crushing computational and communication costs that make them impractical for real-time MEC deployments [1], [10], [11], [29]. Even with clustering-based speedup techniques and improved quasiNewton methods designed to reduce overhead, the computational burden remains too heavy for dynamic systems [24], [29]. In contrast, model manipulation methods directly modify trained models using gradient ascent [30]–[33], knowledge distillation [34], [35], parameter pruning [36], [37], and selective parameter updates [38]–[40] to eliminate target knowledge. These approaches include model transformation methods that adjust parameters to counteract forgotten sample influence, pruning techniques that remove client-specific components, and knowledge distillation frameworks that use historical models as teachers [23], [41]–[43]. While these methods show better efficiency than retraining, existing FU approaches still present issues in MEC systems, including low precision that affects non-target clients and neglect of temporal information freshness that leads to suboptimal forgetting performance, especially in time-critical MEC systems [20], [44]–[47]. To address these issues, we design a novel unlearning mech-
II. S YSTEM M ODEL AND P ROBLEM F ORMULATION A. System Model We consider a Mobile Edge Computing (MEC) system that comprises distributed edge devices deployed at different geographical locations. Edge devices, including IoT sensors, mobile phones, and vehicular units, are connected to edge servers via a wireless federated learning paradigm with limited bandwidth and varying channel conditions. In this system, edge devices collect local data and collaboratively train a model while maintaining data locality and privacy. Fig. 1 illustrates the process of client unlearning in the MEC system. Assume that there are N clients and a central server, where N ≥ 1 denotes the total number of clients. Each client is indexed by n ∈ {1, 2, . . . , N }. Client n is an edge device that n maintains a local dataset Dn = {(xm,n , ym,n )}M m=1 , where xm,n ∈ X represents the m-th data sample, ym,n ∈ Y denotes the corresponding label, m ∈ {1, 2, . . . , Mn } is the sample index, and Mn = |Dn | is the size of client n’s dataset. The
MEC Edge Server
FL Server
Yes Deletion? No
Store all clients' parameters
Unlearn Client K Stop Request
Mobile Devices
anism. Our contributions made in this paper are summarized as follows: • Unlike existing FU approaches that ignore temporal features, we introduce AoI into FU to prioritize parameter modifications based on information freshness. We design SCALE, a dual-level Sensitivity-aware Client unlearning via Adaptive Layer agE-of-information framework combining historical contribution-based layer sensitivity analysis with reinforcement learning-driven adaptive parameter sparsification. • We develop a theoretical framework for SCALE by demonstrating the convergence properties and acceleration advantages of SCALE. The analytical results p show that the designed framework achieves an O( L/|Ls |) convergence advantage over the uniform parameter modification framework, as formal proofs demonstrate. • In sharp contrast to existing studies that rely on simulation, we implement a hardware testbed intergrating USRP software-defined radios as the communication backhaul and MentorPi robotic vehicles as FU clients to validate SCALE. Extensive evaluations across three neural network architectures (LeNet, MobileNetV3, ResNet18) and three unlearning scenarios (client, class, and sample unlearning) demonstrate that SCALE outperforms stateof-the-art baselines in both unlearning effectiveness and communication efficiency. The rest of this paper is organized as follows. Section II presents the system model for MEC-based federated learning and formulates the unlearning problem with different cases. Section III introduces our proposed SCALE framework, including the Historical Contribution-based Layer Sensitivity Analysis and AoI-driven adaptive sparsification algorithm. Section IV provides a comprehensive theoretical analysis with convergence properties and acceleration advantages. Section V evaluates the framework through extensive experiments and real-world tests. Finally, Section VI concludes the paper.
Client 1
Client 2
Client K ...
... D1
D2
DK
Fig. 1: Client unlearning in the MEC system. SN global dataset is D = n=1 Dn with total size M = |D| = PN n=1 Mn . We assume that the global model consists of L layers, where layer l ∈ {1, 2, . . . , L} contains parameters Wl ∈ Rdl with dimension dl . The complete global model PLparameters are Θ = {W1 , W2 , . . . , WL } ∈ Rd , where d = l=1 dl . We define the local loss function for client n as: Mn 1 X ℓ(Θ; xm,n , ym,n ). Fn (Θ) = Mn m=1
(1)
We formulate the global loss function as: F (Θ) =
N X Mn n=1
M N
Fn (Θ) M
n 1 XX = ℓ(Θ; xm,n , ym,n ), M n=1 m=1
(2)
where the contribution of each client is weighted by its data n proportion M M , and ℓ(·) is the per-sample loss function. Training proceeds over T communication rounds, where t ∈ {1, 2, . . . , T } denotes the round index. At round t, the server selects a subset of clients C[t] ⊆ {1, 2, . . . , N } to participate in training, where C[t] denotes the set of selected clients in round t. The server broadcasts the current global model Θ[t] to these selected clients. Each participating client n ∈ C[t] initializes its local model with the global model Θn,0 [t] = Θ[t] and performs E local updates, where e ∈ {0, 1, . . . , E − 1} denotes the local step index: Θn,e+1 [t] = Θn,e [t] − η∇Fn (Θn,e [t]),
(3)
where η > 0 is the learning rate. After local training, each client n ∈ C[t] uploads its final local model Θn,E [t] to the server. The server aggregates these updates via weighted averaging: Θ[t + 1] =
Mn
X P n∈C[t]
n′ ∈C[t]
M n′
Θn,E [t].
(4)
L1
W1
L2
W2
WN
Server Layer
L2
LN
Server Layer
L1
L2
LN ...
Client K Layer
...
L1
L2
LN
Sensitivity Scoring:
Sparsification Execution: Parameter Modification Timestamp Update
Select Top-M sensitive layers
W1
L2
W2
LN
WN
Server Layer
4
Rank by sensitivity scores:
L1
Li
Local-Global Comparison
2
AoI (Information Freshness) Parameter Statistics
Selected Layer Set
3 Layer Selection:
Original Parameter Matrix PPO Agent State Observation:
1
Weighted Combination Generate sensitivity scores
...
Local Model (Client K)
Selected Layers
...
Client Side
2 L1
...
Layer Selection & Grouping
1 Client 1 Layer
Lj Wp
Lp
Parameter Matrix
Li
Wj
Lj
Unlearning Layer ID
Parameter Alignment & Distributional Impact
Wi
...
LN
Sparse
Li
...
...
Global Model
AoI-Driven Adaptive Sparsification
...
Server Side
Sensitive Layer Identification
Lj Lp
Lp
Sparse Parameter Matrix Updated Layers Action Sampling: Layer Selection Sparsification Ratio
3
Reward Calculation: Forgetting Effectiveness Information Freshness
Wi
Target Client Information
Wj Wp
Update Sensitive Local Model Layer (Remember Client K)
Parameter Matrix
Local Unlearning
Update Local Model
Fig. 2: An overview of the SCALE.
B. Problem Formulation Clients can request the removal of their data from the trained global model, which creates challenges due to distributed data ownership, historical contributions over training rounds, and the computational costs associated with full retraining. Assume an unlearning request is specified by U = (Cu , Du , τ ), where S Cu ⊆ {1, 2, . . . , N } is the set of requesting clients, Du = n∈Cu Du,n is the data to be forgotten with Du,n ⊆ Dn , and τ ∈ {client, class, sample} specifies the unlearning granularity. Specifically, we consider three unlearning cases: Client unlearning (τ = client): Client n ∈ {1, 2, . . . , N } requests complete withdrawal from the federation. We set Cu = {n}, Du,n = Dn , and Du,m = ∅ for all m ̸= n; • Class unlearning (τ = class): Client n requests forgetting all samples belonging to specific classes Yu ⊆ Y. We set Cu = {n}, Du,n = {(x, y) ∈ Dn : y ∈ Yu }, and Du,m = ∅ for all m ̸= n; and • Sample unlearning (τ = sample): Client n requests forgetting specific samples Du,n ⊂ Dn while remaining in the federation. We set Cu = {n} and Du,m = ∅ for all m ̸= n. •
SNNote that after unlearning, the remaining PNdataset is Dr = (D \ D ) with size M = |D | = n u,n r r n=1 Mr,n , where n=1 Mr,n = |Dn \ Du,n | is the remaining data size for client n. Our objective is to obtain an unlearned model Θ′ that approximates the ideal retrained model Θ∗ obtained by training from scratch on the remaining data: Θ∗ = arg min F r (Θ) = arg min Θ
Θ
N X Mr,n
Mr n=1
Fnr (Θ),
(5)
P where Fnr (Θ) = M1r,n (x,y)∈Dn \Du,n ℓ(Θ; x, y) is the local loss on client n’s remaining data. III. T HE PROPOSED FRAMEWORK : SCALE To address the above problem, we propose a dual-level FU framework that integrates historical contribution analysis with AoI-aware sparsification to achieve efficient and effective client data removal in mobile edge computing systems. Our approach, coined Sensitivity-aware Client unlearning via Adaptive Layer agE-of-information (SCALE), first identifies the most sensitive layers through historical contribution analysis, then performs fine-grained parameter sparsification guided by information freshness to balance unlearning effectiveness with system responsiveness. The workflow of SCALE is illustrated in Fig. 2. A. Sensitive Layers Identification To efficiently perform unlearning for client n, we must first identify which layers in the global model have been most influenced by the data of this client. Our key observation is that different layers exhibit varying degrees of sensitivity to individual clients; local patterns of client n may heavily influence some layers, while others remain relatively unaffected. We quantify this layer-wise sensitivity by comparing the parameters of each layer between client n’s local model and the global model from Eq. (4). To be specific, we analyze two complementary aspects: parameter alignment (how similar the client’s layer parameters are to the global layer parameters) and distributional impact (how much the global layer would change if client n’s contribution were removed). 1) Parameter Alignment: For each layer l ∈ {1, 2, . . . , L}, we first measure how well the local model parameters of client n align with the corresponding global model parameters. Let
Wl,n ∈ Rdl denote client n’s parameters for layer l, and Wl ∈ Rdl denote the global model’s parameters for the same layer from Θ = {W1 , W2 , . . . , WL }, where dl is the number of parameters in layer l. We calculate the Pearson correlation coefficient [48] by ρl,n = Corr(Wl,n , Wl ),
(6)
where Corr(·, ·) denotes the Pearson correlation function. The alignment-based sensitivity is thus: 1 a Sl,n = − log(1 − ρ2l,n ). 2
(7)
Note that a higher correlation value indicates that the parameters of client n are well-aligned with the global layer. In this case, the parameters of client n have a strong influence on the global model. 2) Distributional Impact: Parameter alignment alone may not capture the full picture of client influence, as a client can exhibit moderate alignment but still cause significant changes when removed. To resolve this, we measure global layer parameters without client n’s contribution: PN j̸=n,j=1 Mj Wl,j , (8) Wl,−n = PN j̸=n,j=1 Mj where Mj is the number of data samples at client j as defined in Eq. (2), and Wl,j denotes client j’s parameters of layer l. We then measure the distributional difference between the current global layer and the hypothetical layer without client n using Kullback–Leibler (KL) divergence [49]: d Sl,n =
dl X i=1
Pi log
Pi , Qi
(9)
where Pi and Qi are the normalized probability values derived from Wl and Wl,−n respectively for the i-th parameter. The KL divergence quantifies how much the parameter distribution would change if client n were excluded, with larger values indicating greater distributional impact. 3) Weighted Score: We combine both measures to obtain a comprehensive sensitivity score for each layer: a d Sl,n = λSl,n + (1 − λ)Sl,n ,
(10)
where λ ∈ [0, 1] is a hyperparameter balancing the importance of parameter alignment versus distributional impact. 4) Sensitive Layer Identification: We rank all layers by their sensitivity scores and select the top-M layers: Ls = {l1 , l2 , . . . , lM } such that Sl1 ,n ≥ Sl2 ,n ≥ · · · ≥ SlM ,n ,
(11)
where M is determined based on computational budget and unlearning requirements. The identified sensitive layers Ls are the primary targets for sparsification to effectively remove the knowledge associated with client n.
B. AoI-Driven Adaptive Sparsification Given the identified sensitive layers Ls from Eq. (11), we need to determine which specific parameters within these layers should be modified for effective unlearning. Our trick is to prioritize parameters based on their information freshness, parameters with higher AoI values contain relatively stale information, making them safer targets for sparsification while achieving effective removal of historical influence of client n. We model this as a sequential decision process where we adaptively select parameter groups to sparsify based on their AoI from Eq. (32) and structural properties. Such a formulation allows us to make informed decisions about which parameters to modify while optimizing information freshness. 1) State Space: Each sensitive layer l ∈ Ls is partitioned into Gl parameter groups for fine-grained control. The system state at time t captures both temporal and structural information: H[t] = {Al,j [t], µl,j [t], σl,j [t]}, (12) where Al,j [t] is the AoI for P parameter group j in layer l from Eq. (32), µl,j [t] = |W1l,j | m∈Wl,j Wl,m,j [t] is the mean parameter value of group j) with |Wl,j | denoting the group q (l, P 1 2 size, and σl,j [t] = m∈Wl,j (Wl,m,j [t] − µl,j [t]) is |Wl,j | the standard deviation of parameters in group (l, j), providing structural information about parameter distributions. 2) Action Space: At each decision step, we choose: A[t] = (ℓ[t], G[t], s[t]),
(13)
where ℓ[t] ∈ Ls is the target sensitive layer to modify, G[t] ⊆ {1, . . . , Gℓ[t] } specifies the selected parameter groups within layer ℓ[t], and s[t] ∈ [0, 1] is the sparsification ratio determining the fraction of parameters to remove from the selected groups. 3) Reward Function: The reward function combines two key components–forgetting effectiveness reward and information freshness reward: R[t] = wf Rf [t] + wc Rc [t],
(14)
where wf , wc ≥ 0 are weighting coefficients controlling the relative importance of each objective. The forgetting effectiveness reward is formulated as: X Sℓ[t],n Rf [t] = s[t], (15) maxl∈Ls Sl,n j∈G[t]
where Sℓ[t],n is the sensitivity score of the chosen layer from Eq. (10), normalized by the maximum sensitivity score across all sensitive layers. The information freshness reward is defined as: Aℓ[t],j [t] 1 X Rc [t] = s[t], (16) |G[t]| maxl,i Al,i [t] j∈G[t]
where |G[t]| is the number of selected parameter groups, and maxl,i Al,i [t] is the maximum AoI across all parameter groups for normalization.
Algorithm 1 SCALE: Sensitivity-aware Client unlearning via Adaptive Layer agE-of-information Input: Global model Θ = {W1 , . . . , WL }, unlearning request U = (Cu , Du , τ ), target client n ∈ Cu , hyperparameters λ, wf , wc , sensitive layers M . 1: Initialize PPO policy πθ , value function Vϕ , buffer B ← ∅; 2: ▷ Phase I: Sensitive Layers Identification 3: for Layer l = 1, 2, . . . , L do a 4: Compute ρl,n = Corr(Wl,n , Wl ) and Sl,n = − 12 log(1 − ρ2l,n ); P d j̸ =n Mj Wl,j P 5: Compute Wl,−n = and Sl,n = j̸=n Mj Pdl Pi P log ; i=1 i Qi a d 6: Compute Sl,n = λSl,n + (1 − λ)Sl,n ; 7: end for 8: Select Ls = {l1 , . . . , lM } with top-M sensitivity scores; 9: for t = 1, 2, . . . , T do 10: ▷ Phase II: AoI-Driven Sparsification 11: for step = 1, 2, . . . , Tcollect do 12: Partition layers in Ls into parameter groups; 13: Update Al,j [t] = t − Tl,j and observe H[t] = {Al,j [t], µl,j [t], σl,j [t]}; 14: Sample A[t] = (ℓ[t], G[t], s[t]) ∼ πθ (·|H[t]); P Sℓ[t],n 15: Compute Rf [t] = j∈G[t] maxl∈Ls Sl,n s[t], Rc [t] = P Aℓ[t],j [t] 1 j∈G[t] maxl,i Al,i [t] s[t]; |G[t]| 16: Calculate R[t] = wf Rf [t] + wc Rc [t]; 17: Apply Θ′ ← Sparsify(Θ′ , ℓ[t], G[t], s[t]) and store (H[t], A[t], R[t], H[t + 1]) in B; 18: end for 19: if |B| ≥ batch size then 20: Compute advantages  from B; 21: for epoch = 1, 2, . . . , Kepoch do 22: Update πθ ← πθ + ηθ ∇θ LCLIP (θ) and Vϕ ← Vϕ − ηϕ ∇ϕ ∥Vϕ (H) − R∥2 ; 23: end for 24: Clear buffer B ← ∅; 25: end if 26: end for 27: return Θ∗
IV. T HEORETICAL A NALYSIS A. Assumptions and Preliminaries Assumption 1. For each layer l ∈ {1, 2, . . . , L}, the global parameters can be decomposed by: X Wl = αl,n Wl,n + αl,j Wl,j + ξl , (17) j̸=n
where αl,n = PNMnM represents client n’s data proportion, k k=1 and ∥ξl ∥2 ≤ β for some small aggregation error bound β > 0. Assumption 2. For parameter group (l, j) with Age of Information Al,j [t] = t − Tl,j where Tl,j is the timestamp of the most recent update, the unlearning effectiveness satisfies: Ul,j (Al,j [t]) = γ0 + γ1 Al,j [t] + ϵl,j ,
(18)
where γ0 > 0 is the baseline unlearning effectiveness, γ1 > 0 indicates that higher AoI improves unlearning effectiveness, and E[ϵl,j ] = 0. Assumption 3. The sensitivity scores satisfy Sl,n ∈ [0, Smax ] for all l ∈ {1, 2, . . . , L} and n ∈ {1, 2, . . . , N }, where Smax is the theoretical upper bound determined by maximum correlation and divergence values. B. Layer Sensitivity Analysis Theorem 1. Under Assumption 1, the combined sensitivity score from Eq. (10) satisfies: Sl,n ≥
2 λαl,n 2 ) + (1 − λ)DKL (Pl ∥Ql ) − O(β), 2(1 − αl,n
(19)
where Pl and Ql are the normalized parameter distributions of Wl and Wl,−n respectively. The overall unlearning procedure of the SCALE approach is illustrated in Algorithm 1. The computational complexity of Algorithm 1 consists of five functional modules. The initialization module establishes Proximal Policy Optimization (PPO) networks πθ , Vϕ and buffer with O(|θ| + |ϕ|) complexity. The sensitive layers identification module computes parameter alignment and distributional impact for all L layers with O(L · d¯ + L log L) complexity, where d¯ is the average layer dimension. The AoI update module partitions M sensitive layers into parameter groups and updates age information with O(M · Ḡ) complexity, where Ḡ is the average number of groups per layer. The sparsification module applies pruning to selected parameter groups with O(|G[t]| · d¯g ) complexity, where d¯g is the average group size. The PPO update module optimizes policies with O(Kepoch · B · (|θ| + |ϕ|)) complexity. The overall complexity is thus O(L log L + Ttotal · (Tcollect · M · Ḡ · d¯g + Kepoch · B · (|θ| + |ϕ|))). Note that the dominant ¯ and iterative factors are sensitive layer identification O(L · d) ¯ sparsification O(Ttotal · Tcollect · M · Ḡ · dg ).
Proof. For parameter alignment, when ρl,n ≈ αl,n : 2 ρ2l,n αl,n 1 a Sl,n = − log(1 − ρ2l,n ) ≥ ≈ 2 ) . (20) 2 2(1 − ρ2l,n ) 2(1 − αl,n
For distributional impact, let Pl and Ql be the normalized parameter distributions derived from Wl and Wl,−n from Eq. (8). The KL divergence DKL (Pl ∥Ql ) quantifies the distributional shift when removing client n’s contribution. The error term O(β) follows from Assumption 1. Corollary 1. The sensitivity analysis correctly ranks layers by client influence, with top-M selected layers Ls satisfying: X l∈Ls
αl,n ≥ (1 − δ)
L X
αl,n ,
(21)
l=1
where δ = L−M represents the fraction of excluded layers. L
TABLE I: Parameter settings.
C. AoI-Driven Convergence Analysis Theorem 2. Under Assumptions 1-3, the AoI-driven sparsification achieves unlearning error: ∥Θ′ − Θ∗ ∥2 ≤
Gl C1 X 1 X s[t]2 , |Ls | Gl j=1 γ0 + γ1 Al,j [t]
(22)
l∈Ls
where C1 = 2Smax is a constant depending on the sensitivity upper bound. Proof. For each parameter group (l, j), the expected parameter change is: E[∥∆Wl,j ∥22 ] = s[t]2 · Ul,j (Al,j [t]) ≤
s[t]2 . (23) γ0 + γ1 Al,j [t]
The reward function from Eq. (14) prioritizes high-AoI groups. Summing over all selected groups and applying MDP optimality completes the proof. Theorem 3. The dual-level framework achieves faster convergence compared to uniform parameter modification: ∥Θ′uni − Θ∗ ∥ ≥ ∥Θ′dua − Θ∗ ∥
s
Ā L · , |Ls | Āsen
(24)
where Āsen is the average AoI of sensitive layer parameters and Ā is the global average AoI. Proof. For uniform parameter modification, the unlearning effort is distributed equally across all L layers: ∥Θ′uni − Θ∗ ∥22 =
L X
∥Wl′ − Wl∗ ∥22
Parameter Number of clients Local training epochs Global communication rounds Learning rate Dirichlet parameter Parameter groups per layer Sensitivity balance factor PPO episodes PPO epochs Clip range Discount factor GAE lambda Learning rate (actor/critic)
Symbol
Value
N E T η α Gl λ − K ϵ γ
100 10 100 0.001 1.0 8 0.5 800 10 0.2 0.99 0.95 3.0 × 10−4
λGAE ηa /ηc
Taking the ratio with equal total budget suni = sdua = s, the above equation becomes: PL 1 PGl 1 ∥Θ′uni − Θ∗ ∥22 |Ls |2 j=1 γ0 +γ1 Al,j [t] l=1 Gl = · P P Gl 1 1 ∥Θ′dua − Θ∗ ∥22 L2 l∈L j=1 s
Gl
γ0 +γ1 Al,j [t]
L · Ā1−1 |Ls |2 = · 1 L2 |Ls | · −1
Āsen
L Āsen |Ls | Āsen · · = , = L |Ls | Ā Ā (27) PL 1 1 −1 where Ā−1 = = and Ā sen l=1 γ0 +γ1 Al,j [t] L P 1 1 . l∈Ls γ0 +γ1 Al,j [t] |Ls | Since sensitive layers are selected based on high client influence and typically have higher AoI values, Āsen ≥ Ā. Taking the square root and incorporating the layer selection advantage, we thus have: s ∥Θ′uni − Θ∗ ∥ Ā L ≥ · . (28) ′ ∗ ∥Θdua − Θ ∥ |Ls | Āsen
l=1
=
Gl L X 1 X l=1
s2uni 2 Gl j=1 L (γ0 + γ1 Al,j [t])
(25) V. P ERFORMANCE E VALUATION
Gl L s2 X 1 X 1 = uni , 2 L Gl j=1 γ0 + γ1 Al,j [t] l=1
A. Experimental Setup
where suni is the total sparsification budget distributed uniformly. For our dual-level approach focusing on sensitive layers Ls : ∥Θ′dua − Θ∗ ∥22 =
X
∥Wl′ − Wl∗ ∥22
l∈Ls
=
Gl X 1 X s2dua Gl j=1 |Ls |2 (γ0 + γ1 Al,j [t])
l∈Ls
Gl s2dua X 1 X 1 = . |Ls |2 Gl j=1 γ0 + γ1 Al,j [t] l∈Ls
(26)
The experimental setup follows a systematic approach to validate both the effectiveness of SCALE and the impact of information freshness on forgetting performance. Table I summarizes the key parameters used throughout our experiments. More specifically, we evaluate SCALE under varying federation scales with 100 client , and test different levels of data heterogeneity using Dirichlet parameter α = 0.3. The Proximal Policy Optimization (PPO) algorithm is configured with standard hyperparameters, while the reward function balances forgetting effectiveness and information freshness with weights wf and wc , respectively. We chose FashionMNIST as our primary dataset due to its complexity and suitability for federated (un)learning evaluation. We additionally adopt CIFAR-100 (100 fine-grained classes) to test generalizability
Client 1 Request Unlearning
Client 2 Server
Request Unlearning
Client 3
Application scenarios USRP 2955 RF
Tx Rx Channel
1
GNU Radio 2
Cable
USRP 2900 1
USRP 2900
RF 1 RF 2
USRP 2955 3
2
USRP 2900
RF 3
3
USRP 2900
Control PC
Client
Client
Client
USRP System Connection
Unlearning Client 1
2
3
Fig. 3: Testbed setup for real-time MEC FU validation.
on a harder task. The choice of three distinct model architectures (MobileNetV3, ResNet18, and LeNet) allows us to evaluate the adaptability of SCALE with different computational complexities. To validate SCALE in practical MEC systems, we implement a testbed using software-defined radios as the communication backhaul for model distribution and synchronization, as shown in Fig. 3. The testbed consists of one USRP 2955 device serving as the communication gateway for the edge server and three USRP 2900 devices deployed as communication interfaces for mobile clients. The USRP 2955 is configured with a transmit power of 20 dBm and operates with a 100 MHz bandwidth, while each USRP 2900 is equipped with dual RF channels, supporting a maximum data rate of 10 Mbps. Each USRP 2900 is mounted on a MentorPi robotic vehicle, which integrates an ARM Cortex-A72 quad-core processor running at 1.5 GHz with 4GB RAM for local model training. A Control PC orchestrates the entire system through GNU Radio, which manages the wireless communication protocols and coordinates the federated learning workflow. The communication operates in the 2.4 GHz frequency band, using OFDM and QPSK modulation to transmit model parameters between clients and the server. During testing, three MentorPi clients first collaborate with the edge server to perform federated learning multiple communication rounds, with the USRP devices handling the transmission of model weights. When an unlearning request is triggered for a target client, the edge server executes the SCALE algorithm, performing sensitive-layer identification and AoIdriven adaptive sparsification to remove the target data of the client from the global model. After the unlearning process converges at the server side, the resulting sparsified model is
pruned and manually deployed to the MentorPi clients for edge inference. We evaluate the unlearning performance directly on these resource-constrained edge devices by measuring the remaining accuracy and forgetting accuracy to validate the practical effectiveness of SCALE under hardware constraints. B. Performance Metrics We evaluate the federated unlearning performance in terms of the following metrics. Remaining Accuracy (RA): This metric reflects whether the unlearned model maintains high accuracy on the remaining data: X 1 RA = I[fΘ′ (x) = y], (29) Mr (x,y)∈Dr
Forgetting Accuracy (FA): This metric indicates whether the unlearned model demonstrates poor performance on erased data: X 1 FA = I[fΘ′ (x) = y], (30) Mu (x,y)∈Du
where fΘ′ (x) = arg maxc∈Y P (y = c|x; Θ′ ) is the prediction function, I[·] is the indicator function that returns 1 if the condition is true and 0 otherwise, Mr = |Dr | and Mu = |Du | are the sizes of remaining and unlearned datasets, respectively. Forgetting Rate: This metric measures the forgetting efficiency through prediction confidence degradation [50]: X PΘ′ (y|x) 1 , (31) FR = 1 − Mu PΘ (y|x) (x,y)∈Du
where PΘ (y|x) is the prediction confidence of the original model Θ for the true label y given input x, PΘ′ (y|x) is the prediction confidence of the unlearned model Θ′ for the same
700
750
650
650
700
600
600
550
550 500
500
LeNet MobileNetV3 ResNet18
450 400
650
Reward
Reward
Reward
700
0
200
400
600
400
800
0
200
400
600
550 500
LeNet MobileNetV3 ResNet18
450
600
LeNet MobileNetV3 ResNet18
450 400
800
0
200
(a) Client Unlearning Reward.
400
600
800
Episode
Episode
Episode
(b) Class Unlearning Reward.
(c) Sample Unlearning Reward.
Fig. 4: PPO training convergence for different unlearning scenarios.
0.16
Class Unlearning Client Unlearning Sample Unlearning
0.5
0.12 0.1
1 0.8 0.6
0.08
0.4
0.06
0.2 0
1
2
3
4
Sensitivity
1.2
0.14
0.04
0.6
Class Unlearning Client Unlearning Sample Unlearning
1.4
Sensitivity
Sensitivity
1.6
Class Unlearning Client Unlearning Sample Unlearning
0.18
0
0.3 0.2 0.1
0
Layer Index
(a) LeNet layer sensitivity.
0.4
5
10
15
Layer Index
(b) MobileNetV3 layer sensitivity.
0
0
10
20
30
40
Layer Index
(c) ResNet18 layer sensitivity.
Fig. 5: Layer sensitivity analysis over different model architectures and unlearning scenarios. true label, and Mu = |Du | is the size of the unlearned dataset. A higher F R value indicates more effective forgetting as the model’s confidence on the target data decreases. Communication Overhead: For fine-grained parameter manipulation, each layer l into Gl is decomposed into subparameter groups, where j ∈ {1, 2, . . . , Gl } denotes the subgroup index. Each sub-group contains parameters Wl,j that can be independently modified. The Age of Information (AoI) for parameter sub-group j in layer l at time t is defined as [12]: Al,j [t] = t − Tl,j ,
(32)
where Tl,j is the timestamp of the most recent update for subgroup j in layer l. The communication overhead is the combination of transmission costs with information freshness: min
C(Θ → Θ′ ) = α · Ct + β · E[Ag [t]],
(33)
where Ct is the total communication cost, Ag [t] = PL PGl P1 L l=1 j=1 Al,j [t] is the global system AoI, M M G l l=1 is the number of layers selected for unlearning, and α, β > 0 are weighting coefficients. Layer Sensitivity: This metric examines whether the selected top-M sensitive layers from Eq. (11) effectively capture the most client-influenced components.
C. Results Fig. 4 shows the PPO training over three unlearning scenarios. All models achieve stable convergence with consistent reward improvements. Notably, client unlearning exhibits the most stable convergence with minimal oscillations, while sample unlearning shows higher volatility due to its fine-grained nature. LeNet demonstrates the fastest initial convergence, whereas ResNet18 exhibits a distinctive two-phase pattern with rapid initial gains followed by fine-tuning stabilization. Fig. 5 presents a layer-wise sensitivity analysis across three model architectures (LeNet, MobileNetV3, ResNet18) under three unlearning scenarios (client unlearning, class unlearning, and sample unlearning). Remember that higher sensitivity values indicate that the corresponding layer contains parameters that are more significantly affected by the client-server data exchange and target information, thus requiring prioritized sparsification to achieve effective unlearning while minimizing impact on the remaining model performance. In particular, Fig. 5a reveals that LeNet exhibits the highest sensitivity in the initial convolutional layers; Fig. 5b demonstrates that MobileNetV3 shows localized sensitivity peaks within the intermediate depthwise separable convolution layers; and Fig. 5c illustrates that ResNet18 displays a more distributed sensitivity pattern across multiple residual layers. Table II presents comprehensive comparisons between SCALE and the state-of-the-art methods, including Retrain [11], FedEraser (F-E) [51], Projected Gradient Ascent
TABLE II: Performance comparison of SCALE on FashionMNIST and CIFAR-100 across three unlearning scenarios. Metrics are reported as value (∆), where ∆ is the gap with Retrain baseline. Underlined values denote SCALE results and bold values denote baseline method results. ↑ indicates higher is better, ↓ indicates lower is better. RA (↑) Model
Retrain
F-E
PGA
FA (↓) FUSED
SCALE
Retrain
F-E
PGA
Avg. AoI (s) (↓) FUSED
SCALE
Retrain
F-E
PGA
FUSED
SCALE
Client Unlearning FashionMNIST LeNet 0.92 MobileNetV3 0.92 ResNet18 0.96
0.74 (0.18) 0.78 (0.14) 0.99 (0.07) 0.82 (0.10) 0.81 (0.11) 0.83 (0.09) 0.81 (0.11) 0.86 (0.06) 0.81 (0.15) 0.84 (0.12) 0.83 (0.13) 0.87 (0.09)
0.00 0.01 0.04
0.12 (0.12) 0.09 (0.09) 0.00 (0.00) 0.00 (0.00) 0.11 (0.10) 0.12 (0.11) 0.04 (0.03) 0.00 (0.01) 0.21 (0.17) 0.18 (0.14) 0.08 (0.04) 0.01 (0.03)
2.59 2.66 6.75
3.00 (0.41) 2.66 (0.07) 1.65 (0.94) 2.01 (0.58) 2.91 (0.25) 2.71 (0.05) 1.74 (0.92) 2.16 (0.50) 7.18 (0.43) 6.82 (0.07) 4.39 (2.36) 5.99 (0.76)
CIFAR-100 LeNet MobileNetV3 ResNet18
0.30 (0.08) 0.32 (0.06) 0.42 (0.04) 0.36 (0.02) 0.51 (0.11) 0.54 (0.08) 0.55 (0.07) 0.58 (0.04) 0.58 (0.13) 0.60 (0.11) 0.61 (0.10) 0.65 (0.06)
0.01 0.02 0.05
0.18 (0.17) 0.14 (0.13) 0.02 (0.01) 0.01 (0.00) 0.16 (0.14) 0.13 (0.11) 0.06 (0.04) 0.02 (0.00) 0.24 (0.19) 0.21 (0.16) 0.10 (0.05) 0.04 (0.01)
2.85 2.93 7.28
3.21 (0.36) 2.93 (0.08) 2.13 (0.72) 2.21 (0.64) 3.18 (0.25) 2.99 (0.06) 1.89 (1.04) 2.34 (0.59) 7.71 (0.43) 7.34 (0.06) 4.71 (2.57) 6.41 (0.87)
0.38 0.62 0.71
Class Unlearning FashionMNIST LeNet 0.90 MobileNetV3 0.91 ResNet18 0.96
0.73 (0.17) 0.76 (0.14) 0.99 (0.09) 0.81 (0.09) 0.78 (0.13) 0.81 (0.10) 0.73 (0.18) 0.87 (0.04) 0.78 (0.18) 0.80 (0.16) 0.73 (0.23) 0.86 (0.10)
0.08 0.04 0.00
0.27 (0.19) 0.20 (0.12) 0.04 (0.04) 0.00 (0.08) 0.17 (0.13) 0.15 (0.11) 0.21 (0.17) 0.02 (0.02) 0.19 (0.19) 0.16 (0.16) 0.08 (0.08) 0.00 (0.00)
2.72 2.79 6.89
3.07 (0.35) 2.64 (0.08) 1.68 (1.04) 2.09 (0.63) 3.14 (0.35) 2.88 (0.09) 1.73 (1.06) 2.19 (0.60) 7.08 (0.19) 6.78 (0.11) 4.52 (2.37) 6.02 (0.87)
CIFAR-100 LeNet MobileNetV3 ResNet18
0.28 (0.08) 0.30 (0.06) 0.40 (0.04) 0.34 (0.02) 0.49 (0.12) 0.52 (0.09) 0.46 (0.15) 0.57 (0.04) 0.55 (0.14) 0.57 (0.12) 0.49 (0.20) 0.62 (0.07)
0.10 0.05 0.01
0.32 (0.22) 0.24 (0.14) 0.05 (0.05) 0.01 (0.09) 0.21 (0.16) 0.19 (0.14) 0.25 (0.20) 0.03 (0.02) 0.23 (0.22) 0.20 (0.19) 0.11 (0.10) 0.01 (0.00)
2.97 3.04 7.42
3.31 (0.34) 2.85 (0.12) 1.60 (1.37) 2.29 (0.68) 3.39 (0.35) 3.13 (0.09) 1.87 (1.17) 2.37 (0.67) 7.62 (0.20) 7.31 (0.11) 4.84 (2.58) 6.47 (0.95)
0.36 0.61 0.69
Sample Unlearning FashionMNIST LeNet 0.89 MobileNetV3 0.88 ResNet18 0.92
0.73 (0.16) 0.74 (0.15) 0.99 (0.10) 0.77 (0.12) 0.76 (0.12) 0.79 (0.09) 0.75 (0.13) 0.82 (0.06) 0.80 (0.12) 0.81 (0.11) 0.54 (0.38) 0.84 (0.08)
0.02 0.01 0.01
0.15 (0.13) 0.13 (0.11) 0.05 (0.03) 0.09 (0.07) 0.18 (0.17) 0.14 (0.13) 0.18 (0.17) 0.10 (0.09) 0.22 (0.21) 0.17 (0.16) 0.13 (0.12) 0.10 (0.09)
2.69 2.81 6.92
3.03 (0.34) 2.75 (0.06) 1.65 (1.04) 2.06 (0.63) 3.29 (0.48) 2.92 (0.11) 1.71 (1.10) 2.68 (0.13) 7.12 (0.20) 6.85 (0.07) 4.52 (2.40) 5.99 (0.93)
CIFAR-100 LeNet MobileNetV3 ResNet18
0.27 (0.07) 0.28 (0.06) 0.39 (0.05) 0.31 (0.03) 0.47 (0.11) 0.50 (0.08) 0.45 (0.13) 0.53 (0.05) 0.54 (0.12) 0.55 (0.11) 0.32 (0.34) 0.58 (0.08)
0.04 0.03 0.02
0.19 (0.15) 0.16 (0.12) 0.07 (0.03) 0.12 (0.08) 0.22 (0.19) 0.18 (0.15) 0.21 (0.18) 0.13 (0.10) 0.26 (0.24) 0.21 (0.19) 0.16 (0.14) 0.13 (0.11)
2.89 3.07 7.46
3.27 (0.38) 2.96 (0.07) 1.53 (1.36) 2.24 (0.65) 3.54 (0.47) 3.18 (0.11) 2.21 (0.86) 2.86 (0.21) 7.66 (0.20) 7.39 (0.07) 4.15 (3.31) 6.43 (1.03)
0.34 0.58 0.66
40 1.1
0.6
1.05 1
0.4
30
35
0
0
10
20
43 42 41 40 39
30
20
30
40
Client Unlearning Class Unlearning Sample Unlearning
0.2
80
30
40
Communication Round
(a) Sum AoI with LeNet.
50
Sum AoI
0.8
Sum AoI
Sum AoI
1
35
0
0
10
20
40
40
Client Unlearning Class Unlearning Sample Unlearning
10
90 88 86 84 82 30
60
30
40
50
0
40
Client Unlearning Class Unlearning Sample Unlearning 0
10
Communication Round
(b) Sum AoI with MobileNetV3.
35
20
20
30
40
50
Communication Round
(c) Sum AoI with ResNet18.
Fig. 6: Sum AoI comparison for different unlearning scenarios across model architectures.
(PGA) [52], and FUSED [53] across three unlearning scenarios on both FashionMNIST and CIFAR-100. As observed, SCALE demonstrates superior performance compared to F-E and PGA while maintaining good efficiency compared to FUSED. For client unlearning on FashionMNIST, SCALE achieves RA values of 0.82, 0.86, and 0.87 across LeNet, MobileNetV3, and ResNet18, substantially outperforming F-E (0.74, 0.81, 0.81) and PGA (0.78, 0.83, 0.84). The Average AoI for SCALE (2.01s, 2.16s, 5.99s) is in the between Retrain (2.59s, 2.66s, 6.75s) and FUSED (1.65s, 1.74s, 4.39s), while F-E exhibits the highest overhead at 3.00s, 2.91s, and 7.18s. Similar observations can be made in class and sample unlearning scenarios, where SCALE consistently maintains better RA and competitive AoI compared to F-E and PGA. Although
FUSED occasionally demonstrates better RA and lower AoI, SCALE achieves a more balanced trade-off between unlearning effectiveness and model utility across diverse architectures, which makes it more suitable for real-world deployment. The same pattern appears on CIFAR-100: SCALE stays within 0.02–0.08 of Retrain in RA across all scenarios, while FUSED collapses on ResNet18 sample unlearning, where its RA drops to 0.32 while Retrain holds at 0.66. D. Testbed Results and Discussion Fig. 6 illustrates the temporal evolution of Sum AoI to the overall system AoI for each communication round throughout the unlearning process on three architectures. The results demonstrate consistent optimization dynamics under SCALE
LeNet (RA)
MobileNetV3 (RA)
ResNet18 (RA)
0.05
0.1
0.15
0.2
MobileNetV3 (FA)
0.25
0
0.3
0.05
0.1
0.15
0.2
Forgetting Accuracy (FA) 0.25
0.3
0
0.5
0.5
0.5
0.3
0.3
0.3
0.1
0.1
0.1
0.4
0.5
0.6
0.7
ResNet18 (FA)
Forgetting Accuracy (FA)
Forgetting Accuracy (FA) 0
LeNet (FA)
0.8
0.4
0.5
0.6
0.7
0.8
0.4
0.05
0.1
0.5
0.15
0.6
0.2
0.25
0.7
Remaining Accuracy (RA)
Remaining Accuracy (RA)
Remaining Accuracy (RA)
(a) Client unlearning.
(b) Class unlearning.
(c) Sample unlearning.
0.3
0.8
Fig. 7: Impact of data heterogeneity on testbed unlearning performance in three scenarios.
implementation, with LeNet achieving the fastest AoI stabilization, followed by MobileNetV3 and ResNet18. From the AoI performance perspective, LeNet demonstrates the lowest communication overhead with 2.013s average AoI, making it the most suitable for deployment in time-sensitive federated environments. Notably, the AoI fluctuations in physical experiments exhibit reduced variance compared to simulation results. This can be attributed to the limited client configuration, which minimizes network congestion and synchronization overhead. Fig. 7 presents the impact of data heterogeneity on unlearning performance across three scenarios in our hardware testbed. To validate SCALE under realistic hardware constraints, we deploy a subset of FashionMNIST images onto MentorPi robotic vehicles and evaluate unlearning performance using pruned model architectures optimized for resource-constrained edge devices. The Dirichlet parameter α controls the degree of non-IID data distribution, where smaller α values indicate higher data heterogeneity among clients. As shown in Fig. 7a, client unlearning performance shows strong dependence on data heterogeneity. Under highly heterogeneous conditions (α = 0.1), LeNet, MobileNetV3, and ResNet18 achieve RA values of 0.56, 0.59, and 0.58, respectively, while FA of 0.19, 0.18, and 0.20 indicate substantial forgetting difficulty with degraded model utility. As α increases to 0.3, RA improves to 0.62-0.67 as FA decreases to 0.14-0.16 across models. At α = 0.5, a more balanced data distribution, all models achieve RA of 0.63-0.68 with FA reduced to 0.14-0.15, which demonstrates substantially improved unlearning effectiveness. ResNet18 shows the most significant improvement, with RA increasing from 0.58 to 0.68 as α increases from 0.1 to 0.5. These observations reveal that balanced data distributions facilitate more effective client removal while preserving model utility, as reduced overlap in client-specific feature representations enables cleaner parameter separation. Fig. 7b shows similar patterns for class unlearning. The consistent improvement across α values validates that uniform class distribution across clients supports more targeted class
removal with minimal impact on remaining knowledge. Note that the sample unlearning results in Fig. 7c exhibit similar heterogeneity sensitivity but with overall degraded performance compared to coarser-grained scenarios. At α = 0.1, LeNet, MobileNetV3, and ResNet18 achieve RA of 0.51, 0.53, and 0.52 with FA of 0.25, 0.26, and 0.24, respectively. Even at α = 0.5, RA only reaches 0.58-0.65 while FA remains at 0.22-0.24, substantially higher than the 0.14-0.15 FA observed in client and class unlearning. ResNet18 demonstrates the largest RA improvement (from 0.52 to 0.65), but its FA shows minimal reduction. This performance gap stems from the finegrained nature of sample removal, in which individual data influences are diffusely distributed across parameter space, making precise localization and elimination significantly more challenging regardless of data distribution characteristics. Results across all three scenarios consistently show that data heterogeneity is a key factor influencing the effectiveness of federated unlearning in MEC environments. Our testbed validation indicates that SCALE achieves stable performance under different levels of heterogeneity on resource-constrained edge devices. In particular, balanced data distributions (α = 0.5) provide the best trade-off between forgetting effectiveness (F A = 0.14 − 0.24) and model utility preservation. VI. C ONCLUSION In this paper, we have presented SCALE: a novel dual-level federated unlearning framework that addresses the issues of low unlearning precision and lack of temporal information freshness in mobile edge computing systems. By integrating historical contribution-based layer sensitivity analysis with Age-of-Information-driven adaptive sparsification, SCALE effectively balances unlearning effectiveness with information freshness while maintaining computational efficiency. Our theoretical analysis demonstrates the convergence properties and acceleration advantages of the proposed framework. At the same time, experiments based on multiple neural network architectures validate the outstanding performance of SCALE compared to existing baselines. The testbed implementation further confirms the practical feasibility of our framework
in real-world MEC environments. We believe that SCALE offers a powerful solution for federated unlearning in MEC that meets stringent regulatory requirements while preserving model utility and system performance. R EFERENCES [1] B. Wu, J. Huang, Q. Duan, L. Dong, and Z. Cai, “Enhancing Vehicular Platooning With Wireless Federated Learning: A Resource-Aware Control Framework,” IEEE Transactions on Networking, pp. 1–16, 2025. [2] R. Chourasia and N. Shah, “Forget Unlearning: Towards True DataDeletion in Machine Learning,” in International Conference on Machine Learning. PMLR, 2023, pp. 6028–6073. [3] N. Xu, Z. Liu, X. Han, Q. Guan, H. Sun, X. Huang, and J. Ma, “An Efficient and Multi-Dimensional Privacy-Preserving Platoon Communication Scheme in Vehicular Networks,” IEEE Transactions on Intelligent Transportation Systems, vol. 26, no. 5, pp. 6831–6847, 2025. [4] B. Wu, Z. Ding, and J. Huang, “RELIEF: Turning Missing Modalities into Training Acceleration for Federated Learning on Heterogeneous IoT Edge,” arXiv preprint arXiv:2604.04243, 2026. [5] L. Dong, J. Huang, and R. W. Heath, “Transformer-Based Dynamic Resource Allocation for Multi-Carrier NOMA Systems,” IEEE Transactions on Cognitive Communications and Networking, vol. 12, pp. 4926– 4941, 2026. [6] N. Romandini, A. Mora, C. Mazzocca, R. Montanari, and P. Bellavista, “Federated Unlearning: A Survey on Methods, Design Guidelines, and Evaluation Metrics,” IEEE Transactions on Neural Networks and Learning Systems, vol. 36, no. 7, pp. 11 697–11 717, 2025. [7] B. Wu, J. Huang, and Q. Duan, “FedTD3: An Accelerated Learning Approach for UAV Trajectory Planning,” in International Conference on Wireless Artificial Intelligent Computing Systems and Applications (WASA). Springer, 2025, pp. 13–24. [8] Z. Ding, J. Huang, Q. Duan, C. Zhang, Y. Zhao, and S. Gu, “A DualLevel Game-Theoretic Approach for Collaborative Learning in UAVAssisted Heterogeneous Vehicle Networks,” in 2025 IEEE International Performance, Computing, and Communications Conference (IPCCC). IEEE, 2025, pp. 1–8. [9] W. Wang, Z. Tian, C. Zhang, A. Liu, and S. Yu, “BFU: Bayesian Federated Unlearning With Parameter Self-Sharing,” in Proc. ACM Asia Conf. on Computer and Communications Security (AsiaCCS), 2023, pp. 567–578. [10] G. Liu, X. Ma, Y. Yang, C. Wang, and J. Liu, “Federaser: Enabling Efficient Client-Level Data Removal From Federated Learning Models,” in Proc. IEEE/ACM 29th Int. Symp. on Quality of Service (IWQoS). IEEE, 2021, pp. 1–10. [11] Y. Liu, L. Xu, X. Yuan, C. Wang, and B. Li, “The Right to Be Forgotten in Federated Learning: An Efficient Realization With Rapid Retraining,” in Proc. IEEE Conf. on Computer Communications (INFOCOM). IEEE, 2022, pp. 1749–1758. [12] B. Wu, J. Huang, and Q. Duan, “Real-Time Intelligent Healthcare Enabled by Federated Digital Twins With AoI Optimization,” IEEE Network, vol. 40, no. 2, pp. 184–191, 2026. [13] B. Wu, Z. Ding, L. Ostigaard, and J. Huang, “Reinforcement LearningBased Energy-Aware Coverage Path Planning for Precision Agriculture,” in 2025 ACM Research on Adaptive and Convergent Systems (RACS). ACM, 2025, pp. 1–8. [14] B. Wu, Z. Cai, W. Wu, and X. Yin, “AoI-aware Resource Management for Smart Health via Deep Reinforcement Learning,” IEEE Access, 2023. [15] Y. Cui, N. Ding, and M. H. Cheung, “AoI-Aware Federated Unlearning for Streaming Data With Online Client Selection and Pricing,” in IEEE INFOCOM 2025–IEEE Conference on Computer Communications. IEEE, 2025, pp. 1–10. [16] C.-C. Xing, Z. Ding, and J. Huang, “A Stochastic GeometryBased Analysis of SWIPT-Assisted Underlaid Device-to-Device Energy Harvesting,” SIGAPP Appl. Comput. Rev., vol. 25, no. 4, p. 18–34, Jan. 2026. [Online]. Available: https://doi.org/10.1145/3787594.3787596 [17] R. R. Depa, Y. Zhang, and D. Xu, “Hybrid Edge Intelligence for RealTime Intrusion Detection in Advanced Traffic Management Systems,” in 2025 IEEE Security and Privacy Workshops (SPW). IEEE, 2025, pp. 352–354.
[18] Z. Ding, J. Huang, and J. Qi, “Learning to Defend: A Multi-Agent Reinforcement Learning Framework for Stackelberg Security Game in Mobile Edge Computing,” in 2026 International Conference on Computing, Networking and Communications (ICNC), 2026, pp. 769– 774. [19] C. Zhou, C. Pan, M. Li, and P. Wang, “Federated Unlearning With Fast Recovery,” IEEE Transactions on Mobile Computing, pp. 1–18, 2025. [20] Y. Zhao, J. Yang, Y. Tao, L. Wang, X. Li, D. Niyato, and H. Vincent Poor, “Exploring Federated Unlearning: Review, Comparison, and Insights,” IEEE Network, pp. 1–1, 2025. [21] X. Sheng, W. Bao, and L. Ge, “Robust Federated Unlearning,” in Proceedings of the 33rd ACM International Conference on Information and Knowledge Management (CIKM), 2024, pp. 2034–2044. [22] T. T. Huynh, T. B. Nguyen, T. T. Nguyen, P. L. Nguyen, H. Yin, Q. V. H. Nguyen, and T. T. Nguyen, “Certified Unlearning for Federated Recommendation,” ACM Transactions on Information Systems, vol. 43, no. 2, pp. 1–29, 2025. [23] W. Zeng, S. Chen, X. Li, and S. Chen, “FU-PA: Federated Unlearning via Parameters Adjustment,” IEEE Transactions on Emerging Topics in Computational Intelligence, pp. 1–12, 2025. [24] M. Ameen, P. Wang, W. Su, X. Wei, and Q. Zhang, “Speed Up Federated Unlearning With Temporary Local Models,” IEEE Transactions on Sustainable Computing, pp. 1–16, 2025. [25] L. Bourtoule, V. Chandrasekaran, C. A. Choquette-Choo, H. Jia, A. Travers, B. Zhang, D. Lie, and N. Papernot, “Machine Unlearning,” in 2021 IEEE Symposium on Security and Privacy (SP). IEEE, 2021, pp. 141–159. [26] J. Wang, S. Guo, X. Xie, and H. Qi, “Federated Unlearning via ClassDiscriminative Pruning,” in Proceedings of the ACM Web Conference 2022, 2022, pp. 622–632. [27] J. Weng, S. Yao, Y. Du, J. Huang, J. Weng, and C. Wang, “Proof of Unlearning: Definitions and Instantiation,” IEEE Transactions on Information Forensics and Security, vol. 19, pp. 3309–3323, 2024. [28] E. Chien, H. Wang, Z. Chen, and P. Li, “Langevin Unlearning: A New Perspective of Noisy Gradient Descent for Machine Unlearning,” Advances in Neural Information Processing Systems (NeurIPS), vol. 37, pp. 79 666–79 703, 2024. [29] J. Huang, B. Wu, Q. Duan, L. Dong, and S. Yu, “A Fast UAV Trajectory Planning Framework in RIS-Assisted Communication Systems With Accelerated Learning via Multithreading and Federating,” IEEE Transactions on Mobile Computing, pp. 1–16, 2025. [30] Z. Huang, X. Cheng, J. Zheng, H. Wang, Z. He, T. Li, and X. Huang, “Unified Gradient-Based Machine Unlearning With Remain Geometry Enhancement,” Advances in Neural Information Processing Systems (NeurIPS), vol. 37, pp. 26 377–26 414, 2024. [31] S. Lin, X. Zhang, W. Susilo, X. Chen, and J. Liu, “GDR-GMA: Machine Unlearning via Direction-Rectified and Magnitude-Adjusted Gradients,” in Proceedings of the 32nd ACM International Conference on Multimedia (ACM MM), 2024, pp. 9087–9095. [32] Z. Pan, Z. Wang, C. Li, K. Zheng, B. Wang, X. Tang, and J. Zhao, “Federated Unlearning With Gradient Descent and Conflict Mitigation,” in Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), vol. 39, no. 19, 2025, pp. 19 804–19 812. [33] B. Wu, Z. Ding, and J. Huang, “PRISM: Exposing and Resolving Spurious Isolation in Federated Multimodal Continual Learning,” arXiv preprint arXiv:2605.01061, 2026. [34] B. Wang, Y. Zi, Y. Sun, Y. Zhao, and B. Qin, “Balancing Forget Quality and Model Utility: A Reverse KL-Divergence Knowledge Distillation Approach for Better Unlearning in LLMs,” in Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 1306–1321. [35] H. Kim, S. Lee, and S. S. Woo, “Layer Attack Unlearning: Fast and Accurate Machine Unlearning via Layer Level Attack and Knowledge Distillation,” in Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), vol. 38, no. 19, 2024, pp. 21 241–21 248. [36] H. Xu, T. Zhu, L. Zhang, W. Zhou, and P. S. Yu, “Update Selective Parameters: Federated Machine Unlearning Based on Model Explanation,” IEEE Transactions on Big Data, vol. 11, no. 2, pp. 524–539, 2025. [37] M. M. U. Zaman, X. Sun, and J. Yao, “Sky of Unlearning (SoUL): Rewiring Federated Machine Unlearning via Selective Pruning,” arXiv preprint arXiv:2504.01705 (arXiv), 2025.
[38] S. Ye, J. Lu, and G. Zhang, “Towards Safe Machine Unlearning: A Paradigm That Mitigates Performance Degradation,” in Proceedings of the ACM on Web Conference 2025, 2025, pp. 4635–4652. [39] Z. Zuo, Z. Tang, K. Li, and A. Datta, “Machine Unlearning Through Fine-Grained Model Parameters Perturbation,” IEEE Transactions on Knowledge and Data Engineering, vol. 37, no. 4, pp. 1975–1988, 2025. [40] B. Wu, Z. Ding, and J. Huang, “A Review of Continual Learning in Edge AI,” IEEE Transactions on Network Science and Engineering, vol. 13, pp. 6571–6588, 2026. [41] H. Zhang, B. Wu, X. Yang, X. Yuan, X. Liu, and X. Yi, “Dynamic Graph Unlearning: A General and Efficient Post-Processing Method via Gradient Transformation,” in Proceedings of the ACM on Web Conference 2025, 2025, pp. 931–944. [42] B. Wu and J. Huang, “Lifecycle-Aware Federated Continual Learning in Mobile Autonomous Systems,” arXiv preprint arXiv:2604.20745, 2026. [43] Z. Ding, B. Wu, and J. Huang, “EASE: Federated Multimodal Unlearning via Entanglement-Aware Anchor Closure,” arXiv preprint arXiv:2605.00733, 2026. [44] N. Ding, Z. Sun, E. Wei, and R. Berry, “Incentivized Federated Learning and Unlearning,” IEEE Transactions on Mobile Computing, pp. 1–17, 2025. [45] Z. Ding, B. Wu, J. Huang, and S. Mao, “Application-Aware Twin-inthe-Loop Planning for Federated Split Learning over Wireless Edge Networks,” arXiv preprint arXiv:2604.26105, 2026. [46] B. Wu, J. Huang, and S. Yu, ““X of Information” Continuum: A Survey on AI-Driven Multi-Dimensional Metrics for Next-Generation Net-
worked Systems,” IEEE Communications Surveys & Tutorials, vol. 28, pp. 5307–5344, 2026. [47] U. Pudasaini, Z. Ding, and J. Huang, “Securing Smart Agriculture with Communication-Efficient Federated Unlearning,” in 2026 IEEE International Conference on High Performance Switching and Routing (HPSR). IEEE, 2026, pp. 1–8. [48] P. Sedgwick, “Pearson’s correlation coefficient,” Bmj, vol. 345, 2012. [49] T. van Erven and P. Harremos, “Rényi Divergence and Kullback-Leibler Divergence,” IEEE Transactions on Information Theory, vol. 60, no. 7, pp. 3797–3820, 2014. [50] Z. Ma, Y. Liu, X. Liu, J. Liu, J. Ma, and K. Ren, “Learn to Forget: Machine Unlearning via Neuron Masking,” IEEE Transactions on Dependable and Secure Computing, vol. 20, no. 4, pp. 3194–3207, 2023. [51] G. Liu, X. Ma, Y. Yang, C. Wang, and J. Liu, “FedEraser: Enabling Efficient Client-Level Data Removal from Federated Learning Models,” in 2021 IEEE/ACM 29th International Symposium on Quality of Service (IWQOS), 2021, pp. 1–10. [52] A. Halimi, S. Kadhe, A. Rawat, and N. Baracaldo, “Federated Unlearning: How to Efficiently Erase a Client in FL?” 2023. [Online]. Available: https://arxiv.org/abs/2207.05521 [53] Z. Zhong, W. Bao, J. Wang, S. Zhang, J. Zhou, L. Lyu, and W. Y. B. Lim, “Unlearning Through Knowledge Overwriting: Reversible Federated Unlearning via Selective Sparse Adapter,” in Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR), 2025, pp. 30 661–30 670.