Zihao Ding
Jun Huang
Liang Dong
Electrical Engr & Computer Science South Dakota State University Brookings, USA [email protected]
Electrical Engr & Computer Science South Dakota State University Brookings, USA [email protected]
Electrical & Computer Engineering Baylor University Waco, USA liang [email protected]
I. I NTRODUCTION
F
EDERATED learning (FL) is increasingly deployed as a long-running network service rather than a one-time training job [1]. A central coordinator trains a model with many clients by exchanging model updates instead of collecting raw data, which is important in edge computing, Internet of Things (IoT), and cross-silo learning systems where data is private, distributed, and costly to move [2]–[4]. Such services already run over mobile and vehicular networks, where clients differ in link quality, compute, and energy [5], [6]. Modern edge agents further turn FL into a self-improving loop: they deploy the current model, interact with the environment, keep verified trajectories, and feed those data back into later rounds [7], [8]. We call this setting a self-improving federated agent network. Keeping raw data local does not settle the privacy obligation of such a network. A data owner may request that a trajectory, a client dataset, or a task source no longer influence the trained model. Regulations such as the General Data Protection Regulation (GDPR) give users a right to be forgot-
ga
l withdraw my consent. Delete my demonstration.
tio n
Server
Ag gre
Abstract—Self-improving federated agent networks keep training after deployment by collecting new trajectories with the current policy and feeding them back into later rounds. This closed loop makes unlearning harder than a one-time model repair. When a data owner requests deletion, the target data may have already shaped later retained trajectories, so retraining or model-side unlearning can leave an influence echo that returns as the network continues to operate. We show that this echo survives retained-data retraining, grows with the amount of forget-shaped retained data, and can be traced from deployment, collection, and aggregation records. To address this problem, we propose MUTE, a Muting Unlearned Trajectories’ Echoes method for reliable deletion in self-improving federated agent networks. MUTE estimates downstream influence from a lightweight server ledger, removes the current residue through a forget-retain update, contains high-influence retained trajectories through quarantine or down-weighting, and audits later behavior to schedule additional erasure under an uplink budget. Experiments on LIBERO with two vision-language-action backbones, three deletion granularities, and a physical Jetson-based edge testbed show that MUTE keeps behavioral leakage and influence regeneration low while preserving task utility and using much less communication than full retraining. Index Terms—Federated unlearning, self-improving networks, reliable data deletion, network privacy, influence tracing, edge agents
Influence Echo
Agent 1
Deploy
t as dc oa Br
arXiv:2607.28829v1 [cs.NI] 30 Jul 2026
When Unlearning Fails: Reliable Data Deletion under Post-Training in Agent Networks
Agent N
Agent 2
Mute Erase Round
Self-improving Loop
Round Unlearning Request
Round
Fig. 1. Influence echo in a self-improving federated agent network. Forget data can affect later retained trajectories through deployment and collection, so deletion must remove both current model residue and future data-side echo.
ten [9]. Federated unlearning (FU) addresses this requirement by removing the effect of requested data from a federated model [10], [11]. Existing FU methods reduce deletion cost through rollback, model editing, or verification [12]–[15], and recent work makes the deletion itself cheap enough for edge deployment [16], [17]. They mainly solve the request-time problem: once a request arrives, they update the current model to behave as if the target data had not been used. This assumption breaks in a self-improving network. The deployed policy shapes what clients collect next, and those new trajectories are reused for later training [18], [19]. If the forget data shaped the policy before deletion, its influence may already have moved into retained data. Retraining on the observed retain set can therefore still train on data causally downstream of the forget set. After operation resumes, the forgotten behavior may return and its membership signal may become visible again [20], [21]. We call this phenomenon the influence echo of forgotten data. As shown in Fig. 1, reliable deletion must remove not only the direct data residue, but also the policy-driven echo that can regenerate later. This leads to our central question: how can a federated network keep a deletion request valid while its agents continue collecting and training on new data?
1 0.7
0.2
, Traj
MiniVLA, Traj
0
MiniVLA, Client
0
MiniVLA, Task
0
3
5
0.6 0.5
0
MiniVLA, Client
0
, Task
MiniVLA, Task
0
2
4
3
5
Rounds after deletion
(a) FSR
6
0
MiniVLA, Client
0
MiniVLA, Task
, Task 0
, Client
0.6 0.4
, Traj
MiniVLA, Traj
, Client
, Client
Task Retrain
0.2
, Task
0.4 1
Trajectory Client
, Traj
MiniVLA, Traj
t-SNE 2
0.4
0.8
IRR
BLI
FSR
0.6
0 1
2
4
6
Rounds after deletion
(b) BLI
Fig. 2. Retraining fails during continued operation. Both forget-set FSR and BLI rise after deletion, showing that self-improvement revives the forgotten behavior.
0
0.2
0.4
0.6
0.8
1
t-SNE 1
Influence fraction
(a) Influence scaling
(b) Influence traceability
Fig. 3. Influence is scalable and traceable. IRR grows with the influence fraction, while t-SNE shows that server-side influence scores are needed to locate forget-shaped retained data.
II. P RELIMINARY R ESULTS AND O BSERVATIONS Prior work addresses only parts of this requirement. FU methods usually assume a fixed retain set and do not model policy-driven future data [22]–[26]. Continual and edge learning preserve useful knowledge across long-running tasks, and lifecycle-aware designs even schedule what to keep and what to drop, but their goal is retention rather than sustained deletion [1], [27], [28]. Machine unlearning methods for large models can suppress target behavior with preference or gradient-based updates [29]–[31], but model-side updates alone do not stop contaminated retained trajectories from reintroducing the behavior. Model-side deletion can also be undone by quantization or relearning attacks [32], [33], so privacy in an operating network is a standing defense rather than a one-time repair [34]. Existing methods lack a unified mechanism that traces downstream influence, removes both model-side and data-side residues, and audits whether the influence reappears during continued operation. To address these limitations, we propose Muting Unlearned Trajectories’ Echoes (MUTE), a reliable deletion method for self-improving federated agent networks. MUTE has three modules. It first replays a lightweight server ledger to estimate how the forget data’s influence propagates through aggregation and policy-driven collection without moving raw trajectories off clients. It then performs influence-aware erasure, where a model-side update removes direct residue and a data-side step quarantines or down-weights high-influence retained trajectories. Finally, it audits behavioral leakage and influence regeneration, and schedules later erasure or containment actions under an uplink budget. The main contributions of this paper are summarized as follows. We formulate reliable deletion for self-improving federated agent networks, where retain data may itself be shaped by forget data and deletion must remain valid after continued operation. We show that retraining can fail, regeneration scales with forget-shaped retained data, and the echo can be traced from network-side records. We develop MUTE to combine influence provenance, influence-aware erasure, and behavioral audit with scheduling. We validate MUTE on LIBERO with two VLA backbones and three deletion granularities, and further check its practicality on a physical Jetson-based edge testbed.
Before formalizing MUTE, we examine unlearning in a selfimproving federated network, where deployed policies keep collecting new trajectories for later training. This closed loop lets the forget data influence retained data, so retraining on retained data may not remove its effect. A. Observation 1: The Echo Survives Retraining Observation 1: Retraining on retained data does not silence the influence echo. As the network continues operating, the forgotten behavior is gradually revived. Does retraining actually remove the data? We build a centralized federated network where clients train a shared policy, deploy it to act, and append collected trajectories to local training. After several rounds, a data owner requests deletion of the forget data. We retrain from scratch on the retained data and continue operating the network. Since residual influence appears through behavior, we measure the forget success rate (FSR), which records how often the forgotten behavior is reproduced, and the behavioral leakage index (BLI), the AUC of a behavioral membership inference attack (MIA) [20]. Figs. 2a and 2b show that both metrics rise after deletion across two backbones, MiniVLA and π0 , and three deletion granularities, trajectory, client, and task. Retraining therefore does not recover the clean forget-free behavior once policydriven collection continues. B. Observation 2: Influence Scaling Observation 2: The influence echo strengthens with the influence fraction. More forget-shaped retained data leads to stronger regeneration after continued learning. What controls how badly deletion fails? We sweep the influence fraction (IF), the share of retained data collected under a policy shaped by the forget data. When IF is zero, collection is exogenous and matches classical deletion. As IF increases, more retained data is generated through the shaped policy. For each IF level, we delete the forget data, retrain on the retained data, continue self-improvement, and measure
the influence regeneration rate (IRR), the fraction of forgotten behavior that returns. Fig. 3a shows that IRR stays near zero under exogenous collection and rises toward one as IF grows. The failure therefore comes from the learn-and-act loop itself, not from a specific deletion algorithm. C. Observation 3: Influence Traceability Observation 3: The echo is downstream of the forget data. It is not always separable in feature space, but can be traced from collection and aggregation records. Why does retained data still carry the influence of the forget data? Under a policy shaped by the forget data, an agent visits different states and records different actions than it would under the counterfactual policy. The collected trajectories therefore carry leftover influence even when they are labeled as retained data. This influence propagates through the collection edge, where a deployed policy shapes future data, and the aggregation edge, where shaped updates bias later global models and data collection. We assign each retained trajectory an influence score from server-side records, including the deployed model version, collection window, and aggregation weight. Fig. 3b shows that client-level carriers form a separable cluster, while trajectory- and task-level carriers overlap with ordinary retained data. The embedding alone is therefore insufficient; the influence score is needed to locate the echo. This traceability motivates MUTE’s influence-aware deletion design. III. S YSTEM M ODEL AND P ROBLEM F ORMULATION A. System Model We consider a centralized federated network with one server and N heterogeneous clients, as in collaborative learning over vehicular and aerial networks [35], [36]. Each client k holds a local dataset Dk that never leaves the client. The server and clients exchange model parameters instead of raw data, which preserves data locality and reduces network traffic. Let D = SN D denote the union of all local datasets. The federation k=1 k trains a policy π(·; θ) with parameters θ by minimizing min θ
N X |Dk | k=1
|D|
Ex∼Dk [ℓ(x; θ)] ,
(1)
where ℓ(·) is the per-sample loss. At round t, the server sends the global parameter θ[t] to selected clients. Each selected client k updates it locally into θ k [t], and the server aggregates the uploaded parameters by θ[t + 1] =
N X
w̃k [t] θ k [t],
(2)
k=1
where w̃k [t] = |Dk |/|D| is the normalized aggregation weight of client k. Unlike a static federation, the network keeps improving while it operates. At round t, client k deploys π(·; θ[t]), acts in its environment, keeps the trajectories accepted by a verifier,
and appends them to Dk . The next local update in (2) then trains on the enlarged dataset. Thus, data collection depends on the deployed policy, while the deployed policy depends on all data collected so far. We now introduce unlearning into this process. A data owner issues a deletion request that marks a forget set Df ⊆ D, leaving the retain set Dr = D \ Df . We consider three granularities: a trajectory-level request removes selected sample trajectories, a client-level request removes the dataset of one client, and a task-level request removes all trajectories of one task across clients. The goal is to erase the effect of Df while preserving the utility supported by Dr . In a static federation, this goal is well defined because the retain set is fixed: one removes Df and retrains on Dr . In our setting, this reference no longer holds. Since data collection is policy-driven, and the policy has been shaped by Df , the retain set Dr may also carry the influence of Df . Retraining on Dr may therefore fail to reach the counterfactual network θ ⋆ that would have been produced had Df never appeared, as shown in Section II. Moreover, requests arrive while training continues, so erased influence may re-enter the model through later collection. We call this effect the regeneration of forgotten influence. Deletion must therefore be sustained over network operation, rather than enforced only once. B. Problem Formulation We formulate deletion over the network above. Since the influence of Df appears through behavior, not only through labeled test error, we measure current leakage by the behavioral leakage index BLI(θ). It is the area under the ROC curve of a membership inference attack that distinguishes forget data from behavioral traces. We measure whether the leakage returns after continued learning by the influence regeneration rate IRR. We write U (θ) for task utility and τ for the response time of a request. At each round, a client either keeps computation local or uploads an adapter update. Let sk [t] denote the uplink volume of client k at round t, which equals one adapter size when the client transmits and zero otherwise. The counterfactual network θ ⋆ is the ideal clean-deletion reference, but it cannot be reproduced in a live network because the past cannot be re-collected without Df . A deletion method must instead drive leakage to the chance level and keep it there while the network continues to learn. At each round, the method chooses client actions a[t] for tracing, erasure, containment, and audit. We minimize the communication cost of deleting Df while preserving utility, meeting the response deadline, and preventing regeneration: (P) :
min
{a[t]}
T X N X
sk [t]
t=1 k=1
s.t. (C1 ) : BLI(θ[t]) ≤ 0.5 + ϵ, (C2 ) : U (θ[t]) ≥ (1 − δ) U0 , (C3 ) : τ ≤ τmax , (C4 ) : IRR ≤ ϱ,
(3)
where ϵ is the chance-level margin, U0 is the pre-request utility, δ is the tolerated utility drop, τmax is the response deadline, and ϱ is the tolerated regeneration after T rounds. Constraint (C1 ) limits current behavioral leakage, and (C2 ) preserves task utility. Constraint (C3 ) enforces timely response, following the timeliness metrics used in networked learning systems [37], [38]. Constraint (C4 ) requires deletion to remain valid after continued learning. IV. S USTAINING N ETWORK P RIVACY WITH MUTE To solve the problem in (3), we design Muting Unlearned Trajectories’ Echoes (MUTE), an unlearning method for selfimproving networks. MUTE follows three steps. It traces how the forget data’s influence spreads through aggregation and data collection, erases the influence from both the model and high-risk retained trajectories, and sustains deletion through behavioral audit and scheduling. A. Influence Provenance The preliminary results show that retained trajectories may inherit the influence of Df when they are collected under a shaped policy. This module converts that influence into a network-side score. Classical data attribution estimates influence from per-sample gradients or Hessians [39], [40], which the server cannot compute without the raw trajectories. During normal operation, the server records each round’s deployed model version, participating clients, aggregation weights, and deployment windows. We denote this lightweight ledger by L[t]. When an unlearning request arrives, the server replays the ledger to estimate where the influence has propagated. The influence spreads through two edges. The aggregation edge carries influence from local training data to the global model and then to other clients. We set γ[t] = 1 at the round where Df first enters training, and update the influence of later global versions by N X w̃k [t]γkd [t], γ[t+1] = 1 − β[t] γ[t] + β[t]
(4)
k=1
where γkd [t] ∈ [0, 1] is the mean influence of client k’s current training data, w̃k [t] is its normalized aggregation weight, and β[t] is the relative size of the round update. The collection edge carries influence from a deployed model to the trajectories collected under it. If trajectory x is collected under version θ[tx ], its influence score is γ(x) = η γ[tx ],
(5)
where η ∈ [0, 1] captures how strongly the deployed policy shapes data collection. Since the ledger follows the round order, one replay costs O(T N ). The server only broadcasts the scalar sequence {γ[t]}, and each client scores its own trajectories from local collection logs using (5). Raw trajectories and per-trajectory metadata remain local. Two properties make this score useful in a live network. The version influence γ[t] decays on its own. A round trained only on clean data adds nothing to the second term of Eq.
(4), so β[t] pulls γ[t] down. The factor η is an upper bound rather than a fitted constant. Over-scoring clean trajectories costs utility that later collection can win back. Under-scoring misses a carrier, which no later step can undo. For scheduling, we summarize the current influence as a network state. Let x[t] collect the model-side and dataside influence of all clients. Aggregation and collection move influence across this state, while containment and erasure reduce it. We write x[t+1] = M(a[t]) x[t],
(6)
where a[t] denotes the scheduling decisions and M(a[t]) is the transition matrix induced by aggregation, collection, containment, and erasure. This compact state lets the scheduler track whether influence is decaying over time. B. Influence-Aware Erasure Given the influence scores, this module removes the current residue from the model and blocks future residue from retained trajectories. For model-side erasure, the client holding Df updates its adapter with a forget-retain objective: L = Lforget Df ; θ, θ ref + λ Lretain Dr ; θ, θ ref , (7) where Lforget suppresses behavior on forget trajectories, Lretain preserves behavior on retained data relative to the pre-request reference model θ ref , and λ balances unlearning and utility. In implementation, Lforget follows negative preference optimization (NPO) [29], [30]. The update touches only the lowrank adapter, so each upload remains small. To avoid later recovery, subsequent local updates are projected away from the dominant forget-gradient subspace [41]. We use NPO rather than gradient ascent on Df . Ascent has no lower bound, so it breaks the retained behavior before the forget behavior is gone. Projection is needed for the same reason the echo exists. Retained trajectories are correlated with Df , so an unprojected update can re-fit the erased direction within a few rounds. For data-side containment, each client grades its trajectories by γ(x). Trajectories with γ(x) ≥ τ q are quarantined and removed from later training. Trajectories with τ d ≤ γ(x) < τ q are down-weighted by 1 − γ(x), while those with γ(x) < τ d are used normally. The thresholds are chosen from the leakage target. Each client sorts local trajectories by influence and contains the highest-scored ones until the predicted residue meets the target. When spare duty cycle is available, the client re-collects replacement trajectories under the erased model, subject to the same energy limits that govern edge agents in the field [42], [43]. Containment is not deletion. A quarantined trajectory stays on its client and only leaves the training stream, so the server can release it once its score drops below τ q. C. Behavioral Audit and Deletion Scheduling A one-time erasure is not enough because the network continues to learn after the request. Behavior planted in a federated
Algorithm 1 MUTE: Muting Unlearned Trajectories’ Echoes Input: initial model θ[0]; forget set Df ; thresholds τ d , τ q ; uplink budget m. Ledger L ← ∅; influence scores inactive for round t = 1, . . . , T do Clients deploy π(·; θ[t]) and collect trajectories; server aggregates into θ[t+1] and appends L[t] ▷ Phase I: Trace the forget data’s influence if an unlearning request for Df arrives then Server replays L for {γ[ℓ]} by (4) and broadcasts it Each client scores its trajectories γ(x) by (5) end if ▷ Phase II: Erase the influence from model and data if unlearning is active then Client holding Df minimizes L of (7) on its adapter, then projects later updates off the forget subspace Each client contains its trajectories by γ(x), τ d , and τ q end if ▷ Phase III: Sustain the guarantee over later rounds Server audits BLI and IRR if BLI or IRR exceeds its target then Set per-client erasure and containment targets that keep ρ(M∞ ) < 1 Assign actions {ak [t]} toward these targets within uplink budget m end if end for Output: unlearned model θ[T ]
model is known to persist over many later rounds [44]. This module audits the remaining influence and schedules future erasure, containment, audit, and normal training actions under resource limits. We use two audit signals. The behavioral leakage index (BLI) measures current leakage using a membership inference attack on canary queries in the normal task stream. It passes when it stays below 0.5 + ϵ. The influence regeneration rate (IRR) measures whether high-influence retained data can revive the forgotten behavior after continued learning. BLI checks the current model state, while IRR checks whether the influence can return. Each audited client returns one score rather than its trajectories. At each round, the server decides how clients participate: some continue normal training, some run adapter-side unlearning, some contain high-influence trajectories, some join the audit, and some may pause collection. These decisions are made under the uplink budget, response deadline, and utility floor. The scheduler prioritizes clients with large model-side or data-side influence, while low-influence clients continue normal work. The feasibility of sustained unlearning follows from (6). The influence decays, and (C4 ) becomes satisfiable, when the closed-loop transition matrix M∞ under the long-run schedule satisfies ρ(M∞ ) < 1. The server therefore selects the lowestcommunication schedule that meets the deadline and utility constraints while keeping this stability condition, and applies the selected actions within the uplink budget m.
V. E XPERIMENT A. Base Models and Datasets We evaluate MUTE on two vision-language-action (VLA) policies with different architectures, so the results do not rest on one action representation. The first is MiniVLA, a compact one-billion-parameter variant of OpenVLA [45]; it reads an image and a language instruction and emits the next action as discrete tokens with an autoregressive decoder. The second is π0 [46], a three-billion-parameter model that produces continuous actions by flow matching. Together they cover the two main action representations in current VLA policies, discrete tokens and continuous flow. Both are small enough to run the full self-improving loop on edge clients, which fits the network we study. We run the whole network once for each backbone. For data, we use LIBERO [47], a benchmark of languageconditioned manipulation tasks with human demonstrations and a simulator. It has four task suites, LIBERO-Spatial, LIBERO-Object, LIBERO-Goal, and LIBERO-Long, which vary the scene layout, the objects, the goal, and the task horizon. We spread the tasks across the N clients so that each client holds a different mix, drawn by a Dirichlet distribution over task labels, which gives the heterogeneous local datasets Dk of Section III. Each client seeds Dk with the benchmark demonstrations and then grows it through the self-improving loop: it deploys π(·; θ[t]), rolls out in the simulator, keeps the trajectories that the verifier accepts, and appends them to Dk . The verifier is the task-success check of the benchmark, the same v that defines the success rate in Section V-C1. We form the forget set Df at the three granularities of Section III: a trajectory-level request removes selected sample trajectories, a client-level request removes one client’s dataset, and a tasklevel request removes all trajectories of one LIBERO task across clients. Because the collected data is policy-driven, a deleted trajectory, client, or task still shapes what the rest of the network collects later. B. Parameter Setup Table I lists the parameter settings. We freeze each backbone and train only a low-rank adapter (LoRA), and run the whole setup once for each backbone. We compare against two internal references: the counterfactual network θ ⋆ computed in simulation, and a no-deletion network that keeps Df . The large-scale runs use a multi-GPU server, while the physical testbed of Section VI runs on the Jetson hardware. C. Performance Metrics We evaluate deletion from four aspects: task utility, current leakage, future regeneration, and communication overhead, which follows the multi-dimensional view of performance in networked learning systems [48]. A valid deletion should remove the requested behavior after the request and keep it removed as the network continues to learn. Success rate (SR) measures task utility. Behavioral leakage index (BLI) measures current leakage from the forget data. Influence regeneration rate (IRR) measures how much deleted behavior returns after
where the behavior is still present; the unlearned model θ unl right after erasure; and θ unl [K], the shadow model obtained by continuing self-improvement on high-influence retained data for K rounds. IRR is the recovered fraction:
TABLE I PARAMETER SETTINGS . Parameter Dirichlet concentration LoRA rank Local epochs Batch size Self-improving rounds Deletion request round Post-request window NPO balance weight Collection-shaping calibration Quarantine threshold Uplink budget
Symbol α r E B T td K λ η τq m
Value 0.4 8 20 64 10 4 6 5.0 0.6 0.45 4
Robotic Arm Controller
CAN Protocol
D405
IRR =
Jetson Thor
Follower Control
D435
Leader USB
Jetson AGX Orin 1
Jetson AGX Orin2
Fig. 4. Physical testbed for system-level validation: a Jetson Thor server and two Jetson AGX edge clients run MUTE on a robotic-arm VLA platform over a real network.
continued self-improvement. Communication cost (Comm) measures the extra uplink traffic introduced by deletion. We report each metric at the deletion round and over the following K rounds. The counterfactual network θ ⋆ , where the forget data never appeared, is the ideal reference. Since it cannot be reproduced in a live network, we compare BLI with the chance level 0.5 and IRR with zero. 1) Success Rate: Success rate (SR) measures task utility as the accepted fraction of rollouts, defined by SR(θ[t]) = Ex∼π(·;θ[t]) [v(x)], where the verifier v(x) = 1 accepts a successful trajectory; a good deletion method keeps SR close to its pre-request value while reducing leakage. 2) Behavioral Leakage Index: Behavioral leakage index (BLI) measures how much the forget data still leaks through behavior. Following membership inference attack (MIA) [20], we use the per-sample loss ℓ(x; θ) as the membership score on canary queries in the task stream. A memorized forget trajectory tends to have lower loss than a held-out trajectory. Let Do be a held-out non-member set never used in training. BLI is defined as {(x, x′ ) ∈ Df × Do : ℓ(x; θ[t]) < ℓ(x′ ; θ[t])} BLI(θ[t]) = . |Df | |Do | (8) This is the area under the ROC curve of the attack. The ideal value is the chance level 0.5, and a higher BLI indicates stronger leakage. BLI corresponds to the leakage constraint (C1 ). 3) Influence Regeneration Rate: Influence regeneration rate (IRR) measures whether deleted behavior returns as learning continues. Let FSR(θ) denote the forget success rate, the rate at which the deleted behavior is reproduced under model θ. We evaluate it at three states: the pre-request reference model θ ref ,
FSR(θ unl [K]) − FSR(θ unl ) , FSR(θ ref ) − FSR(θ unl )
(9)
clipped to [0, 1]. IRR is 0 when the behavior stays erased and 1 when it fully returns. Unlike BLI, which measures current leakage, IRR measures whether future learning can revive the deleted behavior. IRR corresponds to the durability constraint. 4) Communication Cost: Communication cost (Comm) measures the extra uplink traffic introduced by deletion. Normal federated training already uploads adapter updates, so we count only the additional uploads required for tracing, erasure, containment, and audit after the request. This cost is summed over clients and rounds, matching the deletion-related part of the uplink objective in (3). A lower Comm means that deletion consumes less network bandwidth. D. Results Table II reports the main results. IRR grows from trajectory to client to task deletion in every block. MUTE lowers IRR against retraining in all settings, keeps BLI closer to the chance level, and holds SR within 0.04. Under trajectory deletion, uplink drops from 234.6 to 54 on MiniVLA and from 512.8 to 100 on π0 . In Fig. 5, a larger uplink budget m lowers both IRR and BLI. VI. S YSTEM -L EVEL A LGORITHM VALIDATION We validate MUTE on the physical testbed shown in Fig. 4. A Jetson Thor acts as the central server, and two Jetson AGX act as edge clients that time-share one robotic arm. The clients run the self-improving loop on a subset of the LIBERO tasks performed on the real arm, and exchange low-rank adapters with the server over a real network. On this hardware we measure the actual uplink bytes, the deletion response time, and the behavioral recurrence after continued learning, and check that they follow the trends of Section V. A. Overall Results Table II reports MUTE against the Retrain reference across the whole config space. On every backbone, suite, and granularity, MUTE keeps SR close to Retrain and pulls BLI toward the chance level. It also holds IRR below Retrain, so the deletion stays in force as the network keeps learning. The gain that matters for the network is Comm. Retrain replays the entire self-improving schedule, while MUTE adds only light erasure and audit traffic, so its uplink cost is a small fraction of Retrain. The gap grows with the backbone size, since Retrain must re-run the larger model. Deletion also gets harder as the granularity coarsens, from trajectory to client to task, and as the horizon lengthens, from LIBERO-Spatial to LIBEROLong.
TABLE II M AIN LIBERO RESULTS OVER TWO BACKBONES AND THREE DELETION GRANULARITIES . E ACH BLOCK REPORTS MUTE ON FOUR LIBERO SUITES , WITH GAPS TO Retrain IN PARENTHESES . (↑GREEN ) / (↓RED ) MARK CLOSER / FARTHER TO THE IDEAL FOR SR AND IRR, WHILE ( GRAY ) REPORTS PLAIN DIFFERENCES FOR BLI, IRR, AND C OMM . Comm↓
0.532 ( 0.050) 0.534 ( 0.048) 0.529 ( 0.053) 0.545 ( 0.037) 0.582
0.175 (↑0.079) 57.48 ( 184.86) 0.139 (↑0.115) 54.42 ( 187.92) 0.152 (↑0.102) 60.01 ( 182.33) 0.166 (↑0.088) 56.94 ( 185.40) 0.254 242.34
0.739 (↓0.024) 0.739 (↓0.024) 0.740 (↓0.023) 0.726 (↓0.037) 0.763
0.587 ( 0.006) 0.542 ( 0.051) 0.574 ( 0.019) 0.587 ( 0.006) 0.593
0.244 (↑0.094) 61.91 ( 193.82) 0.229 (↑0.109) 61.42 ( 194.31) 0.207 (↑0.131) 55.58 ( 200.15) 0.248 (↑0.090) 54.55 ( 201.18) 0.338 255.73
0.526 ( 0.019) 0.510 ( 0.035) 0.512 ( 0.033) 0.530 ( 0.015) 0.545
0.102 (↑0.058) 0.062 (↑0.098) 0.085 (↑0.075) 0.093 (↑0.067) 0.160
0.804 (↑0.012) 0.795 (↑0.003) 0.806 (↑0.014) 0.780 (↓0.012) 0.792
0.518 ( 0.050) 0.549 ( 0.019) 0.537 ( 0.031) 0.530 ( 0.038) 0.568
0.141 (↑0.100) 0.160 (↑0.081) 0.126 (↑0.115) 0.149 (↑0.092) 0.241
0.754 (↓0.007) 0.759 (↓0.002) 0.765 (↑0.004) 0.747 (↓0.014) 0.761
0.543 ( 0.044) 0.571 ( 0.016) 0.576 ( 0.011) 0.570 ( 0.017) 0.587
0.158 (↑0.149) 0.196 (↑0.111) 0.212 (↑0.095) 0.153 (↑0.154) 0.307
MiniVLA, Traj
0
MiniVLA, Traj
0
MiniVLA, Traj
0
, Traj
MiniVLA, Traj
0
MiniVLA, Client
, Client 0
MiniVLA, Client
, Client 0
MiniVLA, Client
, Client 0
MiniVLA, Client
0
MiniVLA, Task
, Task 0
MiniVLA, Task
, Task 0
MiniVLA, Task
, Task 0
MiniVLA, Task
0
100.68 ( 412.11) 96.85 ( 415.94) 100.72 ( 412.07) 100.38 ( 412.41) 512.79
0.45
IRR
0.55
2
4
0.65
0
0.8
0.05 1.0
1.0
0.8
0.6
0.4
, Traj
, Traj
, Traj
MiniVLA, Traj
0
MiniVLA, Traj
0
MiniVLA, Traj
0
0
, Client
MiniVLA, Client
0
, Client
MiniVLA, Client
0
, Client
MiniVLA, Client
0
, Task 0
MiniVLA, Task
, Task 0
MiniVLA, Task
, Task 0
MiniVLA, Task
0
0.45
0.6
BLI
0.6 0.55
0.55
0.6
0.55 0.5
1
2
4
6
8
10
, Traj , Client , Task
0.55 0.5
0
0.2
0.4
0.6
0.8
1.0
1.0
0.8
0.6
0.4
0.2
0.1
m
(f) m, BLI , Traj
0
MiniVLA, Client
0
MiniVLA, Task
, Task 0
0.85
, Client
SR
0.8 0.75
(g) η, BLI , Traj
MiniVLA, Traj
0
MiniVLA, Client
0
MiniVLA, Task
, Task 0
0.85
, Client
0.8
SR
MiniVLA, Traj
0.75
0.7 0.45
0.55
(h) α, BLI , Traj
MiniVLA, Traj
0
MiniVLA, Client
0
MiniVLA, Task
, Task 0
0.85
, Client
0.8 0.75
0.7 0.35
0.1
(d) α, IRR 0.65
0
0.35
, Task
0.2
MiniVLA, Task
(e) τ q , BLI
0.25
0.6
MiniVLA, Client
q
0.85
0.4
(c) η, IRR 0.65
0.5 0.25
0.2
MiniVLA, Traj
0.5
0.15
10
(b) m, IRR
0.55
0.15
8
, Client
m
BLI
0.6
6
, Traj
0.15
0.05 1
(a) τ q , IRR 0.65
0.25
0.15
q
BLI
0.25
IRR
IRR
0.25
115.28 ( 439.17) 110.47 ( 443.98) 105.01 ( 449.44) 105.84 ( 448.61) 554.45
0.35
, Traj
0.05 0.35
106.58 ( 415.18) 111.69 ( 410.07) 105.63 ( 416.13) 102.23 ( 419.53) 521.76
0.35
, Traj
0.15
0.25
Comm↓
0.752 (↓0.023) 0.784 (↑0.009) 0.745 (↓0.030) 0.785 (↑0.010) 0.775
0.15 0.05 0.15
Task Unlearning BLI IRR↓
SR↑
0.085 (↑0.093) 53.39 ( 181.21) 0.099 (↑0.079) 57.89 ( 176.71) 0.106 (↑0.072) 52.04 ( 182.56) 0.123 (↑0.055) 53.41 ( 181.19) 0.178 234.60
0.35
0.25
Comm↓
0.513 ( 0.038) 0.539 ( 0.012) 0.520 ( 0.031) 0.531 ( 0.020) 0.551
0.35
SR
Client Unlearning BLI IRR↓
SR↑
IRR
MiniVLA Spatial 0.777 (↓0.024) Object 0.788 (↓0.013) Goal 0.782 (↓0.019) Long 0.793 (↓0.008) 0.801 Retrain π0 Spatial 0.817 (↓0.022) Object 0.828 (↓0.011) Goal 0.828 (↓0.011) Long 0.800 (↓0.039) Retrain 0.839
Trajectory Unlearning BLI IRR↓
BLI
SR↑
SR
Suite
2
4
6
q
m
(i) τ q , SR
(j) m, SR
8
10
0
MiniVLA, Client
0
MiniVLA, Task
0
, Traj , Client , Task
0.8 0.75
0.7 1
MiniVLA, Traj
0.7 0
0.2
0.4
0.6
0.8
1.0
1.0
0.8
(k) η, SR
0.6
0.4
0.2
0.1
(l) α, SR
Fig. 5. Sensitivity of MUTE on LIBERO with two backbones, MiniVLA and π0 . The three rows report IRR, BLI, and SR. The four columns each sweep one endogenous parameter: the quarantine threshold τ q , the uplink budget m, the collection-shaping calibration η, and the Dirichlet concentration α. Every panel draws six curves, one for each backbone and deletion granularity, and each parameter varies around its default in Table I.
B. Sensitivity to Endogenous Parameters Fig. 5 sweeps the four endogenous parameters of MUTE around their defaults. Two of them are control knobs. A looser quarantine threshold τ q keeps more high-influence trajectories, so IRR and BLI rise while SR barely moves. A larger uplink budget m buys more cleaning uploads, so IRR and BLI fall. The other two describe the setting. The calibration η works best at its fitted value; IRR and BLI form a shallow valley near η = 0.6 and grow when η is set too low or too high. A smaller Dirichlet concentration α makes the local data
more heterogeneous, which raises IRR and BLI and lowers SR. In every panel, task-level deletion stays the hardest and trajectory-level the easiest, matching the granularity order in Table II. The two backbones stay close, with π0 a little lower and flatter on average. VII. C ONCLUSION We studied reliable data deletion in self-improving federated agent networks, where a deletion request must remain effective while deployed policies continue collecting new training data.
One-time unlearning is insufficient because the forget data can leave an influence echo in later retained trajectories and revive forgotten behavior during continued operation. MUTE addresses this problem by tracing downstream influence from server-side records, erasing model-side residue, containing high-influence retained trajectories, and auditing later regeneration under an uplink budget. Experiments in simulation and on a physical edge testbed show that MUTE suppresses behavioral leakage and influence regeneration while preserving utility with much less communication than retraining. R EFERENCES [1] B. Wu, Z. Ding, and J. Huang, “A Review of Continual Learning in Edge AI,” IEEE Transactions on Network Science and Engineering, vol. 13, pp. 6571–6588, 2026. [2] C. Huang, M. Tang, Q. Ma, J. Huang, and X. Liu, “Promoting collaboration in cross-silo federated learning: Challenges and opportunities,” IEEE Commun. Mag., vol. 62, no. 4, pp. 82–88, 2024. [3] Z. Ding, J. Huang, Y. Zhao, and Z. Cai, “Combating knowledge diversity and catastrophic forgetting in uav-assisted collaborative vehicular learning: A game-theoretic approach,” ACM Trans. Auton. Adapt. Syst., Jun. 2026, just Accepted. [4] B. Wu, Z. Ding, and J. Huang, “RELIEF: Turning Missing Modalities into Training Acceleration for Federated Learning on Heterogeneous IoT Edge,” arXiv preprint arXiv:2604.04243, 2026. [5] J. Huang, B. Wu, Q. Duan, L. Dong, and S. Yu, “A Fast UAV Trajectory Planning Framework in RIS-Assisted Communication Systems With Accelerated Learning via Multithreading and Federating,” IEEE Transactions on Mobile Computing, pp. 1–16, 2025. [6] B. Wu, J. Huang, Q. Duan, L. Dong, and Z. Cai, “Enhancing Vehicular Platooning With Wireless Federated Learning: A Resource-Aware Control Framework,” IEEE/ACM Transactions on Networking, pp. 1–1, 2025. [7] B. Wu, Z. Ding, J. Huang, and Y. Zhao, “Forget to Improve: On-Device LLM-Agent Continual Learning via Budget-Curated Memory,” arXiv preprint arXiv:2606.25115, 2026. [8] L. Dong, J. Huang, and G. Ye Li, “Transmission Games in RIS-Aided MIMO Interference Channels With Nonlinear Energy Harvesting,” IEEE Transactions on Wireless Communications, vol. 25, pp. 20 353–20 369, 2026. [9] European Parliament and Council of the European Union, “Regulation (EU) 2016/679 of the european parliament and of the council of 27 april 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing directive 95/46/EC (general data protection regulation),” Official Journal of the European Union, OJ L 119, pp. 1–88, 2016. [10] N. Romandini, A. Mora, C. Mazzocca, R. Montanari, and P. Bellavista, “Federated unlearning: A survey on methods, design guidelines, and evaluation metrics,” IEEE Trans. Neural Networks Learn. Syst., vol. 36, no. 7, pp. 11 697–11 717, 2025. [11] Z. Ding and J. Huang, “Toward trustworthy federated unlearning for mobile autonomous systems,” IEEE Network, pp. 1–9, 2026. [12] G. Liu, X. Ma, Y. Yang, C. Wang, and J. Liu, “Federaser: Enabling efficient client-level data removal from federated learning models,” in 29th IEEE/ACM International Symposium on Quality of Service, IWQOS 2021, Tokyo, Japan, June 25-28, 2021. IEEE, 2021, pp. 1–10. [13] L. Zhang, T. Zhu, H. Zhang, P. Xiong, and W. Zhou, “Fedrecovery: Differentially private machine unlearning for federated learning frameworks,” IEEE Trans. Inf. Forensics Secur., vol. 18, pp. 4732–4746, 2023. [14] X. Gao, X. Ma, J. Wang, Y. Sun, B. Li, S. Ji, P. Cheng, and J. Chen, “Verifi: Towards verifiable federated unlearning,” IEEE Trans. Dependable Secur. Comput., vol. 21, no. 6, pp. 5720–5736, 2024. [15] Y. Fraboni, M. V. Waerebeke, K. Scaman, R. Vidal, L. Kameni, and M. Lorenzi, “SIFU: sequential informed federated unlearning for efficient and provable client unlearning in federated optimization,” in International Conference on Artificial Intelligence and Statistics, 2-4 May 2024, Palau de Congressos, Valencia, Spain, ser. Proceedings of Machine Learning Research, S. Dasgupta, S. Mandt, and Y. Li, Eds., vol. 238. PMLR, 2024, pp. 3457–3465.
[16] Z. Ding, B. Wu, and J. Huang, “SCALE: Sensitivity-Aware Federated Unlearning with Information Freshness Optimization for Mobile Edge Computing,” in Proceedings of the IEEE International Conference on Distributed Computing Systems (ICDCS), 2026. [17] U. Pudasaini, Z. Ding, and J. Huang, “Securing Smart Agriculture with Communication-Efficient Federated Unlearning,” in Proceedings of the IEEE International Conference on High Performance Switching and Routing (HPSR). IEEE, 2026, pp. 1–8. [18] R. Taori and T. Hashimoto, “Data feedback loops: Model-driven amplification of dataset biases,” in International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, ser. Proceedings of Machine Learning Research, A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, Eds., vol. 202. PMLR, 2023, pp. 33 883–33 920. [19] I. Shumailov, Z. Shumaylov, Y. Zhao, N. Papernot, R. J. Anderson, and Y. Gal, “AI models collapse when trained on recursively generated data,” Nat., vol. 631, no. 8022, pp. 755–759, 2024. [20] N. Carlini, S. Chien, M. Nasr, S. Song, A. Terzis, and F. Tramèr, “Membership inference attacks from first principles,” in 43rd IEEE Symposium on Security and Privacy, SP 2022, San Francisco, CA, USA, May 22-26, 2022. IEEE, 2022, pp. 1897–1914. [21] H. Hu, X. Zhang, Z. Salcic, L. Sun, K. R. Choo, and G. Dobbie, “Source inference attacks: Beyond membership inference attacks in federated learning,” IEEE Trans. Dependable Secur. Comput., vol. 21, no. 4, pp. 3012–3029, 2024. [22] F. Wang, B. Li, and B. Li, “Federated unlearning and its privacy threats,” IEEE Netw., vol. 38, no. 2, pp. 294–300, 2024. [23] Z. Ma, H. Tu, L. Zhou, P. Ji, X. Yan, H. Xu, Z. Wang, and S. Chen, “Hier-fun: Hierarchical federated learning and unlearning in heterogeneous edge computing,” IEEE Internet Things J., vol. 12, no. 7, pp. 8653–8668, 2025. [24] Y. Yuan, B. Wang, C. Zhang, Z. Xiong, C. Li, and L. Zhu, “Toward efficient and robust federated unlearning in iot networks,” IEEE Internet Things J., vol. 11, no. 12, pp. 22 081–22 090, 2024. [25] X. Xia, Z. Wang, R. Sun, B. Liu, I. Khalil, and M. Xue, “Edge unlearning is not ”on edge”! an adaptive exact unlearning system on resourceconstrained devices,” in IEEE Symposium on Security and Privacy, SP 2025, San Francisco, CA, USA, May 12-15, 2025, M. Blanton, W. Enck, and C. Nita-Rotaru, Eds. IEEE, 2025, pp. 2546–2563. [26] Z. Ding, B. Wu, and J. Huang, “EASE: Federated Multimodal Unlearning via Entanglement-Aware Anchor Closure,” arXiv preprint arXiv:2605.00733, 2026. [27] B. Wu and J. Huang, “Lifecycle-Aware Federated Continual Learning in Mobile Autonomous Systems,” arXiv preprint arXiv:2604.20745, 2026. [28] B. Wu, J. Huang, and Y. Zhao, “From Alpha to Omega: LifecycleAware Forgetting Defense in Federated Continual Learning for Planetary Exploration,” in Proceedings of the IEEE International Conference on Distributed Computing Systems (ICDCS), 2026. [29] R. Zhang, L. Lin, Y. Bai, and S. Mei, “Negative preference optimization: From catastrophic collapse to effective unlearning,” in First Conference on Language Modeling, 2024. [30] C. Fan, J. Liu, L. Lin, J. Jia, R. Zhang, S. Mei, and S. Liu, “Simplicity prevails: Rethinking negative preference optimization for LLM unlearning,” CoRR, vol. abs/2410.07163, 2024. [31] C. Gao, L. Wang, K. Ding, C. Weng, X. Wang, and Q. Zhu, “On large language model continual unlearning,” in The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. [32] Z. Zhang, F. Wang, X. Li, Z. Wu, X. Tang, H. Liu, Q. He, W. Yin, and S. Wang, “Catastrophic failure of LLM unlearning via quantization,” in The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. [33] T. Huang, Q. Chen, B. Hu, Y. Zhao, H. Xu, Z. Chen, Y. Chen, and X. Su, “ROVER: robust generative continual identity unlearning against relearning attacks,” in Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, S. Koenig, C. Jenkins, and M. E. Taylor, Eds. AAAI Press, 2026, pp. 5122–5130. [34] Z. Ding, J. Huang, and J. Qi, “Learning to Defend: A Multi-Agent Reinforcement Learning Framework for Stackelberg Security Game in Mobile Edge Computing,” in Proceedings of the International Conference on Computing, Networking and Communications (ICNC), 2026, pp. 769–774.
[35] Z. Ding, J. Huang, Q. Duan, C. Zhang, Y. Zhao, and S. Gu, “A Dual-Level Game-Theoretic Approach for Collaborative Learning in UAV-Assisted Heterogeneous Vehicle Networks,” in Proceedings of the IEEE International Performance, Computing, and Communications Conference (IPCCC), 2025, pp. 1–8. [36] B. Wu, J. Huang, and Q. Duan, “FedTD3: An Accelerated Learning Approach for UAV Trajectory Planning,” in International Conference on Wireless Artificial Intelligent Computing Systems and Applications (WASA). Springer, 2025, pp. 13–24. [37] B. Wu, Z. Cai, W. Wu, and X. Yin, “AoI-Aware Resource Management for Smart Health via Deep Reinforcement Learning,” IEEE Access, 2023. [38] B. Wu, J. Huang, and Q. Duan, “Real-Time Intelligent Healthcare Enabled by Federated Digital Twins With AoI Optimization,” IEEE Network, vol. 40, no. 2, pp. 184–191, 2025. [39] P. W. Koh and P. Liang, “Understanding black-box predictions via influence functions,” in Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017, ser. Proceedings of Machine Learning Research, D. Precup and Y. W. Teh, Eds., vol. 70. PMLR, 2017, pp. 1885–1894. [40] G. Pruthi, F. Liu, S. Kale, and M. Sundararajan, “Estimating training data influence by tracing gradient descent,” in Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., 2020. [41] G. Saha, I. Garg, and K. Roy, “Gradient projection memory for continual learning,” in 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. [42] C.-C. Xing, Z. Ding, and J. Huang, “A Stochastic Geometry-Based Analysis of SWIPT-Assisted Underlaid Device-to-Device Energy Harvesting,” SIGAPP Appl. Comput. Rev., vol. 25, no. 4, pp. 18–34, 2026. [43] J. Huang, B. Wu, Z. Ding, and L. Ostigaard, “Reinforcement LearningBased Energy-Aware Coverage Path Planning for Precision Agriculture,”
in Proceedings of the International Conference on Research in Adaptive and Convergent Systems (RACS). Association for Computing Machinery, 2026. [44] E. Bagdasaryan, A. Veit, Y. Hua, D. Estrin, and V. Shmatikov, “How to backdoor federated learning,” in The 23rd International Conference on Artificial Intelligence and Statistics, AISTATS 2020, 26-28 August 2020, Online [Palermo, Sicily, Italy], ser. Proceedings of Machine Learning Research, S. Chiappa and R. Calandra, Eds., vol. 108. PMLR, 2020, pp. 2938–2948. [45] M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn, “Openvla: An open-source vision-language-action model,” in Proceedings of The 8th Conference on Robot Learning, ser. Proceedings of Machine Learning Research, P. Agrawal, O. Kroemer, and W. Burgard, Eds., vol. 270. PMLR, 06–09 Nov 2025, pp. 2679–2713. [46] K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky, “π 0 : A vision-language-action flow model for general robot control,” CoRR, vol. abs/2410.24164, 2024. [47] B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone, “LIBERO: benchmarking knowledge transfer for lifelong robot learning,” in Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., 2023. [48] B. Wu, J. Huang, and S. Yu, ““X of Information” Continuum: A Survey on AI-Driven Multi-Dimensional Metrics for Next-Generation Networked Systems,” IEEE Communications Surveys & Tutorials, vol. 28, pp. 5307–5344, 2026.