Conceptio › Archive › arXiv CS
arXiv CSopen access

CATA: Continual Machine Unlearning via Conflict-Averse Task Arithmetic

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

CATA: Continual Machine Unlearning via Conflict-Averse Task Arithmetic Shen Lin1∗ , Junhao Dong2∗ , Rongjie Chen1 , Xiaoyu Zhang3 , Li Xu1† , and Xiaofeng Chen3

arXiv:2605.18610v1 [cs.CV] 18 May 2026

1

Fujian Normal University, 2 Nanyang Technological University, 3 Xidian University

Abstract Vision-language models (VLMs) have shown remarkable ability in aligning visual and textual representations, enabling a wide range of multimodal applications. However, their large-scale training data inevitably raises concerns about privacy, copyright, and undesirable content, creating a strong need for machine unlearning. While existing studies mainly focus on single-shot unlearning, practical VLM deployment often involves sequential removal requests over time, giving rise to continual machine unlearning. In this work, we make the first attempt to study continual unlearning for VLMs and identify three key challenges in this setting: effectiveness in removing target knowledge, fidelity in preserving retained model utility, and persistence in preventing knowledge re-emergence under sequential updates. To address these challenges, we propose CATA, a conflict-averse task arithmetic method that represents each forget request as an unlearning task vector. By maintaining historical task vectors and performing sign-aware conflict-averse aggregation, CATA suppresses conflicting update components that may weaken previous forgetting effects. Extensive experiments under both single-shot and continual settings show that CATA outperforms baselines in terms of forgetting effectiveness, model fidelity, and forgetting persistence.

1

Introduction

Recent advances in vision-language models (VLMs), such as CLIP [1], have achieved remarkable success in aligning visual and textual representations, enabling a wide range of downstream applications. Despite their success, the large-scale data used for training VLMs inevitably introduces concerns regarding data privacy, copyright protection, the presence of undesirable or sensitive content, and several trustworthy AI challenges [2, 3, 4, 5]. To address these issues, machine unlearning [6, 7, 8, 9, 10, 11] has emerged as a promising direction for removing the influence of specific data from trained models without retraining from scratch. However, existing studies primarily focus on single-shot unlearning, where the forget request is assumed to be given at once. In practical VLM deployment, removal requests may arrive sequentially over time, as users, data owners, or regulators continuously request specific data, concepts, or associations to be removed. This motivates the study of continual machine unlearning, which aims to process sequential removal requests in an online manner. Although significant progress has been made in single-shot unlearning for VLMs [12, 13, 14, 15, 16, 17], continual machine unlearning for VLMs remains largely underexplored. Directly extending single-shot methods to the continual setting is non-trivial. In single-shot unlearning, the model only needs to remove a fixed target set while maintaining its utility on retained data. In contrast, continual unlearning requires the model to repeatedly process new forgetting requests while preserving the effects of all previous unlearning steps. Since different forget sets may induce different, or even conflicting, parameter updates, a later unlearning step can weaken or partially reverse earlier forgetting ∗ These authors contribute equally. † Corresponding author.

Preprint.

effects. This leads to the knowledge re-emergence problem, where previously forgotten information becomes accessible again after subsequent updates. This problem is particularly important for VLMs. Unlike conventional classification models, VLMs encode visual and textual information in a shared cross-modal representation space. Different images, concepts, categories, and text prompts may be semantically correlated through this space. As a result, removing the influence of one forget set may affect related visual-textual associations, while later unlearning requests may unintentionally restore previously removed knowledge. Therefore, continual unlearning for VLMs should satisfy three key requirements: (i) effectiveness, requiring the model to remove the influence of target data at each step; (ii) fidelity, requiring the model to preserve performance on retained knowledge and downstream tasks; and (iii) persistence, requiring previously forgotten information to remain forgotten after subsequent updates. To address these challenges, we propose a continual machine unlearning method via Conflict-Averse Task Arithmetic, namely CATA. The key idea is to represent each forget request as an unlearning task vector and aggregate historical task vectors in a conflict-aware manner. Specifically, for each incoming forget set, CATA first estimates the parameter direction associated with the target knowledge and takes its negative direction as the corresponding unlearning task vector. To reduce noisy and redundant updates, the task vector is sparsified by retaining the most influential parameter components. Then, CATA performs sign-aware aggregation over historical task vectors, preserving direction-consistent components while suppressing conflicting ones. In this way, CATA mitigates conflicts among sequential unlearning requests, reduces the risk of knowledge re-emergence, and naturally scales to longer sequences of unlearning requests by updating the aggregated task vector rather than repeatedly retraining from scratch. We summarize our contributions as follows: • We introduce the problem of continual machine unlearning for vision-language models, where sequential removal requests must be handled without compromising retained knowledge. We highlight knowledge re-emergence as a core challenge in this setting. • We propose Conflict-averse Task Arithmetic, a task-vector-based unlearning framework that resolves conflicts among sequential forgetting directions through sign-aware aggregation, enabling stable continual unlearning. • We conduct extensive evaluations across single-shot, continual, and long-sequence settings, showing that CATA improves forgetting effectiveness, model fidelity, forgetting persistence, and scalability over strong baselines.

2

Related Work

Continual learning for VLMs. Continual learning aims to enable models to learn from a sequence of tasks while mitigating catastrophic forgetting of previously acquired knowledge. Recent studies have extended continual learning to vision-language models (VLMs), focusing on sequential adaptation of multimodal representations and preservation of zero-shot generalization [18, 19, 20, 21, 22, 23]. Early work explored visually grounded continual learning by incrementally learning compositional phrase representations from streaming visual scenes [18]. Subsequent studies investigated continual fine-tuning of CLIP and proposed strategies such as regularization to reduce zero-shot transfer degradation [19], representation constraints to preserve multimodal representations [21, 23], and parameter-efficient adapters to mitigate task interference [20]. Although these works are closely related to sequential adaptation in VLMs, their objective is fundamentally different from ours. Continual learning aims to acquire new knowledge without forgetting previous tasks, whereas continual unlearning aims to remove specified knowledge while preserving retained knowledge and preventing previously forgotten information from re-emerging. Therefore, methods designed for continual learning cannot be directly applied to continual unlearning. Machine unlearning for VLMs. Machine unlearning aims to remove the influence of specified data or concepts from trained models without retraining from scratch. This is particularly important for VLMs such as CLIP [1], whose large-scale training data may contain private, copyrighted, or undesirable content. Recent works have explored unlearning in VLMs by selectively removing target concepts or associations while preserving general multimodal representations [12, 24, 14, 15, 16, 17]. For example, MultiDelete [24]. separated embeddings associated with the forget set while maintaining unimodal representations for retained data. Other methods constrained unlearning gradients to 2

selected layers [14], disentangled and removed cross-modal associations in CLIP [15], or leveraged regularization and synthetic samples for zero-shot class unlearning [16]. Despite their effectiveness, existing VLM unlearning methods primarily focus on single-shot settings, where the forget request is processed once. In practical deployment, however, removal requests may arrive sequentially over time. Naively applying single-shot methods to each request can cause conflicts among unlearning updates, leading to degraded model utility and knowledge re-emergence. In contrast, we propose Conflict-averse Task Arithmetic to aggregate sequential forgetting directions while suppressing conflicting updates for continual unlearning tasks.

3

Preliminaries and Problem Analysis

3.1

Problem Formulation

Let D = {(xi , si , yi )}N i=1 denote a multimodal dataset, where xi ∈ X represents an image, si ∈ S denotes the associated text input (e.g., prompt or caption), and yi ∈ Y is the supervision signal (e.g., label or target response). A vision-language model (VLM) fθ parameterized by θ ∈ Rd is trained via: θ∗ = arg min E(x,s,y)∼D ℓ(fθ (x, s), y), (1) θ

where ℓ(·) denotes the task-specific loss function. We denote the trained model as the initial model before unlearning, i.e., θ(0) = θ∗ . In this paper, we consider a continual machine unlearning setting, where unlearning requests arrive sequentially in an online manner. Specifically, we define a sequence of multimodal forget sets: {Du(1) , Du(2) , . . . , Du(T ) },

(2)

(t) where each Du

⊂ D corresponds to the forget set specified at time step t. At each step, the unlearning operator updates the current model according to the incoming forget set:

θu(t) = U(θu(t−1) , Du(t) ). (3) This process requires the model to incrementally remove newly specified knowledge while preserving retained knowledge and previously achieved forgetting effects. For each step t, we define the cumulative forget set and retained set as: t [ Du(≤t) = Du(j) , Dr(t) = D \ Du(≤t) . (4) j=1

Ideally, the unlearned model at each step should approximate the model retrained from scratch on the retained data: θu(t) ≈ arg min E(x,s,y)∼D(t) ℓ(fθ (x, s), y). (5) r

θ

Knowledge Re-emergence Problem

In continual machine unlearning, a critical challenge is the knowledge re-emergence problem, where knowledge removed in earlier unlearning steps becomes accessible again after subsequent requests, as illustrated in Fig. 1. Specifically, the original model exhibits high accuracy on several target classes before unlearning. After each class is unlearned, the accuracy on the corresponding class drops significantly, indicating that the target knowledge has been effectively removed at that step. However, after subsequent unlearning requests are processed, the accuracy of previously unlearned classes partially recovers at the final step. This recovery suggests that later updates may unintentionally restore previously removed knowledge, demonstrating the knowledge re-emergence problem in continual machine unlearning. 3

100 80

Accuracy (%)

3.2

Original Model Unlearned Step Final Step

60 40 20 0 bed

bib

castle

shrew

bicycle

Unlearned Class

Figure 1: Illustration of knowledge reemergence. Accuracy on target classes decreases after unlearning but partially recovers at the final step, indicating that previously forgotten knowledge may be restored by later updates.

4

Proposed Method

4.1

Overview

As discussed in Section 3, a key challenge in continual machine unlearning is the re-emergence of previously forgotten knowledge after subsequent unlearning requests. We attribute this phenomenon to conflicts among unlearning updates induced by different forget sets. Inspired by task arithmetic for model editing [25], we represent each unlearning request as a task vector in parameter space, which provides an efficient way to compose multiple unlearning operations. However, directly aggregating these task vectors can introduce conflicting update directions, causing different unlearning effects to cancel or interfere with each other and potentially leading to knowledge re-emergence. Based on this observation, we propose CATA, a conflict-averse task arithmetic framework for continual unlearning in VLMs. As illustrated in Fig. 2, each incoming forget set is first converted into an unlearning task vector by reversing the parameter direction associated with the target knowledge. To reduce noisy and redundant updates, we sparsify each task vector through top-k% masking, retaining only the most influential parameter components. Given the historically sparse task vectors, CATA estimates the dominant update direction at each parameter dimension and aggregates only direction-consistent components, while discarding conflicting ones that may weaken previous forgetting effects. The aggregated task vector is then applied to the pretrained model to obtain the unlearned model for the current step.

Figure 2: An overview of our proposed CATA method. Each incoming forget set is converted into a sparse unlearning task vector through direction reversal and top-k% masking. Historical sparse task vectors are then aggregated by sign-aware conflict filtering, which suppresses conflicting components to mitigate the knowledge re-emergence issue. Unlearning task vector. For the t-th unlearning request, we estimate the parameter direction (t) associated with the target knowledge in Du . Specifically, starting from the pretrained VLM θ(0) , we fine-tune the model on the current forget set: (t)

θf = arg min E(x,s,y)∼D(t) ℓ(fθ (x, s), y), u

θ

(6)

where the optimization is initialized from θ(0) , and ℓ(·) denotes the task-specific VLM objective. The corresponding adaptation direction is defined as: (t)

∆θ(t) = θf − θ(0) .

(7)

Since this direction captures the parameter displacement associated with the current forget set, we define the unlearning task vector as its negative: (t)

τ (t) = −∆θ(t) = −(θf − θ(0) ). 4

(8)

Although unlearning requests arrive sequentially, each task vector is constructed with respect to the same pretrained model θ(0) . This makes task vectors from different requests comparable in a shared parameter space, which is essential for subsequent conflict detection and aggregation. Sparse task vector masking. Dense task vectors may contain many low-magnitude components that have limited impact on forgetting but can introduce noise during continual aggregation. To reduce such redundancy, we sparsify each task vector by retaining only the components with the largest magnitudes. For the unlearning task vector τ (t) , we define a binary mask m(t) ∈ {0, 1}d . Let qk (|τ (t) |) denote the magnitude threshold corresponding to the top-k% largest absolute values in τ (t) . The mask is defined as:   (t) (t) mi = I |τi | ≥ qk (|τ (t) |) , (9) where I(·) is the indicator function. The sparse task vector is then obtained by: τ̂ (t) = m(t) ⊙ τ (t) ,

(10)

where ⊙ denotes element-wise multiplication. This masking operation preserves the dominant parameter changes associated with the current unlearning request while filtering out low-impact components. Conflict-averse aggregation. In continual unlearning, task vectors induced by different forget sets may conflict with each other. For two unlearning requests a and b, a sign conflict occurs at parameter dimension i when: (a) (b) τ̂i · τ̂i < 0. (11) Such conflicting components may cancel each other during aggregation, thereby weakening previous forgetting effects and causing knowledge re-emergence. To address this issue, we maintain a historical task-vector memory: Mt = {τ̂ (1) , τ̂ (2) , . . . , τ̂ (t) }, (12) which stores the sparse task vectors observed up to step t. For each parameter dimension i, we estimate the dominant unlearning direction through a magnitude-weighted sign vote:   t X (j) (t) γi = sgn  τ̂i  , (13) j=1 (t)

where sgn(·) is the sign function. If the summed update is zero, we set γi = 0, indicating that no dominant direction exists at this dimension. We then select the components that are consistent with the dominant direction: n o (j) (j) (t) (t) Ai = j ∈ {1, . . . , t} | τ̂i ̸= 0, sgn(τ̂i ) = γi . (14) (t)

Components with signs inconsistent with γi are treated as conflict-inducing updates and excluded from aggregation. The aggregated task vector at step t is computed as:  (j) (t) (t)  1(t) P , γi ̸= 0 and |Ai | > 0, (t) τ̂i i,(t) τagg = |Ai | j∈Ai (15) 0, otherwise. By aggregating only sign-consistent components, CATA preserves the dominant unlearning direction while suppressing updates that may interfere with previous forgetting effects. (t)

Model update. After obtaining the aggregated task vector τagg , we derive the unlearned model at step t as: (t) θu(t) = θ(0) + λτagg , (16) where λ controls the strength of unlearning. Although unlearning requests arrive sequentially, all (t) task vectors are defined relative to the pretrained model θ(0) . Therefore, applying τagg to θ(0) ensures that the current and historical unlearning directions are integrated in a shared parameter space. The (t) (≤t) resulting model θu is expected to forget the cumulative forget set Du while preserving the knowledge associated with the retained data. The overall procedure of CATA is summarized in Algorithm 1. 5

Algorithm 1: Conflict-averse Task Arithmetic Initialize task-vector memory M0 ← ∅; 2 for t = 1 to T do (t) (t) 3 Fine-tune θ(0) on Du to obtain θf ; 1

(t)

Compute τ (t) = −(θf − θ(0) ) and apply top-k% masking to obtain τ̂ (t) = m(t) ⊙ τ (t) ; Update memory Mt ← Mt−1 ∪ {τ̂ (t) }; for each parameter dimension i do P  (t) (j) t Compute dominant sign γi = sgn ; j=1 τ̂i

4 5 6 7

(t)

(j)

Select consistent indices Ai = {j ≤ t | τ̂i Aggregate non-conflicting components: ( 1 P

8 9

(t)

τagg,i =

(t)

|Ai |

(j)

(t)

̸= 0, sgn(τ̂i ) = γi }; (j)

(t)

j∈Ai

τ̂i ,

(t)

|Ai | > 0, (t)

|Ai | = 0.

0,

end (t) (t) 11 Update θu = θ(0) + λτagg ; 12 end 10

5

Experiments

Datasets, models, and baselines. In the main experiments, we evaluate CATA on ImageNet1K [26] using four representative CLIP backbones, including ViT-B-32, ViT-L-14, RN50, and RN101. Additional results on CIFAR-10 and CIFAR-100 [27] are provided in the supplementary material. Furthermore, we compare CATA with representative machine unlearning baselines, including FT [28], GA [29], Fisher [30], LIP [31], EMMN [32], CLIP-LIP [16], and TIFS [17]. Implementation details. To obtain the task vector, we fine-tune the CLIP visual encoder from the original pre-trained weights for 4 epochs using the AdamW optimizer with a learning rate of 1 × 10−5 and a weight decay of 0.1. The scaling factor λ is set to 0.7 and the top-k trimming ratio is set to 0.3 across all backbones. Additional implementation details and baseline configurations are provided in the supplementary material. Metrics. Following previous work [14], we report zero-shot classification accuracy on the targeted set (Target), the retaining set (Retain), and the full dataset (All). To further evaluate model fidelity and generalization, we test on several unseen datasets, including Food [33], STL [34], and ObjectNet [35]. We compute a normalized score as min(Accunlearn /Accoriginal , 1), where the score is capped at 1 to indicate full preservation of the original model performance. For the Target, we use 100 − ScoreTarget when computing the average score to reflect the degree of forgetting. We report the Avg. Score by averaging the normalized scores across all evaluated datasets. For continual unlearning, we use ∆ to measure the difference between the final accuracy and that at the corresponding forgetting step for each target class and report the Avg. ∆ by averaging ∆ over all forgetting classes. 5.1

Main Results

Comparison experiments in continual unlearning. As shown in Table 1, we compare CATA with baselines under the continual unlearning setting on ImageNet-1K. The results show that CATA achieves the best balance among these objectives on both ViT-B/32 and ViT-L/14. It consistently suppresses the accuracy of forgotten classes to near zero and maintains the lowest Avg. ∆, showing that the removed knowledge remains stably forgotten after subsequent updates. Meanwhile, CATA obtains the highest Avg. Score on both backbones, indicating better preservation of retained knowledge and downstream generalization. In contrast, the baselines typically fail in at least one aspect. FT preserves high retain and downstream performance but shows weak forgetting effectiveness, as the target accuracy often remains high after the corresponding forgetting step. GA achieves stronger forgetting than FT, but causes noticeable degradation in retained and downstream performance, especially under longer unlearning sequences. LIP aggressively suppresses target classes, but severely 6

damages model utility, leading to very low retain and transfer performance. These results demonstrate that CATA achieves more effective, persistent, and high-fidelity continual unlearning by better balancing target removal and utility preservation. Additional results on CIFAR-100 are provided in the supplementary material. Table 1: Performance comparison on ImageNet-1K under continual unlearning. Underlined values indicate the class forgotten at the corresponding step. Step

Target↓ Retain↑ Class 1 Class 2 Class 3 Class 4 Class 5

All↑

Original

Step 0

64.00

24.00

32.00

22.00

90.00

59.36

59.29 82.05 97.36

30.27

88.79

–

–

GA [29]

Step 1 Step 2 Step 3 Step 4 Step 5

2.00 0.00 2.00 0.00 0.00

16.00 0.00 0.00 0.00 0.00

30.00 34.00 2.00 0.00 0.00

24.00 44.00 36.00 12.00 14.00

86.00 84.00 74.00 62.00 10.00

57.69 55.25 52.04 48.34 44.01

57.56 55.13 51.89 48.17 43.81

81.80 79.30 77.00 74.47 70.10

96.96 95.80 95.47 95.16 94.74

28.43 27.79 25.92 23.56 20.02

81.80 80.08 78.58 78.20 95.70

– – – – 1.20

– – – – 84.54

FT [28]

Step 1 Step 2 Step 3 Step 4 Step 5

56.00 52.00 42.00 46.00 40.00

14.00 12.00 10.00 14.00 10.00

50.00 46.00 50.00 46.00 46.00

26.00 28.00 30.00 30.00 34.00

88.00 90.00 88.00 82.00 88.00

62.45 63.26 63.50 63.63 63.85

62.37 63.17 63.40 63.53 63.75

82.17 81.22 79.93 78.72 78.02

97.74 97.92 97.60 97.55 97.47

30.46 30.55 30.58 30.71 30.78

89.61 88.73 87.24 86.92 85.91

– – – – 5.20

– – – – 55.32

LIP [31]

Step 1 Step 2 Step 3 Step 4 Step 5

0.00 0.00 0.00 32.00 74.00

0.00 0.00 0.00 0.00 0.00

0.00 0.00 0.00 0.00 0.00

0.00 0.00 0.00 0.00 0.00

0.00 0.00 0.00 0.00 0.00

0.71 0.15 0.17 0.10 0.01

0.71 0.15 0.17 0.14 0.09

12.84 41.09 1.06 9.97 1.10 7.19 0.91 10.28 0.94 10.40

2.27 0.31 0.29 0.26 0.25

35.16 11.97 12.10 10.44 10.55

– – – – 14.80

– – – – 37.19

Step 1 Step 2 CATA (ours) Step 3 Step 4 Step 5

0.00 0.00 0.00 0.00 0.00

20.00 0.00 0.00 0.00 0.00

26.00 24.00 0.00 0.00 0.00

28.00 36.00 22.00 0.00 0.00

78.00 70.00 58.00 66.00 2.00

56.22 55.17 54.73 54.07 54.33

56.09 55.02 54.54 53.86 54.06

78.18 78.42 78.06 78.04 77.20

96.83 95.89 96.20 96.38 96.69

27.23 27.17 27.06 26.85 26.42

82.44 81.19 81.64 85.13 84.01

– – – – 0.00

– – – – 95.98

Backbone

Method

ViT-B/32

Food↑

STL↑

ObjectNet↑ CIFAR-10↑ Avg. ∆ ↓ Avg. Score↑

Original

Step 0

38.00

46.00

34.00

52.00

88.00

71.72

71.62 92.05 99.41

51.86

95.32

–

–

GA [29]

Step 1 Step 2 Step 3 Step 4 Step 5

6.00 2.00 0.00 0.00 0.00

82.00 12.00 4.00 2.00 2.00

32.00 36.00 2.00 0.00 0.00

56.00 48.00 34.00 6.00 10.00

76.00 26.00 0.00 0.00 2.00

69.95 64.96 43.76 23.35 6.04

69.85 64.76 43.58 23.24 6.02

91.50 88.91 46.30 11.69 5.30

99.35 97.88 73.88 44.90 42.27

51.90 50.45 45.70 34.15 26.29

90.48 87.30 73.81 48.08 43.86

– – – – 4.40

– – – – 57.81

FT [28]

Step 1 Step 2 Step 3 Step 4 Step 5

58.00 28.00 62.00 54.00 66.00

18.00 46.00 16.00 14.00 56.00

32.00 50.00 52.00 50.00 48.00

62.00 64.00 62.00 68.00 56.00

78.00 66.00 68.00 96.00 88.00

75.36 75.81 76.40 76.29 76.69

75.23 75.68 76.28 76.19 76.62

91.72 90.36 90.52 90.08 89.23

99.28 99.30 99.41 99.29 99.25

52.39 50.75 52.09 51.79 51.87

93.87 87.99 92.16 91.54 91.94

– – – – 6.80

– – – – 42.08

LIP [31]

Step 1 Step 2 Step 3 Step 4 Step 5

0.00 0.00 18.00 36.00 48.00

0.00 0.00 0.00 0.00 0.00

0.00 0.00 0.00 0.00 0.00

0.00 0.00 0.00 0.00 0.00

0.00 0.00 0.00 0.00 0.00

2.30 0.17 0.07 0.07 0.04

2.29 0.17 0.08 0.11 0.09

18.55 55.75 1.81 10.86 1.14 9.07 0.82 9.22 0.71 9.22

14.27 2.90 2.54 2.52 2.50

31.16 10.06 9.75 9.71 9.75

– – – – 9.60

– – – – 36.27

Step 1 Step 2 CATA (ours) Step 3 Step 4 Step 5

0.00 0.00 0.00 0.00 0.00

66.00 0.00 0.00 0.00 0.00

36.00 40.00 0.00 0.00 0.00

46.00 50.00 52.00 0.00 8.00

88.00 90.00 86.00 90.00 10.00

70.66 70.69 70.71 69.92 70.22

70.55 70.52 70.49 69.66 69.89

91.22 91.23 91.13 90.99 91.11

51.19 50.42 50.59 50.17 50.15

89.91 90.89 92.37 93.65 93.26

– – – – 1.60

– – – – 96.56

ViT-L/14

99.41 99.36 99.40 99.22 99.31

Table 2: Performance comparison on ImageNet-1K in single-shot unlearning. Backbone

ImageNet

Method Target↓

RN50

RN101

Retain↑

Food↑

STL↑

ObjectNet↑

CIFAR-10↑

Avg. Score↑

All↑

Original

54.10

66.61

65.37

76.49

93.75

25.83

68.84

–

FT [28] GA [29] Fisher [30] LIP [31] EMMN [32] CLIP-LIP [16] TIFS [17]

4.005.88 14.5021.32 0.200.29 0.951.40 10.0014.71 37.0054.41 0.000.00

0.260.53 35.4072.73 0.340.70 1.753.60 36.6775.34 52.77100.00 44.6791.78

24.8047.75 35.3968.14 1.092.10 1.142.19 32.2562.09 53.78100.00 47.3491.14

30.3439.67 51.3967.18 0.230.30 0.230.30 47.4862.07 76.4799.97 79.00100.00

67.0571.52 77.7682.94 10.6911.40 10.5011.20 65.2169.56 93.80100.00 83.0588.59

9.5837.09 14.1854.90 1.144.41 0.592.28 9.6537.36 25.8099.88 16.9565.62

12.2617.81 14.0820.45 10.1214.70 10.4615.19 11.7517.07 67.7798.45 83.80100.00

44.07 63.58 19.04 19.06 58.40 91.98 91.01

CATA (ours)

2.002.94

51.73100.00

51.6899.50

76.2699.70

94.03100.00

25.2297.64

65.9495.79

98.53

Original

80.00

54.67

54.33

81.16

96.46

29.16

73.82

–

FT [28] GA [29] Fisher [30] LIP [31] EMMN [32] CLIP-LIP [16] TIFS [17]

0.000.00 32.0040.00 1.121.40 0.240.30 68.0085.00 2.002.50 0.000.00

0.510.93 38.7670.90 0.380.70 0.270.49 29.3353.65 50.8492.99 50.1791.77

27.6250.84 38.7571.32 0.270.50 1.743.20 26.4348.65 51.4994.77 49.5191.13

38.5447.49 58.4472.01 0.240.30 0.160.20 42.4452.29 78.5096.72 70.5086.87

74.7577.49 84.4587.55 9.9410.30 11.1911.60 72.4675.12 96.4299.96 85.4588.59

13.2645.47 18.6263.85 0.411.41 0.792.71 11.9741.05 27.6694.86 19.1365.60

15.7821.38 18.7325.37 13.4418.21 11.8916.11 17.3123.45 70.0894.93 89.86100.00

49.09 64.43 18.57 19.14 44.17 95.96 89.14

CATA (ours)

0.000.00

53.0697.06

52.9697.48

78.9897.31

96.45100.00

27.2793.52

72.0497.59

97.56

Comparison experiments in single-shot unlearning. Table 2 compares CATA with representative unlearning methods on ImageNet-1K under the single-shot setting. The results show that CATA 7

achieves the best Avg. Score on both RN50 and RN101, indicating that it preserves the original model utility more effectively across in-domain and transfer benchmarks after unlearning. Compared with FT, GA, Fisher, and LIP, which often reduce target accuracy at the cost of severe degradation on retain classes and downstream datasets, CATA maintains substantially stronger retention and generalization performance. CLIP-LIP preserves model utility well but leaves higher target accuracy in some settings, suggesting insufficient forgetting. TIFS achieves strong target suppression, but its average score is lower due to performance drops on retain or transfer benchmarks. In contrast, CATA achieves low target accuracy while maintaining model performance, demonstrating a better balance between forgetting effectiveness and model fidelity in single-shot unlearning. Additional results on CIFAR-10 are provided in the supplementary material. Class 1

Target

100

Class 3

100

Class 4

100

80

80

80

80

60

60

60

60

60

40

40

40

40

40

20

20

20

20

20

0

2

4

6

8

100

0

10

Class 6

0

2

4

6

8

100

0

10

Class 7

0

2

4

6

8

10

Class 8

100

0

0

2

4

6

8

0

10

Class 9

100

80

80

80

80

60

60

60

60

60

40

40

40

40

40

20

20

20

20

20

0

2

4 6 8 Unlearning Step

0

10

0

2

4 6 8 Unlearning Step

10

0

0

2

4 6 8 Unlearning Step

10

0

0

2

0

2

4

4 6 8 Unlearning Step

10

0

6

8

10

4 6 8 Unlearning Step

10

Class 10

100

80

0

Class 5

100

80

0

Target

Class 2

100

0

2

Figure 3: Scalability evaluation on ImageNet-1K in continual unlearning. Each subplot shows the target accuracy of one removed class across continual unlearning steps. Scalability evaluation. We evaluate scalability by progressively increasing the number of unlearning steps, with each step removing one target class. Fig. 3 reports the target accuracy of each removed class across the unlearning sequence. The accuracy of most target classes quickly drops to near zero after the corresponding unlearning step and remains low in subsequent steps, indicating stable forgetting without obvious knowledge re-emergence. These results suggest that CATA can maintain effective target removal as the number of unlearning requests increases. Besides, the detailed stepwise results in the supplementary material (Table 6) indicate that the retain and overall accuracies remain stable as the number of unlearning steps increases. This demonstrates that CATA scales to longer unlearning sequences without significant degradation in model utility. Table 3: Ablation study of aggregation strategies on ImageNet-1K in continual unlearning. Underlined values indicate the class forgotten at the corresponding step. Class 1 Class 2

Target↓ Class 3 Class 4 Class 5

Retain↑

All↑

Food↑

STL↑

ObjectNet↑

CIFAR-10↑

Avg. ∆ ↓

Avg. Score↑

Step 0

64.00

24.00

32.00

22.00

90.00

59.36

59.29

82.05

97.36

30.27

88.79

–

–

Naive Avg.

Step 1 Step 2 Step 3 Step 4 Step 5

0.00 0.00 0.00 4.00 16.00

18.00 0.00 0.00 0.00 2.00

20.00 28.00 0.00 0.00 0.00

32.00 34.00 26.00 2.00 8.00

72.00 72.00 68.00 76.00 34.00

55.84 56.57 56.87 57.25 57.75

55.70 56.42 56.68 57.05 57.53

77.52 80.29 80.53 81.28 81.34

96.54 96.20 96.71 97.08 97.19

27.14 28.32 28.52 28.82 28.82

81.28 82.22 83.72 86.66 86.39

– – – – 4.80

– – – – 88.94

Ours

Step 1 Step 2 Step 3 Step 4 Step 5

0.00 0.00 0.00 0.00 0.00

20.00 0.00 0.00 0.00 0.00

26.00 24.00 0.00 0.00 0.00

28.00 36.00 22.00 0.00 0.00

78.00 70.00 58.00 66.00 2.00

56.22 55.17 54.73 54.07 54.33

56.09 55.02 54.54 53.86 54.06

78.18 78.42 78.06 78.04 77.20

96.83 95.89 96.20 96.38 96.69

27.23 27.17 27.06 26.85 26.42

82.44 81.19 81.64 85.13 84.01

– – – – 0.00

– – – – 95.98

Aggregation

Step

Original

5.2

Ablation Study

Impact of the aggregation strategy. Table 3 evaluates the role of conflict-averse aggregation. Naive averaging directly combines task vectors and therefore suffers from conflicting update directions, 8

leading to incomplete forgetting and noticeable knowledge re-emergence. In contrast, our conflictaverse aggregation suppresses inconsistent components, achieving more complete forgetting and reducing Avg. ∆ to zero. Although it introduces a mild drop in retained performance, it yields a higher Avg. Score, demonstrating a better balance between effectiveness, fidelity, and persistence. These results show that direct aggregation is insufficient for continual unlearning, and conflict-aware aggregation is essential for stable sequential forgetting. Impact of the scaling factor λ. As shown in Fig. 4, we evaluate the effect of the scaling factor λ in continual unlearning. As λ increases, the Target accuracy and Avg.∆ are progressively reduced, indicating stronger forgetting and less knowledge re-emergence. Meanwhile, the overall accuracy stays relatively stable across a wide range of λ, whereas the Avg. Score drops when λ becomes too large, suggesting a loss of model fidelity. These results show that λ governs the balance between forgetting effectiveness and performance preservation. We therefore use λ = 0.7 in our experiments.

(a) Forgetting effectiveness

(b) Model fidelity

Figure 4: Ablation study of the scaling factor λ on ImageNet-1K in continual unlearning. Impact of the top-k% selection. Fig. 5 shows the effect of the top-k% selection in task vector trimming. When k is too small, the task vector becomes overly sparse, leading to insufficient unlearning and higher target accuracy. As k increases, both Target and Avg.∆ quickly decrease and remain close to zero, while All and Avg. Score stay relatively stable. These results indicate that effective continual unlearning can already be achieved with a relatively small top-k%, implying that the task vectors are highly sparse and incur very low storage overhead in practice. Based on this observation, we set k = 0.3 in our evaluation experiments.

(a) Forgetting effectiveness

(b) Model fidelity

Figure 5: Ablation study of the top-k% selection on ImageNet-1K in continual unlearning.

6

Conclusion

In this paper, we study continual machine unlearning for vision-language models, where sequential forget requests introduce unique challenges in effectiveness, fidelity, and persistence. Unlike singleshot unlearning, this setting requires the model to remove newly specified target knowledge while preserving retained utility and preventing previously forgotten knowledge from re-emerging in later steps. To address these challenges, we propose CATA, a framework based on Conflict-averse Task Arithmetic. By representing each forget set as an unlearning task vector and performing conflict-aware aggregation over historical vectors, CATA suppresses inconsistent update directions and produces stable unlearning updates. Experiments on CLIP-based models demonstrate that CATA achieves a strong balance among forgetting effectiveness, model fidelity, and persistence against knowledge re-emergence. These results highlight the importance of explicitly resolving inter-step conflicts for reliable continual unlearning in vision-language models. 9

References [1] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al., “Learning transferable visual models from natural language supervision,” in International Conference on Machine Learning, 2021, pp. 8748–8763. [2] B. Li, P. Qi, B. Liu, S. Di, J. Liu, J. Pei, J. Yi, and B. Zhou, “Trustworthy ai: From principles to practices,” ACM Computing Surveys, vol. 55, no. 9, pp. 1–46, 2023. [3] J. Dong, R. Z. Moayedi, Y.-S. Ong, and S.-M. Moosavi-Dezfooli, “Allies teach better than enemies: Inverse adversaries for robust knowledge distillation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026. [4] J. Dong, J. Liu, X. Qu, and Y.-S. Ong, “Confound from all sides, distill with resilience: Multi-objective adversarial paths to zero-shot robustness,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 624–634. [5] J. Dong, X. Qu, C. Zhang, S. Q. Rong, N. D. Thai, W. Pan, X. Li, T. Liu, P. Koniusz, and Y.-S. Ong, “Tug-of-war no more: Harmonizing accuracy and robustness in vision-language models via stability-aware task vector merging,” in The Fourteenth International Conference on Learning Representations, 2026. [6] V. S. Chundawat, A. K. Tarun, M. Mandal, and M. Kankanhalli, “Can bad teaching induce forgetting? unlearning in deep networks using an incompetent teacher,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, 2023, pp. 7210–7217. [7] S. Lin, X. Zhang, C. Chen, X. Chen, and W. Susilo, “Erm-ktp: Knowledge-level machine unlearning via knowledge transfer,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 20 147–20 155. [8] M. Chen, W. Gao, G. Liu, K. Peng, and C. Wang, “Boundary unlearning: Rapid forgetting of deep networks via shifting the decision boundary,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 7766–7775. [9] S. Lin, X. Zhang, W. Susilo, X. Chen, and J. Liu, “Gdr-gma: Machine unlearning via directionrectified and magnitude-adjusted gradients,” in Proceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 9087–9095. [10] Y. Tong, Y. Wang, J. Yuan, and C. Hu, “Robust machine unlearning for quantized neural networks via adaptive gradient reweighting with similar labels,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 20 603–20 612. [11] Y. Xiao, Q. Ye, L. Hu, H. Zheng, H. Hu, Z. Liang, H. Li, and Y. Jiao, “Reminiscence attack on residuals: Exploiting approximate machine unlearning for privacy,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 3058–3068. [12] S. Poppi, T. Poppi, F. Cocchi, M. Cornia, L. Baraldi, and R. Cucchiara, “Safe-clip: Removing nsfw concepts from vision-and-language models,” in European Conference on Computer Vision, 2024, pp. 340–356. [13] J. Li, Q. Wei, C. Zhang, G. Qi, M. Du, Y. Chen, S. Bi, and F. Liu, “Single image unlearning: Efficient machine unlearning in multimodal large language models,” in Proceedings of the 38th International Conference on Neural Information Processing Systems, 2024, pp. 35 414–35 453. [14] Z. Cai, Y. Tan, and M. S. Asif, “Targeted unlearning with single layer unlearning gradient,” in International Conference on Machine Learning, 2025, pp. 6257–6290. [15] T. Yang, L. Dai, X. Wang, M. Cheng, Y. Tian, and X. Zhang, “Cliperase: Efficient unlearning of visual-textual associations in clip,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 2025, pp. 30 438–30 452. [16] A. Kravets and V. P. Namboodiri, “Zero-shot class unlearning in clip with synthetic samples,” in 2025 IEEE/CVF Winter Conference on Applications of Computer Vision, 2025, pp. 6456–6464. [17] Z. Zhang, G. Liu, C. Fleming, R. R. Kompella, and C. Xu, “Targeted forgetting of image subgroups in clip models,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 9870–9880. [18] X. Jin, J. Du, A. Sadhu, R. Nevatia, and X. Ren, “Visually grounded continual learning of compositional phrases,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, 2020, pp. 2018–2029. 10

[19] Z. Zheng, M. Ma, K. Wang, Z. Qin, X. Yue, and Y. You, “Preventing zero-shot transfer degradation in continual learning of vision-language models,” in 2023 IEEE/CVF International Conference on Computer Vision, 2023, pp. 19 068–19 079. [20] J. Yu, Y. Zhuge, L. Zhang, P. Hu, D. Wang, H. Lu, and Y. He, “Boosting continual learning of vision-language models via mixture-of-experts adapters,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 23 219–23 230. [21] Z. Ni, L. Wei, S. Tang, Y. Zhuang, and Q. Tian, “Continual vision-language representation learning with off-diagonal information,” in Proceedings of the 40th International Conference on Machine Learning, 2023, pp. 26 129–26 149. [22] J. Dong, P. Koniusz, X. Qu, and Y.-S. Ong, “Stabilizing modality gap & lowering gradient norms improve zero-shot adversarial robustness of vlms,” in Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1, 2025, pp. 236–247. [23] W. Liu, F. Zhu, L. Wei, and Q. Tian, “C-CLIP: Multimodal continual learning for visionlanguage model,” in The Thirteenth International Conference on Learning Representations, 2025. [Online]. Available: https://openreview.net/forum?id=sb7qHFYwBc [24] J. Cheng and H. Amiri, “Multidelete for multimodal machine unlearning,” in European Conference on Computer Vision, 2024, pp. 165–184. [25] G. Ilharco, M. T. Ribeiro, M. Wortsman, L. Schmidt, H. Hajishirzi, and A. Farhadi, “Editing models with task arithmetic,” in The Eleventh International Conference on Learning Representations, 2023. [Online]. Available: https://openreview.net/forum?id=6t0Kwf8-jrj [26] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE Conference on Computer Vision and Pattern Recognition, 2009, pp. 248–255. [27] A. Krizhevsky, “Learning multiple layers of features from tiny images,” Master’s thesis, University of Tront, 2009. [28] A. Warnecke, L. Pirch, C. Wressnegger, and K. Rieck, “Machine unlearning of features and labels,” in Proceedings 2023 Network and Distributed System Security Symposium, 2023. [29] A. Thudi, G. Deza, V. Chandrasekaran, and N. Papernot, “Unrolling sgd: Understanding factors influencing machine unlearning,” in 2022 IEEE 7th European Symposium on Security and Privacy, 2022, pp. 303–319. [30] A. Golatkar, A. Achille, and S. Soatto, “Eternal sunshine of the spotless net: Selective forgetting in deep networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 9304–9312. [31] J. Foster, K. Fogarty, S. Schoepf, Z. Dugue, C. Öztireli, and A. Brintrup, “An information theoretic approach to machine unlearning,” 2024. [Online]. Available: https://arxiv.org/abs/2402.01401 [32] V. S. Chundawat, A. K. Tarun, M. Mandal, and M. Kankanhalli, “Zero-shot machine unlearning,” IEEE Transactions on Information Forensics and Security, vol. 18, pp. 2345–2354, 2023. [33] L. Bossard, M. Guillaumin, and L. Van Gool, “Food-101–mining discriminative components with random forests,” in European Conference on Computer Vision, 2014, pp. 446–461. [34] A. Coates, A. Ng, and H. Lee, “An analysis of single-layer networks in unsupervised feature learning,” in Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, 2011, pp. 215–223. [35] A. Barbu, D. Mayo, J. Alverio, W. Luo, C. Wang, D. Gutfreund, J. Tenenbaum, and B. Katz, “Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models,” in Proceedings of the 33rd International Conference on Neural Information Processing Systems, 2019, pp. 9453–9463.

11

A

Theoretical Analysis

As a simple analysis, we formulate the regret bound of the proposed CATA in the convex setting. Firstly, we introduce convexity, smoothness, and Lipschitz assumptions. Assumption 1 (Convexity). Each ℓt : Rd → R is convex. Assumption 2 (µ-smoothness). Each ℓt is µ-smooth, i.e., its gradient is µ-Lipschitz: ∥∇ℓt (θ) − ∇ℓt (θ′ )∥ ≤ µ∥θ − θ′ ∥,

∀θ, θ′ .

Assumption 3 (L-Lipschitzness). Each ℓt is L-Lipschitz, i.e., |ℓt (θ) − ℓt (θ′ )| ≤ L∥θ − θ′ ∥,

∀θ, θ′ .

Equivalently, ∥∇ℓt (θ)∥ ≤ L whenever differentiable. Then, we assume that the task vector is uniformly similar to the genuine gradient. Assumption 4 (Gradient approximation error). For each step t, the sparse task vector satisfies ∥τ̂ (t) + ∇ℓt (θ(0) )∥ ≤ δ, where δ ≥ 0 is a constant that depends on the fine-tuning procedure, sparsification ratio, and the Lipschitz constant of the loss. Moreover, we upper bound the accumulated discrepancy between Gt and the genuine gradient. Pt Lemma 1 (Cumulative gradient approximation error). Let Gt = j=1 τ̂ (j) and let ∇ℓj (θj ) be the true gradient at the point θj chosen by CATA. Then Gt −

t X

∇ℓj (θj ) ≤ µRt + tδ.

j=1

Proof. By the triangle inequality, t X j=1

τ̂ (j) −

t X

∇ℓj (θj ) ≤

j=1

t X

∥τ̂ (j) + ∇ℓj (θ(0) )∥ +

j=1

t X

∥∇ℓj (θ(0) ) − ∇ℓj (θj )∥.

j=1

The first sum is bounded by tδ by Assumption 4. Finally, we derive the Regret upper bound of CATA by 1) upper-bounding the accumulated discrepancy between Gt and the genuine gradient (Lemma 1), 2) upper-bounding the regret bound induced by the genuine gradient. Theorem 1 (Regret bound of CATA). Under Assumptions 1–4, if the step-size is chosen as η = R√ , then the cumulative regret of CATA satisfies (L+δ) T RegretT =

T X

ℓt (θt ) − min

t=1

θ

T X

√ ℓt (θ) ≤ R(L + δ) T + 2µR2 T + 2RδT.

t=1

Proof. We decompose the regret into three parts: RegretT =

T X

 ℓt (θt ) − ℓt (θ∗ ) .

t=1

|

{z

true regret

}

By convexity (Assumption 1), ℓt (θt ) − ℓt (θ∗ ) ≤ ⟨∇ℓt (θt ), θt − θ∗ ⟩. Besides,

⟨∇ℓt (θt ), θt − θ∗ ⟩ = ⟨τ̂ (t) , θt − θ∗ ⟩ + ⟨∇ℓt (θt ) − τ̂ (t) , θt − θ∗ ⟩. 12

The first term is the linear regret with respect to the approximate gradients τ̂ (t) . For the update Pt−1 θt = θ(0) − ηGt−1 with Gt−1 = j=1 τ̂ (j) , standard analysis gives T X

⟨τ̂ (t) , θt − θ∗ ⟩ ≤

t=1 ∗

T ∥θ∗ − θ(0) ∥2 η X (t) 2 ∥τ̂ ∥ . + 2η 2 t=1

(0)

(t) Using ∥θ − √ θ ∥ ≤ R and ∥τ̂ ∥ ≤ L + δ (by Assumptions 3 and 4), and choosing η = R/((L + δ) T ), we obtain T X

⟨τ̂ (t) , θt − θ∗ ⟩ ≤

t=1

√ η R2 + T (L + δ)2 = R(L + δ) T . 2η 2

Using Cauchy–Schwarz and the bound ∥θt − θ∗ ∥ ≤ 2R, T X

⟨∇ℓt (θt ) − τ̂ (t) , θt − θ∗ ⟩ ≤

t=1

Now, Hence

T X

∥∇ℓt (θt ) − τ̂ (t) ∥ · ∥θt − θ∗ ∥ ≤ 2R

t=1

T X

∥∇ℓt (θt ) − τ̂ (t) ∥.

t=1

∥∇ℓt (θt ) − τ̂ (t) ∥ ≤ ∥∇ℓt (θt ) − ∇ℓt (θ(0) )∥ + ∥∇ℓt (θ(0) ) + τ̂ (t) ∥ ≤ µR + δ. T X ⟨∇ℓt (θt ) − τ̂ (t) , θt − θ∗ ⟩ ≤ 2R(µR + δ)T. t=1

Putting everything above together, √ √ RegretT ≤ R(L + δ) T + 2RµRT + 2RδT = R(L + δ) T + 2µR2 T + 2RδT.

B

Additional Implementation Details

All experiments are conducted on a single NVIDIA RTX 4090 GPU with 48GB memory, using Python 3.11.14 and PyTorch 2.9.1. All CLIP backbones are initialized from the official OpenAI released pretrained weights3 . In the main continual unlearning experiments on ImageNet-1K, we sequentially remove five target classes: tailed frog, diamondback, barn spider, briard, and dungeness crab. For the implementation details of baselines, we follow the official implementations or commonly adopted settings from prior work as described below. • FT [28] updates the image encoder using only the retain set Dr , while keeping the text encoder frozen, and optimizes the standard CLIP contrastive loss. For ImageNet, we subsample approximately 50k retain samples per step to reduce computation. We train for 2 epochs per step using Adam with a learning rate of 1 × 10−6 and a batch size of 128. (t)

• GA [29] performs gradient ascent on the current forget set Du to maximize the CLIP loss and disrupt image–text alignment for target classes. Only the image encoder is updated, with gradient norm clipping applied for stability. We use Adam with a learning rate of 1 × 10−6 , a batch size of 128, and train for 2 epochs per step with gradient clipping set to 1.0. • Fisher [30] estimates the Fisher Information Matrix on the forgetting dataset and perturbs the model parameters by adding Gaussian noise whose variance is inversely proportional to the Fisher information, thereby reducing the model’s sensitivity to the target data. In the (t) continual setting, we apply this procedure sequentially at each step based on Du . • LIP [31] freezes the image encoder and updates the text projection matrix via a LoRA adapter, optimizing a combination of forget loss, retain loss, and a regularization term. To reduce computation on ImageNet, up to 512 retain classes are sampled per step. After optimization, the LoRA weights are merged back into the model. We use Adam with a learning rate of 0.01, LoRA rank 5, and run 2,000 iterations per step with λ1 = 0.3, dynamically adjusted λ2 , and λ3 = 1.0. 3 https://github.com/openai/CLIP

13

• EMMN [32] adopts a teacher–student framework with a learnable pseudo-image generator. The teacher model is fixed, while the student is updated using pseudo-images filtered by confidence, and training alternates between generator and student updates in a 1:5 ratio. We use Adam for both components, with a student learning rate of 1 × 10−6 , a generator learning rate of 1 × 10−5 , batch size 128, temperature T = 1.0, and confidence threshold δ = 0.5. • CLIP-LIP [16] extends LIP to the CLIP framework by applying low-rank adaptation to the text projection while preserving the pretrained image encoder. Similar to LIP, it updates the text projection using a low-rank LoRA module to suppress representations of the forgetting classes while maintaining alignment for retained classes. In the continual setting, we apply CLIP-LIP sequentially, where each step updates the model using the cumulative forgetting set. We follow the original implementation and use Adam with a learning rate of 0.01, LoRA rank 5, and 2,000 iterations per step. • TIFS [14] performs targeted forgetting by suppressing the model’s responses on the forgetting dataset while preserving performance on retained data. Specifically, it minimizes prediction confidence on target classes while maintaining alignment on non-target samples. In the continual setting, we apply TIFS sequentially, where at step t the model is updated (t) using the current forgetting set Du together with retained data. We follow the original implementation and update only the image encoder using Adam with a learning rate of 1 × 10−6 , batch size 128, and train for 2 epochs per step.

C

Additional Experimental Results

Additional comparison experiments in single-shot unlearning. Table 4 reports additional singleshot unlearning results on CIFAR-10. Overall, CATA achieves the best Avg. Score on both RN50 and RN101, showing strong utility preservation after unlearning. Compared with FT, GA, Fisher, and LIP, which substantially degrade retained and transfer performance, CATA maintains much higher accuracy on retained classes and downstream datasets. Although CLIP-LIP and TIFS preserve model utility in some settings, they often retain higher target accuracy or show weaker overall balance. These results further confirm that CATA achieves effective target removal while preserving model fidelity in single-shot unlearning. These observations are consistent with the main experimental results on ImageNet-1K in Table 2, further validating the robustness of CATA across different datasets and backbones. Table 4: Performance comparison on CIFAR-10 in single-shot unlearning. Backbone

CIFAR-10

Method Target↓

RN50

RN101

Retain↑

Food↑

STL↑

ObjectNet↑

All↑

ImageNet Target↓

Avg. Score↑

All↑

Original

54.10

66.61

65.37

76.49

93.75

25.83

71.33

53.81

–

FT [28] GA [29] Fisher [30] LIP [31] EMMN [32] CLIP-LIP [16] TIFS [17]

22.4041.40 5.9010.91 0.000.00 0.000.00 0.300.55 67.80100.00 5.039.30

63.2995.02 27.0140.55 12.3918.60 14.8522.30 11.3417.02 68.44100.00 61.6192.50

58.2089.03 24.9038.09 12.1618.60 14.5822.30 10.2415.66 68.34100.00 60.4792.50

1.702.22 1.822.38 1.221.60 1.612.10 49.7064.98 76.78100.00 65.1785.20

45.2048.21 34.1636.44 15.4716.50 15.1916.20 46.8649.98 93.6699.90 87.3893.20

0.291.12 0.883.41 0.100.40 0.080.30 10.2139.53 25.91100.00 21.7084.00

0.100.14 0.000.00 0.000.00 0.140.20 48.0067.29 57.3380.37 66.2792.90

0.000.00 0.400.74 0.110.20 0.160.30 36.3267.50 53.85100.00 45.0983.80

36.79 26.34 19.49 20.46 52.68 81.87 89.35

CATA (ours)

3.306.10

69.46100.00

62.8496.13

74.4397.31

88.8494.76

24.5194.89

31.5544.24

50.2593.38

96.52

Original

70.60

75.39

74.91

81.10

96.45

29.16

66.67

55.25

–

FT [28] GA [29] Fisher [30] LIP [31] EMMN [32] CLIP-LIP [16] TIFS [17]

14.9021.10 4.506.37 0.000.00 0.000.00 78.30100.00 3.605.10 8.9712.70

63.8484.68 25.8634.30 15.0820.00 16.1321.40 6.478.58 75.67100.00 73.9698.10

57.8977.28 23.7231.66 14.9820.00 16.0321.40 13.6518.22 68.5091.44 73.4998.10

1.702.10 3.474.28 1.381.70 2.272.80 45.2955.84 79.1197.55 69.1085.20

66.9069.36 56.3358.40 17.6518.30 16.4017.00 42.8544.43 90.8894.22 91.5394.90

1.876.41 1.976.76 0.120.40 0.090.30 10.7136.73 26.8392.01 24.4984.00

0.000.00 0.000.00 0.130.20 0.000.00 62.6794.00 21.3331.99 58.7488.10

1.292.33 1.642.97 0.170.30 0.280.50 36.5466.14 55.25100.00 36.5266.10

40.13 29.00 20.11 20.43 39.13 87.76 87.73

CATA (ours)

3.054.31

69.4192.07

70.7794.47

78.9797.37

89.7193.01

28.2796.95

27.3340.99

51.0292.34

87.86

Additional comparison experiments in continual unlearning. Table 5 reports additional stepwise continual unlearning results on CIFAR-100. The observations are consistent with the main results on ImageNet-1K. FT generally preserves high retain and downstream performance but fails to effectively forget the target classes. GA achieves stronger target suppression than FT, but its retained and overall performance degrades as the number of unlearning steps increases. LIP aggressively suppresses many target classes, but severely damages model utility across retained and 14

transfer datasets. In contrast, CATA consistently reduces the accuracy of the forgotten classes while maintaining stronger downstream performance, demonstrating a better balance among effectiveness, fidelity, and irreversibility across different datasets and backbones. Table 5: Performance comparison on CIFAR-100 in continual unlearning. Underlined values indicate the class forgotten at the corresponding step. Backbone

ViT-B/32

rose

wardro

Retain↑

All↑

Original

Step 0 66.00 61.00 72.00 61.00

80.00

61.20

61.54 82.05 97.36

30.27

88.79

–

–

GA [29]

Step 1 Step 2 Step 3 Step 4 Step 5

11.00 14.00 15.00 8.00 15.00

38.00 22.00 21.00 11.00 17.00

66.00 59.00 40.00 23.00 38.00

56.00 48.00 44.00 19.00 23.00

65.00 79.00 78.00 63.00 52.00

57.77 53.15 52.13 39.12 43.58

57.24 52.71 51.50 38.40 42.85

81.79 81.90 81.66 81.03 81.41

97.41 97.41 97.33 97.03 97.21

29.85 29.85 29.66 29.23 29.52

87.28 84.10 81.63 69.93 75.68

– – – – 3.00

– – – – 74.23

FT [28]

Step 1 Step 2 Step 3 Step 4 Step 5

70.00 65.00 66.00 58.00 51.00

74.00 60.00 53.00 51.00 48.00

85.00 82.00 79.00 77.00 73.00

90.00 89.00 91.00 87.00 73.00

95.00 96.00 96.00 96.00 96.00

83.17 84.62 85.27 84.93 84.98

83.15 84.31 84.86 84.37 84.14

79.32 77.40 76.35 75.60 75.14

97.38 96.81 96.50 96.28 95.74

29.53 28.81 28.33 27.81 27.49

91.66 90.37 89.49 88.47 87.28

– – – – 10.20

– – – – 59.78

LIP [31]

Step 1 Step 2 Step 3 Step 4 Step 5

0.00 0.00 0.00 69.00 54.00 0.00 25.00 23.00 13.00 52.00 10.00 10.00 42.00 1.00 3.00

0.00 0.00 0.00 0.00 0.00

0.00 0.00 0.00 0.00 9.00

2.46 0.22 0.63 0.51 0.61

2.34 1.42 1.19 1.21 1.10

10.50 34.60 1.06 9.71 1.10 7.24 1.06 8.14 1.14 9.85

4.51 0.33 0.19 0.11 0.10

10.89 8.88 9.30 9.40 9.51

– – – – 21.00

– – – – 40.42

45.00 68.00 44.00 0.00 64.00 38.00 2.00 7.00 38.00 2.00 9.00 3.00 5.00 8.00 7.00

62.00 76.00 53.00 44.00 5.00

55.69 45.38 47.43 46.61 50.39

55.10 44.91 46.07 44.87 48.13

81.67 81.68 82.02 81.74 81.72

97.38 97.06 97.12 97.09 97.16

29.95 29.46 29.90 29.61 29.78

85.15 75.19 77.30 76.31 79.16

– – – – 2.20

– – – – 91.72

Original

Step 0 90.00 82.00 71.00 91.00

62.00

74.88

75.10 92.05 99.41

51.86

95.32

–

–

GA [29]

Step 1 Step 2 Step 3 Step 4 Step 5

26.00 64.00 32.00 78.00 10.00 9.00 7.00 10.00 15.00 20.00 8.00 16.00 14.00 16.00 4.00 11.00 13.00 7.00 0.00 5.00

40.00 15.00 14.00 3.00 2.00

61.21 19.29 24.77 10.69 6.01

60.55 18.84 24.26 10.64 5.98

91.90 91.72 91.71 91.68 91.58

99.40 99.26 99.24 99.19 99.15

51.61 51.36 51.41 51.24 51.33

93.42 71.29 71.01 57.93 31.49

– – – – 5.80

– – – – 74.14

FT [28]

Step 1 Step 2 Step 3 Step 4 Step 5

87.00 69.00 74.00 71.00 80.00

88.00 80.00 75.00 69.00 61.00

89.00 90.00 89.00 87.00 89.00

93.00 94.00 91.00 88.00 76.00

96.00 98.00 98.00 97.00 97.00

89.55 90.12 90.73 91.18 91.38

89.60 89.92 90.46 90.74 90.84

90.99 90.19 89.53 89.53 89.18

98.94 98.84 98.05 97.90 98.06

50.93 49.82 48.82 47.95 47.93

95.78 95.31 94.05 95.09 93.66

– – – – 7.60

– – – – 54.60

LIP [31]

Step 1 Step 2 Step 3 Step 4 Step 5

0.00 24.00 13.00 11.00 10.00

37.00 94.00 93.00 93.00 84.00

0.00 0.00 2.00 2.00 2.00

0.00 0.00 0.00 0.00 0.00

0.00 1.00 1.00 0.00 8.00

8.89 0.03 0.02 0.02 0.01

8.80 1.22 1.11 1.08 1.03

39.62 1.01 0.79 0.82 0.82

67.38 18.15 11.18 12.00 13.35

21.74 3.96 3.07 3.13 3.31

59.95 12.13 12.14 11.90 11.98

– – – – 4.00

– – – – 36.85

Step 1 0.00 77.00 61.00 81.00 Step 2 26.00 0.00 44.00 36.00 CATA (ours) Step 3 24.00 2.00 4.00 43.00 Step 4 25.00 2.00 6.00 2.00 Step 5 4.00 4.00 5.00 3.00

23.00 40.00 28.00 28.00 1.00

65.39 49.46 50.74 46.12 50.55

64.39 48.45 49.21 44.44 50.41

92.12 91.85 91.99 91.79 91.79

99.40 99.40 99.38 99.36 99.33

51.76 51.40 51.34 51.26 50.59

94.12 84.64 87.38 86.42 88.41

– – – – 2.00

– – – – 91.21

bed

0.00 2.00 1.00 1.00 1.00

crab

Food↑

STL↑

ObjectNet↑ CIFAR-10↑ Avg. ∆ ↓ Avg. Score↑

Step

Step 1 Step 2 CATA (ours) Step 3 Step 4 Step 5

ViT-L/14

Target↓ house

Method

Detailed results for scalability evaluation. Table 6 provides the detailed numerical results corresponding to the scalability evaluation in the main paper, where the results are visualized in figure form. We evaluate CATA on ImageNet-1K with ViT-B/32 by increasing the number of unlearning steps to 10, with each step forgetting one target class. The underlined value indicates the class forgotten at the corresponding step. The results show that CATA remains effective as the unlearning sequence becomes longer. Once a class is forgotten, its target accuracy generally stays close to zero in subsequent steps, indicating that the proposed method can suppress knowledge re-emergence over longer sequences. Meanwhile, the retain and overall accuracy remain relatively stable across 10 steps, further supporting the scalability of CATA for continual unlearning.

D

Naive averaging aggregation

Naive averaging aggregation is used as an ablation variant of CATA by replacing the proposed conflict-averse aggregation with direct averaging over all accumulated task vectors. Given the task vectors {τ (i) }N i=1 computed from the fixed original CLIP anchor, the averaged task vector is: N

τavg =

1 X (i) τ . N i=1

The unlearned model is then obtained as: (t) θu(t) = θ(0) + λτavg . 15

(17)

(18)

Table 6: Detailed numerical results for the scalability evaluation on ImageNet-1K. Each step forgets one target class, and underlined values indicate the class forgotten at the corresponding step. briard

Target↓ dungeness crab

bib

loupe

radio telescope

dock

32.00

22.00

90.00

46.00

30.00

50.00

22.00

14.00

34.00

66.00

48.00

34.00

52.00

24.00

0.00

18.00

38.00

62.00

44.00

34.00

48.00

20.00

0.00

0.00

0.00

24.00

50.00

44.00

28.00

36.00

0.00

0.00

0.00

0.00

58.00

44.00

26.00

34.00

Step 5

0.00

0.00

0.00

0.00

2.00

42.00

22.00

Step 6

0.00

0.00

0.00

0.00

2.00

0.00

16.00

Step 7

2.00

0.00

0.00

0.00

2.00

0.00

Step 8

4.00

0.00

0.00

0.00

2.00

Step 9

4.00

0.00

0.00

0.00

Step 10

4.00

0.00

0.00

0.00

tailed frog

diamondback

barn spider

Step 0

64.00

24.00

Step 1

0.00

10.00

Step 2

0.00

Step 3 Step 4

Retain↑

All↑

52.00

59.46

59.29

60.00

53.94

53.74

58.00

53.88

53.67

22.00

54.00

53.78

53.50

28.00

54.00

53.14

52.86

30.00

28.00

56.00

53.44

53.09

34.00

28.00

56.00

54.42

54.01

2.00

46.00

26.00

56.00

54.87

54.45

0.00

2.00

2.00

26.00

60.00

54.93

54.48

6.00

2.00

2.00

2.00

4.00

52.00

55.27

54.79

6.00

2.00

2.00

4.00

4.00

4.00

55.16

54.64

streetcar

All other settings, including task-vector construction, top-k% trimming, scaling factor, and evaluation protocol, are kept the same as CATA. Additional ablation results for aggregation strategies. Table 7 provides additional step-wise results for the aggregation ablation on CIFAR-100. Naive average aggregation directly averages historical task vectors without considering sign conflicts across unlearning steps. As a result, several previously forgotten classes show increased target accuracy in later steps, indicating potential knowledge reemergence caused by conflicting updates. In contrast, CATA consistently keeps the accuracy of forgotten classes close to zero across subsequent steps and achieves a lower Avg. ∆ with a higher Avg. Score. These results further support the observations in the main paper: directly averaging task vectors is insufficient for stable continual unlearning, while conflict-averse aggregation better balances effectiveness, fidelity, and persistence. Table 7: Ablation study of aggregation strategies on Imagenet in continual unlearning. Underlined values indicate the class forgotten at the corresponding step. Aggregation

Step

Original

Step 0

Naive Avg.

Step 1 Step 2 Step 3 Step 4 Step 5

Ours

Step 1 Step 2 Step 3 Step 4 Step 5

Class 2

Target↓ Class 3 Class 4 Class 5

Retain↑

All↑

Food↑

38.00

46.00

34.00

52.00

88.00

71.72

71.62

0.00 0.00 0.00 0.00 2.00

64.00 0.00 0.00 4.00 6.00

36.00 34.00 2.00 4.00 4.00

46.00 50.00 52.00 16.00 20.00

88.00 90.00 90.00 90.00 46.00

70.93 71.23 71.36 71.19 71.33

0.00 0.00 0.00 0.00 0.00

66.00 0.00 0.00 0.00 0.00

36.00 40.00 0.00 0.00 0.00

46.00 50.00 52.00 0.00 8.00

88.00 90.00 86.00 90.00 10.00

70.66 70.69 70.71 69.92 70.22

70.55 70.52 70.49 69.66 69.89

Class 1

E

Broader Impact and Limitations

E.1

Broader Impact

Avg. ∆ ↓

STL↑

ObjectNet↑

CIFAR-10↑

92.05

99.41

51.86

95.32

–

–

70.81 91.348 71.05 91.60 71.15 91.62 70.95 91.62 71.05 91.78

99.42 99.40 99.42 99.38 99.41

51.02 51.43 51.54 51.57 50.61

90.62 93.24 93.95 94.58 94.49

– – – – 2.80

– – – – 88.57

99.41 99.36 99.40 99.22 99.31

51.19 50.42 50.59 50.17 50.15

89.91 90.89 92.37 93.65 93.26

– – – – 1.60

– – – – 96.56

91.22 91.23 91.13 90.99 91.11

Avg. Score↑

This work studies continual machine unlearning for vision-language models, aiming to support responsible deployment of large-scale multimodal systems. By enabling models to process sequential removal requests, CATA can help address practical concerns related to privacy protection, copyright compliance, and the removal of undesirable or sensitive content. Compared with retraining from scratch, the proposed method provides a more efficient alternative, which may reduce computational cost and make unlearning more accessible in real-world deployments. At the same time, machine unlearning should be applied with care. Improved unlearning techniques may be used to remove harmful or unauthorized knowledge, but they may also be misused to selectively erase information for improper purposes, such as hiding model behavior or removing evidence of biased training data. Therefore, continual unlearning should be accompanied by transparent evaluation protocols and appropriate auditing mechanisms to ensure that deletion requests are handled faithfully and responsibly. 16

E.2

Limitations

Although CATA shows strong performance under both single-shot and continual unlearning settings, several limitations remain. First, our experiments are mainly conducted on CLIP-based models and classification-oriented benchmarks. Extending the method to broader VLM architectures, such as generative multimodal models, remains an important direction. Second, our current setting focuses on class-level forgetting, where each unlearning step removes one target class. More fine-grained scenarios, such as instance-level, concept-level, or text-only forgetting, may introduce additional challenges. Third, the proposed method relies on task vectors obtained by fine-tuning on forget sets. Although the top-k% trimming strategy reduces storage overhead, maintaining task vectors for long unlearning sequences may still require additional memory. Finally, while our conflict-averse aggregation mitigates knowledge re-emergence, it does not provide formal deletion guarantees. Developing verifiable and theoretically grounded continual unlearning methods for VLMs remains an important future direction.

17

Record · ID 200499 · SHA-256 8952c5db42019c3d
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.