Conceptio › Archive › arXiv CS
arXiv CSopen access

CLARE: Scalable Class-Incremental Continual Learning via a Sparsity-Based Framework

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

CLARE: Scalable Class-Incremental Continual Learning via a Sparsity-Based Framework

1

Yunxiang Fu1 Meng Lou1 Zicheng Liao2 Yizhou Yu1 School of Computing and Data Science, The University of Hong Kong 2 Hong Kong Generative AI Research and Development Center

arXiv:2609.17026v1 [cs.LG] 15 Sep 2026

[email protected], [email protected] [email protected], [email protected]

Abstract

nition capacity on all seen classes. Traditional CIL methods can often be categorized into three main paradigms: regularization-based methods [15, 18], replay-based methods [21], and optimization-based methods [5]. Recent advancements leverage strong pretrained models (PTM) to further improve performance instead of training models from scratch, as pretrained models contain rich prior knowledge learned from large-scale datasets [4, 41, 48]. In particular, two prominent directions include leveraging taskspecific parameters with a routing mechanism to use the most relevant parameters during inference [8, 35, 41, 46, 48] and merging task-specific parameters into a single set of parameters for all tasks [19, 31, 43]. While the former achieves stronger performance, it usually requires the number of stored task-specific parameters to increase linearly with the number of tasks and rely on an accurate routing mechanism to predict the task identity of inputs during inference [25]. In this work, we focus on the latter paradigm since it is efficient when there are many tasks to be learned, as it does not require storing task-specific parameters.

Continual learning must balance the learning of new knowledge with the retention of previously learned knowledge to incrementally learn tasks from a data stream without catastrophic forgetting. While leveraging pretrained models has significantly advanced continual learning, existing methods exhibit a scalability bottleneck when trained sequentially on many tasks, suffering from performance degradation due to inter-task interference and loss of plasticity. Inspired by evidence that sparse fine-tuning achieves performance comparable to full fine-tuning, this paper presents a novel sparsity-driven continual learning framework. Our continual learning method, termed CLARE, operates in two stages: it first identifies a sparse, taskcritical parameter mask via a sparsity-inducing objective, then performs mask-constrained fine-tuning by only optimizing parameters selected by the mask. This two-stage sparse adapter mechanism enables all tasks to be accumulated within a shared adapter space while reducing destructive interference across tasks. Extensive experiments demonstrate the scalability of CLARE. On the long tasksequence benchmark Omnibenchmark-1k, CLARE outperforms strong baselines in final accuracy by a large margin, e.g, improving EASE by 4.64% and 13.34% after learning 100 tasks, respectively.

Although recent PTM-based CIL methods (InfLoRA [19], SD-LoRA [43]) that do not store task-specific parameters have shown promising performance on short task sequences (e.g., 10 or 20 tasks), scaling these methods to longer task sequences typically means substantial performance sacrifice [25]. This decline stems primarily from an imbalance between interference and plasticity. As the task sequence lengthens, effective new task learning causes catastrophic forgetting of earlier knowledge as new updates overwrite or conflict with parameters crucial for previous tasks. However, effective earlier knowledge preservation restricts a model’s capacity to integrate new information, resulting in progressively poorer performance on new tasks.

1. Introduction The core challenge of continual learning (CL) lies in achieving a balance between the capacity to learn new diverse tasks (learning plasticity) and the ability to retain previously learned knowledge without catastrophic forgetting (memory stability). Class-incremental learning (CIL) stands as one of the most challenging settings in CL, requiring a model to incrementally learn new classes over time without accessing previous task data, while maintaining recog-

In this paper, we hypothesize that strategically learning a small number of parameters for each task can already maintain sufficient plasticity while dramatically reducing the likelihood of destructive interference across tasks. This 1

85

80

81.5

81.3

81.0

78.5

81.5

81.4

81.0

80.0

81.4

79.6

Top-1 Accuracy (%)

75

70 68.1

65

60

55.3

55

Magnitude-based top-k% parameter updates Full Fine-tune Baseline 50

0

10

20

30

40

50

60

70

80

90

100

Percentage of Parameters Updated (%)

Figure 1. Sparse parameter update analysis. Left: Long-tail distribution of parameter update magnitudes shows most parameters experience tiny updates (< 0.01), while only 8.5K parameters have updates ≥ 0.05. Right: Sparse parameter updates achieve performance close to full fine-tuning on ImageNet-R using ImageNet-1K pretrained ViT-B/16.

hypothesis is supported by existing literature on learning sparse neural networks [27, 28, 42] as well as empirical evidence showing that the magnitude of parameter updates during fine-tuning follows a long-tailed distribution (Figure 1 (left)), with substantial updates being confined to a tiny subset of parameters. More importantly, it is only necessary to update a small proportion of the model parameters to achieve competitive task-specific performance, as illustrated in Figure 1 (right). On the basis of this insight, we propose a sparsity-driven continual learning framework that learns a sparse subset of parameters for each task, enabling effective scaling to extended sequences while balancing the plasticity-stability trade-off in continual learning models. Our sparsity-driven framework manages parameter allocation across task sequences through a two-stage learning process. Starting from a pretrained base model, we first identify task-critical parameters by optimizing a sparsityinducing objective, which produces a binary mask identifying the most relevant parameters for the task. We then perform mask-constrained fine-tuning, updating only these relevant parameters while keeping the remainder frozen. This enables the model to achieve promising performance by updating only a sparse subset of the total parameters, thereby facilitating targeted knowledge acquisition with minimal interference and preserved plasticity. In practice, the parameters learned for new tasks are incrementally fused into the base model via simple accumulation for computational efficiency. Distinct from InfLoRA [19] and SD-LoRA [43], our method explicitly learns the subset of parameters that are most important to each task, instead of choosing them based on heuristics such as orthogonal constraints. We evaluate CLARE through extensive experiments spanning both long and standard task-sequence settings. On long-sequence benchmarks, CLARE achieves the highest final accuracy of 66.88% after 100 tasks on

OmniBenchmark-1k [25], surpassing the strongest adapterbased baseline in final accuracy by 4.64% and the LoRAbased SD-LoRA [43] by over 38%. On 50-task splits of ImageNet-R and ImageNet-A, CLARE outperforms InfLoRA[19] by 13.9% and 19.28% points in final accuracy, respectively. On standard class-incremental benchmarks with 10 and 20 tasks, CLARE attains the highest final accuracy across all six evaluated settings, with particularly large margins on datasets with distribution shift: on ImageNet-A and ObjectNet, it improves final accuracy by 5.40% and 8.40% points over the strongest baseline, respectively. These results demonstrate that capacity-aware sparse adapter learning scales effectively from short to long task sequences and provides a unified solution that does not require task-specific routing or task identity at inference time. In summary, our contributions are: • We propose CLARE, a sparsity-driven adapter-based continual learning framework that learns task-critical sparse adapter masks before mask-constrained task learning. • We explicitly learn the task-critical mask using an L1 regularized objective to induce natural sparsity in taskspecific parameter updates, allowing each task to discover a compact set of task-critical parameters. • We provide extensive evaluations and ablations showing how learned sparsity, two-stage optimization, and adapter capacity affect performance across short, medium, and long task sequences.

2. Related work Pretrained Model-Based CIL Traditional CIL methods address catastrophic forgetting through regularization constraints [15, 18], rehearsal strategies that retain exemplars from previous tasks [33], gradient constraints [21], and parameter-isolation mechanisms [1, 29]. Inspired by the de2

P single shared adapter αshared = α0 + t Mt ⊙(αt − α0 ), every input uses that same adapter with all coordinates active, and no task identity or routing is used. The mask Mt only decides which currently-unused coordinates a new task t may modify. A direct consequence is that earlier tasks’ coordinates remain active for every later input, so prior knowledge is reused rather than gated. This gives CLARE high learning plasticity (average incremental accuracy), not merely forgetting mitigation. This also makes sparse parameter selection part of the learning objective rather than a hand-designed or purely post-hoc pruning rule.

velopment of strong representations learned by pretrained models (PTM) for vision [4, 7, 22–24, 32, 34] and their applicability to different tasks [2, 6, 11, 17, 26], recent CIL methods increasingly build on frozen or partially tuned pretrained vision transformers [25, 41, 48]. Prompt-based methods such as L2P [41], DualPrompt [40], and CODAPrompt [35] learn task-related prompt tokens and retrieve them at inference. Adapter-based methods further improve plasticity by learning lightweight task-specific modules, as in EASE [48] and SEMA [38]. However, these methods often store task-specific parameters and depend on accurate retrieval or routing during inference, which becomes increasingly challenging as the task sequence grows. In contrast, our work focuses on learning sparse task updates that can be merged into a single task-agnostic model.

3. Method 3.1. Problem Definition We study exemplar-free class-incremental learning. The model receives a sequence of image classification tasks {D1 , . . . , DT }. Task t contains training samples Dt = t {(xi , yi )}ni=1 whose labels belong to a new class set Ct . The class sets are disjoint for different tasks, so Ci ∩ Cj = ∅ for i ̸= j. During task t, the model can only access Dt and cannot replay samples from previous tasks. After learning task t, the model must classify test samples from all seen classes St Yt = i=1 Ci . We build CLARE as a parameter-efficient CIL method on top of a pretrained vision transformer (ViT-B/16 [4]). The pretrained backbone parameters are denoted by θ0 and kept frozen. We insert lightweight adapters [3] into each transformer block and denote all adapter parameters by α. The classifier after task t is denoted by ϕt , and the prediction function is written as f (x; θ0 , α, ϕt ). CLARE only updates the adapters and classifier. It does not fine-tune the full backbone. Figure 2 gives an overview. CLARE keeps one shared adapter across the whole task sequence. For each task, it first explicitly learns a parameter-importance sparse mask for the current task using an L1 -regularized objective. Subsequently, it trains only the selected adapter parameters. After task learning, the masked task update is added to the shared adapter, and the next task starts from this updated adapter.

Parameter-Efficient CIL Another line of PTM-based CIL seeks to avoid task-specific routing by combining taskspecific updates into one model for inference. Works like InfLoRA [19], SD-LoRA [43], and LoDA [10] learn taskspecific LoRA modules and impose constraints to reduce interference among tasks. MagMax [31] merges task-specific parameter updates into a shared model, using predefined or post-hoc rules such as random pruning or magnitudebased selection. There are also conceptually works that focus on LLM tasks, including OA-Adapter [37], Share [14], CSBoRA [20], OPLoRA [44]. These methods are free of inference-time routing, but do not explicitly learn which parameters should be updated for each task before task learning, leading to suboptimal learning plasticity. In contrast, our method uses an L1 -regularized objective to induce naturally sparse task updates, then uses the discovered mask for constrained training and merging. Sparse and Mask-Based Continual Learning Sparsity has also been widely studied in continual learning through subnetwork selection, pruning, and mask learning. Piggyback [30] learns task-specific binary masks over a fixed backbone, while PackNet [29] progressively prunes and allocates parameters for new tasks. Other sparse or subnetwork-based methods similarly reduce forgetting by assigning different parameter subsets to different tasks [1, 36, 39, 45]. These methods usually maintain task-specific masks or sparse subnetworks and often require task identity or task-specific selection during inference. For example, PackNet [29] and Piggyback [30] keep a per-task mask/subnetwork and, at test time, select the right one using the task identity or an arg-max over stored masks. They also rely on knowing the number of tasks in advance. Distinctively, CLARE’s core novelty is that sparsity is a training-time allocation mechanism, not an inference-time selection mechanism. CLARE keeps no per-task mask at inference. After each task its sparse update is merged into a

3.2. Capacity-Aware Sparse Adapter Learning CLARE uses one shared adapter to learn all tasks. This design is parameter-efficient, but it may create a clear source of forgetting. Specifically, if a new task changes adapter parameters that are important for old tasks, the feature representation of historical classes can shift. The classifier was trained on the historical representations, so this shift can damage old-task predictions and cause catastrophic forgetting. The problem becomes more severe in long task sequences because more tasks compete for the same limited trainable parameters. 3

Figure 2. (a) A shared adapter contains parameters that have been frozen after previous tasks (blue) and parameters that remain available for future tasks (orange). (b) For a new task, CLARE first selects a sparse set of parameters from the available pool (Stage 1), then resets them to their original values and optimizes only those selected parameters (Stage 2). After training, the updated parameters are frozen and added to the frozen set.

To control this interference, CLARE explicitly manages which scalar entries of the adapter parameters can be updated by each task. We call each scalar entry an adapter coordinate. Let Ω = {1, . . . , N } index all adapter coordinates, where N is the total number of adapter parameters. After several tasks have been learned, some coordinates have already beenSselected by previous task masks. t−1 For task t, let Ω<t = i=1 Ωi be the set of coordinates used by previous tasks, and let At = Ω \ Ω<t

the method does not suffer from interference where each parameter coordinate is used by 5% × 100 = 5 different tasks on average when assigning a fixed 5% budget to every task. This geometric schedule slows capacity consumption and explains why a single adapter can support a 100-task sequence. Let α0 denote the initial adapter parameters before any continual task is learned. For task t, CLARE learns an adapter state αt and defines the task-specific adapter update as ∆αt = αt − α0 . (3)

(1)

be the set of coordinates still available for task t. CLARE selects a sparse set Ωt ⊆ At for the current task. In this way, new tasks are guided toward unused adapter parameters instead of freely overwriting parameters already assigned to earlier tasks. Importantly, since knowledge of all historical tasks are learned in Ω<t , we encourage Ωt to build on the accumulated adapter state from previous tasks and use minimal sparse updates for the new task t. Specifically, the percentage of At that can be updated for task t (sparsity ratio) is set to a fixed value. For example, ρ = 0.95 means that task t only uses 1 − ρ = 5% of At to learn, not 5% of the original adapter capacity for every task. If Ft = |At | is the number of available coordinates before task t, then Ft+1 ≈ ρFt ,

Ft ≈ ρt−1 N.

Sparsity is imposed on ∆αt , not on the full pretrained backbone. Since a task modifies only free coordinates At in Equation 1 while used ones are frozen. In this trainable subspace, the previous shared adapter equals α0 , so defining ∆αt w.r.t. α0 or the previous shared state is equivalent. After each task, the sparse update is immediately added to the shared adapter. For tasks after the first one, CLARE starts from this learned shared adapter instead of returning to an independent adapter. The next task can therefore use the representation accumulated so far, while the freecoordinate mask limits direct overwriting of earlier task updates.

3.3. Two-Stage Sparse Update Optimization We now describe how the sparse adapter coordinate set Ωt and the adapter update ∆αt are learned for each task t. CLARE uses a two-stage approach: the first stage selects Ωt from the currently available coordinates, and the second stage optimizes ∆αt while restricting adapter updates to Ωt .

(2)

After T tasks, the used capacity is approximately 1 − ρT . For ρ = 0.95 and T = 100, this value is 1 − 0.95100 ≈ 0.994. The adapter is nearly saturated after 100 tasks, but 4

Stage 1: L1 -induced mask discovery. For task t, the first stage decides which available adapter coordinates should be assigned to the current task. CLARE starts from the current shared adapter and updates only the coordinates in At for a few epochs using the current task data. The objective is

Excluding Ω<t further prevents direct reuse of coordinates already assigned to previous tasks. Stage 2: mask-constrained learning. The adapter values obtained in Stage 1 are not used as the final task parameters. CLARE restores the adapter to the state used at the beginning of Stage 1: the initial adapter for the first task and the shared adapter learned so far for later tasks. It then learns task t while updating only the coordinates in Ωt . Equivalently, for any adapter coordinate j not selected (j) in Stage 1, CLARE sets the gradient of αt to zero before each optimizer step:

min E(x,y)∼Dt [ℓ (f (x; θ0 , α, ϕt ), y)]+λ ∥PAt (α − α0 )∥1 , α (4) where PAt (·) keeps only currently free adapter coordinates and λ controls the strength of the sparsity term. The classifier provides the task loss, while the purpose of this stage is to reveal which adapter coordinates respond to the new task. The L1 penalty encourages most free-coordinate updates to stay close to zero. As a result, a coordinate that still changes by a large amount is likely to be important for learning task t, while coordinates with near-zero changes can be left unused. Using the adapter values reached at the end of this short optimization, CLARE scores each available coordinate by the magnitude of this movement: (j)

st

(j)

= α(j) − α0

,

j ∈ At .

min E(x,y)∼Dt [ℓcls (f (x; θ0 , αt , ϕt ), y)] ,

αt ,ϕt

∇α(j) = 0 t

Here ∇α(j) denotes the gradient with respect to the jt th scalar adapter coordinate after the task loss is backpropagated. Only coordinates in Ωt receive adapter updates, while all other adapter coordinates remain fixed. This reset is important: Stage 1 is designed to select coordinates under an L1 pressure, whereas Stage 2 is designed to learn the task well once the coordinate budget has been fixed. Removing the L1 term in Stage 2 prevents the sparsity objective from weakening classification learning. In practice, the classifier is expanded for the new classes and trained together with the active adapter coordinates.

(5)

Larger scores mean that the coordinate moved more even under the L1 penalty, so CLARE treats them as more useful for the current task. The number of coordinates selected for task t follows the remaining-capacity schedule in the previous subsection. The target budget is k̄t = round ((1 − ρ)|At |) .

(9)

if j ∈ / Ωt .

3.4. Sequential Sparse Update and Inference After task learning, CLARE stores the masked adapter delta ∆αtsparse = Mt ⊙ (αt − α0 ).

(6)

(10)

(j)

where Mt = I[j ∈ Ωt ]. Thus, only the coordinates selected for task t are written into the task update. All other coordinates contribute zero. The shared adapter used after task t is obtained by adding the sparse task deltas to the initial adapter state:

Since k̄t may be zero after rounding when few coordinates remain, CLARE clips it to a valid integer budget:   kt = min |At |, max 1, k̄t , when |At | > 0. (7) If no available coordinate remains, CLARE assigns an (j) empty mask. Otherwise, CLARE ranks all scores {st : j ∈ At } globally and sets τt to the kt -th largest score. The selected coordinate set is n o (j) Ωt = j ∈ At : st ≥ τt . (8)

(t)

αshared = α0 +

t X

∆αisparse .

(11)

i=1

Because each task selects coordinates from the remaining free capacity, the shared adapter accumulates task knowledge while reducing direct overlap among task updates. The update is performed sequentially after each task, so the next task always starts from the adapter learned so far. At inference time, CLARE uses the frozen backbone θ0 , (t) the shared adapter αshared , and a single classifier ϕt over all seen classes:

(j)

Equivalently, the binary mask is Mt = I[j ∈ Ωt ]. If several coordinates have exactly the same score as τt , they are all included, so the mask size can differ slightly from kt in the rare case of ties. This global top-kt rule has two purposes. First, it gives each task only a fixed fraction of the coordinates still available, which preserves capacity for future tasks. Second, it lets the task use that capacity wherever the learned update indicates it is most useful, instead of forcing every adapter layer to receive the same budget.

(t)

ŷ = arg max fc (x; θ0 , αshared , ϕt ). c∈Yt

(12)

Throughout continual learning, we maintain only one set of adapter weights and do not require task identity for inference. We use a cosine classifier whose class weights 5

Metrics. We report the average incremental accuracy Ā and the final accuracy AT . Here, AT is the accuracy over all seen classes after training the final task, and Ā is the mean of the accuracies measured after each incremental task. Higher values are better for both metrics.

grow as new classes arrive following standard implementations [48]. CLARE differs from existing sparse and parameterefficient continual learning methods in two ways: (1) Its sparsity is induced during learning through a L1 regularization. (2) The mask is capacity-aware since each task selects from the remaining free coordinates. These properties allow CLARE to use a single adapter to scale to long task sequences.

4.2. Long Task Sequence Evaluation Table 1 evaluates long task sequences, a setting where single adapter-based models that do not leverage task identity at inference struggle (e.g., InfLoRA, SD-LoRA). CLARE consistently achieves the highest final accuracy and average incremental accuracy. On the longest OmniBenchmark-1k benchmark comprising 100 tasks, CLARE attains a final accuracy of 66.88%, which represents a substantial 137% relative improvement over SD-LoRA [43]. This gain suggests that the capacity-aware sparsity introduced by CLARE can mitigate catastrophic forgetting even for 100 tasks. Compared to Aper [49], which trains the adapter only on the first task, freezes it for subsequent tasks, and relies on class prototypes for prediction, CLARE yields significant improvements of 11.69% on 50-task ImageNet-R and 9.4% on 50task ImageNet-A, respectively. This demonstrates stronger learning plasticity of CLARE. These results indicate that CLARE can maintain a balance between learning plasticity and knowledge retention even over long task sequences.

4. Experiments We evaluate CLARE in exemplar-free class-incremental learning (CIL) under two settings ranging from 10 tasks to 100 tasks. The main setting studies the challenging long task sequence setup, where methods must preserve performance as the number of incremental tasks grows to 100. The second setting leverages standard CIL protocols on established datasets to verify that the same sparse shared-adapter design in CLARE remains competitive under shorter sequences. We then comprehensively investigate the effect of each component of CLARE through controlled ablation studies.

4.1. Experimental Setup Datasets. Long-sequence evaluation uses ImageNet-R [12] and ImageNet-A [13] split into 50 tasks with 4 classes per task, and OmniBenchmark-1k [25] split into 100 tasks with 10 classes per task. We denote IncN as the number of novel classes learned per task. Following previous works [19, 43, 48, 49], standard benchmark evaluation uses ImageNet-R [12] with 10 tasks (Inc20) and 20 tasks (Inc10), CIFAR-100 [16] with 10 tasks (Inc10) and 20 tasks (Inc5), ImageNetA [13] with 10 tasks (Inc20), and ObjectNet with 10 tasks (Inc20). Baselines. We compare with representative pretrainedmodel CIL methods from different families, including prompt-based (L2P [41], DualPrompt [40], CODA-Prompt [35]), adapter-based [8, 38, 48, 49], classifier-based [9, 47] and LoRA-based [19, 43]. These baselines cover methods that use task-specific parameters or prompts, and methods that maintain a single model at inference time. All adapterbased and LoRA-based baselines are run from their official implementations. Implementation details. CLARE follows prior works [19, 43, 48, 49] to use a ViT-B/16 pretrained on ImageNet-21K as the frozen backbone and trains lightweight adapters together with the classifier. During training, we use a learning rate of 0.02 with cosine learning rate decay, a batch size of 32, and a SGD optimizer with weight decay of 0.0005. The adapter bottleneck dimension is 64, the first-stage mask discovery is trained for 5 epochs with an L1 coefficient of 10−4 , and the second-stage masked training is run for 20 epochs. The sparsity ratio is set to 95%.

4.3. Standard Benchmark Evaluation Table 2 evaluates CLARE under the commonly used CIL settings with 10 and 20 tasks [19, 35, 38, 41, 43, 48]. Across all six reported settings over four datasets, CLARE achieves the highest final accuracy among the compared methods. The gains are especially clear on the datasets with larger distribution shifts: on ImageNet-A Inc20, CLARE improves AT from 55.96% for the strongest baseline (SD-LoRA) to 61.36%, and on ObjectNet Inc20, it improves AT from 59.37 to 67.77. The results on ImageNet-R and CIFAR100 show a similar pattern under both 10 and 20 task splits. On ImageNet-R, CLARE improves final accuracy over the strongest baseline by 2.47% in the 20-task Inc10 setting. On CIFAR-100 with 10 tasks, the improvement is 3.91%. Together with Table 1, these consistent gains across different number of tasks indicate that constraining adapter updates to sparse, selected coordinates in CLARE is a scalable and can generalize to various ranges of tasks.

4.4. Ablation Studies We investigate each design component of CLARE, which are the capacity-aware sparse adapter learning for determining how many parameters to use for the current task and two-stage sparse update optimization for maximizing learning plasticity using the amount of parameters available. Table 3 shows results on ImageNet-R with 10 tasks and Omnibenchmark-1k with 100 tasks. In the second row 6

Table 1. Performance comparison on benchmarks with long task sequences. We report average incremental accuracy Ā and final accuracy AT .

Method

ImageNet-R 50 Tasks (Inc4) Ā AT

ImageNet-A 50 Tasks (Inc4) Ā AT

OmniBenchmark-1k 100 Tasks (Inc10) Ā AT

L2P(CVPR’22) DualPrompt(ECCV’22) CODA-Prompt(CVPR’23) EASE(CVPR’24) SEMA(CVPR’25) APER-Adapter(IJCV’25) InfLoRA(CVPR’24) SD-LoRA(ICLR’25) CLARE (Ours)

69.16 64.00 62.43 78.11 67.80 72.43 71.68 68.40 83.00

49.89 43.85 38.24 59.86 52.99 61.29 50.17 54.72 68.83

60.91 62.18 64.16 65.00 56.55 73.23 51.53 53.97 78.32

63.45 56.33 57.57 70.63 59.32 64.83 62.62 63.28 76.52

36.41 29.95 26.60 47.53 40.68 48.58 38.70 41.28 57.98

48.87 49.45 51.75 53.54 33.96 62.24 27.01 28.15 66.88

Table 2. Performance comparison with state-of-the-art methods on standard CIL benchmarks. We report average incremental accuracy Ā and final accuracy AT .

Method

ImageNet-R 10 Tasks (Inc20) 20 Tasks (Inc10) Ā AT Ā AT

CIFAR-100 10 Tasks (Inc10) 20 Tasks (Inc5) Ā AT Ā AT

ImageNet-A 10 Tasks (Inc20) Ā AT

ObjectNet 10 Tasks (Inc20) Ā AT

L2P(CVPR’22) DualPrompt(ECCV’22) CODA-Prompt(CVPR’23) EASE(CVPR’24) SEMA(CVPR’25) APER-Adapter(IJCV’25) InfLoRA(CVPR’24) SD-LoRA(ICLR’25) CLARE (Ours)

75.46 73.10 77.97 81.74 81.39 75.82 80.82 82.04 83.88

85.92 89.65 91.05 92.11 91.60 92.22 91.70 92.54 94.73

49.39 53.71 53.54 65.34 63.83 60.53 58.50 64.95 69.46

66.77 64.31 66.53 71.04 67.95 69.24 70.67 70.37 76.67

69.77 67.18 72.27 76.17 77.84 67.95 75.65 77.34 79.73

63.75 66.52 70.45 81.18 77.84 72.35 77.28 80.22 82.67

55.78 61.77 64.68 74.62 69.60 64.33 71.01 75.26 77.73

of Table 3, we enforce a constant predefined portion of parameters for each task. We report the best results from over four different sparsity ratios {10%, 5%, 2.5%, 1%} for both datasets. It can be observed that on the 100 task setting, the final accuracy drops significantly by 13.71% to 53.17%, demonstrating the importance of capacity-aware sparsity. In the third row (w/o two-stage), we remove the two-stage learning process and directly train the adapter, and select the parameters for each task based on its magnitude. It can be seen that final accuracy decreased for both settings, suggesting that the two stage optimization is beneficial. We further investigate design choices and important hyperparameters within each component. Specifically, we study the effect of the sparsity ratio, the training strategies for the two-stage sparse update optimization, and the effect of the L1 normalization coefficient λ.

79.19 84.89 86.44 87.72 86.75 87.45 86.51 88.01 91.92

85.94 87.87 89.11 91.51 92.23 90.65 89.13 90.90 94.96

79.93 81.15 81.96 85.80 87.84 85.15 81.46 85.18 91.73

41.71 41.67 42.73 55.04 52.21 49.57 46.28 55.96 61.36

55.16 52.99 56.80 59.37 54.92 57.41 58.04 58.54 67.77

the sparsity ratio to be smaller (90%) can improve performance on the short 10 task setting, but lead to substantial decline in performance in the 100 task setting. This decline is due to insufficient number of parameters to learn new task when the number of tasks increases to 100.

Training Strategies for Two-Stage Learning . Table 5 examines the effect of different mask selection and optimization strategies, as well as the sensitivity to the sparsity coefficient λ. Removing the sparse mask entirely (no sparse mask) causes severe forgetting, with final accuracy collapsing to near zero on both benchmarks, confirming that restricting adapter updates is essential for retaining previous knowledge. Replacing the learned mask with a random selection of the same number of available coordinates (random mask) reduces final accuracy by 0.51% on ImageNet-R and 5.65% on OmniBenchmark-1k, indicating that the L1 guided mask identifies task-relevant coordinates are beneficial. When each task is trained independently from the initial pretrained adapter, without leveraging prior accumulated knowledge in the shared adapter (independent tuning), performance drops drastically to 32.33% and 11.03% final accuracy, respectively. This collapse shows that knowledge

Impact of the Sparsity Ratio . Table 4 shows the impact of the sparsity ratio for 10 tasks (ImageNet-R) and 100 tasks (Omnibenchmark-1k). We observe that CLARE is robust to the sparsity ratio when choosing from 95% to 97.5%, where the declines in final accuracy are not substantial (-0.38% for 10 tasks and -0.41% for 100 tasks). We note that choosing 7

Table 3. Impact of capacity-aware sparsity and two stage learning on ImageNet-R with 10 tasks and Omnibenchmark-1K with 100 tasks.

Ā

Variant CLARE w/o capacity-aware sparsity w/o two-stage learning

ImageNet-R AT ∆AT

83.88 83.94 83.25

79.73 79.60 78.96

Table 4. Impact of the sparsity ratio.

Variant

Ā

90% 95% 97.5% 99%

83.88 83.88 82.38 79.62

ImageNet-R AT ∆AT 84.04 79.73 79.35 73.14

+0.45 0.00 -0.38 -6.59

59.25 66.88 66.47 65.59

78.32 64.20 75.19

66.88 53.17 63.09

0.00 -13.71 -3.79

quarter of the default width) changes AT by only −0.11 points, while doubling it to d = 128 improves AT by only +0.13 points. Across this eight-fold range, the spread in final accuracy is 0.27 points, and even d = 16 remains well above the strongest baselines on the same split in Table 1 (EASE 70.63, InfLoRA 62.62). This 50-task setting is a stringent test of capacity: with ρ = 0.95, about 1 − 0.9550 ≈ 92% of the adapter has already been assigned, so extra width would help if the method were limited by the number of parameters. The nearly unchanged AT indicates that it is not. CLARE’s improvements instead come from assigning each task a sparse subset of the remaining coordinates, which reduces destructive interference even when the adapter itself is small.

Omnibenchmark-1K Ā AT ∆AT 73.49 78.32 79.03 77.62

0.00 -0.13 -1.08

Omnibenchmark-1K Ā AT ∆AT

-7.63 0.00 -0.41 -1.29

accumulation across tasks is critical and that sparse masking alone is insufficient without a shared representation. We further study alternatives for the first stage of the two-stage optimization. Using L2 regularization instead of L1 (Stage-1 L2 ) yields a modest decline, notably a 1.52% lower final accuracy on the 100-task setting, suggesting that the sparsity-inducing property of L1 is better suited for selecting a compact set of coordinates. Removing the sparsity term entirely (λ = 0) and selecting coordinates solely by update magnitude after unregularized probing reduces final accuracy by 0.88% and 1.91%, respectively, confirming that the L1 penalty helps identify the most essential adapter parameters for each task. Varying λ around the default value of 0.0001 (λ = 0.00001 and λ = 0.001) yields stable final accuracies, demonstrating that the method is robust to the choice of this hyperparameter.

4.5. Efficiency CLARE keeps a single shared adapter and does not use task identity or routing at inference, so it adds no inference-time overhead relative to a standard adapter model. Although Stage 1 spends a few epochs on mask discovery, Stage 2 updates only the selected sparse coordinates, which keeps training efficient. On 50-task ImageNet-A, the total training time of CLARE is 4.37 hours on a single NVIDIA L40 GPU, slightly faster than SD-LoRA (4.41 hours) under the same setting.

5. Limitations

Robustness to Seeds In the main experiments and ablations of the paper, we partition each dataset by randomly shuffling the class order using the seed 1993, following EASE. In Table 6, we evaluate the performance of CLARE and a representative baseline SD-LoRA on three random seed and report the mean±std. We observe that variance is small and final accuracy is close to that of the seed 1993 so CLARE is robust to different orders across different benchmarks.

CLARE is designed so that the sequence length T need not be known in advance. Rather than splitting the adapter into a predefined number of task slots, each new task selects a sparse subset of the coordinates that have not yet been assigned. The consumed capacity therefore grows as 1 − ρT . With the default ρ = 0.95, a single adapter is only about 99.4% full after 100 tasks, which is enough to remain accurate from the standard 10–20 task setting through the 100-task OmniBenchmark-1k protocol. The limitation is that this pool is finite and cannot scale indefinitely to very large number of tasks like 10k tasks. On substantially longer streams the remaining free coordinates will run out, and new tasks would then have little unused capacity left to learn. Expanding capacity at that point does not require changing the allocation principle. One can freeze a saturated adapter and continue with a fresh one, or increase the bottleneck when free coordinates become scarce, then ap-

Impact of Adapter Capacity . The adapter bottleneck dimension d controls the total number of trainable adapter coordinates, i.e., the size of the pool from which each task selects its sparse mask. We investigate the impact of d on CLARE’s gains on long task sequences. Specifically, we vary d ∈ {16, 32, 64, 128} on ImageNet-R with 50 tasks. Table 7 shows that final accuracy is nearly insensitive to adapter width. Reducing the bottleneck to d = 16 (one 8

Table 5. Impact of mask selection strategies, two-stage training, and L1 coefficient λ on ImageNet-R with 10 tasks and Omnibenchmark1K with 100 tasks. independent tuning trains each task from the initial (not shared) adapter, then merges.

ImageNet-R AT ∆AT

Ā

Variant

Omnibenchmark-1K Ā AT ∆AT

CLARE (λ = 0.0001)

83.88

79.73

0.00

78.32

66.88

0.00

No sparse mask Random mask Independent tuning

24.36 83.16 58.32

0.23 79.22 32.33

-79.50 -0.51 -47.40

8.73 70.10 23.95

0.04 61.23 11.03

-66.84 -5.65 -55.85

Stage-1 L2 regularization Stage-1 λ = 0 Stage-1 λ = 0.00001 Stage-1 λ = 0.001

83.96 83.05 84.08 84.29

79.59 78.85 79.92 79.61

-0.20 -0.88 +0.19 -0.12

76.17 75.84 77.34 76.92

65.36 64.97 65.69 66.16

-1.52 -1.91 -1.19 -0.58

Table 6. Final accuracy AT (mean±std over 3 random seeds).

Method

ImageNet-R 10 Task

ImageNet-A 50 Task

Omnibenchmark-1k 100 Task

CLARE SD-LoRA

79.48±0.45 77.50±0.38

57.49±0.62 41.33±1.91

67.10±0.46 13.27±12.83

Table 7. Impact of adapter bottleneck dimension d on ImageNet-R with 50 tasks (Inc4). Relative width is measured with respect to the default d = 64.

Variant

Relative width

AT

∆AT

d = 16 d = 32 d = 64 d = 128

×0.25 ×0.50 ×1.00 ×2.00

76.41 76.38 76.52 76.65

-0.11 -0.14 0.00 +0.13

saturates the adapter as the number of tasks increases, limiting the maximum sequence length that the current method can support without expanding the adapter capacity. Developing dynamic capacity allocation strategies that adjust the sparsity budget based on task difficulty or inter-task similarity could extend the method to substantially longer sequences. Additionally, while CLARE is evaluated on vision tasks with a ViT backbone, extending the capacity-aware sparse allocation principle to other modalities, architecture, and domains remains an open question.

ply the same remaining-capacity schedule. We leave such expansion to future work. Notably, prior works like SDLoRA mostly focus on 10 to 20 tasks, while our proposed CLARE can learn effectively for 100 tasks.

References [1] Rahaf Aljundi, Punarjay Chakravarty, and Tinne Tuytelaars. Expert gate: Lifelong learning with a network of experts. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3366–3375, 2017. 2, 3 [2] Chaoqi Chen, Jiongcheng Li, Xiaoguang Han, Xiaoqing Liu, and Yizhou Yu. Compound domain generalization via metaknowledge encoding. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 7109–7119. IEEE, 2022. 3 [3] Shoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang, Yibing Song, Jue Wang, and Ping Luo. Adaptformer: Adapting vision transformers for scalable visual recognition. Advances in Neural Information Processing Systems, 35:16664–16678, 2022. 3 [4] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 1, 3 [5] Mehrdad Farajtabar, Navid Azizan, Alex Mott, and Ang Li. Orthogonal gradient descent for continual learning. In Inter-

6. Conclusion In this paper, we have addressed the challenge of scaling class-incremental continual learning to long task sequences by proposing CLARE, a sparsity-driven adapterbased framework. CLARE is built on a simple yet effective principle, where each task updates only a sparse subset of parameters, selected from parameters not already assigned to earlier tasks. The sparse updates are accumulated as it learns new tasks. This design eliminates the need for taskspecific routing or stored task-specific parameters at inference, making it particularly suitable for long task sequences where storing per-task modules becomes impractical. Extensive experiments on both long-sequence and standard benchmarks demonstrate that CLARE consistently outperforms representative continual learning methods. Several limitations point to directions for future work. For instance, the geometric capacity schedule inevitably 9

[18] Zhizhong Li and Derek Hoiem. Learning without forgetting. IEEE transactions on pattern analysis and machine intelligence, 40(12):2935–2947, 2017. 1, 2 [19] Yan-Shuo Liang and Wu-Jun Li. Inflora: Interference-free low-rank adaptation for continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 23638–23647, 2024. 1, 2, 3, 6 [20] Yuyang Liu, Lai-Man Po, Farrell Hung, Zhuohan Wang, Haoxuan Wu, Zeyu Jiang, Kun Li, Xuyuan Xu, and KwokWai Cheung. Csbora: A continual learning method for large language models with true orthogonality and reduced forgetting. Pattern Recognition, 179:113782, 2026. 3 [21] David Lopez-Paz and Marc’Aurelio Ranzato. Gradient episodic memory for continual learning. Advances in neural information processing systems, 30, 2017. 1, 2 [22] Meng Lou and Yizhou Yu. Overlock: An overview-firstlook-closely-next convnet with context-mixing dynamic kernels. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 128–138. IEEE, 2025. 3 [23] Meng Lou, Yunxiang Fu, and Yizhou Yu. Sparx: A sparse cross-layer connection mechanism for hierarchical vision mamba and transformer networks. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 19104– 19114, 2025. [24] Meng Lou, Shu Zhang, Hong-Yu Zhou, Sibei Yang, Chuan Wu, and Yizhou Yu. Transxnet: learning both global and local dynamics with a dual dynamic token mixer for visual recognition. IEEE Transactions on Neural Networks and Learning Systems, 36(6):11534–11547, 2025. 3 [25] Meng Lou, Yunxiang Fu, and Yizhou Yu. Scaling continual learning to 300+ tasks with bi-level routing mixture-ofexperts. In Forty-third International Conference on Machine Learning, 2026. 1, 2, 3, 6 [26] Meng Lou, Hanzhong Guo, Linwei Chen, and Yizhou Yu. Overcoming catastrophic forgetting in visual continual learning with reinforcement fine-tuning. arXiv preprint arXiv:2605.09640, 2026. 3 [27] Christos Louizos, Max Welling, and Diederik P Kingma. Learning sparse neural networks through l 0 regularization. International Conference on Learning Representations, 2018. 2 [28] Rongrong Ma, Jianyu Miao, Lingfeng Niu, and Peng Zhang. Transformed l 1 regularization for learning sparse deep neural networks. Neural Networks, 119:286–298, 2019. 2 [29] Arun Mallya and Svetlana Lazebnik. Packnet: Adding multiple tasks to a single network by iterative pruning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 7765–7773, 2018. 2, 3 [30] Arun Mallya, Dillon Davis, and Svetlana Lazebnik. Piggyback: Adapting a single network to multiple tasks by learning to mask weights. In Proceedings of the European conference on computer vision (ECCV), pages 67–82, 2018. 3 [31] Daniel Marczak, Bartłomiej Twardowski, Tomasz Trzciński, and Sebastian Cygert. Magmax: Leveraging model merging for seamless continual learning. In European Conference on Computer Vision, pages 379–395. Springer, 2024. 1, 3

national conference on artificial intelligence and statistics, pages 3762–3773. PMLR, 2020. 1 [6] Yunxiang Fu, Chaoqi Chen, Yu Qiao, and Yizhou Yu. Dreamda: Generative data augmentation with diffusion models. arXiv preprint arXiv:2403.12803, 2024. 3 [7] Yunxiang Fu, Meng Lou, and Yizhou Yu. Segman: Omniscale context modeling with state space models and local attention for semantic segmentation. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 19077–19087. IEEE, 2025. 3 [8] Zijian Gao, Wangwang Jia, Xingxing Zhang, Dulan Zhou, Kele Xu, Feng Dawei, Yong Dou, Xinjun Mao, and Huaimin Wang. Knowledge memorization and rumination for pretrained model-based class-incremental learning. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 20523–20533, 2025. 1, 6 [9] Dipam Goswami, Yuyang Liu, Bartłomiej Twardowski, and Joost Van De Weijer. Fecam: Exploiting the heterogeneity of class distributions in exemplar-free continual learning. Advances in Neural Information Processing Systems, 36:6582– 6595, 2023. 6 [10] Lingfeng He, De Cheng, Huaijie Wang, Xi Yang, Nannan Wang, and Xinbo Gao. Task-driven subspace decomposition for knowledge sharing and isolation in lora-based continual learning. arXiv preprint arXiv:2603.00191, 2026. 3 [11] Xiang He, Sibei Yang, Guanbin Li, Haofeng Li, Huiyou Chang, and Yizhou Yu. Non-local context encoder: Robust biomedical image segmentation against adversarial attacks. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 8417–8424, 2019. 3 [12] Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, et al. The many faces of robustness: A critical analysis of out-of-distribution generalization. In Proceedings of the IEEE/CVF international conference on computer vision, pages 8340–8349, 2021. 6 [13] Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song. Natural adversarial examples. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 15262–15271, 2021. 6 [14] Prakhar Kaushik, Ankit Vaidya, Shravan Chaudhari, Rama Chellappa, and Alan Yuille. Shared lora subspaces for almost strict continual learning. arXiv preprint arXiv:2602.06043, 2026. 3 [15] James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka GrabskaBarwinska, et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 114(13):3521–3526, 2017. 1, 2 [16] Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images.(2009), 2009. 6 [17] Jiayuan Li, Zhen Wang, Nan Xu, and Zhuhong You. Semantic segmentation with scale alignment and contextual information fusion for multimodal remote sensing images. Information Fusion, page 103671, 2025. 3

10

and Ying Wei. Sd-lora: Scalable decoupled low-rank adaptation for class incremental learning. In International Conference on Learning Representations, 2025. 1, 2, 3, 6 [44] Yifeng Xiong and Xiaohui Xie. Oplora: Orthogonal projection lora prevents catastrophic forgetting during parameterefficient fine-tuning. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 34088–34096, 2026. 3 [45] Murat Onur Yildirim, Elif Ceren Gok, Ghada Sokar, Decebal Constantin Mocanu, and Joaquin Vanschoren. Continual learning with dynamic sparse training: Exploring algorithms for effective model updates. In Conference on parsimony and learning, pages 94–107. PMLR, 2024. 3 [46] Jiazuo Yu, Yunzhi Zhuge, Lu Zhang, Ping Hu, Dong Wang, Huchuan Lu, and You He. Boosting continual learning of vision-language models via mixture-of-experts adapters. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 23219–23230, 2024. 1 [47] Gengwei Zhang, Liyuan Wang, Guoliang Kang, Ling Chen, and Yunchao Wei. Slca: Slow learner with classifier alignment for continual learning on a pre-trained model. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 19148–19158, 2023. 6 [48] Da-Wei Zhou, Hai-Long Sun, Han-Jia Ye, and De-Chuan Zhan. Expandable subspace ensemble for pre-trained modelbased class-incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 23554–23564, 2024. 1, 3, 6 [49] Da-Wei Zhou, Zi-Wen Cai, Han-Jia Ye, De-Chuan Zhan, and Ziwei Liu. Revisiting class-incremental learning with pre-trained models: Generalizability and adaptivity are all you need. International Journal of Computer Vision, 133(3): 1012–1032, 2025. 6

[32] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PmLR, 2021. 3 [33] Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H Lampert. icarl: Incremental classifier and representation learning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 2001–2010, 2017. 2 [34] Cheng Shi, Yizhou Yu, and Sibei Yang. Vision transformers need more than registers. In 2026 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 26328–26337. IEEE, 2026. 3 [35] James Seale Smith, Leonid Karlinsky, Vyshnavi Gutta, Paola Cascante-Bonilla, Donghyun Kim, Assaf Arbelle, Rameswar Panda, Rogerio Feris, and Zsolt Kira. Coda-prompt: Continual decomposed attention-based prompting for rehearsal-free continual learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11909–11919, 2023. 1, 3, 6 [36] Fengqiang Wan and Yang Yang. Probabilistic group mask guided discrete optimization for incremental learning. In Forty-second International Conference on Machine Learning, 2025. 3 [37] Zhiyi Wan, Wanrou Du, Liang Li, Miao Pan, and Xiaoqi Qin. Budget-adaptive adapter tuning in orthogonal subspaces for continual learning in llms. arXiv e-prints, pages arXiv–2505, 2025. 3 [38] Huiyi Wang, Haodong Lu, Lina Yao, and Dong Gong. Selfexpansion of pre-trained models with mixture of adapters for continual learning. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 10087–10098, 2025. 3, 6 [39] Zifeng Wang, Zheng Zhan, Yifan Gong, Geng Yuan, Wei Niu, Tong Jian, Bin Ren, Stratis Ioannidis, Yanzhi Wang, and Jennifer Dy. Sparcl: Sparse continual learning on the edge. Advances in Neural Information Processing Systems, 35:20366–20380, 2022. 3 [40] Zifeng Wang, Zizhao Zhang, Sayna Ebrahimi, Ruoxi Sun, Han Zhang, Chen-Yu Lee, Xiaoqi Ren, Guolong Su, Vincent Perot, Jennifer Dy, et al. Dualprompt: Complementary prompting for rehearsal-free continual learning. In European conference on computer vision, pages 631–648. Springer, 2022. 3, 6 [41] Zifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang, Ruoxi Sun, Xiaoqi Ren, Guolong Su, Vincent Perot, Jennifer Dy, and Tomas Pfister. Learning to prompt for continual learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 139–149, 2022. 1, 3, 6 [42] Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li. Learning structured sparsity in deep neural networks. Advances in neural information processing systems, 29, 2016. 2 [43] Yichen Wu, Hongming Piao, Long-Kai Huang, Renzhen Wang, Wanhua Li, Hanspeter Pfister, Deyu Meng, Kede Ma,

11

Record · ID 919413 · SHA-256 312c07ce38425b6a
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.