CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation Fengming Yu Haiwei Pan∗ Kejia Zhang Chunling Chen Jian Guan Baoying Ma College of Computer Science and Technology, Harbin Engineering University, Harbin, China
arXiv:2607.27054v1 [cs.LG] 29 Jul 2026
Abstract Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model compression. The emergence of diverse model architectures has extended KD from homogeneous to heterogeneous settings. However, differences in architectural inductive biases between the teacher and student models often result in substantial representation discrepancies, limiting the effectiveness of direct knowledge transfer. Recently, redundancy suppression has offered a new perspective on heterogeneous KD by preserving cross-architecture invariance and reducing feature redundancy through decorrelation of teacher-student feature correlations. Nevertheless, this formulation may weaken useful structural information through uniform decorrelation, while a fixed coefficient may make the effective contribution of redundancy suppression sensitive to teacher-student pairs and training stages. To address these problems, Correlation Calibration-based Redundancy Suppression (CoCaRS) is proposed to better retain structural information while suppressing redundancy and reduce sensitivity to coefficient settings across teacher-student pairs and training stages. Specifically, CoCaRS calibrates feature decorrelation through Confusion Evidence Estimation (CEE) and Strength Allocation Control (SAC), which respectively capture reliable semantic relations for correlation estimation and preserve discriminative structure during decorrelation. Adaptive Coefficient Regulation (ACR) further regulates the contribution of the calibrated redundancy suppression objective according to its relative loss scale, reducing sensitivity to coefficient settings. Extensive experiments on CIFAR-100 and ImageNet-1K validate the effectiveness of CoCaRS in improving distillation performance and reducing sensitivity to coefficient settings. Code will be released soon.
Introduction Recent progress in visual recognition has been largely driven by advanced neural architectures, such as CNNs (He et al. 2016; Sandler et al. 2018; Liu et al. 2022b), ViTs (Dosovitskiy et al. 2021; Touvron et al. 2021; Liu et al. 2021b) and MLP-based models (Tolstikhin et al. 2021; Touvron et al. 2023). While these models achieve strong performance, their computational and storage costs often hinder deployment in resource-constrained scenarios. Knowledge distillation (KD) (Hinton, Vinyals, and Dean 2015) has therefore become ∗
Corresponding author.
a widely used approach for model compression, where a lightweight student model is trained with guidance from a strong teacher to reduce inference cost with minimal performance degradation. Existing KD methods can be categorized into responsebased, feature-based, and relation-based approaches (Gou, Yu, and Maybank 2021; Pan et al. 2026). Response-based methods transfer probability distributions via soft targets (Hinton, Vinyals, and Dean 2015; Yang et al. 2019; Son et al. 2021), while feature-based methods use intermediate representations as additional supervision (Romero et al. 2015; Heo et al. 2019b; Lin et al. 2022). Relation-based methods further distill structural information, such as relations among samples or feature representations (Yim et al. 2017; Park et al. 2019; Tian, Krishnan, and Isola 2020). These distillation methods have demonstrated their effectiveness in homogeneous settings. Nevertheless, restricting distillation to such homogeneous scenarios limits the flexibility of teacher selection, since a high-performing teacher with the same architecture as the student may not always be available in practice. For heterogeneous model pairs, differences in architectural inductive biases can lead to substantial discrepancies between teacher and student representations, reducing the compatibility of transferred knowledge and resulting in mismatched supervision and suboptimal performance (Raghu et al. 2021; Liu et al. 2022a; Hao et al. 2023). Recent heterogeneous KD studies have explored different ways to adapt teacher supervision across architectures. For example, OFA (Hao et al. 2023) projects intermediate representations into the logit space, FBT (Li et al. 2025) fuses heterogeneous features through an auxiliary model, and PAT (Lin et al. 2025) adapts teacher representations through feature prompting and region-aware attention. Different from these adaptation strategies, RSD (Zhang et al. 2025) formulates heterogeneous KD from the perspective of redundancy suppression. It constructs a feature correlation matrix between teacher and student representations, where the diagonal entries encourage invariance across architectures, while the off-diagonal entries are constrained by a feature decorrelation objective to suppress redundancy, with a fixed coefficient controlling the contribution of the resulting RSD term. However, RSD applies a uniform decorrelation constraint to feature correlations, even though some of them may also encode useful structural information. Consequently, such infor-
mation may be weakened together with redundancy. Moreover, as the RSD term varies in scale relative to the task loss across teacher-student pairs and training stages, a fixed coefficient may yield varying effective contributions and make distillation performance sensitive to coefficient selection. To resolve these problems, Correlation Calibration-based Redundancy Suppression (CoCaRS) is proposed to calibrate feature decorrelation for better preservation of structural information encoded in feature correlations, while adaptively regulating the effective contribution of redundancy suppression to reduce sensitivity to coefficient selection. Specifically, CoCaRS introduces Semantic Correlation Calibration (SCC) to retain cross-architecture invariance while calibrating feature decorrelation through Confusion Evidence Estimation (CEE) and Strength Allocation Control (SAC). CEE derives confusion weights from teacher responses to capture reliable semantic relations for correlation estimation, whereas SAC constructs a semantic strength map from a discriminative subspace induced by the teacher classifier to preserve discriminative structure during decorrelation. Adaptive Coefficient Regulation (ACR) further regulates the contribution of SCC according to its loss scale relative to the task objective, thereby reducing sensitivity to coefficient settings. The main contributions of this work are summarized as follows: • CoCaRS is proposed to refine redundancy suppression for heterogeneous KD through correlation calibration, retaining structural information while adaptively regulating its effective contribution. • In SCC, CEE captures reliable semantic relations for correlation estimation, while SAC preserves discriminative structure during decorrelation. • ACR adaptively regulates the effective contribution of SCC to reduce sensitivity to coefficient settings. • Extensive experiments on CIFAR-100 and ImageNet-1K demonstrate the effectiveness of CoCaRS across diverse heterogeneous teacher-student pairs.
problem from the perspective of redundancy suppression, preserving architecture invariant knowledge while suppressing redundant correlations. Our work follows this direction and further calibrates the decorrelation process.
Related Work
Redundancy Suppression Distillation (RSD) introduces a redundancy suppression perspective for heterogeneous distillation. Given teacher representations f t ∈ RB×D and student s representations fori ∈ RB×ds , an adaptor h(·) maps the student representations into the teacher feature space, yielding f s ∈ RB×D . Here, B is the batch size, while D and ds are the teacher and student feature dimensions. Based on f t and f s , RSD constructs a Pearson correlation matrix P ∈ RD×D between teacher and student feature units: PB t ¯t s ¯s k=1 (fki − fi )(fkj − fj ) Pij = qP , (1) B t ¯t 2 PB (f s − f¯s )2 j k=1 (fki − fi ) k=1 kj
Knowledge Distillation Knowledge distillation transfer knowledge from a teacher model to a smaller student model through soft labels (Hinton, Vinyals, and Dean 2015). Subsequent methods improve student learning with richer teacher supervision (Yang et al. 2019; Son et al. 2021; Lin et al. 2022; Lao et al. 2023; Tian, Krishnan, and Isola 2020; Xu et al. 2022). Knowledge distillation across heterogeneous architectures has also been explored. (Touvron et al. 2021; Ren et al. 2022; Liu et al. 2022a; Zhao, Song, and Liang 2023). These methods usually target specific architecture pairs or fixed transfer directions, limiting their applicability to diverse heterogeneous pairs. Recent studies therefore explore general frameworks for heterogeneous KD. OFA (Hao et al. 2023) projects intermediate representations into logits to alleviate the semantic mismatch. FBT (Li et al. 2025) integrates teacher and student representations through an auxiliary model. PAT (Lin et al. 2025) adapts teacher representations through prompt tuning and aligns student features through region-aware attention. In contrast, RSD (Zhang et al. 2025) formulates this
Semantic-Aware Redundancy Suppression Redundancy suppression has been widely studied in representation learning. Barlow Twins (Zbontar et al. 2021) reduces feature redundancy by driving the cross-correlation matrix between two augmented views toward the identity matrix, while VICReg (Bardes, Ponce, and LeCun 2022) regularizes covariance to reduce dependencies among embedding variables. RSD extends this principle to heterogeneous KD through a teacher-student correlation matrix, whose diagonal terms preserve cross-architecture invariance and off-diagonal terms suppress redundancy through decorrelation. Relational structures in teacher representations have also been explored in KD. SPKD (Tung and Mori 2019) preserves pairwise similarities between samples, CC (Peng et al. 2019) transfers instance correlations, RKD (Park et al. 2019) distills distance and angle relations, and ICKD (Liu et al. 2021a) matches inter-channel correlations. These methods show that correlations in teacher representations can encode structural information, which is relevant to the decorrelation term in redundancy suppression. Beyond these relational structures, SemCKD (Chen et al. 2021a) addresses teacher-student semantic mismatch through adaptive cross-layer calibration. SimKD (Chen et al. 2022) reuses the teacher classifier to guide student feature learning. Neural Collapse (Papyan, Han, and Donoho 2020) reveals the geometric alignment between classifier weights and class means, and NCKD (Zhang, Song, and He 2025) exploits such class geometry for KD. These works motivate semantic calibration and classifier structure in redundancy suppression.
Method Preliminaries
where f¯ denotes the mini-batch mean. Using ID as the target, diagonal entries are driven toward one to preserve cross-architecture invariance, whereas off-diagonal entries are driven toward zero to suppress feature redundancy. The RSD objective can be written as an off-diagonal reweighted mean-square error: D X 1 X P2ij ], (2) LRSD = 2 [ (1 − Pii )2 + κ D i=1 i̸=j
where κ controls the strength of the decorrelation objective. With β weighting the RSD term, the overall objective is L = LCE + βLRSD .
(3)
Overall Framework CoCaRS revisits redundancy suppression from the perspective of correlation calibration, as shown in Fig. 1. Semantic Correlation Calibration (SCC) calibrates feature decorrelation in P while retaining cross-architecture invariance, rather than imposing a uniform constraint. Confusion Evidence Estimation (CEE) derives confusion weights from positive and reciprocal negative evidence to calibrate correlation estimation, while Strength Allocation Control (SAC) constructs a semantic strength map to allocate decorrelation strength. Adaptive Coefficient Regulation (ACR) further regulates the SCC contribution using its loss scale relative to the task objective, reducing sensitivity to coefficient settings. Overall, CoCaRS combines semantic calibration of feature decorrelation with adaptive regulation of the SCC contribution.
Confusion Evidence Estimation In RSD, feature redundancy is suppressed by penalizing the off-diagonal entries of the teacher-student correlation matrix. CEE introduces sample awareness into correlation estimation by estimating semantic confusion evidence for each training sample. Samples with stronger semantic confusion reflected in teacher responses are treated as more informative for correlation estimation. In this way, sample-level importance is incorporated into correlation estimation in SCC through teacher dark knowledge aggregated over retrieved samples, while the original redundancy suppression objective is preserved. Non-target teacher responses are used as semantic cues for confusion evidence, since they encode dark knowledge about semantic relations beyond the ground-truth label (Zhao et al. 2022). Accordingly, given the training set D = {(xi , yi )}N i=1 and a collection of pretrained teachers {Tm }M m=1 , a teacher response bank B is constructed before distillation: B = { ki , yi , zti,1 , . . . , zti,M }N (4) i=1 , where ki = ϕk (xi ) is the retrieval key extracted by a pretrained key encoder ϕk , and zti,m ∈ RC denotes the logits produced by teacher Tm . For each sample xi , a query embedding qi = ϕk (xi ) is used to retrieve a knowledge set Ki from B. Let Ki+ denote the positive knowledge set containing retrieved samples labeled yi . Aggregating teacher logits over Ki+ provides a local estimate of the non-target responses around xi . Given the normalized teacher logits z̃tj,m for each retrieved sample xj , the positive evidence is formulated as + e+ i = Ψyi (
M X X 1 z̃tj,m ), + M |Ki | m=1 +
(5)
xj ∈Ki
where Ψ+ yi (·) masks the ground-truth entry and rescales the C remaining non-target entries. The obtained e+ ini ∈ R dicates which non-target classes are supported as potential semantic competitors to yi in the local neighborhood of xi .
To assess whether these semantic competitors reflect credible confusion with yi , negative evidence is further estimated from Ki− , which contains retrieved samples with labels different from yi . For each non-target class c, let − Ki,c = {xj ∈ Ki− | yj = c}. The corresponding negative evidence entry evaluates the credibility of the semantic confusion between yi and c by measuring the teacher responses − to yi for samples in Ki,c . Formally, it is defined as ( P P M − 1 t − z̃ − j,m,yi , |Ki,c | > 0, m=1 xj ∈Ki,c − M |K | i,c ẽi,c = − 0, |Ki,c | = 0. (6) The entries are assembled according to their class coordinates and calibrated as − − − − e− i = Ψ ([ẽi,1 , ẽi,2 , . . . , ẽi,C ]),
(7)
where Ψ− (·) denotes the calibration of the assembled negC ative evidence vector. The obtained e− serves as i ∈ R calibrated reciprocal evidence for assessing the credibility of the semantic confusion indicated by e+ i . Given the positive and negative evidence defined above, the final confusion weight is obtained under an asymmetric evidence model. The positive evidence determines the support of possible non-target confusion, whereas the negative evidence calibrates the credibility of this support. This design reflects the assumption that reciprocal responses should reinforce an existing confusion pattern rather than create an independent supervision signal. Accordingly, the calibrated confusion weight is defined as − wiconf = e+ i ⊙ (1 + γei ),
(8)
where γ controls the strength of reciprocal calibration. The resulting wiconf ∈ RC is a calibrated confusion weight vector that captures reliable semantic relations between yi and its non-target competitors for subsequent correlation estimation.
Strength Allocation Control SAC calibrates the off-diagonal decorrelation term through a semantic strength map that allocates decorrelation strength according to associations between feature dimensions. Neural collapse theory relates classifier weight vectors to classlevel feature prototypes at convergence (Papyan, Han, and Donoho 2020), and NCKD (Zhang, Song, and He 2025) further exploits this theory for classifier construction in distillation. Motivated by this, SAC uses the teacher classifier weights to induce a discriminative subspace for constructing the semantic strength map. Given the normalized teacher classifier weights W̃, an orthonormal basis Q for the discriminative subspace of the teacher classifier is obtained from W̃⊤ through reduced QR decomposition. When W̃⊤ has full row rank, rank reduction is applied before QR decomposition. The target rank is estimated from the stable rank, which measures the effective rank of a matrix. The semantic matrix is therefore defined as Msem = |QQ⊤ |, which captures the projection structure of the induced discriminative subspace over feature dimensions. Since the diagonal entries serve as invariance anchors
T-Block 1
❄️
T-Block 2
❄️
T-Block 3
❄️
❄️
T-FC
T. Feature
T. Logits
T. Feature
🔥
Proj.
Correlation Estimation
SCC
T-FC Weight
Aligned S. Feature
Aligned S. Feature
🔥
🔥
S-Block 1
S-Block 2
🔥
🔥
S-Block 3
SAC
S-FC S. Feature
Confusion Weight
S. Logits Pressure Calibration
❄️
Strength Map
Key Encoder
Query Embedding
Strength Map
CEE
Response Bank
Retrieved Knowledge
❄️ Frozen 🔥 Learnable
Confusion Weight
(a) Framework Architecture
c1 yi
c1
c2
c3
c4
Semantic Correlation Matrix
(b) Semantic Correlation Calibration (SCC)
c5
c2 c1 c4
Non-GT Mask
Neg. Knowledge
GT Extract
Neg. Evidence Pool
Neg. Evidence
Retrieved Knowledge
Transpose
T-FC Weight
Calib.
Transposed Weight
QR Decomp.
Orthonormal Basis Q
R Factor
Confusion Weight
yi Agg.
Q
GT Mask
Pos. Knowledge
Aggregate
Pos. Evidence Pool
MatMul
Pos. Evidence
Diag. Mask Transpose
(c) Confusion Evidence Estimator (CEE)
QT
Semantic Matrix
Strength Map
(d) Strength Allocation Controller (SAC)
Figure 1: Overview of the proposed CoCaRS framework. (a) CoCaRS refines redundancy suppression for heterogeneous distillation through CEE and SAC; (b) The resulting confusion weights and strength map jointly calibrate feature decorrelation in SCC; (c) CEE derives confusion weights from positive and reciprocal negative evidence in retrieved knowledge; (d) SAC constructs the strength map from a discriminative subspace induced by the teacher classifier through QR decomposition. rather than decorrelation terms, they are excluded to obtain the off-diagonal component Mof f = Msem ⊙ (1 − I). The strength map is defined as Mκij =
of f exp(−τκ Mij ) , i ̸= j, P of f 1 a̸=b exp(−τκ Mab ) D(D−1)
(9)
where τκ controls the semantic modulation strength. A larger f Mof indicates a stronger association between the correij sponding feature dimensions within the induced discriminative subspace, thereby reducing the decorrelation strength applied to the corresponding feature correlation to preserve discriminative structure.
Distillation Formulation Given the confusion weights and the semantic strength map defined above, SCC integrates them into the decorrelation objective while preserving the diagonal invariance term. Let f̂ t and f̂ s denote the normalized teacher and student features. The SCC objective is formulated as " D 1 X LSCC = 2 (1 − Pii )2 D i=1 (10) !2 # B X X conf ˆt ˆs κ . +κ Mij Sk fk,i fk,j i̸=j
k=1
In this objective, κ controls the overall decorrelation strength, while the strength map Mκ specifies its relative coefficients.
Skconf is derived from the confusion weight as Skconf = Norm 1 + α∥wkconf ∥1 ,
(11)
where α controls the effect of confusion evidence on sample weighting in correlation estimation. A basic training objective combines the cross-entropy task loss with the SCC term as L = LCE + λLSCC ,
(12)
where λ controls the strength of the SCC term. However, the relative scale of the SCC term can vary across training stages and teacher-student pairs, making a static coefficient less suitable for maintaining a balanced optimization process. To address this, Adaptive Coefficient Regulation (ACR) regulates the effective contribution of SCC according to its relative loss scale with respect to the task objective: rt =
LtSCC . LtCE + ϵ
(13)
The SCC coefficient is regulated according to the deviation of rt from the target ratio ρ. To reduce fluctuations in loss magnitudes, the coefficient is updated through EMA as ρ λt = ηλt−1 + (1 − η)G , (14) rt + ϵ where η denotes the EMA decay factor, and G(·) is a bounded modulation function. The final objective is defined as L = LCE + λt LSCC ,
(15)
Table 1: Top-1 accuracy (%) on CIFAR-100. The best and second best results are in bold and underlined. From Scratch
Feature-based
Response-based
Heterogeneous-KD
Teacher
Student
T.
S.
FitNet
CC
RKD
CRD
KD
DKD
DIST
OFA
PAT
RSD
CoCaRS
CNN-based students
Swin-T ViT-S Mixer-B/16 Swin-T ViT-S Mixer-B/16
ResNet18 ResNet18 ResNet18 MobileNetV2 MobileNetV2 MobileNetV2
89.26 92.04 87.29 89.26 92.04 87.29
74.01 74.01 74.01 73.68 73.68 73.68
78.87 77.71 77.15 74.28 73.54 73.78
74.19 74.26 74.26 71.19 70.67 70.73
74.11 73.72 73.75 69.00 68.46 68.95
77.63 76.60 76.42 79.80 78.14 78.15
78.74 77.26 77.79 74.68 72.77 73.33
80.26 78.10 78.67 71.07 69.80 70.20
77.75 76.49 76.36 72.89 72.54 73.26
80.54 80.15 79.39 80.98 78.45 78.78
81.22 80.11 80.07 78.78 78.87 78.62
83.92 81.50 81.85 83.68 81.68 81.74
85.42 85.22 83.85 85.50 85.62 84.46
ViT-based students
ConvNeXt-T Mixer-B/16 ConvNeXt-T Mixer-B/16
DeiT-T DeiT-T Swin-P Swin-P
88.41 87.29 88.41 87.29
68.00 68.00 72.63 72.63
60.78 71.05 24.06 75.20
68.01 68.13 72.63 73.32
69.79 69.89 71.73 70.82
65.94 65.35 67.09 67.03
72.99 71.36 76.44 75.93
74.60 73.44 76.80 76.39
73.55 71.67 76.41 75.85
75.76 73.90 78.32 76.65
79.59 74.66 80.74 78.44
82.46 78.50 82.21 81.28
84.09 81.54 85.08 84.05
MLP-based students
ConvNeXt-T Swin-T
ResMLP-S12 ResMLP-S12
88.41 89.26
66.56 66.56
45.47 63.12
67.70 68.37
65.82 64.66
63.35 61.72
72.25 71.89
73.22 72.82
71.93 11.05
75.21 73.58
83.50 80.94
84.21 82.67
86.63 84.99
88.85
71.45
66.25
71.12
70.06
71.44
74.62
74.61
69.15
77.64
79.63
82.14
84.70
Average
Experiments Experimental Setup Models Teacher-student pairs are constructed from models with different architectures. For CNN-based models, ResNet (He et al. 2016), MobileNetV2 (Sandler et al. 2018), and ConvNeXt (Liu et al. 2022b) are adopted. Transformer-based models cover ViT (Dosovitskiy et al. 2021), DeiT (Touvron et al. 2021), Swin Transformer (Liu et al. 2021b), and its lightweight variants, Swin-Pico and Swin-Nano (Hao et al. 2023). MLP-based models include MLP-Mixer (Tolstikhin et al. 2021) and ResMLP (Touvron et al. 2023). Datasets Experiments are conducted on CIFAR-100 (Krizhevsky, Hinton et al. 2009) and ImageNet-1K (Deng et al. 2009). CIFAR-100 contains 60,000 images from 100 classes, with 50,000 images used for training and 10,000 images used for testing. ImageNet-1K is a large-scale dataset containing approximately 1.28 million training images and 50,000 validation images from 1,000 classes. Baselines Several representative KD methods are selected for comparison. Feature-based methods include FitNet (Romero et al. 2015), CC (Peng et al. 2019), RKD (Park et al. 2019), and CRD (Tian, Krishnan, and Isola 2020), which transfer knowledge through intermediate representations or feature relations. Response-based methods include KD (Hinton, Vinyals, and Dean 2015), DKD (Zhao et al. 2022), and DIST (Huang et al. 2022), which align the output responses of teacher and student models. Heterogeneous KD methods, including OFA (Hao et al. 2023), PAT (Lin et al. 2025), and RSD (Zhang et al. 2025), are also compared.
Main Results Results on CIFAR-100 Experiments are conducted on CIFAR-100 using heterogeneous teacher-student pairs involving CNNs, ViTs, and MLPs. As reported in Table 1, CoCaRS achieves the best results across all evaluated pairs, with an average accuracy of 84.70%. Feature-based and responsebased methods obtain average Top-1 accuracies of 69.72% and 72.79%, respectively, and the former is below the scratch baseline of 71.45%. These results indicate the limited applicability of homogeneous methods to heterogeneous teacher-
student pairs, whereas methods designed for heterogeneous KD achieve higher average accuracies. RSD achieves an average accuracy of 82.14% through redundancy suppression, while CoCaRS raises the average accuracy to 84.70% by calibrating feature decorrelation, yielding a 2.56% gain. This result supports the benefit of calibrating feature decorrelation rather than applying a uniform constraint. Results on ImageNet-1K CoCaRS further achieves the highest average Top-1 accuracy of 74.46% on ImageNet1K, outperforming RSD by 0.43% points, as reported in Table 2. Feature-based and response-based methods obtain lower average accuracies than methods developed for heterogeneous KD, reflecting the difficulty of transferring knowledge across heterogeneous architectures. The performance of RSD demonstrates the effectiveness of redundancy suppression in this setting, while the further improvement achieved by CoCaRS supports the benefit of correlation calibration beyond the original formulation. Notably, CoCaRS yields its largest improvement over RSD on Swin-T → ResMLP-S12, with a gain of 0.85%. These results further validate correlation calibration for large-scale heterogeneous distillation.
Ablation Study Effect of Core Components Table 3 shows that removing CEE, SAC, or ACR consistently degrades performance. The performance drops caused by removing CEE or SAC support calibrating both correlation estimation and decorrelation strength rather than applying uniform decorrelation. Meanwhile, the degradation caused by removing ACR supports the need to regulate the effective contribution of SCC according to its loss scale relative to the task objective. Together, these results validate semantic calibration in SCC and adaptive regulation of its contribution. Effect of CEE Formulation The construction of confusion evidence from retrieved knowledge and the reciprocal calibration provided by negative evidence are examined in Table 4. Removing CEE entirely reduces performance, confirming the contribution of confusion evidence to correlation estimation in SCC. Removing negative evidence (w/o NE) also leads to lower performance, supporting its contribution
Table 2: Top-1 accuracy (%) on ImageNet-1K. The best and second best results are in bold and underlined. Teacher
Student
From Scratch
Feature-based
Response-based
Heterogeneous-KD
T.
S.
FitNet
CC
RKD
CRD
KD
DKD
DIST
OFA
PAT
RSD
CoCaRS
81.35 76.55
69.75 68.87
71.18 71.59
70.07 70.79
68.89 69.86
69.09 68.89
71.14 71.92
71.10 70.93
70.91 71.74
71.85 72.12
71.54 72.22
72.13 71.90
72.63 72.15
82.05
72.17
70.45
73.12
71.47
69.18
74.00
73.95
74.07
74.41
74.44
74.46
74.59
81.35
76.65
76.48
76.15
75.10
73.40
76.67
76.99
77.25
77.31
77.59
77.61
78.46
80.33
71.86
72.43
72.53
71.33
70.14
73.43
73.24
73.49
73.92
73.95
74.03
74.46
CNN-based students Swin-T Mixer-B/16
ResNet18 MobileNetV2
ViT-based students ConvNeXt-T
DeiT-T
MLP-based students Swin-T
ResMLP-S12 Average
Table 3: Ablation study on CIFAR-100: effect of the core components of CoCaRS. w/ w/ w/ Swin-T Mixer-B/16 ViT-S ConvNeXt-T Mixer-B/16 ConvNeXt-T CEE SAC ACR ResNet18 ResNet18 MobileNetV2 DeiT-T Swin-P ResMLP-S12 RSD Baseline
83.92
81.85
81.68
82.46
81.28
84.21
✓ –
– ✓
– –
84.38 84.52
82.45 82.60
83.72 83.96
82.95 82.93
83.04 83.35
85.20 85.50
– ✓ ✓
✓ – ✓
✓ ✓ –
84.81 84.71 84.50
83.60 83.37 83.62
85.09 84.94 84.91
83.54 83.65 83.62
83.50 83.44 83.65
85.69 85.53 85.94
✓
✓
✓
85.42
83.85
85.62
84.09
84.05
86.63
to calibrating the credibility of the semantic confusion indicated by positive evidence. Replacing the retrieved samples with randomly selected samples (Random) results in a clear performance decrease, suggesting that an effective local estimate of the non-target responses relies on retrieved samples associated with the query sample. These results support the use of retrieved knowledge for positive evidence estimation and negative evidence for reciprocal calibration. Table 4: Ablation study on CIFAR-100: effects of retrieved knowledge and negative evidence in CEE. CEE Setting
Swin-T Mixer-B/16 ViT-S ConvNeXt-T Mixer-B/16 ConvNeXt-T ResNet18 ResNet18 MobileNetV2 DeiT-T Swin-P ResMLP-S12
w/o CEE w/o NE Random
84.81 84.50 84.60
83.60 83.60 83.22
85.09 85.04 84.64
83.54 83.35 83.20
83.50 83.33 82.55
85.69 85.76 86.05
CoCaRS
85.42
83.85
85.62
84.09
84.05
86.63
Effect of SAC Formulation The construction of the semantic strength map from the induced discriminative subspace and the direction of decorrelation strength allocation are examined in Table 5. Removing SAC entirely reduces performance, showing the contribution of the semantic strength map to decorrelation strength allocation. When the allocation direction is inverted (Inverted), larger values of Mof f lead to greater decorrelation strength, while the resulting performance remains close to that without SAC. This contrast supports allocating lower decorrelation strength to stronger
associations in the induced discriminative subspace. Replacing the induced discriminative subspace with a random orthogonal subspace (Random) also performs worse than the complete formulation, supporting the use of teacher classifier weights as the structural basis of the semantic strength map. These results further validate the proposed direction of decorrelation strength allocation. Table 5: Ablation study on CIFAR-100: effects of strength allocation and discriminative subspace in SAC. SAC Setting
Swin-T Mixer-B/16 ViT-S ConvNeXt-T Mixer-B/16 ConvNeXt-T ResNet18 ResNet18 MobileNetV2 DeiT-T Swin-P ResMLP-S12
w/o SAC Inverted Random
84.71 84.68 84.96
83.37 83.31 83.57
84.94 85.03 84.90
83.65 83.24 83.78
83.44 83.58 83.33
85.53 85.73 86.03
CoCaRS
85.42
83.85
85.62
84.09
84.05
86.63
Effect of Diagonal Invariance The role of diagonal invariance is examined in Table 6. Extending the modulation introduced by CEE, SAC, or both to the diagonal term consistently degrades performance. The diagonal term is intended to preserve cross-architecture invariance between heterogeneous representations. Applying confusion weights or the semantic strength map to this term may weaken its role in preserving cross-architecture invariance. These results support preserving diagonal invariance while restricting CEE and SAC to off-diagonal redundancy suppression. Table 6: Ablation study on CIFAR-100: effect of CEE and SAC on the invariance term of SCC. Invariance Term Swin-T Mixer-B/16 ViT-S ConvNeXt-T Mixer-B/16 ConvNeXt-T CEE SAC ResNet18 ResNet18 MobileNetV2 DeiT-T Swin-P ResMLP-S12 ✓ – ✓
– ✓ ✓
84.48 85.02 84.81
83.46 83.51 83.11
84.99 84.78 84.60
83.01 83.19 82.60
82.95 83.41 83.40
85.28 85.81 85.94
–
–
85.42
83.85
85.62
84.09
84.05
86.63
Further Analysis Performance in Homogeneous Settings To further evaluate CoCaRS under homogeneous settings, experiments are
Table 7: Top-1 accuracy (%) on ImageNet-1K for homo. settings. The best and second best results are in bold and underlined.
ResNet18 MobileNet
Homogeneous-KD
T.
S.
KD
OFD
CRD
RKD
CAT
SimKD
Review
DKD
SDD
DIST
OFA
RSD
CoCaRS
73.31 80.36
69.75 68.58
70.66 68.58
70.81 71.25
71.17 71.37
71.34 71.32
71.26 72.24
71.59 72.25
71.61 72.56
71.70 72.05
71.14 72.24
72.07 73.24
72.10 73.28
72.18 73.08
72.58 74.91
conducted on two teacher-student pairs on ImageNet-1K. The comparison additionally includes OFD (Heo et al. 2019a), CAT (Guo et al. 2023), SimKD (Chen et al. 2022), Review (Chen et al. 2021b), and SDD (Wei, Luo, and Luo 2024). CoCaRS achieves the best performance and outperforms RSD on both pairs, as reported in Table 7. These results further support the effectiveness of introducing correlation calibration into redundancy suppression. Together with the results on heterogeneous teacher-student pairs, this additional evaluation demonstrates the robustness of CoCaRS across different distillation settings.
32.57
1.19
0
5
3.55
0
60
60
40
30
30
30
30
40
0 120 0
30
60
90
Swin-T vs. ResMLP-S12 (RSD)
0 120 0
30
60
90
120
Swin-T vs. ResMLP-S12 (CoCaRS)
Figure 2: Intermediate representation similarity between Swin-T and ResMLP-S12 measured by CKA, where brighter regions indicate higher similarity. Computational Cost A comparison of additional trainable parameters and peak memory usage across methods is presented in Fig. 3. CoCaRS requires fewer additional trainable parameters than OFA and PAT while matching that of RSD, since its proposed components introduce no additional learnable parameters. Its peak memory usage is comparable to that of OFA and substantially lower than that of PAT. Compared to RSD, CoCaRS achieves a better overall balance between distillation performance and training memory, despite its moderate increase in memory overhead. Since the additional components are used only during distillation, the original inference cost of the student is preserved. ACR for SCC Balance As shown in Fig. 4(a), the relative scale of SCC to CE differs across the three heterogeneous model pairs and varies during training. Consequently, without ACR, the same fixed coefficient may produce different
SCC / CE
60
90
Mixer-B/16 -> Swin-P
8.12 3.74
6.84
3.99 4.65
8.16
Mixer-B/16 -> ResNet18 ConvNeXt-T -> ResMLP-S12
(b) Effective SCC Contribution CX ResMLP SwinT MNV2 Mixer-DeiT
0.40
80
60
60
6.70 7.23
Before ACR
90
30
14.56
(a) Relative Loss Scale
120
Swin-T vs. ResMLP-S12 (w/o KD)
0.89 0.89
19.79
effective contributions of SCC, making it difficult to maintain a consistent balance with the task objective across model pairs and training stages. Figure 4(b) further shows that the SCC coefficient is regulated to different extents for the three model pairs according to their relative loss scales. This regulation maintains adaptive control over the contribution of SCC despite variations in loss scale, thereby reducing sensitivity to coefficient selection.
90
0 120 0
3.45 3.80
1.31 1.31
Figure 3: Comparison of additional trainable parameters and peak memory usage on CIFAR-100.
120
90
7.94
ViT-S -> ResNet18
90
60
1.63
Mixer-B/16 -> ResNet18 ConvNeXt-T -> ResMLP-S12
17.23 10.56
10
120
Swin-T vs. Swin-T
0.89 0.89
Mixer-B/16 -> Swin-P
15
90
30
0.92 0.92
OFA PAT RSD CoCaRS
14.94 20.77
14.86
ViT-S -> ResNet18
20
120
0
87.82
84.63
80 60 40 20
SCC / CE
Intermediate Representation Similarity The similarity between the intermediate representations of a Swin-T teacher and a ResMLP-S12 student is visualized using CKA (Kornblith et al. 2019) in Fig. 2. Without KD, the student exhibits relatively low feature similarity with the teacher, reflecting the representation discrepancy between heterogeneous architectures. Both RSD and CoCaRS improve the similarity between teacher and student intermediate features. Compared with RSD, CoCaRS exhibits higher similarity, particularly in the shallow and deep regions. The reduced representation discrepancy is consistent with the improved distillation performance and further supports the effectiveness of CoCaRS.
0
Heterogeneous-KD
0.35
60
After ARC
20
SCC Contribution
ResNet34 ResNet50
From Scratch
Extra Params (M)
Student
Peak Memory (GB)
Teacher
0.30 0.25 0.20 0.15 0.10
0
50
100
150
200
250
300 epoch
0.05
50
100
150
200
250
300 Epoch
Figure 4: ACR regulates the effective contribution of SCC across heterogeneous model pairs during training.
Conclusion In this work, redundancy suppression in heterogeneous knowledge distillation was revisited by considering the limitation of uniform decorrelation, which may weaken useful structural information encoded in feature correlations. CoCaRS addresses this limitation through semantic calibration of feature decorrelation while preserving cross-architecture invariance, together with adaptive regulation of the calibrated objective during optimization. Experimental results across heterogeneous model pairs consistently support the effectiveness of this formulation. Overall, these findings demonstrate the effectiveness of semantic correlation calibration in improving redundancy suppression for heterogeneous knowledge distillation.
References Bardes, A.; Ponce, J.; and LeCun, Y. 2022. VICReg: Variance-Invariance-Covariance Regularization for SelfSupervised Learning. In The Tenth International Conference on Learning Representations, ICLR 2022. OpenReview.net. Chen, D.; Mei, J.; Zhang, H.; Wang, C.; Feng, Y.; and Chen, C. 2022. Knowledge Distillation with the Reused Teacher Classifier. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, 11923–11932. IEEE. Chen, D.; Mei, J.; Zhang, Y.; Wang, C.; Wang, Z.; Feng, Y.; and Chen, C. 2021a. Cross-Layer Distillation with Semantic Calibration. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, 7028–7036. AAAI Press. Chen, P.; Liu, S.; Zhao, H.; and Jia, J. 2021b. Distilling Knowledge via Knowledge Review. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, 5008–5017. Computer Vision Foundation / IEEE. Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and FeiFei, L. 2009. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, 248–255. Ieee. Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In 9th International Conference on Learning Representations, ICLR 2021. OpenReview.net. Gou, J.; Yu, B.; and Maybank, S. J. 2021. Knowledge distillation: A survey. International Journal of Computer Vision, 129(6): 1789–1819. Guo, Z.; Yan, H.; Li, H.; and Lin, X. 2023. Class Attention Transfer Based Knowledge Distillation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, 11868–11877. IEEE. Hao, Z.; Guo, J.; Han, K.; Tang, Y.; Hu, H.; Wang, Y.; and Xu, C. 2023. One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge Distillation. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023. He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep Residual Learning for Image Recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, 770–778. IEEE Computer Society. Heo, B.; Kim, J.; Yun, S.; Park, H.; Kwak, N.; and Choi, J. Y. 2019a. A Comprehensive Overhaul of Feature Distillation. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, 1921–1930. IEEE. Heo, B.; Lee, M.; Yun, S.; and Choi, J. Y. 2019b. Knowledge Transfer via Distillation of Activation Boundaries Formed by Hidden Neurons. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, 3779–3787. AAAI Press. Hinton, G. E.; Vinyals, O.; and Dean, J. 2015. Distilling the Knowledge in a Neural Network. CoRR, abs/1503.02531.
Huang, T.; You, S.; Wang, F.; Qian, C.; and Xu, C. 2022. Knowledge Distillation from A Stronger Teacher. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022. Kornblith, S.; Norouzi, M.; Lee, H.; and Hinton, G. 2019. Similarity of Neural Network Representations Revisited. In Proceedings of the 36th International Conference on Machine Learning, volume 97, 3519–3529. PMLR. Krizhevsky, A.; Hinton, G.; et al. 2009. Learning multiple layers of features from tiny images. Lao, S.; Song, G.; Liu, B.; Liu, Y.; and Yang, Y. 2023. Masked Autoencoders Are Stronger Knowledge Distillers. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, 6361–6370. IEEE. Li, G.; Wang, Q.; Yan, K.; Ding, S.; Gao, Y.; and Xia, G. 2025. Fuse Before Transfer: Knowledge Fusion for Heterogeneous Distillation. In IEEE/CVF International Conference on Computer Vision, ICCV 2025, 3445–3454. IEEE. Lin, J.; Yao, Y.; Hsu, C.; Xie, H.; Shuai, H.; and Cheng, W. 2025. Perspective-Aware Teaching: Adapting Knowledge for Heterogeneous Distillation. In IEEE/CVF International Conference on Computer Vision, ICCV 2025, 4178–4187. IEEE. Lin, S.; Xie, H.; Wang, B.; Yu, K.; Chang, X.; Liang, X.; and Wang, G. 2022. Knowledge Distillation via the Target-aware Transformer. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, 10905–10914. IEEE. Liu, L.; Huang, Q.; Lin, S.; Xie, H.; Wang, B.; Chang, X.; and Liang, X. 2021a. Exploring Inter-Channel Correlation for Diversity-preserved Knowledge Distillation. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, 8251–8260. IEEE. Liu, Y.; Cao, J.; Li, B.; Hu, W.; Ding, J.; and Li, L. 2022a. Cross-Architecture Knowledge Distillation. In Computer Vision - ACCV 2022, volume 13845, 179–195. Springer. Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021b. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, 9992–10002. IEEE. Liu, Z.; Mao, H.; Wu, C.; Feichtenhofer, C.; Darrell, T.; and Xie, S. 2022b. A ConvNet for the 2020s. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, 11966–11976. IEEE. Pan, H.; Yu, F.; Zhang, K.; Lan, H.; Meng, Q.; and Li, Z. 2026. Knowledge Distillation in Visual Algorithms: A Survey. Journal of Computer Research and Development, 63(1): 90–122. Papyan, V.; Han, X. Y.; and Donoho, D. L. 2020. Prevalence of Neural Collapse during the terminal phase of deep learning training. CoRR, abs/2008.08186. Park, W.; Kim, D.; Lu, Y.; and Cho, M. 2019. Relational Knowledge Distillation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, 3967–3976. Computer Vision Foundation / IEEE.
Peng, B.; Jin, X.; Li, D.; Zhou, S.; Wu, Y.; Liu, J.; Zhang, Z.; and Liu, Y. 2019. Correlation Congruence for Knowledge Distillation. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, 5006–5015. IEEE. Raghu, M.; Unterthiner, T.; Kornblith, S.; Zhang, C.; and Dosovitskiy, A. 2021. Do Vision Transformers See Like Convolutional Neural Networks? In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, 12116–12128. Ren, S.; Gao, Z.; Hua, T.; Xue, Z.; Tian, Y.; He, S.; and Zhao, H. 2022. Co-advise: Cross Inductive Bias Distillation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, 16752–16761. IEEE. Romero, A.; Ballas, N.; Kahou, S. E.; Chassang, A.; Gatta, C.; and Bengio, Y. 2015. FitNets: Hints for Thin Deep Nets. In 3rd International Conference on Learning Representations, ICLR 2015. Sandler, M.; Howard, A. G.; Zhu, M.; Zhmoginov, A.; and Chen, L. 2018. MobileNetV2: Inverted Residuals and Linear Bottlenecks. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, 4510–4520. Computer Vision Foundation / IEEE Computer Society. Son, W.; Na, J.; Choi, J.; and Hwang, W. 2021. Densely Guided Knowledge Distillation using Multiple Teacher Assistants. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, 9375–9384. IEEE. Tian, Y.; Krishnan, D.; and Isola, P. 2020. Contrastive Representation Distillation. In 8th International Conference on Learning Representations, ICLR 2020. OpenReview.net. Tolstikhin, I. O.; Houlsby, N.; Kolesnikov, A.; Beyer, L.; Zhai, X.; Unterthiner, T.; Yung, J.; Steiner, A.; Keysers, D.; Uszkoreit, J.; Lucic, M.; and Dosovitskiy, A. 2021. MLP-Mixer: An all-MLP Architecture for Vision. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, 24261–24272. Touvron, H.; Bojanowski, P.; Caron, M.; Cord, M.; El-Nouby, A.; Grave, E.; Izacard, G.; Joulin, A.; Synnaeve, G.; Verbeek, J.; and Jégou, H. 2023. ResMLP: Feedforward Networks for Image Classification With Data-Efficient Training. IEEE Trans. Pattern Anal. Mach. Intell., 45(4): 5314–5321. Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and Jégou, H. 2021. Training data-efficient image transformers & distillation through attention. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, volume 139, 10347–10357. PMLR. Tung, F.; and Mori, G. 2019. Similarity-Preserving Knowledge Distillation. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, 1365–1374. IEEE. Wei, S.; Luo, C.; and Luo, Y. 2024. Scale Decoupled Distillation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, 15975–15983. IEEE. Xu, H.; Fang, J.; Zhang, X.; Xie, L.; Wang, X.; Dai, W.; Xiong, H.; and Tian, Q. 2022. Bag of Instances Aggregation
Boosts Self-supervised Distillation. In The Tenth International Conference on Learning Representations, ICLR 2022. OpenReview.net. Yang, C.; Xie, L.; Qiao, S.; and Yuille, A. L. 2019. Training Deep Neural Networks in Generations: A More Tolerant Teacher Educates Better Students. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, 5628– 5635. AAAI Press. Yim, J.; Joo, D.; Bae, J.; and Kim, J. 2017. A Gift from Knowledge Distillation: Fast Optimization, Network Minimization and Transfer Learning. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, 7130–7138. IEEE Computer Society. Zbontar, J.; Jing, L.; Misra, I.; LeCun, Y.; and Deny, S. 2021. Barlow Twins: Self-Supervised Learning via Redundancy Reduction. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, volume 139, 12310–12320. PMLR. Zhang, S.; Song, Z.; and He, K. 2025. Neural Collapse Inspired Knowledge Distillation. In Thirty-Ninth AAAI Conference on Artificial Intelligence, AAAI 2025, 22542–22550. AAAI Press. Zhang, W.; Liu, Y.; Ran, W.; and Ma, C. 2025. CrossArchitecture Distillation Made Simple with Redundancy Suppression. In IEEE/CVF International Conference on Computer Vision, ICCV 2025, Honolulu, HI, USA, October 19-25, 2025, 23256–23266. IEEE. Zhao, B.; Cui, Q.; Song, R.; Qiu, Y.; and Liang, J. 2022. Decoupled Knowledge Distillation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, 11943–11952. IEEE. Zhao, B.; Song, R.; and Liang, J. 2023. Cumulative Spatial Knowledge Distillation for Vision Transformers. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, 6123–6132. IEEE.