OCT-FedSIR: Toward Trustworthy Federated Ophthalmic
arXiv:2609.14734v1 [cs.LG] 13 Sep 2026
Learning under Annotation Noise Sina Gholami1 , Abdulmoneam Ali1 , Tania Haghighi1 , Rashadul H. Badhon1 , Behafarin Emam1 , Sally S. Y. Ong2 , Atalie C. Thompson2 , Theodore Leng4 , Ahmed Arafa1 , Jennifer I. Lim3 , and Minhaj Nur Alam*1 1 Department of Electrical and Computer Engineering, University of North Carolina at
Charlotte, Charlotte, NC, USA 2 Department of Ophthalmology, Wake Forest School of Medicine, Winston-Salem, NC, USA 3 Department of Ophthalmology and Visual Sciences, University of Illinois Chicago, Chicago,
IL, USA 4 Byers Eye Institute at Stanford, Stanford University School of Medicine, Stanford, CA, USA
Abstract Federated learning enables collaborative model development without centralizing patient data, but the reliability of annotations retained at participating institutions cannot always be assumed. This creates a particular challenge in ophthalmic imaging, where legitimate differences in disease prevalence and local class composition may resemble changes produced by corrupted supervision. We introduce OCT-FedSIR, a reliability-aware spectral framework for federated OCT classification under client-dependent annotation noise and heterogeneous data distributions. OCT-FedSIR combines class-balanced spectral estimation, Stage-I logit adjustment, and complementary spectral descriptors to distinguish annotation-related disruption from statistical heterogeneity, followed by selective spectral relabeling and noise-aware federated optimization. The framework was evaluated on three ophthalmic classification tasks using the Kermany, University of Illinois Chicago, and Wake Forest datasets under symmetric and structured asymmetric annotation noise and three levels of non-IID heterogeneity. Across 117 experimental conditions, OCT-FedSIR achieved a mean accuracy of 86.73%, compared with 79.94% for RoFL and 78.75% for FedCorr. Under the controlled corruption settings, OCT-FedSIR correctly separated clients with original and corrupted annotations across all evaluated conditions, whereas the original FedSIR identification procedure showed reduced performance, particularly under asymmetric noise. Spectral relabeling recovered 77.2% of intentionally corrupted annotations on average, with a mean correction precision of 91.3% and a mean false-correction rate of 3.5%. Retaining corrected clients also outperformed spectral pruning by 9.30 percentage points on average. These findings identify annotation reliability as an important consideration for trustworthy ophthalmic federated learning and show that corrupted supervision can often be corrected without discarding potentially informative client data. * Corresponding author: [email protected]
1
Introduction Optical coherence tomography (OCT) has become a cornerstone of retinal care by enabling detailed visualization of retinal morphology that supports disease detection, characterization, and longitudinal monitoring. Deep learning has shown promise for automated OCT interpretation, including disease classification from retinal B-scans, detection and quantification of pathological features, and clinically oriented diagnosis and referral from volumetric OCT imaging.1–4 Together, these studies, including our recent studies,5–7 have established that diagnostically relevant information can be learned directly from OCT and have motivated the development of models capable of operating across broader and more heterogeneous clinical settings. Translation across institutions, however, requires models to accommodate differences in patient populations, disease prevalence, imaging systems, acquisition protocols, and clinical workflows. Because ophthalmic data are naturally distributed across hospitals and imaging centers, centralizing sufficiently diverse datasets may be constrained by patient-privacy requirements, including compliance with the Health Insurance Portability and Accountability Act (HIPAA), as well as data-ownership and institutional-governance considerations. Federated learning (FL)8 provides an alternative by enabling multiple institutions to optimize a shared model while retaining patient data locally. In ophthalmology, FL has been investigated across retinal image segmentation, diabetic retinopathy (DR) classification, retinopathy of prematurity, OCT-based age-related macular degeneration (AMD) classification, and glaucoma detection.9–12 More recent studies have extended federated and distributed learning to multi-disease retinal classification, self-supervised representation learning, masked-autoencoder pretraining, and personalized OCT classification.13–16 Collectively, these studies demonstrate the feasibility of collaborative ophthalmic model development without centralizing raw images, but have largely focused on statistical and domain heterogeneity rather than whether the supervision available at participating clients is itself reliable. This distinction is important for trustworthy FL. Retaining patient images locally addresses where data are stored, but does not ensure that the annotations used to train a shared model are reliable. Retinal labels may be assigned prospectively by specialist graders or derived retrospectively from clinical documentation, diagnostic codes, and electronic health records, and these sources do not provide uniform annotation quality.17 Agreement may vary with grader expertise, diagnostic criteria, image quality, disease severity, coexisting pathology, and the availability of complementary clinical information. Intergrader variability has been documented throughout medical and ophthalmic image interpretation, particularly for conditions represented along a severity continuum.18, 19 OCT presents an additional challenge because diagnostically distinct retinal disorders may share overlapping structural findings, and grading decisions may become uncertain near disease or severity boundaries. Annotation errors may therefore be structured rather than uniformly distributed across samples or diagnostic categories. The consequences of unreliable supervision can be substantial. Label noise is known to impair generalization in medical image analysis.20, 21 In OCT classification, compensating for label-noise rates of 10%, 15%, and 20% has been estimated to require approximately four-, nine-, and fourteen-fold increases in training data, respectively, to recover performance obtained with clean annotations.22 Retinal imaging studies have similarly demonstrated progressive performance degradation with increasing annotation 2
noise and partial recovery following automated label cleaning.23 Correction itself, however, introduces another risk: corrupted annotations may remain undetected, while labels that were originally correct may be modified incorrectly.23 Reliable learning under annotation noise therefore requires more than improving downstream accuracy; it requires determining which supervision is unreliable, whether corrupted labels can be recovered, and whether correction can be performed without substantially damaging annotations that were already correct. This problem becomes particularly challenging in FL because annotation reliability must be inferred from heterogeneous local datasets without centralizing the underlying images. Existing noisy-label FL methods have addressed unreliable supervision through reliable-sample selection, confidence- or lossbased weighting, noisy-client identification, knowledge distillation (KD), robust aggregation, and optimization of corrected label distributions.24–30 FedSIR31 introduced a complementary perspective by exploiting the spectral structure of class-specific feature representations. Rather than relying primarily on prediction confidence or training loss, FedSIR uses changes in feature geometry to identify clients affected by label corruption and constructs spectral references from identified clean clients to guide selective relabeling. This representation-level strategy demonstrated strong robustness to severe label noise and non-IID heterogeneity on federated benchmark datasets. Its translation to ophthalmic data, however, exposes a fundamental challenge: an unusual client is not necessarily an unreliable client. Differences in local disease prevalence, class imbalance, incomplete representation of diagnostic categories, and strongly non-IID client distributions can alter learned feature geometry even when annotations are entirely correct. Consequently, spectral deviations caused by legitimate differences in clinical case mix may resemble those caused by annotation corruption. This creates an important challenge for trustworthy ophthalmic FL: distinguishing genuine clinical heterogeneity from unreliable supervision before using that distinction to modify labels or client contributions. To address this problem, we introduce OCT-FedSIR, a spectral-guided framework for federated OCT classification under client-dependent annotation noise and heterogeneous local data distributions. OCTFedSIR strengthens client characterization through class-balanced spectral estimation, Stage-I logit adjustment (LA), and complementary similarity descriptors designed to reduce sensitivity to differences in local class composition. Rather than automatically excluding clients identified as unreliable, spectral information learned from reliable clients is used to guide selective relabeling, allowing potentially informative local data to remain available for federated optimization. Importantly, the framework evaluates annotation recovery itself in addition to downstream classification performance, enabling both successful recovery and harmful correction of originally correct labels to be quantified. The main contributions of this study are: • We introduce OCT-FedSIR, a reliability-aware spectral framework for ophthalmic FL under annotation noise. OCT-FedSIR combines class-balanced spectral estimation, Stage-I LA, and complementary spectral similarity descriptors to improve discrimination between annotation-related disruption and legitimate variation in local disease composition. • We evaluate spectral relabeling at the individual-sample level by quantifying recovery of corrupted annotations, preservation of originally correct labels, correction precision, and residual label-noise 3
rate. We further compare correction and retention of clients identified as unreliable with spectral pruning to determine whether useful local information can be recovered rather than discarded. • We evaluate OCT-FedSIR across three distinct ophthalmic classification settings: multiclass OCT disease classification using the Kermany dataset, ordinal DR classification using a University of Illinois Chicago (UIC) cohort, and binary geographic atrophy (GA) classification using a Wake Forest (WF) cohort. Experiments span symmetric and structured asymmetric annotation noise, multiple noise ratios, and varying degrees of non-IID client heterogeneity, with comparisons against conventional FL and dedicated noisy-label FL approaches. Through these experiments, we investigate whether spectral characteristics can distinguish annotationrelated disruption from legitimate client heterogeneity, how reliably corrupted supervision can be recovered without damaging originally correct annotations, and whether retaining corrected clients provides greater benefit than excluding them from federated optimization. More broadly, this study examines annotation reliability as an important and previously underexplored requirement for trustworthy ophthalmic FL.
Figure 1: Overview of the OCT-FedSIR framework. a, Global initialization is distributed to participating clients. b, Local training. c–e, Spectral statistics are estimated, communicated to the server, and used to identify clean and noisy clients. f, Clean-client spectral information is used to construct representative and residual references for relabeling. g, Noise-aware federated optimization using LA, KD, and DaAgg.
Results We evaluated OCT-FedSIR across three distinct ophthalmic classification tasks: multiclass OCT disease classification using Kermany, ordinal DR severity classification using UIC, and binary GA classifica4
Table 1: Mean classification accuracy (%) across label-noise levels and client-heterogeneity settings. Symmetric and asymmetric results are averaged over 21 and 18 experimental conditions per dataset, respectively, and overall accuracy is averaged across all 117 conditions. Bold and underline indicate the best- and second-best-performing noisy-label FL methods, respectively. Pruning and all-clean FedAvg are shown as reference configurations. Multi-class (Kermany)
Ordinal DR severity (UIC)
Binary GA (WF)
Method
Sym.
Asym.
Sym.
Asym.
Sym.
Asym.
FedAvg8 FedProx32 RoFL24 RHFL26 FedLSR33 FedCorr27 FedNed29 FedELC30 FedNoRo28 OCT-FedSIR (Ours)
70.13 69.85 87.56 44.53 67.46 86.19 69.21 60.01 64.23 94.76
53.79 52.28 82.24 33.53 62.32 81.76 54.20 50.84 61.48 93.59
68.71 67.40 81.69 52.65 71.55 81.57 67.46 64.36 76.07 88.39
57.10 55.80 81.94 38.39 64.99 83.43 56.77 55.79 68.04 89.57
65.82 64.89 72.45 59.64 69.81 68.88 64.72 61.56 70.24 75.79
66.44 66.32 73.49 45.42 71.86 70.62 59.52 62.36 69.62 78.47
64.02 63.11 79.94 46.20 68.12 78.75 62.38 59.37 68.43 86.73
Pruning FedAvg (all clients clean)
81.23 98.97
84.20 98.97
83.47 94.47
79.97 94.47
67.22 81.86
68.53 81.86
77.43 91.76
Overall
tion using WF. Experiments included controlled symmetric and structured asymmetric annotation noise across multiple noise ratios and three degrees of client heterogeneity, α ∈ {2.0, 0.5, 0.1}, with lower values representing increasingly heterogeneous non-IID client distributions. OCT-FedSIR was compared with conventional FL baselines, including FedAvg8 and FedProx,32 and dedicated noisy-label FL methods, including RoFL,24 RHFL,26 FedLSR,33 FedCorr,27 FedNed,29 FedELC,30 and FedNoRo.28 We additionally evaluated a pruning baseline in which clients identified as having corrupted annotations were excluded from subsequent federated optimization, and an all-clean FedAvg reference in which all participating clients retained their original annotations.
Overall performance Table 1 summarizes classification accuracy across datasets, noise mechanisms, noise levels, and clientheterogeneity settings. For each dataset, symmetric results were averaged across 21 experimental conditions, corresponding to seven noise levels ranging from 30% to 90% with 10% increments and three values of α, whereas asymmetric results were averaged across 18 conditions, corresponding to six noise levels ranging from 40% to 90% and the same three values of α. The overall result therefore summarizes 117 condition-level mean accuracies. Across all 117 evaluated conditions, OCT-FedSIR achieved the highest mean accuracy among methods trained in the presence of noisy clients, reaching 86.73%. The next strongest methods were RoFL at 5
Table 2: Clean-client identification F1 score (%) for FedSIR and OCT-FedSIR, averaged across the evaluated label-noise levels and Dirichlet heterogeneity settings. Bold indicates the best-performing method. Dataset
Method
Symmetric F1 (%)
Asymmetric F1 (%)
Kermany
FedSIR OCT-FedSIR
92.4 100.0
71.8 100.0
UIC
FedSIR OCT-FedSIR
89.7 100.0
68.5 100.0
WF
FedSIR OCT-FedSIR
91.2 100.0
74.6 100.0
79.94% and FedCorr at 78.75%, corresponding to absolute differences of 6.79 and 7.98 percentage points, respectively. On Kermany, OCT-FedSIR achieved mean accuracies of 94.76% and 93.59% under symmetric and asymmetric noise, respectively. The corresponding values were 88.39% and 89.57% for UIC and 75.79% and 78.47% for WF. The all-clean FedAvg reference achieved an overall mean accuracy of 91.76%, a 5.03 percentage-point difference relative to OCT-FedSIR under corrupted supervision. In addition to overall accuracy, we examined macro-F1 to characterize performance across diagnostic classes. Macro-F1 results under symmetric and asymmetric label noise are shown in Figs. 3 and 4, respectively. Across the three ophthalmic tasks, the macro-F1 results generally followed the trends observed for classification accuracy. OCT-FedSIR maintained comparatively strong macro-F1 across increasing label noise levels, with the greatest degradation occurring when severe symmetric label noise was combined with strong client heterogeneity.
Spectral identification and annotation recovery We next evaluated the two stages that precede noise-aware federated optimization: identification of clients with corrupted labels and recovery of their sample-level supervision. We first examined whether the FedSIR spectral identification procedure transferred directly to the ophthalmic setting considered here, particularly under structured asymmetric label noise. For this comparison, FedSIR and OCT-FedSIR procedures were evaluated using the same client partitions, corrupted labels, model architectures, and experimental seeds. The FedSIR identification procedure remained comparatively effective under symmetric noise but showed reduced clean-client F1 under asymmetric noise. The principal failure mode was over-assignment to the clean component: true clean clients were generally retained, but additional clients affected by label noise were also classified as clean. This reduced the purity of the clean reference subset used for subsequent spectral construction. In contrast, the OCT-FedSIR identification procedure completely separated clients retaining their original annotations from clients subjected to simulated label corruption across both noise mechanisms (Table 2). Across all evaluated datasets, noise mechanisms, noise ratios, and client-heterogeneity settings, OCT-FedSIR correctly recovered the simulated annotation6
status partition. Identification accuracy was 100%, with 100% recall for clients subjected to label corruption and 100% precision for clients assigned to the clean reference subset. These results apply to the controlled corruption settings evaluated here. (Fig. 2 shows that complete separation was retained under both symmetric and asymmetric label noise, including the most heterogeneous configuration α = 0.1). Compared with the original identification formulation, the extended procedure incorporated classbalanced spectral estimation, Stage-I LA, and the complementary maximum off-diagonal similarity descriptor smax . The resulting spectral representations showed greater separation between clients with clean k and corrupted labels. Corrupted-label clients exhibited greater cross-class alignment, reflected by the mean off-diagonal similarity µk , off-diagonal energy ek , and maximum off-diagonal similarity smax , which k jointly formed the descriptor space used by the two-component Gaussian Mixture model (GMM). We then examined whether identifying corrupted-label clients resulted in recovery of their local annotations. Across the six dataset–noise summaries, spectral relabeling restored a mean of 77.2% of intentionally corrupted labels to their reference class (Table 3). Mean recovery was 75.7% under symmetric noise and 78.7% under asymmetric noise. Recovery rates under symmetric and asymmetric noise were 81.6% and 84.3% for Kermany, 75.2% and 78.6% for UIC, and 70.4% and 73.1% for WF, respectively. Accepted corrections were predominantly consistent with the reference labels. Correction precision ranged from 86.7% to 95.0%, with the highest values observed on Kermany. Across the six dataset–noise summaries, the mean false-correction rate was 3.5%, corresponding to preservation of approximately 96.5% of annotations that were initially correct. False-correction rates ranged from 2.4% to 4.8%. The proportion of annotations modified by OCT-FedSIR varied according to dataset and noise mechanism. Under symmetric noise, edit rates were 52.0%, 49.7%, and 48.7% for Kermany, UIC, and WF, respectively. Under asymmetric noise, the corresponding values were 57.7%, 55.5%, and 31.1%. The lower edit rate observed for WF under asymmetric noise was accompanied by a recovery rate of 73.1% and correction precision of 89.2%, demonstrating that a lower overall frequency of label changes did not preclude recovery of a substantial fraction of corrupted annotations. Together, the client- and sample-level analyses show that the spectral procedure not only identified clients affected by simulated annotation noise but also recovered a substantial fraction of the corrupted supervision while producing comparatively few harmful changes to labels that were originally correct.
Dataset-specific robustness Kermany OCT-FedSIR remained robust across the Kermany experiments, with differences among methods becoming most apparent under severe label noise (Tables S1 and S2; Figs. 3a and 4a). The macro-F1 curves showed a pattern consistent with the accuracy results, with OCT-FedSIR maintaining strong class-balanced performance across most symmetric and asymmetric noise settings. Under symmetric noise with α = 2, OCT-FedSIR achieved 96.00% accuracy at 80% noise and 93.60% at 90%, compared with 83.50% for RoFL, 20.60% for FedAvg, and 19.70% for FedProx at the 90% noise level. With α = 0.5, OCT-FedSIR remained between 97.30% and 99.70% across the complete 30–90% symmetric noise range. Performance decreased under the most heterogeneous configuration. At α = 0.1, OCT-FedSIR achieved 74.10% and 74.20% ac7
(a) Clean clients
(b) Asymmetric label noise
(d) Distributions of the client-level spectral descriptors µk , ek , and smax for clean k and noisy clients.
(c) Symmetric label noise
(e) Joint spectral descriptor space formed by (µk , ek , smax ). k
(f) Client-identification confusion matrices.
Figure 2: Spectral characterization and identification of clients with corrupted labels in representative Kermany experiments under severe statistical heterogeneity (α = 0.1). a–c, Class-wise spectral similarity matrices for clean, asymmetric-noise, and symmetric-noise clients, respectively. d, Distributions of mean off-diagonal similarity (µk ), off-diagonal energy (ek ), and maximum off-diagonal similarity (smax ). e, Joint k spectral descriptor space. f, Client-identification confusion matrices across the evaluated symmetric and asymmetric noise levels.
8
(a) Kermany.
(b) UIC.
(c) WF.
Figure 3: Macro-F1 performance under symmetric label noise. Results are shown across α ∈ {2.0, 0.5, 0.1} for a, Kermany OCT classification; b, UIC DR classification; and c, WF GA classification. The dashed line denotes the all-clean FedAvg reference.
9
(a) Kermany.
(b) UIC.
(c) WF.
Figure 4: Macro-F1 performance under asymmetric label noise. Results are shown across α ∈ {2.0, 0.5, 0.1} for a, Kermany OCT classification; b, UIC DR classification; and c, WF GA classification. The dashed line denotes the all-clean FedAvg reference.
10
Table 3: Sample-level spectral relabeling performance at clients selected for correction. Values are reported as mean ± standard deviation across three independent seeds and averaged across the evaluated label-noise levels and Dirichlet heterogeneity settings. For WF under asymmetric noise, the target corruption ratio applies only to the GA source class. Dataset
Edit rate (%)
Recovery (%)
Correction precision (%)
False-correction rate (%)
Residual noise (%)
2.7 ± 0.4 3.9 ± 0.6 4.8 ± 0.7
12.1 ± 1.5 16.4 ± 2.2 19.7 ± 2.7
2.4 ± 0.4 3.4 ± 0.5 3.7 ± 0.6
11.0 ± 1.5 15.1 ± 2.0 12.5 ± 1.7
a. Symmetric noise (mean target noise = 60%) Kermany UIC WF
52.0 ± 1.8 49.7 ± 2.2 48.7 ± 2.6
81.6 ± 2.1 75.2 ± 2.7 70.4 ± 3.1
94.1 ± 0.9 90.8 ± 1.2 86.7 ± 1.5
b. Asymmetric noise (mean target noise = 65%) Kermany UIC WF
57.7 ± 2.0 55.5 ± 2.4 31.1 ± 2.8
84.3 ± 1.9 78.6 ± 2.4 73.1 ± 2.9
95.0 ± 0.8 92.1 ± 1.1 89.2 ± 1.4
curacy at 80% and 90% symmetric noise, respectively, compared with 72.60% and 72.90% for RoFL. A similar convergence among the methods was observed in macro-F1 at the most severe symmetric-noise levels (Fig. 3a). OCT-FedSIR showed greater stability under asymmetric noise, which was also reflected in macro-F1 (Fig. 4a). With α = 2, accuracy remained between 96.70% and 98.60% across all evaluated noise levels and reached 97.10% at 90% noise, compared with 91.20% for FedCorr and 71.90% for RoFL. At α = 0.5, OCT-FedSIR retained 95.70% accuracy at 90% noise. Under the strongest heterogeneity (α = 0.1), accuracy decreased from 95.50% at 40% noise to 75.60% at 90%, but remained above RoFL (71.50%), FedNoRo (64.10%), and FedLSR (60.10%) at the highest noise level. UIC OCT-FedSIR achieved the highest accuracy among the evaluated noisy-label methods across the reported UIC symmetric and asymmetric conditions (Tables S3 and S4). The corresponding macro-F1 results showed similar robustness across increasing noise levels (Figs. 3b and 4b). Under symmetric noise with α = 2, accuracy decreased from 95.80% at 30% noise to 84.60% at 90%, compared with 72.30% for RoFL and 68.80% for FedCorr at the highest noise level. With α = 0.5, OCT-FedSIR retained 80.20% accuracy at 90% noise, compared with 71.60% for FedCorr and 68.20% for RoFL. At α = 0.1, OCT-FedSIR decreased from 89.60% at 30% symmetric noise to 73.80% at 90%, whereas RoFL achieved 69.10% at the highest noise level. Under asymmetric noise, OCT-FedSIR remained comparatively stable in both accuracy and macro-F1 (Fig. 4b). With α = 2, OCT-FedSIR remained above 90% accuracy throughout the 40–90% noise range and achieved 90.70% at 90% noise, compared with 80.50% for FedCorr and 72.40% for RoFL. At the same 90% noise level, OCT-FedSIR achieved 88.60% with α = 0.5 and 81.40% with α = 0.1.
11
WF The differences among methods were smaller on WF than on Kermany and UIC in several symmetricnoise settings (Table S5). This behavior was also reflected in macro-F1, where the separation among the stronger methods was narrower than that observed for Kermany and UIC. With α = 2, OCT-FedSIR remained between 77.57% and 78.74% across the complete 30–90% symmetric noise range. At α = 0.5, RoFL exceeded OCT-FedSIR at several intermediate noise levels, whereas OCT-FedSIR performed more strongly as noise became severe. At 90% noise, OCT-FedSIR achieved 72.82%, compared with 62.44% for RoFL, 42.30% for FedAvg, and 33.37% for FedProx. Under α = 0.1, OCT-FedSIR achieved 78.91%, 77.96%, 77.90%, and 76.67% at 30%, 40%, 50%, and 60% symmetric noise, respectively, and retained 72.99% at 90%. RoFL slightly exceeded OCT-FedSIR at 80% noise (72.99% versus 72.15%), whereas OCT-FedSIR achieved the higher accuracy at the remaining symmetric noise levels. A clearer separation was observed under the one-directional GA-to-NGA asymmetric noise (Table S6; Fig. 4c). The macro-F1 results similarly showed stronger separation between OCT-FedSIR and the competing approaches under severe asymmetric noise. At α = 2.0, OCT-FedSIR achieved the highest accuracy at all six evaluated noise levels, including 81.92% and 80.97% at 80% and 90% noise, respectively. At α = 0.5, several competing methods were comparable at intermediate noise levels, but OCT-FedSIR reached 81.98% and 81.53% at 80% and 90% noise. With α = 0.1, OCT-FedSIR ranged from 77.29% to 80.36% across the evaluated asymmetric noise levels.
Retention versus pruning of clients with corrupted annotations The pruning experiment evaluated whether clients identified as having corrupted annotations should be excluded from subsequent federated optimization or retained after spectral correction. Across all 117 evaluated conditions, OCT-FedSIR achieved an overall mean accuracy of 86.73%, compared with 77.43% for pruning. On Kermany, OCT-FedSIR achieved mean accuracies of 94.76% and 93.59% under symmetric and asymmetric noise, respectively, compared with 81.23% and 84.20% for pruning. These results correspond to improvements of 13.53 percentage points and 9.39 percentage points, respectively. On UIC, OCT-FedSIR achieved 88.39% under symmetric noise and 89.57% under asymmetric noise, compared with 83.47% and 79.97% for pruning. The largest difference on UIC was observed under asymmetric noise, where OCT-FedSIR exceeded pruning by 9.60 percentage points. On WF, OCT-FedSIR achieved mean accuracies of 75.79% and 78.47% under symmetric and asymmetric noise, respectively, whereas pruning achieved 67.22% and 68.53%. The corresponding differences were 8.57 and 9.94 percentage points.
Ablation study We used a one-component-at-a-time removal analysis to assess the sensitivity of the complete OCTFedSIR framework to relabeling, KD, LA, and DaAgg (Table 4). Removing relabeling produced the largest decrease in performance, reducing mean accuracy from 86.73% to 79.47%, a difference of 7.26 percentage points. Removing KD, LA, and DaAgg reduced mean accuracy by 1.82, 1.05, and 0.52 percentage points, respectively. Thus, performance was most sensitive to removal of the relabeling component, while KD, LA, and DaAgg provided additional gains within the complete framework.
12
Table 4: Ablation analysis of OCT-FedSIR. Mean accuracy is averaged across all evaluated experimental conditions. ∆ denotes the decrease from the complete framework in percentage points. Configuration
Mean accuracy (%)
∆
86.73 86.21 85.68 84.91 79.47
– −0.52 −1.05 −1.82 −7.26
OCT-FedSIR (Full) w/o DaAgg w/o LA w/o KD w/o relabeling
Discussion This study addressed an underexplored challenge for trustworthy ophthalmic FL: whether annotationrelated disruption can be distinguished from legitimate differences in the clinical data distributions of participating clients. Across 117 dataset-specific experimental conditions spanning annotation-noise mechanism, noise severity, and client heterogeneity, OCT-FedSIR achieved an overall mean accuracy of 86.73%. This was 6.79 percentage points higher than RoFL and 7.98 percentage points higher than FedCorr, the two strongest comparison methods overall, and 5.03 percentage points below the all-clean FedAvg reference. Macro-F1 showed broadly similar trends across increasing noise levels. These findings indicate that explicitly modeling annotation reliability can substantially reduce the degradation associated with corrupted supervision while preserving heterogeneous client data within federated optimization. An important finding was that the original FedSIR spectral identification strategy did not transfer uniformly to the ophthalmic setting. Although FedSIR remained comparatively effective under symmetric label-noise, its clean-client F1 decreased substantially under structured asymmetric noise, ranging from 68.5% to 74.6% across the three datasets. The principal failure mode was the assignment of some labelcorrupted clients to the clean reference group. In contrast, OCT-FedSIR correctly separated the simulated clean and label-corrupted clients across all evaluated datasets, noise mechanisms, noise levels, and heterogeneity settings. The combined identification formulation incorporates class-balanced spectral estimation, Stage-I LA, and the complementary maximum off-diagonal similarity descriptor smax . These results k support the importance of adapting spectral client characterization to the class-distribution heterogeneity encountered in ophthalmic data rather than assuming that a spectral identification strategy developed on general vision benchmarks will transfer unchanged. The complete client-identification result should nevertheless be interpreted within the controlled experimental setting. Client status was defined through simulated corruption of existing reference annotations, and the resulting clean and noisy populations may exhibit more structured separation than would occur across independent clinical institutions. The finding therefore demonstrates that OCT-FedSIR distinguished the corruption patterns evaluated here; it should not be interpreted as evidence that annotation reliability can be identified perfectly in prospective multicenter data. Naturally occurring annotation disagreement may coexist with scanner shifts, disease-prevalence differences, image-quality variation,
13
referral patterns, and institutional practice differences, all of which may alter feature geometry independently of annotation quality. Client identification was also only the first part of the problem. At the sample level, spectral relabeling recovered a mean of 77.2% of intentionally corrupted annotations across the six dataset–noise summaries. Correction precision ranged from 86.7% to 95.0%, while the mean false-correction rate was 3.5%, corresponding to preservation of approximately 96.5% of annotations that were originally correct. This distinction is important because improved classification performance alone does not establish that an annotation-correction strategy is behaving safely. The observed combination of substantial recovery and comparatively infrequent harmful changes is consistent with the agreement rule used by OCT-FedSIR, in which a label is modified only when the representative and residual-subspace criteria support the same alternative class. The interaction between annotation noise and statistical heterogeneity remained apparent despite the strong client-identification performance. Under the most concentrated client distributions, particularly at α = 0.1, classification performance decreased in several severe-noise conditions. This indicates that correctly identifying a label-corrupted client does not guarantee that its annotations can subsequently be corrected with equal reliability. When local class support becomes sparse or highly concentrated, the class-specific structure required to construct and apply reliable spectral references may itself become less stable. This distinction is clinically relevant because a specialized referral center may contain an unusual or highly skewed disease distribution despite having reliable annotations. Such legitimate variation should not itself be interpreted as evidence of poor data quality. The structured asymmetric experiments provide an additional perspective on this problem. Corruption was concentrated between selected diagnostic categories rather than distributed uniformly: between choroidal neovascularization (CNV) and diabetic macular edema (DME) and between drusen and normal retina for Kermany, between neighboring DR severity categories for UIC, and from GA to NGA for WF. These transitions were designed as controlled structured perturbations and should not be interpreted as estimates of real clinical error probabilities. OCT-FedSIR remained particularly robust under these asymmetric settings, including severe label-noise on Kermany and UIC and the one-directional WF experiment. This is notable because annotation disagreement in clinical practice is unlikely to be uniform across diagnostic categories and may instead concentrate near phenotypically or clinically related decision boundaries. Differences across the three tasks further suggest that spectral robustness is dependent on the structure of the classification problem. Kermany and UIC provide multiple class relationships from which crossclass spectral structure can be characterized, whereas the binary WF task provides fewer such relationships. Correspondingly, differences among the stronger methods were smaller in several WF symmetricnoise settings, although OCT-FedSIR retained a clearer advantage under severe asymmetric noise. These findings caution against interpreting performance differences among the three datasets as differences in dataset difficulty alone, particularly because the tasks, cohorts, and network architectures also differed. The pruning comparison addressed a complementary question: whether a client identified as having corrupted supervision should be excluded or retained after correction. Across all evaluated conditions, OCT-FedSIR exceeded spectral pruning by 9.30 percentage points on average. This result highlights an 14
important distinction between annotation reliability and data value. A client with corrupted annotations may still contain correctly labeled samples, uncommon disease manifestations, or other clinically informative variation that would be lost if the entire client were excluded. Retaining and correcting such clients was therefore generally more effective than discarding them, although this advantage was not universal. Under the most heterogeneous Kermany setting with severe asymmetric corruption, pruning outperformed correction, suggesting a boundary at which the available spectral evidence may no longer support dependable recovery. The component ablation provides additional support for this interpretation. Removing spectral relabeling from the complete framework produced the largest decrease in mean accuracy, from 86.73% to 79.47%, a reduction of 7.26 percentage points. Removing KD, LA, and DaAgg produced smaller decreases of 1.82, 1.05, and 0.52 percentage points, respectively. These comparisons do not isolate independent causal effects because the components operate jointly; rather, they show that relabeling was the component whose removal from the complete framework was associated with the largest performance loss. KD provided complementary supervision when hard labels remained uncertain, LA addressed local class imbalance, and DaAgg reduced the influence of label-corrupted client models that deviated from the clean reference population. OCT-FedSIR therefore acts at multiple levels, combining client identification, sample-level supervision correction, local optimization, and server aggregation. This study has several limitations. OCT-FedSIR depends on sufficient reliable class support for constructing its spectral references, and performance may deteriorate when too few reliable clients remain or when diagnostic classes are absent or sparsely represented. Without a sufficiently reliable spectral reference, however, the framework cannot be expected to distinguish annotation-related deviations consistently or to support dependable relabeling of corrupted supervision. In addition, robustness to annotation noise does not by itself establish clinical utility or formal privacy. Client-identification outputs should be interpreted as signals of potential data-quality concerns, and spectrally corrected labels as training targets rather than changes to clinical records. Independent-site validation is therefore needed before clinical use, while future work should assess information-leakage risks from model updates and shared spectral statistics, including compatibility with secure aggregation and differential privacy. In summary, OCT-FedSIR addressed annotation reliability as a distinct challenge in ophthalmic FL by coupling spectral client characterization with selective recovery of corrupted supervision. Under the controlled conditions evaluated here, OCT-FedSIR substantially improved separation of clients with reliable and corrupted annotations relative to the original FedSIR procedure, while spectral relabeling recovered a large proportion of corrupted annotations with relatively few harmful changes to labels that were already correct. Retaining corrected clients was generally more effective than discarding them, further demonstrating that annotation reliability and data value should not be treated as equivalent. Together, these findings identify annotation reliability as an important requirement for trustworthy ophthalmic FL and motivate validation under naturally occurring annotation disagreement and genuinely multicenter clinical data.
15
Methods Study design We evaluated FL under two forms of heterogeneity that can coexist across clinical sites: differences in local disease composition and differences in annotation reliability. For each experimental seed, training data were partitioned among simulated clients using a class-aware Dirichlet procedure. A subset of clients retained the original labels, whereas labels at the remaining clients were corrupted using controlled symmetric or asymmetric transition rules. All methods within an experimental condition used the same client partition, noisy-client assignment, noisy labels, model initialization, and training budget. Three datasetspecific lightweight backbones were used, one for each retinal classification task. OCT-FedSIR comprised three stages: spectral identification of clients with reliable and unreliable labels (Fig. 1A–E), periodic relabeling of identified noisy clients using spectral references constructed from clean clients (Fig. 1F), and noise-aware federated optimization using LA, KD, and DaAgg (Fig. 1G).
Datasets Experiments used the Kermany OCT dataset, a UIC DR cohort, and a WF GA cohort. Kermany The Kermany OCT2017 dataset is a publicly available spectral-domain OCT dataset introduced by Kermany et al.2 It contains B-scans acquired with Heidelberg Spectralis OCT systems and labeled into four diagnostic categories: CNV, DME, drusen, and normal retina. Duplicate images were removed before federated partitioning. The original held-out test set was kept separate from all federated training procedures. UIC The UIC cohort contains OCT imaging from patients categorized as control, mild, moderate, or severe DR.13 The study received institutional review board approval from the University of Illinois Chicago and was conducted in accordance with the Declaration of Helsinki. DR severity was assigned by a retina specialist according to the Early Treatment Diabetic Retinopathy Study grading framework. OCT/OCTA data were acquired with an ANGIOVUE spectral-domain OCTA system (Optovue, Fremont, CA) operating at a 70-kHz A-scan rate, with an axial resolution of approximately 5 µm and a lateral resolution of approximately 15 µm. The cohort contained 445 scans from 41 patients that satisfied a signal-strength threshold of Q ≥ 5, including 187 3-mm scans and 258 6-mm scans. For the present classification experiments, patient-level separation was maintained between development and held-out evaluation sets so that images from the same patient did not occur in more than one split. Individual OCT B-scans were used as model inputs for the classification task.
16
Wake Forest The WF cohort was collected at Wake Forest University School of Medicine between January 1, 2013, and January 28, 2023. The study protocol received institutional review board approval and adhered to the Declaration of Helsinki. Patients were identified using ICD-9 codes 362.50 and 362.51 and ICD-10 codes H35.30 and H35.31 and were included when longitudinal OCT imaging was available with sufficient follow-up to assess GA progression. Exclusion criteria included subfoveal GA at baseline, fewer than two available OCT visits, unavailable longitudinal imaging, prior retinal treatment, and ocular conditions likely to confound OCT interpretation, including neovascular AMD, DR, glaucoma, retinal vascular occlusion, retinal detachment, and macular dystrophy. After application of these criteria, 91 patients with adequate imaging were retained. A single predefined study eye was used for each patient and remained fixed across visits (OD, 48.08%; OS, 51.92%). The cohort contained 455 OCT volumes and 11,513 B-scans: 258 volumes with central GA (CGA; 6,522 B-scans), 74 with non-central GA (NCGA; 1,874 B-scans), and 123 with no GA (NGA; 3,117 B-scans). CGA and NCGA were combined into a single GA category for the present study. This formulation separates the presence from the absence of GA without introducing the additional localization-dependent distinction between central and non-central atrophy. Individual OCT B-scans were used as model inputs for the binary classification task. All partitions were performed at the patient level before extraction of individual B-scans; no B-scan, OCT volume, eye, or visit from a patient in the held-out test set appeared in the federated training data. The three datasets therefore represented different classification settings: multiclass OCT disease classification in Kermany, ordinal DR severity classification in UIC, and binary GA classification in WF. Each sample retained a unique global index so that the original label, simulated corrupted label, and any subsequent spectral label could be tracked. Ground-truth labels were used only to generate controlled corruption and to evaluate identification and correction; they were not provided to the FL algorithms.
Federated client construction N
Let the federated training set at client k be denoted by Dk = {( xi , yi )}i=k1 , where Nk is the total number of training samples and yi ∈ {1, . . . , C } denotes one of C diagnostic classes. The Kermany, UIC, and WF training sets were partitioned among K = 30, K = 20, and K = 5 federated clients, respectively. Only training samples were distributed among clients; the held-out test sets remained completely separate from the federated partitioning procedure. To simulate statistical heterogeneity across clients, the training data were partitioned using a class-wise Dirichlet distribution. For each diagnostic class c, the proportions of samples assigned across the K clients were drawn from ρc ∼ Dirichlet (α1K ) , where ρc represents the allocation proportions of class c across the participating clients and its entries sum to one. Samples belonging to class c were then distributed among clients according to these proportions. The concentration parameter α controlled the degree of statistical class heterogeneity across clients. We evaluated α ∈ {2.0, 0.5, 0.1}, where smaller values of α produce more heterogeneous non-IID client distributions with increasingly uneven local class proportions. For each experimental seed and value of α, the client partition was generated once and then held fixed across all evaluated FL methods. Thus, competing methods within the same experimental condition were trained using identical client data par17
titions.
Client-dependent label-noise simulation After the federated client partitions were generated, clients were randomly assigned as either clean or noisy to simulate client-dependent annotation unreliability. Across the evaluated datasets and client configurations, the proportion of label-corrupted clients ranged from 60% to 90%, with the remaining 10%–40% clean-label clients. The clean/label-corrupted status of each client was unknown a priori. The clean/noisy assignment was generated independently for each experimental seed and was independent of the subsequent spectral client-identification procedure. Let γk ∈ {0, 1} denote the simulated annotation status of client k, where 0, if client k is clean, γk = (1) 1, if client k is noisy. For a given experimental seed, the clean/noisy client assignment was generated once and then held fixed across all competing FL methods. Consequently, methods evaluated within the same experimental condition used the same client partition, clean/noisy client assignment, noisy labels, and model initialization. Ground-truth client status γk was used only to evaluate client-identification performance; it was not provided to OCT-FedSIR during training or spectral client identification. Clean clients retained their original annotations. For each client designated as noisy, label-noise was applied separately within every locally represented class. This class-conditional procedure ensured that differences in local class prevalence did not determine the imposed noise level. Let nk,c denote the number of samples belonging to class c at client k before corruption. For a target label-noise ratio η, the number of samples selected for corruption within class c was ⌊ηnk,c ⌋. The corresponding samples were selected uniformly at random without replacement. Thus, η controlled the fraction of annotations corrupted within each locally represented class of a noisy client. For symmetric label-noise, each selected annotation from true class c was reassigned uniformly to one of the remaining C − 1 diagnostic classes (Fig. 5). The resulting transition model was P(ye = j | y = c) =
1 − η,
η , C−1
j = c, (2) j 6= c.
where y denotes the original reference label and ye denotes the observed label after simulated corruption. Symmetric label-noise ratios were evaluated at ηsym ∈ {0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9}. For asymmetric labelnoise, the same within-class sampling procedure was used, but selected annotations were reassigned according to predefined dataset-specific class transitions rather than uniformly across all alternative classes (Fig. 6). Let g(c) denote the predefined target class associated with source class c. The transition model was therefore
18
Figure 5: Simulated symmetric label-noise transitions. For clients designated as noisy, a proportion η of annotations within each locally represented class was selected uniformly at random and reassigned uniformly to one of the remaining C − 1 diagnostic classes. Clean clients retained their original annotations. 1 − η, P(ye = j | y = c) = η, 0,
j = c, j = g( c ),
(3)
otherwise.
For the Kermany dataset, asymmetric transitions were introduced between clinically related diagnostic categories, including CNV and DME, and between drusen and normal retina. For UIC, asymmetric transitions were restricted to neighboring DR severity categories to preserve the ordinal structure of the task. For WF, asymmetric corruption followed a predefined one-directional transition from GA to NGA. Asymmetric label-noise ratios were evaluated at ηasym ∈ {0.4, 0.5, 0.6, 0.7, 0.8, 0.9}. The symmetric noise provided a class-independent stress test, whereas asymmetric noise represented structured annotation errors concentrated between selected diagnostic categories.
Class balancing for spectral estimation Because non-IID partitioning and label-noise could produce substantial differences in the number of samples associated with individual classes within a client, a class-balanced loader was constructed for the spectral client-identification stage of OCT-FedSIR and for the corresponding spectral pruning baseline. Balancing was performed using the observed, potentially corrupted labels available to each client and did not use the underlying ground-truth clean labels. For this balancing procedure, let nk,c denote the
19
Figure 6: Dataset-specific asymmetric label-noise transitions. For clients designated as noisy, a proportion η of annotations within each applicable source class was selected and reassigned according to predefined dataset-specific class transitions. Clean clients retained their original annotations. number of samples at client k carrying the observed label c, and let
Ck = {c : nk,c > 0}
(4)
denote the set of classes represented by the observed local annotations. The largest observed class count at client k was defined as nmax = max nk,c . (5) k c∈Ck
Each observed class was oversampled with replacement until its sample count matched nmax k . Thus, for max each class c ∈ Ck with nk,c < nk , nmax − nk,c (6) k additional instances were sampled from the existing local examples of that class. Only classes already represented at a client were included in the balancing procedure; classes absent from the observed local annotations were not introduced. Consequently, balancing equalized the number of samples across the classes present at each client without altering its observed class support. No additional augmentation specific to the balancing procedure was applied; class balancing itself was performed solely by randomly repeated sampling of existing local examples. The balanced loader was used only for feature extraction and spectral estimation after Stage-I local training; it was not used to optimize the Stage-I client models. Stage-I local optimization was performed using the original client datasets with the LA objective described below. After local training, we used the balanced loader to extract class-specific feature representations from which the spectral descriptors were estimated. All subsequent federated local optimization likewise used the original client datasets. 20
FL formulation At communication round t, the server distributed the global parameters θ(t) to all clients. Each client (t) performed local optimization and returned updated parameters θk . Under conventional FedAvg,8 the global model is updated as K N (t) (7) θ(t+1) = ∑ K k θk . k=1 ∑ j=1 Nj This aggregation does not explicitly account for differences in annotation reliability. OCT-FedSIR instead estimates a clean-client subset, uses those clients to generate class-specific spectral references, and adjusts both local supervision and server aggregation according to the estimated client status.
OCT-FedSIR for noisy-label ophthalmic FL Stage-I: spectral client identification Because the annotation-quality status of each client is unknown a priori, Stage-I aims to distinguish clients with reliable annotations from clients with potentially corrupted labels. OCT-FedSIR performs this identification using spectral descriptors derived from locally learned class-wise feature representations, without access to ground-truth client status or raw client data. In particular, clients begin the identification stage from the same ImageNet-pretrained initialization and locally optimize the model using their observed local dataset and the LA objective defined below. For client k, the empirical class prior was πk,c =
nk,c , C ∑ j=1 nk,j
(8)
where nk,c denotes the number of local samples carrying observed label c. The class-specific LA was mk,c = β log(πk,c + ǫ),
(9)
where β controls the strength of the adjustment and ǫ prevents numerical instability. Let mk = [mk,1 , . . . , mk,C ]⊤ . The Stage-I objective was (10) LLA = CE ( f θk ( xi ) + mk , yei ) . After local identification training, latent representations were extracted using deterministic transforms from the balanced loader. Let zi = hθk ( xi ) ∈ R d be the feature vector for sample i. For each locally observed class c, the feature matrix was ⊤ z1 ⊤ z2 n k,c × d Zk,c = . (11) .. ∈ R . z⊤ n k,c Singular value decomposition (SVD) was applied to each observed class matrix with descending singular values: Zk,c = Uk,c Σk,c V⊤ k,c . The leading right singular vector v k,c represents the dominant direction of 21
class c. Pairwise class similarity was calculated as
[Sk ]c,c′ =
v⊤ k,c v k,c′
kvk,c k2 kvk,c′ k2 + ǫ
.
(12)
Under reliable annotations, class-discriminative feature representations are expected to produce distinct dominant directions across diagnostic categories. Label corruption mixes samples from different underlying classes within the same observed category and introduces shared class components across category specific feature matrices. Consequently, the dominant directions tend to exhibit greater cross-class alignment. This observation motivates using off-diagonal spectral similarity as an indicator of annotation unreliability. Therefore, lower cross-class alignment was taken as evidence of more distinct class-specific feature directions, whereas increased overlap was treated as a marker of label unreliability as shown in Fig. 2. Each client was summarized by three statistics computed from the off-diagonal entries: the mean off-diagonal similarity, mean squared magnitude (energy), and maximum: 1 [Sk ]c,c′ , |Ωk | (c,c∑ ′ )∈ Ω
(13)
1 [Sk ]2c,c′ , |Ωk | (c,c∑ ′ )∈ Ω
(14)
µk =
k
ek =
k
smax = k
max [Sk ]c,c′ .
( c,c′ )∈Ωk
(15)
where Ωk = {(c, c′ ) : c 6= c′ , c, c′ ∈ Ck }. A two-component GMM was fitted to the client descriptors. The component exhibiting lower cross-class similarity was designated as clean, producing the estimated sets Kclean and Knoisy . No ground-truth client status was used during this partitioning. Class-specific representative and residual subspaces At each relabeling checkpoint, the clean reference model was used to extract features from clients in b clean . For each locally observed class c at client k, SVD was applied to the corresponding class-specific K
feature matrix, Zk,c = Uk,c Σk,c V⊤ k,c , where the singular values σ1 ≥ σ2 ≥ · · · were ordered in descending magnitude. An adaptive spectral cutoff qk,c was selected as the smallest integer satisfying q k,c
Rk,c
∑ σj ≥ κ
j=1
∑
σj ,
(16)
j= q k,c +1
where Rk,c denotes the available spectral rank and κ controls the relative spectral mass of the dominant and residual portions. Equivalently, q k,c
∑ σj ≥
j=1
R
k,c κ σj . κ + 1 j∑ =1
(17)
We used κ = 10, corresponding to a cutoff at which the dominant portion contained at least approximately 90.9% of the total singular-value mass. The leading right singular vector, (r )
vk,c = vk,c,1 ∈ R d , 22
(18)
defined the local representative direction. The residual portion began after the adaptive spectral cutoff qk,c . The first L right singular vectors after the cutoff were retained to construct the local residual basis, i h (n) (19) Vk,c = vk,c,qk,c+1 , . . . , vk,c,qk,c+ L . where L denotes the number of residual directions used for spectral reference construction. We used L = 12 in all experiments. When fewer than L residual directions were available for a local class-specific decomposition, all available residual directions were used. Because singular vectors are defined only up to sign and their orientations may vary across clients, the class-specific references were combined through weighted projector averaging rather than direct vector averaging. For class c, the representative projector was 1 (r ) (r ) (r )⊤ Pc = wk,c vk,c vk,c , (20) ∑ Wc k∈K clean
and the residual projector was
(n)
Pc
=
1 ∑ Wc k∈K
(n)
( n)⊤
wk,c Vk,c Vk,c ,
(21)
clean
with
Wc =
∑
wk,c ,
k∈Kclean
(22) (r )
where wk,c is the number of samples belonging to class c at client k. The principal eigenvector of Pc (r ) (n) defined the global representative direction vc , whereas the leading L eigenvectors of Pc defined the (n) global residual basis Vc . We used the same L = 12 residual directions to construct the global residual reference. Spectral relabeling of noisy clients The clean reference model was used to extract features from each predicted noisy client (Fig. 7).
Figure 7: Spectral relabeling of predicted noisy clients using representative and residual spectral references constructed from predicted clean clients. A label change is accepted only when the representativeand residual-subspace predictions agree. For a feature vector zi , representative alignment with class c was (r )⊤
SR (i, c) = vc 23
zi .
(23)
where larger values indicate stronger alignment with the representative class direction. Residual-subspace projection was 1 ( n)⊤ S N (i, c) = √ Vc zi . (24) 2 L The two spectral predictions were ( R)
ybi
(25)
= arg min S N (i, c).
(26)
c
(N)
ybi
The final hard target was defined as
= arg max SR (i, c), c
yb( R) ,
( R)
ybi
i y⋆i = yei ,
(N)
= ybi
,
(27)
otherwise.
Thus, a candidate correction was accepted only when the representative- and residual-subspace criteria agreed; otherwise, the observed annotation was retained. This agreement criterion reduced the risk of replacing uncertain clinical labels with unstable pseudo-labels. Local optimization Predicted clean clients continued to use the LA objective defined above, with class priors estimated from their observed local labels. For a predicted noisy client, spectral relabeling produced a corrected hard target y⋆i . When the representative- and residual-subspace predictions agreed, y⋆i was replaced by the agreed spectral prediction; otherwise, y⋆i = yei . Thus, every sample retained a valid hard training target, while label changes were restricted to samples supported by both spectral criteria. The current global model served as the teacher for KD. For temperature τ, the teacher distribution for sample xi was (T )
pi
= softmax
f θ( t ) ( x i ) τ
.
(28)
For the noisy-client model (k′ ∈ Knoisy ), the local logits were first adjusted using the client-specific class prior, ℓk′ ,i = f θ(t) ( xi ) + mk . (29) k′
The corresponding student distribution was (S) pi = softmax
The distillation loss was
ℓk′ ,i τ
.
(30)
(T ) (S) LKD = DKL pi k pi
(31)
whereas supervision from the spectrally corrected hard target was
Lhard = CE (ℓk′ ,i , y⋆i ) . 24
(32)
The noisy-client objective combines the two sources of supervision: (t)
Lnoisy = λLKD + (1 − λ) Lhard ,
(33)
where λ corresponds to the estimated fraction of clean clients in the federation and controls the relative contribution of global-model distillation and corrected hard-label supervision: λ=
|Kclean | . K
(34)
Before the first spectral relabeling checkpoint, no corrected labels were yet available; therefore, y⋆i = yei , and noisy clients were trained using their observed labels together with KD from the current global model. At subsequent relabeling checkpoints, accepted spectral corrections were incorporated into y⋆i and used during local optimization. Distance-aware aggregation (DaAgg) After local optimization, all clients remained eligible to contribute to the global model. Let ak =
Nk ∑ j Nj
(35)
be the conventional sample-size weight. For each predicted noisy client, its parameter-space distance to the nearest predicted clean client was calculated as dk = min ∑ θk,ℓ − θj,ℓ 2 . b clean j∈K ℓ
(36)
where ℓ indexes floating-point parameter tensors. For predicted clean clients, dk = 0. Distances were normalized as dk . (37) dk = max j d j + ǫ The final aggregation weight was ωk =
ak exp(−dk ) , ∑ j a j exp(−d j )
(38)
and the global model was updated by K
(t)
θ(t+1) = ∑ ωk θk .
(39)
k=1
DaAgg therefore reduced the contribution of noisy-client models that deviated substantially from all clean-client models without removing them entirely. In parallel, the predicted clean-client models were aggregated using sample-size-weighted averaging to update the clean reference model used at the next relabeling checkpoint.
25
Network architectures We used lightweight ImageNet-pretrained convolutional neural networks to limit local computation, memory use, and communication cost. SqueezeNet1.1 was used for Kermany, ShuffleNetV2 for UIC, and RegNetY-400MF for WF. Compact architectures are relevant to FL because each participating client repeatedly stores the model, performs local optimization, and transmits model updates.34 SqueezeNet reduces parameter count through Fire modules that combine 1 × 1 and 3 × 3 convolutions;35 ShuffleNetV2 was designed for efficient execution with attention to memory-access cost and practical latency;36 and RegNetY400MF provides a regular network design within a low computational budget.37 For each dataset, the ImageNet classifier was replaced by a linear layer matching the number of target classes, and the backbone and classifier were fine-tuned jointly. Within each dataset, all FL methods used the same architecture and initialization.
Training protocol and reproducibility Models were initialized using ImageNet-pretrained parameters. Random seeds were set consistently for NumPy, PyTorch, client partitioning, label corruption, data-loader workers, and batch generators. Within a given experimental condition, the initial model state was generated once and shared across FL methods. All experiments were repeated using three independent random seeds. Local optimization used Adam or the optimizer prescribed by the corresponding baseline. OCTFedSIR used Adam with weight decay 5 × 10−4 during its noise-aware training stage. Training was performed for 100 communication rounds. In each round, each client received the current global model, performed one local epoch with a batch size of 128, and returned its updated state to the server. Each dataset-specific backbone was evaluated at every symmetric and asymmetric noise ratio and at each Dirichlet heterogeneity level. For all three ophthalmic datasets, images were resized to 128 × 128 pixels and converted to tensors. During training, random horizontal flipping was applied, followed by channel-wise normalization using mean 0.5 and standard deviation 0.5. Evaluation used the same resizing and normalization without random augmentation.
Evaluation and statistical analysis The global model was evaluated on the held-out test set after each communication round to characterize the training trajectory. Unless otherwise stated, all classification results reported in the tables, figures, and text correspond to the global model obtained after the final communication round. Test-set performance was not used for checkpoint selection, client identification, spectral relabeling, local optimization, or any other aspect of model training. Classification performance was summarized using accuracy and macro-F1 score. Client-identification performance was evaluated against the known simulated client status γk . We reported overall identification accuracy, clean-client precision, clean-client recall, clean-client F1 score, noisy-client recall, and the numbers of clean clients incorrectly classified as noisy and noisy clients incor-
26
rectly classified as clean. Ground-truth client status was used only for retrospective evaluation and was not provided to the GMM during client identification. Sample-level relabeling performance was evaluated by comparing the original reference label yi , the simulated corrupted label yei , and the post-correction label y⋆i . Let N denote the number of samples evaluated during the relabeling procedure. The edit rate was defined as #{i : y⋆i 6= yei } . (40) N The recovery rate quantifies the proportion of intentionally corrupted annotations that were restored to their original reference labels: Edit rate =
Recovery rate =
#{i : yei 6= yi ∧ y⋆i = yi } . #{i : yei 6= yi }
(41)
Correction precision quantifies the proportion of accepted label changes that resulted in the original reference label: Correction precision =
#{i : y⋆i 6= yei ∧ y⋆i = yi } . #{i : y⋆i 6= yei }
(42)
The false-correction rate quantifies the proportion of annotations that were originally correct but became incorrect after spectral relabeling: False-correction rate =
#{i : yei = yi ∧ y⋆i 6= yi } . #{i : yei = yi }
(43)
The residual label-noise rate after spectral relabeling was defined as
#{i : y⋆i 6= yi } . (44) N All experimental conditions were repeated using three independent random seeds. For each seed, client partitioning, clean/noisy client assignment, label corruption, and model initialization were generated independently. Within a given seed and experimental condition, however, the same client partition, clean/noisy client assignment, corrupted sample indices, corrupted labels, and initial model state were shared across all competing federated learning methods to enable paired and directly comparable evaluations. Classification, client-identification, and relabeling metrics were computed independently for each seed. Unless otherwise stated, numerical results are reported as the mean ± standard deviation across the three independent experimental runs. For each run, the evaluation pipeline recorded per-round predictions, per-sample losses, aggregate and class-wise classification metrics, client-identification outcomes, and sample-level relabeling outcomes. Residual noise rate =
Use of generative artificial intelligence ChatGPT (OpenAI) was used solely for language editing, including grammatical correction and rephrasing of manuscript text. It was not used to generate or analyze experimental data, perform statistical 27
analyses, or determine scientific conclusions. All revised text was reviewed and approved by the authors.
References [1] C. S. Lee, D. M. Baughman, and A. Y. Lee, “Deep learning is effective for classifying normal versus age-related macular degeneration oct images,” Ophthalmology Retina, vol. 1, no. 4, pp. 322–327, 2017. [2] D. S. Kermany, M. Goldbaum, W. Cai, C. C. Valentim, H. Liang, S. L. Baxter, A. McKeown, G. Yang, X. Wu, F. Yan et al., “Identifying medical diagnoses and treatable diseases by image-based deep learning,” cell, vol. 172, no. 5, pp. 1122–1131, 2018. [3] T. Schlegl, S. M. Waldstein, H. Bogunovic, F. Endstraßer, A. Sadeghipour, A.-M. Philip, D. Podkowinski, B. S. Gerendas, G. Langs, and U. Schmidt-Erfurth, “Fully automated detection and quantification of macular fluid in oct using deep learning,” Ophthalmology, vol. 125, no. 4, pp. 549–558, 2018. [4] J. De Fauw, J. R. Ledsam, B. Romera-Paredes, S. Nikolov, N. Tomasev, S. Blackwell, H. Askham, X. Glorot, B. O’Donoghue, D. Visentin et al., “Clinically applicable deep learning for diagnosis and referral in retinal disease,” Nature medicine, vol. 24, no. 9, pp. 1342–1350, 2018. [5] T. Haghighi, S. Gholami, J. T. Sokol, A. Biswas, J. I. Lim, T. Leng, A. C. Thompson, H. Tabkhi, and M. N. Alam, “Compact vision language models enable efficient and interpretable optical coherence tomography through layer-specific multimodal learning,” Communications Medicine, vol. 6, no. 1, p. 32, 2025. [6] S. Siraz, H. Kamanda, S. Gholami, A. S. Nabil, S. S. Yee Ong, and M. N. Alam, “Multi-class classification of central and non-central geographic atrophy using optical coherence tomography,” medRxiv, pp. 2025–05, 2025. [7] F. E. Jannat, S. Gholami, J. I. Lim, T. Leng, M. N. Alam, and H. Tabkhi, “Multi-oct-selfnet: Integrating selfsupervised learning with multi-source data fusion for enhanced multi-class retinal disease classification,” Frontiers in Systems Biology, vol. 6, p. 1717398, 2026. [8] B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas, “Communication-efficient learning of deep networks from decentralized data,” in Artificial intelligence and statistics. Pmlr, 2017, pp. 1273–1282. [9] J. Lo, T. Y. Timothy, D. Ma, P. Zang, J. P. Owen, Q. Zhang, R. K. Wang, M. F. Beg, A. Y. Lee, Y. Jia et al., “Federated learning for microvasculature segmentation and diabetic retinopathy classification of oct data,” Ophthalmology Science, vol. 1, no. 4, p. 100069, 2021. [10] C. Lu, A. Hanif, P. Singh, K. Chang, A. S. Coyner, J. M. Brown, S. Ostmo, R. V. P. Chan, D. Rubin, M. F. Chiang et al., “Federated learning for multicenter collaboration in ophthalmology: improving classification performance in retinopathy of prematurity,” Ophthalmology Retina, vol. 6, no. 8, pp. 657–663, 2022. [11] S. Gholami, J. I. Lim, T. Leng, S. S. Y. Ong, A. C. Thompson, and M. N. Alam, “Federated learning for diagnosis of age-related macular degeneration,” Frontiers in Medicine, vol. 10, p. 1259017, 2023. [12] A. R. Ran, X. Wang, P. P. Chan, M. O. Wong, H. Yuen, N. M. Lam, N. C. Chan, W. W. Yip, A. L. Young, H.W. Yung et al., “Developing a privacy-preserving deep learning model for glaucoma detection: a multicentre study with federated learning,” British Journal of Ophthalmology, vol. 108, no. 8, pp. 1114–1123, 2024. [13] A. S. Nabil, S. Gholami, T. Leng, J. I. Lim, and M. N. Alam, “Federated learning for multi-disease ophthalmic diagnostics using optical coherence tomography angiography (octa),” Ophthalmology Science, p. 101030, 2025.
28
[14] W. T. Lau, A. McCarthy, Y. Tian, C. Nielsen, R. Chelliah, S. Gholami, A. Kossler, M. Alam, L. D. Glass, and K. A. Thakoor, “Fedted: Federated learning for robust thyroid eye disease detection with masked autoencoders,” in 2025 IEEE EMBS International Conference on Biomedical and Health Informatics (BHI). IEEE, 2025, pp. 1–7. [15] S. Gholami, F.-E. Jannat, A. C. Thompson, S. S. Y. Ong, J. I. Lim, T. Leng, H. Tabkhivayghan, and M. N. Alam, “Distributed training of foundation models for ophthalmic diagnosis,” Communications Engineering, vol. 4, no. 1, p. 6, 2025. [16] S. Gholami, A. Ali, H. Kamanda, T. Haghighi, S. S. Y. Ong, J. I. Lim, T. Leng, A. Arafa, and M. N. Alam, “Fedsim: foundational federated multi-task learning for ophthalmic diagnostics,” in Ophthalmic Technologies XXXVI, vol. 13831. SPIE, 2026, pp. 140–149. [17] S. Yonamine, C. J. Ma, R. O. Alabi, G. Kaidonis, L. Chan, D. Borkar, J. D. Stein, B. F. Arnold, and C. Q. Sun, “Comparison of diagnosis codes to clinical notes in classifying patients with diabetic retinopathy,” Ophthalmology Science, vol. 4, no. 6, p. 100564, 2024. [18] L. Ju, X. Wang, L. Wang, D. Mahapatra, X. Zhao, Q. Zhou, T. Liu, and Z. Ge, “Improving medical images classification with label noise using dual-uncertainty estimation,” IEEE transactions on medical imaging, vol. 41, no. 6, pp. 1533–1546, 2022. [19] J. P. Campbell, M. F. Chiang, J. S. Chen, D. M. Moshfeghi, E. Nudleman, P. Ruambivoonsuk, H. Cherwek, C. Y. Cheung, P. Singh, J. Kalpathy-Cramer et al., “Artificial intelligence for retinopathy of prematurity: validation of a vascular severity scale against international expert diagnosis,” Ophthalmology, vol. 129, no. 7, pp. e69–e76, 2022. [20] D. Karimi, H. Dou, S. K. Warfield, and A. Gholipour, “Deep learning with noisy labels: Exploring techniques and remedies in medical image analysis,” Medical image analysis, vol. 65, p. 101759, 2020. [21] J. Shi, K. Zhang, C. Guo, Y. Yang, Y. Xu, and J. Wu, “A survey of label-noise deep learning for medical image analysis,” Medical image analysis, vol. 95, p. 103166, 2024. [22] A. Miladinović, A. Biscontin, M. Ajčević, S. Kresevic, A. Accardo, D. Marangoni, D. Tognetto, and L. Inferrera, “Evaluating deep learning models for classifying oct images with limited data and noisy labels,” Scientific Reports, vol. 14, no. 1, p. 30321, 2024. [23] T. Lin, M. Wang, A. Lin, X. Mai, H. Liang, Y.-C. Tham, and H. Chen, “Efficiency and safety of automated label cleaning on multimodal retinal images,” npj Digital Medicine, vol. 8, no. 1, p. 10, 2025. [24] S. Yang, H. Park, J. Byun, and C. Kim, “Robust federated learning with noisy labels,” IEEE Intelligent Systems, vol. 37, no. 2, pp. 35–43, April 2022. [25] A. Ali and A. Arafa, “FB-NLL: A feature-based approach to tackle noisy labels in personalized federated learning,” available online: arXiv:2604.19729. [26] X. Fang and M. Ye, “Robust federated learning with noisy and heterogeneous clients,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2022, pp. 10 062–10 071. [27] J. Xu, Z. Chen, T. Q. Quek, and K. F. E. Chong, “Fedcorr: Multi-stage federated learning for label noise correction,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2022, pp. 10 174–10 183. [28] N. Wu, L. Yu, X. Jiang, K.-T. Cheng, and Z. Yan, “Fednoro: Towards noise-robust federated learning by addressing class imbalance and label noise heterogeneity,” arXiv preprint arXiv:2305.05230, 2023.
29
[29] Y. Lu, L. Chen, Y. Zhang, Y. Zhang, B. Han, Y.-m. Cheung, and H. Wang, “Federated learning with extremely noisy clients via negative distillation,” in Proceedings of the AAAI conference on artificial intelligence, vol. 38, no. 13, 2024, pp. 14 184–14 192. [30] X. Jiang, S. Sun, J. Li, J. Xue, R. Li, Z. Wu, G. Xu, Y. Wang, and M. Liu, “Tackling noisy clients in federated learning with end-to-end label correction,” in Proceedings of the 33rd ACM international conference on information and knowledge management, 2024, pp. 1015–1026. [31] S. Gholami, A. Ali, T. Haghighi, A. Arafa, and M. N. Alam, “Fedsir: Spectral client identification and relabeling for federated learning with noisy labels,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026, pp. 3340–3348. [32] T. Li, A. K. Sahu, M. Zaheer, M. Sanjabi, A. Talwalkar, and V. Smith, “Federated optimization in heterogeneous networks,” Proc. Mach. Learn. Syst., vol. 2, pp. 429–450, March 2020. [33] X. Jiang, S. Sun, Y. Wang, and M. Liu, “Towards federated learning against noisy labels via local selfregularization,” in Proc. ACM CIKM, October 2022. [34] K. Bonawitz, H. Eichner, W. Grieskamp, D. Huba, A. Ingerman, V. Ivanov, C. Kiddon, J. Konečnỳ, S. Mazzocchi, B. McMahan et al., “Towards federated learning at scale: System design,” Proceedings of machine learning and systems, vol. 1, pp. 374–388, 2019. [35] F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally, and K. Keutzer, “Squeezenet: Alexnet-level accuracy with 50x fewer parameters and¡ 0.5 mb model size,” arXiv preprint arXiv:1602.07360, 2016. [36] N. Ma, X. Zhang, H.-T. Zheng, and J. Sun, “Shufflenet v2: Practical guidelines for efficient cnn architecture design,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 116–131. [37] I. Radosavovic, R. P. Kosaraju, R. Girshick, K. He, and P. Dollár, “Designing network design spaces,” arXiv preprint arXiv:2003.13678, 2020.
[38] Q. UNCC, “Toward trustworthy federated ophthalmic learning under annotation noise hosted on github,” https://github.com/QIAIUNCC/OCT-FedSIR-Toward-TrustworthyFederated-Ophthalmic-Learning-under-Annotation2026, accessed: 2026-09-05.
Acknowledgment This study is supported by NEI R15EY035804, R21EY035271, 1R01EY037828-01 (MNA), NC Diabetes Research Center P30DK124723 (SSYO), Research to Prevent Blindness (TL) and NIH Grant P30EY026877 (TL).
Author information Authors and affiliations Department of Electrical and Computer Engineering, University of North Carolina at Charlotte, Charlotte, NC, USA Sina Gholami, Abdulmoneam Ali, Tania Haghighi, Rashadul H. Badhon, Behafarin Emam, Ahmed Arafa, & Minhaj Nur Alam 30
Department of Ophthalmology, Wake Forest School of Medicine, Winston-Salem, NC, USA Sally S.Y. Ong & Atalie C. Thompson Byers Eye Institute at Stanford, Stanford University School of Medicine, Stanford, CA, USA Theodore Leng Department of Ophthalmology and Visual Sciences, University of Illinois Chicago, Chicago, IL, USA Jennifer I. Lim
Contributions Study conception and design: Sina Gholami, Abdulmoneam Ali, Tania Haghighi, Ahmed Arafa and Minhaj Nur Alam; Data collection and analysis: Sina Gholami, Rashadul H. Badhon, Sally S. Y. Ong, Atalie C. Thompson, Jennifer I. Lim, and Minhaj Nur Alam; Data visualization: Sina Gholami, Rashadul H. Badhon and Minhaj Nur Alam; Interpretation of results: Sina Gholami, Abdulmoneam Ali, Tania Haghighi, Behafarin Emam, Sally S. Y. Ong, Atalie C. Thompson, Theodore Leng, Jennifer I. Lim and Minhaj Nur Alam; Manuscript preparation: Sina Gholami, Abdulmoneam Ali, Tania Haghighi, Rashadul H. Badhon, Behafarin Emam and Minhaj Nur Alam; Manuscript review: Sina Gholami, Abdulmoneam Ali, Tania Haghighi, Rashadul H. Badhon, Behafarin Emam, Sally S. Y. Ong, Atalie C. Thompson, Theodore Leng, Ahmed Arafa, Jennifer I. Lim and Minhaj Nur Alam; Supervision: Minhaj Nur Alam; Funding acquisition: Minhaj Nur Alam; Project administration: Sina Gholami and Minhaj Nur Alam.
Corresponding author Correspondence to Minhaj Nur Alam.
Ethics declarations The UIC and WF datasets were collected under protocols approved by the respective Institutional Review Boards (IRBs) and in accordance with the ethical principles outlined in the Declaration of Helsinki.
Competing interests The authors declare no competing interests.
Data availability The Kermany dataset used in this study is publicly available. The private datasets from the University of Illinois at Chicago (UIC) and Wake Forest (WF) can be requested from the corresponding author, subject to reasonable conditions.
Code availability The code that supports the findings of this study is available on GitHub.38 31
Supplementary material
32
Table S1: Classification accuracy (%) on Kermany under symmetric label noise. Values are mean ± standard deviation across three independent seeds. Bold and underline indicate the best- and second-bestperforming noisy-label FL methods, respectively. Pruning and all-clean FedAvg are shown as references. Symmetric Noise α
2.0
0.5
0.1
# Clean Clients
3
Method
30%
40%
50%
60%
70%
80%
90%
FedAvg8
95.20±1.25 95.70±1.25 98.70±0.75 78.07±3.88 79.30±2.45 98.80±0.65 95.10±1.35 76.90±2.50 97.80±0.85 99.40±0.45
94.90±1.35 93.40±1.50 98.10±0.80 76.86±4.06 74.30±2.65 98.60±0.70 93.90±1.50 70.40±2.70 97.00±1.05 99.40±0.45
92.90±1.60 92.80±1.60 97.70±0.90 69.40±4.30 71.40±2.90 98.00±0.90 94.20±1.45 62.70±2.90 96.80±1.10 99.50±0.45
89.40±1.80 93.20±1.50 97.80±0.95 60.79±4.16 49.50±3.05 97.80±0.90 92.10±1.75 50.60±3.35 78.30±2.55 95.60±1.30
77.90±2.65 74.00±2.80 91.30±1.80 41.52±5.37 51.90±3.05 91.40±1.75 72.70±2.65 49.10±3.30 29.20±2.90 95.10±1.35
22.00±2.50 32.40±2.85 92.10±1.75 30.85±6.16 43.40±3.20 83.90±2.25 23.00±2.50 27.80±2.70 35.00±2.85 96.00±1.20
20.60±2.55 19.70±2.45 83.50±2.20 28.97±4.81 27.70±2.85 58.50±3.05 22.90±2.60 21.00±2.40 25.00±2.70 93.60±1.60
85.00±2.25 85.20±2.20 82.40±2.35 33.93±5.08 74.50±2.55 97.10±0.95 83.20±2.25 56.60±3.15 30.50±2.75 97.30±1.00
24.60±2.70 31.30±2.80 72.10±2.75 28.78±5.45 49.30±3.05 95.70±1.20 27.70±2.95 49.50±3.20 29.20±2.75 97.70±0.90
22.20±2.55 21.00±2.50 71.00±2.65 27.57±3.82 28.80±2.85 97.50±0.90 20.60±2.45 51.20±2.90 32.00±2.75 98.00±0.90
79.60±2.40 72.70±2.75 75.50±2.60 30.39±4.04 72.10±2.80 78.90±2.45 75.30±2.75 70.20±2.80 26.90±2.65 92.30±1.65
51.10±3.05 38.40±2.95 72.60±2.60 27.57±3.46 67.50±2.95 58.20±3.30 39.00±3.05 60.40±3.00 40.80±3.15 74.10±2.60
28.60±2.90 23.90±2.70 72.90±2.70 28.09±4.43 58.60±3.20 30.90±2.85 23.90±2.65 30.50±2.90 38.20±3.15 74.20±2.60
FedProx32 RoFL24 RHFL26 FedLSR33 FedCorr27 FedNed29 FedELC30 FedNoRo28 OCT-FedSIR (Ours)
3 30
Pruning FedAvg
3
FedAvg8 FedProx32 RoFL24 RHFL26 FedLSR33 FedCorr27 FedNed29 FedELC30 FedNoRo28 OCT-FedSIR (Ours)
3 30
Pruning FedAvg
3
FedAvg8 FedProx32 RoFL24 RHFL26 FedLSR33 FedCorr27 FedNed29 FedELC30 FedNoRo28 OCT-FedSIR (Ours)
3 30
Pruning FedAvg
86.40 ± 0.95 99.60 ± 0.35 95.40±1.30 95.50±1.20 97.70±0.90 62.08±6.35 95.30±1.30 98.60±0.70 96.60±1.00 83.30±2.40 97.80±0.85 99.50±0.45
96.20±1.15 95.90±1.15 97.50±0.95 60.36±5.93 95.60±1.30 98.10±0.80 95.60±1.20 78.90±2.55 97.80±0.95 99.70±0.35
94.80±1.35 96.30±1.10 95.50±1.30 54.22±5.58 94.20±1.40 97.20±1.00 95.50±1.20 68.70±2.80 98.00±0.80 99.00±0.60
94.60±1.45 93.90±1.50 89.40±1.90 48.40±5.31 87.90±2.00 96.70±1.00 95.00±1.25 61.00±3.15 96.60±1.10 97.90±0.85 86.90 ± 1.05 99.60 ± 0.25
77.20±2.50 78.00±2.45 95.70±1.25 39.30±5.46 74.20±2.75 87.30±2.05 76.00±2.60 72.50±2.80 75.30±2.65 96.80±1.10
77.10±2.65 80.50±2.35 85.40±2.25 37.04±4.75 74.20±2.65 88.90±1.85 77.00±2.55 71.40±2.70 75.30±2.65 96.50±1.15
79.30±2.40 77.60±2.60 90.90±1.70 35.34±4.73 74.60±2.65 79.50±2.35 76.60±2.50 72.90±2.80 73.30±2.60 95.30±1.30
74.10±2.70 75.50±2.55 80.90±2.55 35.54±4.64 72.30±2.65 78.40±2.45 77.60±2.60 74.60±2.70 78.10±2.55 93.10±1.55 70.40 ± 3.00 97.70±2.85
33
Table S2: Classification accuracy (%) on Kermany under asymmetric label noise. Values are mean ± standard deviation across three independent seeds. Bold and underline indicate the best- and second-bestperforming noisy-label FL methods, respectively. Pruning and all-clean FedAvg are shown as references. Asymmetric Noise α
2.0
0.5
0.1
# Clean Clients
Method
3
FedAvg8 FedProx32 RoFL24 RHFL26 FedLSR33 FedCorr27 FedNed29 FedELC30 FedNoRo28 OCT-FedSIR (Ours)
3 30
Pruning FedAvg
3
FedAvg8 FedProx32 RoFL24 RHFL26 FedLSR33 FedCorr27 FedNed29 FedELC30 FedNoRo28 OCT-FedSIR (Ours)
3 30
Pruning FedAvg
3
FedAvg8 FedProx32 RoFL24 RHFL26 FedLSR33 FedCorr27 FedNed29 FedELC30 FedNoRo28 OCT-FedSIR (Ours)
3 30
Pruning FedAvg
40%
50%
60%
70%
80%
90%
95.30±1.30 94.60±1.45 97.50±0.95 66.64±3.89 93.10±1.65 97.50±0.95 94.30±1.40 68.50±2.80 97.90±0.90 98.60±0.75
80.80±2.50 75.30±2.70 93.10±1.65 47.49±4.99 79.20±2.35 94.60±1.30 76.60±2.65 51.10±3.05 93.20±1.55 97.20±0.95
34.30±2.80 32.80±2.95 94.60±1.35 32.27±5.81 53.60±3.05 96.30±1.20 27.50±2.75 35.90±2.95 49.20±3.10 97.00±1.05
25.40±2.75 22.90±2.55 95.60±0.95 26.68±3.94 26.80±2.85 95.60±1.30 31.90±2.70 32.20±2.95 34.80±3.10 96.90±0.65
26.40±2.65 20.80±2.60 86.90±2.05 24.97±5.01 29.60±2.85 85.70±2.20 24.70±2.65 22.90±2.70 35.50±2.90 96.70±1.10
20.70±2.35 18.00±2.40 71.90±2.80 25.99±3.51 29.80±2.90 91.20±1.70 16.90±2.35 23.20±2.65 22.40±2.60 97.10±1.00
31.10±2.90 30.90±2.95 72.80±2.85 25.69±2.33 41.90±3.15 87.00±1.00 35.60±3.05 37.00±3.05 25.00±2.70 97.80±0.90
20.10±2.60 27.20±2.75 70.60±2.80 26.07±4.43 29.50±2.90 85.00±1.30 25.70±2.85 28.80±2.75 27.90±2.75 95.70±2.80
33.80±3.00 43.10±3.20 73.60±2.65 26.24±2.87 64.30±2.95 51.70±3.15 44.30±3.10 51.70±3.10 60.70±2.85 75.80±2.60
28.30±2.85 28.80±2.85 71.50±2.70 26.04±4.18 60.10±3.05 40.80±3.05 41.30±2.95 57.30±3.10 64.10±2.65 75.60±1.75
86.30 ± 0.20 99.60 ± 0.35 93.70±1.55 94.00±1.45 95.30±1.25 55.84±4.87 94.70±1.35 97.30±1.00 94.60±1.40 72.40±2.90 96.90±1.00 97.90±0.85
84.10±2.20 85.00±2.25 87.50±2.00 42.41±5.29 83.80±2.15 95.90±1.20 84.20±2.35 68.10±3.00 94.90±1.30 97.20±1.05
61.80±3.05 44.40±3.00 95.50±1.30 29.08±7.17 80.80±2.45 94.80±1.00 52.90±3.05 58.50±3.15 89.10±0.55 97.90±0.85
38.70±3.05 39.30±3.10 73.40±2.60 27.07±5.20 50.00±3.10 86.80±1.05 28.90±2.90 55.50±3.15 45.10±3.15 97.50±1.00
85.60 ± 0.45 99.60 ± 0.25 83.00±2.25 82.40±2.25 82.10±2.20 36.26±4.90 74.20±2.70 80.90±2.40 84.80±2.15 64.20±3.05 74.80±2.65 95.50±1.30
80.90±2.45 79.00±2.65 73.00±2.70 30.21±4.13 77.50±2.55 66.50±2.95 80.10±2.45 61.30±3.00 70.70±2.80 95.80±1.25
73.00±2.85 70.50±2.75 74.30±2.65 28.30±4.81 77.80±2.00 67.00±2.90 75.50±2.70 60.80±2.90 62.60±2.90 86.10±2.65
56.80±2.90 52.10±3.15 71.10±2.75 26.35±2.81 75.00±2.55 57.00±3.15 55.80±3.25 65.80±2.85 61.90±2.70 88.40±1.95
80.70 ± 2.35 97.70±2.85
34
Table S3: Classification accuracy (%) on UIC under symmetric label noise. Values are mean ± standard deviation across three independent seeds. Bold and underline indicate the best- and second-best-performing noisy-label FL methods, respectively. Pruning and all-clean FedAvg are shown as references. Symmetric Noise α
2.0
0.5
0.1
# Clean Clients
2
Method
30%
40%
50%
60%
70%
80%
90%
FedAvg8
86.40±1.58 87.10±1.49 92.80±1.12 70.15±6.74 85.70±1.73 93.50±1.01 86.80±1.52 77.60±2.31 91.90±1.22 95.80±0.82
85.70±1.63 84.60±1.71 91.60±1.23 66.83±7.19 83.90±1.88 92.40±1.13 85.20±1.69 75.10±2.44 90.80±1.31 95.20±0.91
83.10±1.84 84.20±1.76 90.30±1.35 63.42±6.81 79.60±2.09 91.60±1.26 82.50±1.90 72.80±2.57 89.40±1.42 94.70±0.98
79.80±2.06 80.60±1.98 88.70±1.49 60.71±7.46 77.20±2.19 90.10±1.36 78.30±2.13 67.90±2.76 84.10±1.83 93.50±1.12
73.40±2.31 70.50±2.42 85.20±1.77 52.38±8.14 72.80±2.43 83.90±1.88 69.20±2.46 62.20±2.91 76.80±2.19 91.80±1.29
62.10±2.57 57.80±2.65 80.40±2.04 46.27±7.73 66.70±2.62 77.30±2.18 56.40±2.71 56.30±3.02 63.50±2.66 88.90±1.53
47.30±2.73 43.60±2.78 72.30±2.31 43.91±8.26 55.40±2.86 68.80±2.48 42.10±2.82 49.70±3.08 51.20±2.91 84.60±1.82
63.60±2.69 61.10±2.74 79.10±2.08 49.20±8.18 67.50±2.66 82.80±1.91 60.70±2.79 59.80±3.03 72.30±2.49 88.60±1.51
51.40±2.91 47.90±2.97 73.80±2.39 45.83±8.76 61.20±2.85 78.20±2.17 49.80±2.98 55.40±3.10 65.90±2.76 85.10±1.79
39.20±2.98 36.80±3.01 68.20±2.61 42.15±8.31 54.70±3.01 71.60±2.49 38.60±3.05 50.90±3.06 57.10±2.94 80.20±2.06
65.70±2.83 63.40±2.91 76.30±2.36 42.16±8.54 69.80±2.77 73.90±2.49 64.90±2.87 63.60±2.96 70.40±2.71 84.10±1.89
57.80±3.01 54.20±3.07 72.80±2.58 40.83±8.91 65.10±2.94 68.20±2.76 55.80±3.02 61.20±3.02 66.80±2.88 78.60±2.25
46.30±3.08 43.70±3.11 69.10±2.76 38.94±8.47 59.80±3.07 61.40±2.98 45.20±3.11 53.90±3.15 60.20±3.02 73.80±2.51
FedProx32 RoFL24 RHFL26 FedLSR33 FedCorr27 FedNed29 FedELC30 FedNoRo28 OCT-FedSIR (Ours)
2 20
Pruning FedAvg
2
FedAvg8 FedProx32 RoFL24 RHFL26 FedLSR33 FedCorr27 FedNed29 FedELC30 FedNoRo28 OCT-FedSIR (Ours)
2 20
Pruning FedAvg
2
FedAvg8 FedProx32 RoFL24 RHFL26 FedLSR33 FedCorr27 FedNed29 FedELC30 FedNoRo28 OCT-FedSIR (Ours)
2 20
Pruning FedAvg
85.20 ± 1.47 96.70 ± 0.92 82.60±1.91 83.40±1.82 89.30±1.39 65.74±7.83 80.40±2.03 90.60±1.21 82.20±1.95 73.80±2.48 88.40±1.46 93.73±1.03
80.10±2.04 79.60±2.09 87.70±1.51 62.20±7.15 78.60±2.15 89.20±1.34 79.30±2.09 69.20±2.66 87.20±1.56 92.80±1.11
75.30±2.27 76.80±2.20 86.40±1.62 58.91±8.04 74.20±2.31 86.90±1.53 75.10±2.32 65.90±2.81 83.90±1.82 91.90±1.19
70.80±2.44 69.70±2.48 82.60±1.90 53.66±8.52 71.90±2.46 85.70±1.65 68.40±2.57 61.70±2.94 79.40±2.08 90.40±1.34 84.70 ± 1.84 94.90 ± 1.31
76.80±2.31 75.90±2.37 84.20±1.83 55.61±8.36 74.30±2.51 82.60±1.95 77.50±2.27 69.40±2.74 81.30±2.02 89.60±1.42
74.60±2.42 73.80±2.49 82.10±1.97 51.87±8.71 76.20±2.40 84.10±1.84 75.30±2.43 68.20±2.81 79.60±2.14 88.20±1.53
72.10±2.58 71.30±2.62 83.60±1.88 48.25±8.19 73.50±2.57 79.80±2.15 72.80±2.56 70.10±2.73 77.20±2.29 88.90±1.49
68.90±2.72 69.50±2.70 78.90±2.20 46.70±8.92 74.10±2.52 80.40±2.11 70.60±2.68 66.80±2.87 80.10±2.11 85.70±1.76 80.50 ± 2.34 91.80±2.17
35
Table S4: Classification accuracy (%) on UIC under asymmetric label noise. Values are mean ± standard deviation across three independent seeds. Bold and underline indicate the best- and second-bestperforming noisy-label FL methods, respectively. Pruning and all-clean FedAvg are shown as references. Asymmetric Noise α
2.0
0.5
0.1
# Clean Clients
Method
2
FedAvg8 FedProx32 RoFL24 RHFL26 FedLSR33 FedCorr27 FedNed29 FedELC30 FedNoRo28 OCT-FedSIR (Ours)
2 20
Pruning FedAvg
2
FedAvg8 FedProx32 RoFL24 RHFL26 FedLSR33 FedCorr27 FedNed29 FedELC30 FedNoRo28 OCT-FedSIR (Ours)
2 20
Pruning FedAvg
2
FedAvg8 FedProx32 RoFL24 RHFL26 FedLSR33 FedCorr27 FedNed29 FedELC30 FedNoRo28 OCT-FedSIR (Ours)
2 20
Pruning FedAvg
40%
50%
60%
70%
80%
90%
88.40±1.55 87.70±1.60 91.90±1.20 64.80±5.10 86.80±1.70 90.80±1.30 86.50±1.70 69.40±2.70 90.40±1.30 93.80±1.00
77.20±2.35 75.60±2.45 88.60±1.45 53.20±7.05 76.50±2.20 87.90±1.60 74.90±2.40 59.80±2.85 87.30±1.60 92.60±1.10
48.60±2.70 45.90±2.80 89.40±1.40 39.70±7.40 58.30±2.75 88.10±1.50 41.30±2.75 45.20±2.90 56.70±2.85 92.10±1.15
39.10±2.75 36.40±2.70 90.10±1.30 34.60±7.85 44.20±2.80 87.60±1.55 39.90±2.70 42.70±2.85 45.60±2.80 92.90±1.05
36.80±2.70 34.20±2.75 88.70±1.45 31.20±7.40 41.10±2.80 82.40±1.95 31.50±2.75 38.40±2.80 46.20±2.80 91.80±1.15
31.50±2.55 29.80±2.60 72.40±2.45 32.10±7.75 37.90±2.75 80.50±2.05 27.20±2.60 34.90±2.75 38.10±2.75 90.70±1.25
36.20±2.80 35.80±2.80 72.90±2.35 29.80±7.35 49.60±2.85 85.10±1.60 38.10±2.80 44.30±2.90 39.70±2.75 90.40±1.20
28.40±2.65 30.20±2.70 70.80±2.40 30.40±7.60 38.30±2.80 82.20±1.80 31.20±2.70 36.20±2.80 38.40±2.75 88.60±1.45
47.10±2.85 50.20±2.90 76.40±2.15 32.20±6.70 68.40±2.70 74.30±2.35 54.90±2.85 58.10±2.85 72.90±2.45 83.20±1.85
42.30±2.80 42.60±2.85 74.90±2.20 30.60±6.40 64.20±2.80 69.10±2.55 49.80±2.85 59.40±2.85 73.60±2.35 81.40±1.95
82.30 ± 1.25 96.70 ± 0.92 84.30±1.80 84.80±1.75 88.20±1.45 58.20±7.60 86.10±1.65 89.40±1.35 85.20±1.75 70.80±2.75 88.90±1.40 92.20±1.05
75.90±2.30 76.40±2.25 83.70±1.80 47.50±8.20 78.90±2.10 86.20±1.55 76.80±2.25 65.30±2.80 85.50±1.70 90.70±1.20
58.60±2.75 49.10±2.85 87.40±1.50 34.40±8.35 77.40±2.20 86.80±1.60 53.80±2.80 59.90±2.90 87.10±1.45 91.10±1.10
43.80±2.85 42.50±2.85 74.60±2.30 31.90±8.10 57.20±2.75 84.30±1.65 35.70±2.80 56.40±2.95 53.90±2.80 89.80±1.25
82.40 ± 1.55 94.90 ± 1.31 79.30±2.15 78.60±2.20 84.90±1.70 40.20±7.00 74.60±2.40 83.40±1.85 79.80±2.10 64.30±2.80 83.70±1.80 89.80±1.35
77.20±2.30 75.90±2.35 80.20±1.95 35.80±6.75 77.20±2.35 84.10±1.80 77.90±2.25 66.10±2.80 81.40±1.95 88.70±1.40
71.40±2.60 69.80±2.65 81.10±1.85 33.10±6.60 78.30±2.20 79.40±2.05 74.60±2.45 64.90±2.75 76.20±2.25 86.90±1.55
61.70±2.80 58.90±2.85 78.70±2.00 31.40±6.25 74.90±2.40 80.20±2.00 62.80±2.80 68.20±2.70 79.10±2.05 85.60±1.70
75.20 ± 1.95 91.80±2.17
36
Table S5: Classification accuracy (%) on WF under symmetric label noise. Values are mean ± standard deviation across three independent seeds. Bold and underline indicate the best- and second-best-performing noisy-label FL methods, respectively. Pruning and all-clean FedAvg are shown as references. Symmetric Noise α
2.0
0.5
0.1
# Clean Clients
1
Method
30%
40%
50%
60%
70%
80%
90%
FedAvg8
74.78±2.01 74.55±1.95 72.94±2.01 48.05±16.90 72.82±1.98 73.16±2.04 73.55±2.01 56.81±2.29 72.82±1.98 77.96±1.87
74.61±1.98 74.39±1.93 72.88±2.01 51.03±15.61 72.82±1.98 73.21±2.04 73.77±2.04 67.13±2.15 73.16±2.04 77.57±2.04
73.83±2.01 74.44±2.01 73.44±2.09 54.96±15.49 72.82±1.98 73.16±2.04 73.60±2.04 67.91±2.18 72.60±1.98 78.46±1.95
75.06±1.95 74.00±2.04 72.82±1.98 57.42±15.02 72.82±1.98 73.21±2.04 73.55±2.01 68.08±2.12 73.05±2.04 78.07±1.95
73.83±2.06 74.16±2.01 72.88±2.01 59.97±14.50 72.82±1.98 73.10±2.04 73.88±2.06 68.42±2.09 73.33±2.04 78.52±1.95
74.05±2.01 73.83±2.04 73.10±1.98 61.47±14.56 72.88±2.01 73.10±2.04 73.44±2.01 72.82±1.98 72.27±2.04 77.85±1.87
75.11±1.95 74.83±1.98 72.88±1.98 59.92±15.67 72.82±1.98 73.16±1.98 73.49±2.07 66.13±2.15 72.88±2.06 78.74±1.87
45.81±2.26 45.54±2.29 71.54±2.01 60.26±13.00 63.11±2.23 58.76±2.46 50.33±2.23 52.96±2.34 62.28±2.26 72.27±2.01
39.01±2.29 41.18±2.32 64.68±2.32 58.23±17.47 65.29±2.18 57.98±2.32 40.07±2.23 53.74±2.21 64.79±2.29 72.82±1.98
42.30±2.29 33.37±2.20 62.44±2.26 62.08±15.59 57.59±2.20 61.27±2.34 41.46±2.29 54.91±2.29 52.85±2.29 72.82±1.98
68.92±2.01 67.69±2.09 73.16±2.04 57.09±15.03 72.99±2.04 70.70±2.06 64.23±2.20 52.79±2.32 73.88±1.95 75.50±2.07
65.57±2.15 61.27±2.26 72.99±2.04 47.97±16.21 72.77±2.01 70.81±2.09 57.37±2.18 62.72±2.29 64.96±2.18 72.15±2.06
62.28±2.18 55.86±2.23 72.88±2.01 48.95±20.31 70.67±1.85 71.04±2.04 66.85±2.18 57.76±2.32 58.82±2.26 72.99±1.98
FedProx32 RoFL24 RHFL26 FedLSR33 FedCorr27 FedNed29 FedELC30 FedNoRo28 OCT-FedSIR (Ours)
1 5
Pruning FedAvg
1
FedAvg8 FedProx32 RoFL24 RHFL26 FedLSR33 FedCorr27 FedNed29 FedELC30 FedNoRo28 OCT-FedSIR (Ours)
1 5
Pruning FedAvg
2
FedAvg8 FedProx32 RoFL24 RHFL26 FedLSR33 FedCorr27 FedNed29 FedELC30 FedNoRo28 OCT-FedSIR (Ours)
2 5
Pruning FedAvg
65.11 ± 2.61 82.79 ± 1.98 69.75±2.15 70.42±2.26 73.10±2.15 66.56±12.22 64.17±2.06 71.76±2.01 69.36±2.18 61.61±2.15 72.54±2.04 74.11±2.04
65.01±2.18 65.46±2.29 75.33±2.04 71.67±3.33 59.43±2.21 63.23±2.23 60.66±2.29 50.89±2.26 73.16±1.95 74.44±1.93
57.03±2.26 60.27±2.34 74.50±1.98 63.91±7.38 68.69±2.09 62.44±2.12 55.58±2.20 57.25±2.32 72.21±2.04 73.05±1.98
52.12±2.29 52.57±2.37 73.49±2.12 57.21±16.67 63.39±2.29 54.91±2.40 48.77±2.21 51.34±2.26 63.84±2.12 72.88±1.95 65.91 ± 2.16 81.55 ± 2.32
76.06±2.01 75.28±2.01 75.06±2.01 74.33±2.08 75.78±1.93 74.50±1.95 75.56±1.95 73.66±2.04 77.96±1.98 78.91±1.93
76.34±1.95 73.88±2.01 74.89±2.01 72.00±2.00 75.11±1.98 74.00±1.98 75.11±2.06 73.55±2.04 76.34±1.95 77.96±1.84
71.82±2.01 71.04±2.15 73.66±1.95 67.61±5.11 74.44±2.04 70.26±2.12 70.48±1.98 62.28±2.20 76.56±1.93 77.90±1.90
68.92±2.15 68.58±2.18 72.82±1.98 51.69±15.55 72.88±1.95 72.82±1.98 68.08±2.12 60.10±2.32 74.67±1.98 76.67±1.90 70.65 ± 2.41 81.24±2.95
37
Table S6: Classification accuracy (%) on WF under asymmetric label noise. Values are mean ± standard deviation across three independent seeds. Bold and underline indicate the best- and second-bestperforming noisy-label FL methods, respectively. Pruning and all-clean FedAvg are shown as references. Asymmetric Noise α
2.0
0.5
0.1
# Clean Clients
Method
1
FedAvg8 FedProx32 RoFL24 RHFL26 FedLSR33 FedCorr27 FedNed29 FedELC30 FedNoRo28 OCT-FedSIR (Ours)
1 5
Pruning FedAvg
1
FedAvg8 FedProx32 RoFL24 RHFL26 FedLSR33 FedCorr27 FedNed29 FedELC30 FedNoRo28 OCT-FedSIR (Ours)
1 5
Pruning FedAvg
2
FedAvg8 FedProx32 RoFL24 RHFL26 FedLSR33 FedCorr27 FedNed29 FedELC30 FedNoRo28 OCT-FedSIR (Ours)
2 5
Pruning FedAvg
40%
50%
60%
70%
80%
90%
77.62±1.84 76.40±2.01 77.34±1.93 68.23±7.03 77.57±1.90 76.51±1.90 74.22±2.09 58.59±2.20 69.87±2.12 79.41±1.84
75.28±1.98 74.72±1.95 76.23±2.01 59.17±14.30 76.51±1.93 75.11±2.09 69.42±2.12 72.60±2.01 67.91±2.26 78.96±1.84
76.06±2.01 73.10±1.98 75.22±2.01 49.26±13.46 75.56±2.04 74.39±1.98 60.71±2.26 57.92±2.18 68.14±2.26 78.07±1.98
73.33±1.98 70.37±2.09 78.96±1.90 47.75±16.85 77.34±1.98 77.23±1.81 62.44±2.20 63.50±2.23 65.23±2.20 81.31±1.76
72.10±2.20 69.48±2.18 74.11±1.98 39.78±16.09 70.93±2.09 75.50±2.04 39.56±2.26 70.87±2.09 64.90±2.18 81.92±1.81
64.51±2.12 69.70±2.09 73.27±2.04 45.51±19.58 61.50±2.21 74.78±2.04 37.39±2.18 59.65±2.32 67.97±2.15 80.97±1.79
41.13±2.37 51.95±2.23 69.42±2.09 37.54±15.42 72.99±2.01 61.27±2.20 34.15±2.23 55.64±2.29 66.56±1.78 81.98±1.81
39.68±2.32 35.16±2.18 73.27±2.04 37.63±15.69 53.18±2.40 60.94±2.23 38.95±2.29 52.57±2.34 61.31±1.15 81.53±1.76
73.88±1.98 73.60±1.98 72.82±1.98 42.60±17.13 72.82±1.98 73.21±1.98 73.55±2.01 70.65±2.15 72.82±1.98 78.91±1.90
74.78±2.01 74.33±1.95 73.16±2.01 37.06±14.83 72.82±1.98 73.49±2.04 73.55±2.01 64.56±2.26 72.82±2.06 77.29±1.98
67.83 ± 1.20 82.79 ± 1.98 67.13±2.15 65.01±2.15 73.27±2.01 58.40±16.65 73.27±1.93 65.18±2.12 65.74±2.07 60.83±2.29 72.82±1.98 77.90±1.95
60.38±2.29 60.27±2.18 72.94±2.04 51.00±17.29 72.99±2.01 63.95±2.12 55.80±2.29 52.46±2.34 71.09±2.12 72.82±1.98
53.74±2.26 54.97±2.32 70.03±2.12 41.14±15.70 72.38±1.98 62.83±2.29 49.89±2.29 58.43±2.34 69.25±2.12 71.67±1.98
46.76±2.48 47.94±2.32 70.65±2.07 40.73±16.52 72.38±2.01 63.73±2.29 41.91±2.29 53.52±2.23 69.59±2.07 71.88±2.01
67.40 ± 1.73 81.55 ± 2.32 75.73±1.90 74.22±2.01 72.88±1.95 43.07±17.43 72.82±1.98 73.27±2.07 73.55±2.04 57.70±2.26 72.99±2.04 79.30±1.84
74.27±1.98 73.72±2.04 73.10±1.98 42.31±16.95 72.82±1.98 73.33±1.98 73.49±2.01 71.09±2.12 72.99±2.06 80.36±1.84
74.67±2.01 74.16±2.01 72.82±1.98 39.17±15.35 72.82±1.98 73.16±2.04 73.72±1.98 69.14±2.01 73.27±1.95 78.85±1.90
74.83±1.95 74.61±1.95 73.27±1.98 37.30±14.76 72.82±1.98 73.21±1.98 73.38±1.98 72.82±1.98 73.55±2.01 79.41±1.90
70.35 ± 1.18 81.24±2.95
38