DETECTING AND REFURBISHING GROUND TRUTH ERRORS DURING TRAINING OF DEEP LEARNING-BASED ECHOCARDIOGRAPHY SEGMENTATION MODELS Iman Islam, Bram Ruijsink, Andrew J. Reader, Andrew P. King School of Biomedical Engineering and Imaging Sciences, King’s College London, London, UK
arXiv:2604.12832v1 [cs.CV] 14 Apr 2026
ABSTRACT Deep learning-based medical image segmentation typically relies on ground truth (GT) labels obtained through manual annotation, but these can be prone to random errors or systematic biases. This study examines the robustness of deep learning models to such errors in echocardiography (echo) segmentation and evaluates a novel strategy for detecting and refurbishing erroneous labels during model training. Using the CAMUS dataset, we simulate three error types, then compare a loss-based GT label error detection method with one based on Variance of Gradients (VOG). We also propose a pseudo-labelling approach to refurbish suspected erroneous GT labels. We assess the performance of our proposed approach under varying error levels. Results show that VOG proved highly effective in flagging erroneous GT labels during training. However, a standard U-Net maintained strong performance under random label errors and moderate levels of systematic errors (up to 50%). The detection and refurbishment approach improved performance, particularly under high-error conditions. Index Terms— echocardiography segmentation, quality control, deep learning 1. INTRODUCTION Deep learning (DL) models have been widely proposed for automating the segmentation of cardiac structures from echocardiography (echo) images [1, 2]. Subsequently, these segmentations are often used to calculate functional biomarkers such as ejection fraction and therefore their reliability is of crucial importance in patient management [3, 4]. The training of DL segmentation models requires access to (preferably) large datasets with associated annotations, or ground truth (GT) labels of the structures of interest. These GT labels are often treated as a ‘gold standard’; however, they are typically created by humans who are prone to making mistakes, particularly in complex and laborious tasks such as medical image segmentation. The assumption is always that GT labels are correct, but in reality minor differences due to inter-annotator variability [1] and more severe annotation errors are common in medical imaging [5]. The impact that such errors can have on model performance when they are
included in the training data has been mostly overlooked in DL-based segmentation of medical images. 1.1. Aims and Contributions First, this study aims to investigate the robustness of a stateof-the-art (SOTA) DL segmentation model to errors in the GT labels used in training. A range of synthetic but realistic errors are introduced into the training data with varying frequencies and their impact on model performance is measured. Second, we propose an automated technique to detect potentially problematic GT labels and a method to ‘refurbish’ them during training. The necessity and utility of this approach is investigated. 1.2. Related Work Most work to date on identifying and dealing with GT label errors has been in classification tasks, and there has been limited work in segmentation tasks. Segmentation introduces unique challenges that are not present in classification tasks, such as spatial continuity and boundary sensitivity. In general, research into GT label errors mainly considers three tasks: (i) identifying which GT labels are noisy, (ii) determining the optimal stage in the training process to intervene and (iii) selecting an appropriate method to ‘clean’ or ‘refurbish’ them. The majority of methods for detecting GT label errors when training classification models assume that samples with persistently high training losses are likely to be mislabelled. Therefore, in order to select the erroneous labels, these works have considered the loss distribution of samples [6, 7]. Some papers have also proposed considering the model’s prediction confidence, and judged low confidence to indicate errors in the GT label [8]. Additionally, a recent paper suggested considering the variance of gradients (VOG) for individual samples, assuming that large changes in gradients are indicative of errors in the GT label [9]. In order to reduce the impact of erroneous labels, several label refurbishment techniques have been proposed. One common approach is pseudo-labelling, where potentially incorrect GT labels are replaced with the current model’s prediction during training [6, 8, 7]. Another approach is
reweighting losses, in which samples suspected to be mislabelled are assigned lower weights during optimisation [10]. In addition, sample filtering methods discard samples during training so that the model learns only from clean samples [11]. These label refurbishment techniques have been implemented from the start of training [8, 10, 7] or after a warm-up period [6]. Fewer studies have explicitly addressed GT label errors in segmentation tasks. One approach has involved using model uncertainty to locate label regions that are likely to be incorrect to inform pixel-level relabelling [12]. Alternatively, an iterative refinement pipeline has been proposed in which the initial segmentation network generated improved annotations that were used to supervise a second stage network to reduce the reliance on imperfect original labels over time [13]. Another approach included applying adaptive early-learning to emphasize clean samples during training [14]. Finally, selfrelabelling techniques have been proposed to refine segmentation labels dynamically based on historical predictions [15]. Currently, significant gaps remain in the literature concerning the impact of GT label errors on SOTA segmentation models. First, prior work has predominantly been restricted to binary segmentation scenarios, ignoring the common task of multi-class segmentation. Second, although there are some methods to mitigate the effects of GT label errors, a comprehensive evaluation of the robustness of models to different types of errors is lacking. This lack of evaluation limits our understanding of how well these models handle real-world annotation errors. Finally, the impact of incorrect GT segmentations in echo images remains unexplored. Given the inherent challenges of echo image segmentation, such as variability in anatomical structures and poor image quality, addressing these limitations is essential to improve model reliability and trust in real-world clinical settings.
2. MATERIALS AND METHODS 2.1. Dataset We used the CAMUS dataset [1], which contains 2D echo images with manual labellings of the left ventricle (LV), left ventricular myocardium (LVM) and left atrium (LA). Specifically, we selected the apical four-chamber (A4C) view images at end-diastole (ED), resulting in a total of 500 annotated samples.
(i) Incomplete label error: Part of a structure is left unlabelled (i.e., as background), simulating interruptions or incomplete annotations. (ii) Boundary distortion error: Labels are dilated or eroded, mimicking confusion at boundaries and different annotation styles. (iii) Merged labels error: One structure is incorrectly labelled as another already labelled structure, simulating unchecked label application. We examined two categories of error: random errors - arbitrary mistakes simulating inattentive annotation where there is an equal chance of every structure being affected, and systematic errors - consistent mislabelling of specific structure(s). Systematic errors were applied as follows: (i) for incomplete labels, the LV was consistently partially annotated; (ii) for boundary distortion, the labels were consistently eroded; and (iii) for merged labels, the LVM was relabelled and merged with the LV class.
Fig. 1. A sample image and GT label with examples of the three types of synthetic errors: incomplete label, boundary distortion and merged label. Blue=LV, Green=LVM, Yellow=LA
2.3. Baseline Segmentation Model A U-Net [16] was used as the baseline model to segment the LV, LVM, and LA. Models were trained with cross-entropy loss using the Adam optimizer (learning rate 0.001, batch size 4) for 100 epochs. Hyperparameters were selected via a grid search based on validation performance. Data were split 80%/10%/10% into training, validation, and test sets. The model from the epoch with the best validation foreground Dice score was used for test evaluation. 2.4. GT Label Error Detection and Refurbishment Pipeline A summary of the GT label error detection and refurbishment pipeline is provided in Algorithm 1. Details of the detection and refurbishment steps are provided below.
2.2. GT Label Error Generation
2.4.1. Label Error Detection
To evaluate the impact of GT label errors, we introduced synthetic but realistic errors into a random subset of the training and validation data. The specific error types used in this study are often found in real-life echo datasets and are summarised below (examples shown in Fig. 1):
To detect GT label errors, we use the VOG method which measures gradient variability of each training sample’s activations over several epochs. Samples with high gradient variability are likely mislabeled or noisy. We adopt the VOG approach by Khanal et al. [9], as defined in Equation 1.
effectiveness of the GT label error detection and refurbishment method.
v j D u X u1 X 1 2 t (Sie − µi ) VOGij = D t e=j−t
(1)
d=1
2.4.2. Refurbishment Adapted from SELFIE [6], we perform label refurbishment every 5 epochs following a 10-epoch warm-up. Each refurbishment replaces the training label with a pseudo-label obtained by averaging the model’s predictions over the preceding 5 epochs (Equation 2). ŷi =
E 1 X (e) pi 5
(2)
e=E−4
(e)
where ŷi is the refined pseudo-label for sample i, pi is the model’s prediction for sample i at epoch e, and E is the current epoch. Algorithm 1 Iterative VOG based Label Refurbishment Require: S: training set data, e: epochs, T : warm-up period, t: iteration interval Ensure: Model parameters, R: refurbished samples 1: for each epoch e in epochs do 2: for each batch B in S do 3: Train model on B 4: Record gradients for every sample 5: end for 6: if e > T and e mod t = 0 then 7: Calculate VOG for every sample 8: for samples with V OG > Q3 + 1.5 × IQR do 9: Apply label refurbishment method 10: Update GT segmentations in D 11: end for 12: end if 13: end for
3. EXPERIMENTS AND RESULTS The purpose of these experiments is to evaluate (i) the impact of GT label errors on segmentation performance and (ii) the
3.1. Experiment 1: GT Label Error Detection To evaluate the effectiveness of our method for identifying GT label errors, we conducted an experiment in which we apply random label errors to 25% of the training and validation sets. This was done for the three label errors detailed in Section 2.2. We compared our VOG based strategy with the more commonly used loss-based approach (i.e. based upon the training loss rather than VOG, with the same outlier detection method). We tested two different proportions of GT label errors. Table 1 summarises the results. In most cases, the VOG method consistently outperformed the loss-based approach, achieving higher accuracy, sensitivity and specificity in identifying GT label errors. Therefore, we use the VOG based strategy in our subsequent experiments. Table 1. Experiment 1 - Comparison of methods for identifying erroneous GT labels. Incomplete Boundary Merged labels distortion labels VOG Loss VOG Loss VOG Loss Accuracy 0.90 0.81 0.99 0.95 0.99 0.90 Sensitivity 0.86 0.32 0.97 0.87 1 1 Specificity 0.91 0.97 0.99 0.97 0.99 0.87
3.2. Experiment 2: Label Refurbishment To assess the refurbishment strategy’s effectiveness in correcting GT label errors, we compared Dice scores of refurbished labels before and after refurbishment, using the original GT labels as reference. This was tested under random and systematic label errors. As shown in Table 2, refurbishment consistently improved Dice scores, especially when random errors were below 50% and systematic errors below 25%. Table 2. Experiment 2 - Evaluation of refurbishment method. Dice scores for the erroneous dataset before and after refurbishment.
Systematic Random
where D is the dimension of the gradient vector, t = 5 is the number of previous epochs used to compute the variance, Sij is the gradient vector Pj of the activation for a sample xi at epoch j, and µi = 1t e=j−t Sie is the mean gradient vector over the last t epochs. To identify potential GT label errors, we apply a threshold based on the upper bound of the interquartile range (IQR): samples with VOG scores greater than Q3 + 1.5 × IQR are flagged as outliers (Q3 is the third quartile value). This robust criterion is commonly used in outlier detection and avoids reliance on arbitrary cutoffs.
12.5% 25% 50% 12.5% 25% 50%
before after after after before after after after
Incomplete labels 0.85 ± 0.17 0.95 ± 0.03 0.96 ± 0.04 0.90 ± 0.14 0.68 ± 0.03 0.94 ± 0.04 0.75 ± 0.10 0.69 ± 0.06
Boundary distortion 0.68 ± 0.06 0.84 ± 0.06 0.75 ± 0.13 0.70 ± 0.10 0.52 ± 0.11 0.87 ± 0.04 0.72 ± 0.15 0.54 ± 0.13
Merge labels 0.53 ± 0.45 0.93 ± 0.04 0.93 ± 0.04 0.69 ± 0.37 0.74 ± 0.04 0.94 ± 0.03 0.79 ± 0.09 0.77 ± 0.08
3.3. Experiment 3: Complete Pipeline
levels of GT label errors.
To evaluate model robustness to random GT label errors, we trained both a baseline U-Net and a model incorporating our GT label error detection and refurbishment strategy. Synthetic random segmentation errors were introduced into the training and validation sets at varying proportions, i.e., for each of the three error types (incomplete label, boundary distortion and merged labels), random class(es) were chosen for each affected sample. Results are shown in Fig. 2. As the proportion of incorrect labels increased, segmentation performance gradually declined. Surprisingly, despite the introduction of GT label errors, the baseline U-Net maintained relatively strong performance. The refurbishment strategy resulted in some improvements over the baseline, especially at higher proportions of errors.
Fig. 3. Experiment 3 - Systematic GT errors. Test set results when training with systematic GT label errors. Box plots display the Dice coefficients for each segmented structure and the overall mean for each model. The orange stars represent cases where the refurbished model performs better than the baseline model with a statistically significant difference by a Wilcoxon signed rank test (p-value<0.05).
4. DISCUSSION AND CONCLUSION
Fig. 2. Experiment 3 - Random GT label errors. Test set results when training with random GT label errors. Box plots display the Dice coefficients for each segmented structure and the overall mean for each model. The orange stars represent cases where the refurbished model performs better than the baseline model with a statistically significant difference by a Wilcoxon signed rank test (p-value<0.05). We repeated the previous experiment to assess model performance under systematic GT label errors. Results are shown in Fig. 3. Systematic errors led to a more pronounced degradation in segmentation quality compared to random errors. In the boundary distortion and merged labels scenarios, segmentation performance deteriorated significantly, with a higher incidence of outlier cases. With incomplete labels, the baseline model continued to perform reasonably well, again showing good robustness. Across all systematic error types, the refurbishment strategy offered small gains over the baseline at low
Across all experiments, we observed that the baseline segmentation model was surprisingly robust to a wide range of GT label errors. Random label errors had surprisingly little effect on final performance. For systematic errors, refurbishment provided some improvement, especially at moderate levels of corruption. This suggests that when the GT label errors follow structured patterns, error detection and refurbishment strategies may have a role to play. In practice, annotation errors typically affect less than 10% of data, making our high-error experiments a worst-case scenario. To our knowledge, the observed robustness of a naively trained U-Net under such conditions has not been previously reported and we believe that these findings offer a valuable contribution to the field. We also found that VOG reliably identified corrupted labels across error types and proportions, making it a practical, lightweight method for flagging incorrect GT labels without retraining or architectural changes. Taken together, these results highlight a promising degree of resilience in segmentation models and point toward selective, rather than routine, use of label refurbishment and correction strategies in clinical practice.
5. COMPLIANCE WITH ETHICAL STANDARDS
tional conference on machine learning. pp. 5907–5915.
PMLR, 2019,
This research study was conducted retrospectively using human subject data made available in open access [1]. Ethical approval was not required as confirmed by the license attached with the open access data.
[7] J. Li, R. Socher, and S. C. Hoi, “Dividemix: Learning with noisy labels as semi-supervised learning,” arXiv preprint arXiv:2002.07394, 2020.
6. ACKNOWLEDGMENTS
[8] L. Huang, C. Zhang, and H. Zhang, “Self-adaptive training: beyond empirical risk minimization,” Advances in neural information processing systems, vol. 33, pp. 19 365–19 376, 2020.
We would like to acknowledge funding from the EPSRC Centre for Doctoral Training in Medical Imaging (EP/L015226/1). 7. REFERENCES [1] S. Leclerc, E. Smistad, J. Pedrosa, A. Ostvik, F. Cervenansky, F. Espinosa, T. Espeland, E. A. R. Berg, P.-M. Jodoin, T. Grenier, C. Lartizien, J. Dhooge, L. Lovstakken, and O. Bernard, “Deep Learning for Segmentation Using an Open Large-Scale Dataset in 2D Echocardiography,” IEEE Transactions on Medical Imaging, vol. 38, no. 9, pp. 2198–2210, Sep. 2019. [2] J. Tromp, P. J. Seekings, C.-L. Hung, M. B. Iversen, M. J. Frost, W. Ouwerkerk, Z. Jiang, F. Eisenhaber, R. S. M. Goh, H. Zhao, W. Huang, L.-H. Ling, D. Sim, P. Cozzone, A. M. Richards, H. K. Lee, S. D. Solomon, C. S. P. Lam, and J. A. Ezekowitz, “Automated interpretation of systolic and diastolic function on the echocardiogram: a multicohort study,” The Lancet Digital Health, vol. 4, no. 1, pp. e46–e54, Jan. 2022. [3] E. Puyol-Antón, B. Ruijsink, B. S. Sidhu, J. Gould, B. Porter, M. K. Elliott, V. Mehta, H. Gu, C. A. Rinaldi, M. cowie et al., “Ai-enabled assessment of cardiac systolic and diastolic function from echocardiography,” in International Workshop on Advances in Simplifying Medical Ultrasound. Springer, 2022, pp. 75–85. [4] A. Ghorbani, D. Ouyang, A. Abid, B. He, J. H. Chen, R. A. Harrington, D. H. Liang, E. A. Ashley, and J. Y. Zou, “Deep learning interpretation of echocardiograms,” npj Digital Medicine, vol. 3, no. 1, p. 10, Jan. 2020. [Online]. Available: https://www.nature.com/articles/s41746-019-0216-8 [5] J. Mariscal-Harana, C. Asher, V. Vergani, M. Rizvi, L. Keehn, R. J. Kim, R. M. Judd, S. E. Petersen, R. Razavi, A. P. King et al., “An artificial intelligence tool for automated analysis of large-scale unstructured clinical cine cardiac magnetic resonance databases,” European Heart Journal-Digital Health, vol. 4, no. 5, pp. 370–383, 2023. [6] H. Song, M. Kim, and J.-G. Lee, “Selfie: Refurbishing unclean samples for robust deep learning,” in Interna-
[9] B. Khanal, T. Dai, B. Bhattarai, and C. Linte, “Active label refinement for robust training of imbalanced medical image classification tasks in the presence of high label noise,” in International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2024, pp. 37–47. [10] P. Chen, J. Ye, G. Chen, J. Zhao, and P.-A. Heng, “Beyond class-conditional assumption: A primary attempt to combat instance-dependent label noise,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 13, 2021, pp. 11 442–11 450. [11] B. Han, Q. Yao, X. Yu, G. Niu, M. Xu, W. Hu, I. Tsang, and M. Sugiyama, “Co-teaching: Robust training of deep neural networks with extremely noisy labels,” Advances in neural information processing systems, vol. 31, 2018. [12] E. Redekop and A. Chernyavskiy, “Uncertainty-based method for improving poorly labeled segmentation datasets,” in 2021 IEEE 18th international symposium on biomedical imaging (ISBI). IEEE, 2021, pp. 1831– 1835. [13] C. Xue, Q. Deng, X. Li, Q. Dou, and P.-A. Heng, “Cascaded robust learning at imperfect labels for chest x-ray segmentation,” in Medical Image Computing and Computer Assisted Intervention. Springer, 2020, pp. 579– 588. [14] S. Liu, K. Liu, W. Zhu, Y. Shen, and C. FernandezGranda, “Adaptive early-learning correction for segmentation from noisy annotations,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 2606–2616. [15] J. Li, R. Li, R. Han, and S. Wang, “Self-relabeling for noise-tolerant retina vessel segmentation through label reliability estimation,” BMC Medical Imaging, vol. 22, no. 1, p. 8, 2022. [16] O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in International Conference on Medical image computing and computer-assisted intervention. Springer, 2015, pp. 234–241.