ConceptioArchiveNCBI PubMed Central
NCBI PubMed Centralopen access

An Evaluation of the Mean Teacher Framework for Semi-Supervised Cataract Surgical Image Segmentation.

Faraji M et al. · ncbi_pmc
NCBI PubMed Central · Papers · License: Open Access
Open Source ↗Direct PDF ↓
distributed systems architecture

Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Transl Vis Sci Technol . 2026 Apr 9;15(4):5. doi: 10.1167/tvst.15.4.5 Search in PMC Search in PubMed View in NLM Catalog Add to search An Evaluation of the Mean Teacher Framework for Semi-Supervised Cataract Surgical Image Segmentation Mahtab Faraji Mahtab Faraji 1 Illinois Eye and Ear Infirmary, Department of Ophthalmology and Visual Sciences, University of Illinois Chicago, Chicago, IL, USA 2 Department of Biomedical Engineering, University of Illinois Chicago, Chicago, IL, USA 3 Artificial Intelligence in Ophthalmology (Ai-O) Center, University of Illinois Chicago, Chicago, IL, USA Find articles by Mahtab Faraji 1, 2, 3 , Darvin Yi Darvin Yi 1 Illinois Eye and Ear Infirmary, Department of Ophthalmology and Visual Sciences, University of Illinois Chicago, Chicago, IL, USA 3 Artificial Intelligence in Ophthalmology (Ai-O) Center, University of Illinois Chicago, Chicago, IL, USA Find articles by Darvin Yi 1, 3 , Michael J Heiferman Michael J Heiferman 1 Illinois Eye and Ear Infirmary, Department of Ophthalmology and Visual Sciences, University of Illinois Chicago, Chicago, IL, USA 3 Artificial Intelligence in Ophthalmology (Ai-O) Center, University of Illinois Chicago, Chicago, IL, USA Find articles by Michael J Heiferman 1, 3 , Homa Rashidisabet Homa Rashidisabet 1 Illinois Eye and Ear Infirmary, Department of Ophthalmology and Visual Sciences, University of Illinois Chicago, Chicago, IL, USA 3 Artificial Intelligence in Ophthalmology (Ai-O) Center, University of Illinois Chicago, Chicago, IL, USA Find articles by Homa Rashidisabet 1, 3, ✉ Author information Article notes Copyright and License information 1 Illinois Eye and Ear Infirmary, Department of Ophthalmology and Visual Sciences, University of Illinois Chicago, Chicago, IL, USA 2 Department of Biomedical Engineering, University of Illinois Chicago, Chicago, IL, USA 3 Artificial Intelligence in Ophthalmology (Ai-O) Center, University of Illinois Chicago, Chicago, IL, USA * Correspondence : Homa Rashidisabet, Illinois Eye and Ear Infirmary, Department of Ophthalmology and Visual Sciences, University of Illinois Chicago, 1855 W. Taylor St., Ste 1, Chicago, IL 60612, USA. e-mail: [email protected] ✉ Corresponding author. Accepted 2026 Feb 16; Received 2025 Jul 25; Collection date 2026 Apr. Copyright 2026 The Authors This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License. PMC Copyright notice PMCID: PMC13077722  PMID: 41954327 Abstract Purpose The purpose of this study was to evaluate a semi-supervised Mean Teacher (MT) framework for semantic segmentation of cataract surgical images, addressing the challenge of limited labeled data in real-world clinical applications. Methods We adapted the MT framework for four-class segmentation of the iris, pupil, intraocular lens, and surgical instruments using a small labeled set and 40,000 unlabeled Cataract-1K frames. Performance was assessed using Dice Similarity Coefficient (DSC) and the 95th percentile Hausdorff Distance (HD95). MT was compared with a fully supervised UNet under varying labeled and unlabeled conditions, and additional baselines provided broader context. Ablation studies evaluated noise types and key hyperparameters, including consistency weight (λ) and exponential moving average (EMA) decay (α). Internal validation used Cataract-1K, and external testing was done on CaDIS and CatInstSeg. Results Our MT model consistently outperformed the supervised UNet baseline. With 100 labeled images and 40,000 unlabeled frames, MT achieved a DSC of 0.76 ± 0.12 compared to UNet with 0.59 ± 0.17 ( t = 21.23, P < 0.05). On CaDIS and CatInstSeg, MT reached DSCs of 0.69 ± 0.17 and 0.71 ± 0.20, outperforming UNet at 0.62 ± 0.08 and 0.64 ± 0.23. Optimal performance was observed with λ = 0.1, α = 0.995, and Gaussian noise σ = 15. Conclusions The MT framework provides an effective semi-supervised solution for surgical image segmentation with limited annotations. Full source code and utility scripts will be released upon acceptance at: https://github.com/mahtabfaraji1/Semi-supervised-segmentation-of-cataract-surgical-images . Translational Relevance Accurate segmentation of ocular anatomy and instruments supports surgical guidance, intraoperative decision making, and training tools in data-limited clinical environments. Keywords: cataract surgical images, deep learning, semantic segmentation, semi-supervised segmentation, artificial intelligence Introduction Computer-assisted surgical systems for cataract procedures rely on video and image analysis to enhance intraoperative guidance and support skill assessment. 1 A key part of these systems is the semantic segmentation of anatomic structures and surgical instruments. 2 This segmentation forms the basis for both real-time tasks, such as ocular structure localization, intraocular lens (IOL) alignment, and corneal incision guidance, 1 as well as offline applications like surgical training. 3 Although recent supervised deep-learning methods 4 – 10 have shown promise in cataract surgical image segmentation, they rely on extensive pixel-level annotations. These annotations are labor-intensive to generate, requiring clinical expertise and detailed labeling of high-resolution frames. 11 , 12 As a result, the number of publicly available annotated datasets for cataract surgical image segmentation remains limited, with only a few, such as Cataract-1K, 5 CaDIS, 11 and CatInstSeg, 12 constraining the training and evaluation of segmentation models. To address the challenge of inadequate labeled data, a growing body of research in medical imaging has begun to focus on alternative supervision techniques in situations where annotations are limited, 13 , 14 such as weakly supervised learning. 15 , 16 Among these, semi-supervised learning (SSL) methods have shown promising performance in magnetic resonance imaging (MRI) 17 and computed tomography (CT), 18 where there is often a lack of labeled data. Unlike fully supervised methods, which depend only on labeled samples, SSL uses both labeled and unlabeled data by applying techniques like consistency regularization and pseudo-labeling. 19 Because unlabeled data can typically be acquired with minimal manual effort, the performance gains achieved by SSL often come at a relatively low additional cost. 20 Therefore, in this paper, we will focus on SSL, which is more suitable for use cases with partially labeled datasets and significant quantities of unlabeled data. SSL has also been used in ophthalmology, 21 showing promising results in tasks like retinal vessel segmentation, 22 optic disc and cup segmentation, 23 , 24 and lesion segmentation and disease severity classification 25 using fundus photographs and optical coherence tomography (OCT) images. However, even with this progress, the use of SSL methods for surgical image segmentation is still relatively limited. 26 – 28 In cataract surgery, limited studies have applied SSL techniques, 29 – 31 and these often lack detailed class-wise performance evaluation, particularly for smaller or under-represented categories, such as IOL and surgical instruments. This is clinically relevant, as segmentation performance can vary across classes due to imbalance and structural complexity. As noted by Grammatikopoulou et al. 32 and Pissas et al., 9 reliable performance across all categories is essential for downstream applications such as surgical guidance and skill assessment. Among the current SSL techniques, the mean teacher (MT) framework 33 has become a popular method for segmentation tasks with minimal supervision. The MT approach is a consistency-based SSL method that aligns outputs between a student model and its exponentially averaged teacher. Unlike other SSL approaches that introduce architectural modifications or multi-branch configurations, 1 , 20 , 34 , 35 MT maintains a relatively simple structure by reusing the same segmentation network architecture for both student and teacher models. As a result of its design, MT has been utilized in multiple areas, such as medical image segmentation. For example, when applied to the left atrium (LA) dataset with 20% of the data labeled, MT reached a Dice similarity coefficient (DSC) of 89.01%. 36 Recent variants of the MT framework incorporate mechanisms such as sharpness-aware optimization (COMT) 37 and cross-attention (UG-CEMT), 38 which improve segmentation performance but increase training complexity. 39 COMT and UG-CEMT reported DSC of 91.02% and 89.73%, respectively. In this study, we present an MT-based SSL approach for multi-class segmentation of cataract surgical images. Motivated by recent studies 39 suggesting that the standard MT framework can yield competitive results, we optimized our MT models through hyperparameter tuning and examined the effect of various noise functions on segmentation under limited labeled data. We compared our MT models primarily against a fully supervised UNet, 40 which serves as a widely used baseline for medical image segmentation. To provide a broader context, we also report results from established supervised architectures, including DeepLabV3+, 41 UPerNet, 42 and SwinUNet, 43 trained using the full labeled dataset. We further assessed generalizability by evaluating all models on external datasets. Our contributions include building upon the MT framework originally proposed by Tarvainen et al. 33 to develop a semi-supervised MT framework for segmenting cataract surgery images into four clinically relevant classes: iris, pupil, IOL, and instrument, under limited annotation settings, following the setup used by Ghamsarian et al. 6 We also identify the minimum amount of unlabeled data required to outperform supervised baselines and conduct ablation studies on core MT hyperparameters (e.g., consistency weight and exponential moving average [EMA] decay) and various noise functions to determine optimal configurations and enhance generalization on both internal and external validation datasets. Methods Dataset Source and Preprocessing This study utilized three publicly available datasets, including Cataract-1K, 6 CaDIS, 32 and CatInstSeg. 44 The Cataract-1K dataset 6 was used as the primary training source for our experiments. It consists of 2256 manually annotated frames extracted from 30 cataract surgery videos recorded at the Klinikum Klagenfurt Eye Clinic, each with a spatial resolution of 1024 × 768 pixels. 6 Whereas the annotations include 13 anatomic and instrument classes, our study focuses on four clinically relevant structures: iris, pupil, IOL, and surgical instruments, as shown in Figure 1 . These structures do not appear in every frame. A detailed breakdown of the proportion of frames containing the structure and average pixel-wise coverage is provided in Supplementary Table S1 ( Section A.1 ). Figure 1. Open in a new tab Randomly selected examples from the Cataract-1K 6 dataset showing segmentation of four clinically relevant structures, including the iris, pupil, intraocular lens, and surgical instruments. In addition to the labeled frames, we extracted 40,000 unlabeled images from the unannotated parts of the Cataract-1K video. We used the same frame extraction rate as the original dataset, taking one frame every 5 seconds. To ensure the data was of good quality, we manually filtered out frames that had severe motion blur or did not show visible eye structures, such as the iris. The complete data retrieval procedures are described in Supplementary Section B.3 . Table 1 shows the labeled and unlabeled samples used for training in our experiments. It also shows how the annotated classes are distributed in the training, validation, and test sets. Table 1. Overview of the Cataract-1K Dataset, Detailing its Source, Resolution, and the Distribution of Labeled Frames by Class Used in this Research Dataset Source Resolution, px Labeled Frames Classes Train Val Test Unlabeled Frames Cataract-1K 6 Surgical video 1024 × 768 2,256 Iris 1,803 240 213 40,000 microscope Pupil 1,803 240 213 — Instr 1,435 174 169 — IOL 403 63 69 — Open in a new tab IOL, intraocular lens; Instr surgical instruments. The chosen classes are a selection from the complete dataset, which encompasses a wider range of detailed surgical instrument categories that are not included in this study. To assess the generalizability of our models, we conducted external validation on publicly available CaDIS 32 and CatInstSeg 44 datasets. The CaDIS dataset consists of 4670 frames annotated into 36 categories, including 4 anatomic tissues, 29 types of surgical instruments, and 3 miscellaneous classes. CatInstSeg consists of 843 frames and focuses on surgical instrument segmentation with annotations for 11 distinct instrument types. In both datasets, not all structures are present in every frame. The distribution of frames in which each anatomic structure appears together with their average pixel-wise extent, and is reported in Supplementary Tables S2 (CaDIS, Sections A.2 ) and S3 (CatInstSeg, Section A.3 ). Table 2 summarizes the test set distributions and key acquisition characteristics for these external datasets used in the external evaluation. Table 2. Distribution of Annotated Test Samples From External Datasets (CaDIS and CatInstSeg) used for the External Evaluation Dataset Source Resolution, px Labeled Frames Classes CaDIS 32 Zeiss OPMI Lumera Microscope 1920 × 1080 4670 Anatomic tissues, surgical instruments, and miscellaneous CatInstSeg 44 Not specified 1280 × 720 843 Surgical instruments Open in a new tab Semi-Supervised Segmentation Framework The general design of our suggested MT semi-supervised semantic segmentation framework is illustrated in Figure 2 . The training process leverages a small labeled dataset D L = { ( x i , y i ) } i = 1 | D L | , where x i ∈ X ⊂ R H × W × 3 denotes an input RGB image of spatial dimensions H × W , and y i ∈ Y ⊂ R H × W is the corresponding pixel-wise segmentation mask. In addition, we utilize a substantially larger unlabeled dataset D U = { ( x i ) } i = 1 | D U | with | D L | < < | D U |. The framework comprises a student network θ s and a teacher network θ t , both utilizing the same architecture. Figure 2. Open in a new tab The proposed framework of the MT semi-supervised semantic segmentation model for cataract surgery images. Total Loss The total training loss for the student model, denoted as L T , combines both supervised and unsupervised components and is defined as follows in Equation 1 : L T = L s + λ × L c (1) where L s denotes the supervised segmentation loss, L c is the consistency loss on unlabeled data, and λ ∈ R is a consistency weight that controls the influence of the unsupervised component during training. Supervised Segmentation Loss The supervised segmentation loss combines cross-entropy and dice losses to jointly optimize pixel-wise classification and region overlap, following standard practice in medical image segmentation as follows in Equation 2 45 : L s u p e r v i s e d = 1 D L ∑ x , y ∈ D L 1 2 ∑ ω ∈ Ω 1 Ω L C E ∑ ω ∈ Ω y ω , p s l x , ω + L D i c e y , p s l x (2) here, Ω represents the set of pixel indices, y (ω) is the ground truth label at the pixel ω, and p s l ( x , ω ) is the student model's predicted class probability at pixel ω given the input image x . The cross-entropy loss L C E is the main classification loss. The dice loss L D i c e focuses on spatial overlap and promotes segmentation consistency across class boundaries. Consistency Loss and Input Noise Strategy The consistency loss enforces alignment of the predictions made by the student and the teacher on unlabeled data, and it is calculated as shown in Equation 3 46 : L c o n s i s t e n c y = 1 D U ∑ x ∈ D U ∑ ω ∈ Ω 1 Ω L M S E p t u x , ω , p s u x , ω (3) here, p t u ( x , ω ) and p s u ( x , ω ) represent the teacher and student predictions at pixel ω given the input image x , respectively. The L M S E represents the mean squared error (MSE) loss, encouraging the student model to mimic the teacher's pseudo-labels. In our setup, the student model receives a corrupted version of the unlabeled image, whereas the teacher processes the clear input. The types of noise include Gaussian noise, Gaussian blur, and impulse noise. Introducing noise to the student promotes consistency under changes. This approach helps students learn robust features while keeping consistent guidance from the teacher. Teacher Model Update The teacher model parameters θ t are updated using an EMA of the student parameters θ s , as shown in Equation 4 : θ t = α × θ t + 1 - α × θ s (4) where α ∈ ( 0 , 1 ) is the EMA decay factor that regulates the update rate. This update mechanism stabilizes the pseudo-labels produced by the teacher, ensuring smoother optimization for the student. Experimental Settings Baseline Methods To evaluate how well the MT framework works, we compared its performance to the state-of-the-art fully supervised models commonly used in research. We selected the UNet 40 architecture with a ResNet-50 backbone pre-trained on ImageNet, as our main baseline model. This network contains approximately 30.0 million trainable parameters. We used the same ImageNet-pretrained UNet-ResNet50 architecture for both the student and teacher networks in our MT framework; however, because the teacher model is implemented as EMA of the student's weights, it does not introduce additional trainable parameters. In our experiments, the supervised UNet model was trained using only the labeled subsets without access to unlabeled data. For consistency, we used identical training configurations, including input resolution, augmentation strategies, and optimization settings, for both the MT and UNet baseline models. In addition, to provide a broader context beyond the UNet baseline, we additionally included results from three widely used supervised architectures: DeepLabV3+, UPerNet, and SwinUNet. For DeepLabV3+ and UPerNet, we report performance values directly from the corresponding reference implementation, 6 as these models were originally trained on the full labeled Cataract-1K dataset under comparable settings. In contrast, SwinUNet was trained as part of this work using the full labeled dataset. SwinUNet 43 uses a hierarchical Swin Transformer encoder and was initialized from the publicly released Swin-Tiny pretrained checkpoint. Although its architecture requires its own recommended training configuration, we ensured that SwinUNet followed the same preprocessing, augmentation, and evaluation protocols used for the other supervised models, allowing for consistent comparisons across baselines. Effect of Labeled and Unlabeled Data Size We systematically evaluated the effectiveness of the MT framework under varying quantities of labeled and unlabeled data from the Cataract-1K dataset. Specifically, we evaluated performance using 100, 300, 900, and the full set of 1803 labeled images while maintaining access to the full pool of 40,000 unlabeled frames. The same labeled subsets were used for the supervised UNet baseline to ensure consistency in comparison. The labeled data sampling strategy is detailed in the Supplementary Section B.1 , and the class distributions for each labeled subset are provided in Supplementary Table S5 ( Section B.2 ). To investigate the role of unlabeled data, we fixed the labeled set to 100 images and varied the size of the unlabeled set, using 5000, 10,000, 20,000, and 40,000 images. These subsets were randomly sampled from the full unlabeled pool. Generalization to External Datasets We performed external validation using the supervised UNet baseline model that was trained on the completely labeled Cataract-1K dataset, along with our semi-supervised MT model that incorporated the same labeled data and an additional 40,000 unlabeled images from the Cataract-1K dataset. Both models were evaluated on two external datasets, CaDIS and CatInstSeg. To ensure consistent evaluation of surgical instrument segmentation across datasets, we reformulated the external datasets into a binary segmentation format, distinguishing foreground (instruments) from background. Specifically, we selected only those instrument classes in CaDIS and CatInstSeg that were also annotated in the Cataract-1K dataset (as detailed in Supplementary Table S4 ( Section A.4 )) and assigned them to the foreground class. All remaining regions were treated as background. Supplementary Figure S1 illustrates representative examples from each dataset, showing the original surgical frame alongside its corresponding ground truth binary segmentation mask. Effect of Core MT Parameters and Noise Function We conducted a series of ablation experiments to evaluate the sensitivity of MT to key design choices. All experiments were performed using a fixed configuration of 100 labeled and 40,000 unlabeled Cataract-1K images, with all other training settings held constant. Specifically, we varied the λ, α, and the type and severity of input noises applied to the student model. Implementation Details All models were developed using PyTorch and trained on a workstation equipped with three NVIDIA GeForce RTX 3090 GPUs (24 GB VRAM each). Performance assessments were carried out based on the results of the student model during inference. The training process involved 40,000 iterations using the stochastic gradient descent (SGD) optimizer. It started with a learning rate of 0.01, a momentum value of 0.9, and a weight decay of 0.0001. The mini-batch size was set at 24. The base value of λ was kept at zero for an initial warm-up period and then increased to 0.1 over the next 1000 iterations using a sigmoid ramp-up schedule to improve training stability during the early phase. The EMA decay rate for the teacher model was set at 0.995. All images for the MT and UNet models were resized to 256 × 256 using nearest-neighbor interpolation, whereas SwinUNet used an input resolution of 224 × 224 to match the requirements of its pretrained Swin-Tiny backbone. All images were augmented with random horizontal flips, vertical flips, and rotations. For every model, the dataset was split by video index to ensure patient-independent partitions, with 80% of videos used for training, 10% for validation, and 10% for testing. For the semi-supervised setup, noises were applied to the student model's input to encourage consistency under noise. Three types of noises were evaluated: (1) Gaussian noise, applied by sampling from a zero-mean normal distribution with standard deviations of σ = 2 (low), 15 (medium), and 30 (high); (2) Gaussian blur, applied using kernel sizes of 3 × 3, 5 × 5, and 7 × 7; and (3) impulse noise, simulated by randomly replacing 1%, 3%, and 5% of pixel intensities with maximum or minimum values. Severity levels were selected based on visual inspection to ensure that anatomic interpretability was qualitatively preserved while still introducing meaningful noises for regularization. A representative image illustrating each noise type at three severity levels is provided in Supplementary Figure S3 ( Section E ). In addition to tuning MT-specific hyperparameters and input noises, we performed a grid search to optimize general training parameters, such as learning rate, batch size, weight decay, and data augmentation, for both the MT model and all supervised baselines. The final parameters were chosen according to their performance on the validation set. A detailed summary of the hyperparameter configurations evaluated during grid search is provided in Supplementary Table S6 ( Section C ). Evaluation Metrics The effectiveness of the segmentation was primarily evaluated using the DSC. 47 The DSC measures the overlap between the predicted segmentations and the true ground truth segmentations. Its definition is as follows in Equation 5 : D S C = 2 × T P 2 × T P + F P + F N (5) where TP refers to true positives, FP denotes false positives, and FN signifies false negatives. A DSC value of 1 indicates perfect overlap, whereas 0 indicates no overlap. In addition to our primary metric, we present the 95th percentile Hausdorff Distance (HD95), 47 which quantifies the spatial difference between the boundaries of predicted segmentations and the actual ground-truth segmentations. The definition of HD is as follows in Equation 6 : H D = max max x ∈ X min y ∈ Y d x , y , max y ∈ Y min x ∈ X d x , y (6) where X and Y are sets of surface points from predicted and actual segmentations and d (., .) denotes the Euclidean distance. HD95 is computed by taking the 95th percentile of these distances rather than the maximum to reduce sensitivity to outliers. Unlike the DSC, HD95 does not have an upper bound and is reported in units of pixels (or millimeters when physical spacing is known). 48 An HD95 of 0 indicates perfect boundary alignment; however, what constitutes “good” depends strongly on image resolution, anatomic size, and task complexity. To assess the statistical significance of performance differences, we conducted two-tailed paired t -tests comparing the MT framework against the supervised baselines. The tests were applied to per-image DSC and HD95 scores for each anatomic class and each labeled-data regime, and we additionally computed macro-averaged t -tests by averaging the per-class metrics for each image before comparison. All paired comparisons were performed on identical test samples, and differences were considered statistically significant at P < 0.05, which is widely used in medical imaging research to indicate statistically meaningful differences. 49 , 50 Results Segmentation Performance Table 3 reports the per-class averaged DSC and HD95 performance, together with the associated statistical significance values, for both supervised baselines and our semi-supervised MT models across different labeled-data regimes. With 100 labeled images, the semi-supervised MT model achieved a DSC and HD95 of 0.76 ± 0.12 and 65.78 ± 49.45, respectively, compared to 0.59 ± 0.17 and 99.82 ±  49.16 (t = 21.23, P < 0.05 for DSC and t = 15.33, P < 0.05 for HD95) for the supervised UNet. When trained on the full set of 1803 labeled images, the MT framework further improved to a DSC of 0.90 ± 0.06 and HD95 of 17.78 ± 24.93, again outperforming the UNet baseline with a statistically significant margin (t = 13.0, P < 0.05 for DSC and t = 4.38, P < 0.05 for HD95). Table 3. Per-Class Averaged DSC and HD95 for Cataract Surgical Image Segmentation Using the Supervised Baselines and our Semi-Supervised MT Models Across Different Labeled Data Sizes Frame Used Performance (Mean ± STD) Statistic (t-Value) Model Unlabeled Labeled DCS↑ HD95↓ DCS HD95 DeepLabV3+ 6 , † 0 Full 0.85 — — — UPerNet 6 , † Full 0.89 — — — SwinUNet 43 Full 0.89  ±  0.06 17.04  ±  17.71 6.37 * 0.41 UNet 40 0 100 0.59  ±  0.17 99.82  ±  49.16 21.23 * 15.33 * 300 0.80  ±  0.11 51.66  ±  45.68 16.04 * 10.68 * 900 0.85  ±  0.10 27.63  ±  33.87 12.22 * 7.86 * Full 0.86  ±  0.08 23.72  ±  26.82 13.0 * 4.38 * Semi-supervised (Ours) 40k 100 0.76  ±  0.12 * 65.78  ±  49.45 * — — 300 0.87  ±  0.07 * 29.87  ±  29.85 * — — 900 0.89  ±  0.07 * 18.14  ±  19.96 * — — Full 0.90  ±  0.06 * 17.78  ±  24.93 * — — Open in a new tab * Statistically significant improvements of MT over the baseline ( P < 0.05, two-tailed paired t -test) are indicated with an asterisk (*). † External results are reproduced from a previously published study. 6 For reference, Table 3 also includes external benchmarks reported in a prior study, 6 where the fully supervised DeepLabV3+ and UPerNet models achieved average DSCs of 0.85 and 0.89, respectively. Statistical values are not provided for these two models because their results were taken directly from the prior publication, 6 and the corresponding per-image predictions required for statistical testing were not available. In our evaluation of fully supervised baselines, SwinUNet 43 achieved a DSC of 0.89 ± 0.06 and an HD95 of 17.04 ± 17.71. Under the full-data setting, our MT model obtained a DSC of 0.90 ± 0.06 and an HD95 of 17.78 ± 24.93 (t = 6.37, P < 0.05 for DSC and t = 0.41, P = 0.681 for HD95). Detailed per-class results are provided in Supplementary Sections D1 ( Tables S7 and S8 ) and D2 ( Tables S9 and Table S10 ). Figure 3 presents qualitative segmentation outputs from both the supervised UNet baseline and our semi-supervised MT model, trained using different quantities of labeled Cataract-1K images. The MT models were trained with 40,000 unlabeled images. Segmented regions are color-coded as follows: pupil (sky blue), iris (orange), IOL (bluish green), and instruments (reddish purple). Visual differences can be observed between models trained with 100 versus full labeled images, particularly in the delineation of small or complex structures, such as the IOL and instruments. Figure 3. Open in a new tab Qualitative comparison of segmentation results from the UNet baseline and our MT models trained with varying labeled data sizes. The MT models were trained with 40,000 unlabeled images. Color coding: pupil ( sky blue ), iris ( orange ), intraocular lens ( bluish green ), and instruments ( reddish purple ). Table 4 reports segmentation performance, measured by DSC and HD95, for our semi-supervised MT model trained with increasing amounts of unlabeled Cataract-1K images while keeping the labeled set fixed at 100 images. The average DSC improved from 0.63 ± 0.17 to 0.76 ± 0.12 when increasing unlabeled data from 5000 to 40,000 frames, with a corresponding reduction in HD95 from 93.07 ± 43.43 to 65.78 ±  49.45. For reference, the UNet baseline trained on the same 100 labeled images without unlabeled data achieved an average DSC of 0.59   ± 0.17 and an HD95 of 99.82 ± 49.16. Per-class DSC and HD95 values for the iris, pupil, IOL, and instruments are also provided in Supplementary Section D.3 ( Tables S11 for DSC and S12 for HD95). Table 4. Per-Class Averaged DSC and HD95 for Cataract Surgical Image Segmentation Using the UNet Baseline and our Semi-Supervised MT Models Across Different Unlabeled Data Sizes Frame Used Performance (Mean ± STD) Model Labeled Unlabeled DCS↑ HD95↓ UNet 40 100 0 0.59  ±  0.17 99.82  ±  49.16 Semi-supervised (Ours) 100 5k 0.63  ±  0.17 93.07  ±  43.43 10k 0.67  ±  0.13 80.51  ±  45.58 20k 0.73  ±  0.14 78.27  ±  50.07 40k 0.76  ±  0.12 65.78  ±  49.45 Open in a new tab Generalization to External Dataset Table 5 reports binary surgical instrument segmentation performance for the UNet baseline and our semi-supervised MT models, both trained on the fully labeled Cataract-1K dataset. On the Cataract-1K test set, MT achieved a DSC of 0.88 ± 0.09 and HD95 of 5.52 ± 9.21, whereas the UNet baseline achieved a DSC of 0.83 ± 0.13 and HD95 of 7.70 ± 11.12. These improvements were statistically significant for both DSC and HD95. Table 5. Binary Surgical Instrument Segmentation Results for the UNet Baseline and our Semi-Supervised MT Models Trained on Cataract-1K, Evaluated on the Internal Test Set and Two External Datasets (CaDIS and CatInstSeg) Performance (Mean ± STD) Statistic (t-Value) Model Labeled Unlabeled Source Dataset Target Dataset Domain DCS↑ HD95↓ DCS HD95 UNet 40 Full 0 Cataract-1K Cataract-1K Internal 0.83  ±  0.13 7.70  ±  11.12 5.79 * 2.92 * CaDIS External 0.62  ±  0.08 36.10  ±  30.55 2.18 * 4.04 * CatInstSeg External 0.64  ±  0.23 54.32  ±  66.34 12.06 * 7.53 * Semi-supervised (Ours) Full 40k Cataract-1K Cataract-1K Internal 0.88  ±  0.09 5.52  ±  9.21 — — CaDIS External 0.69  ±  0.17 19.78  ±  20.87 — — CatInstSeg External 0.71  ±  0.20 36.30  ±  53.92 — — Open in a new tab * Statistically significant improvements of MT over the baseline ( P < 0.05, two-tailed paired t -test) are indicated with an asterisk (*). In external evaluations, the UNet baseline achieved DSCs of 0.62 ±   0.08 on CaDIS and 0.64 ± 0.23 on CatInstSeg, with HD95 values of 36.10 ± 30.55 and 54.32 ± 66.34, respectively. The MT obtained DSCs of 0.69 ± 0.17 (CaDIS) and 0.71 ± 0.20 (CatInstSeg), and HD95 values of 19.78 ± 19.02 and 36.30 ± 53.92, respectively. Paired t -tests confirmed that the improvements were statistically significant for both DSC and HD95. Effect of Core MT Parameters and Noise Function Additional results from ablation experiments on MT hyperparameters and input noise are provided in Supplementary Section E . The highest average DSC was observed at λ = 0.1 and α = 0.995. For input noise, medium Gaussian noise (σ = 15) produced an average DSC of 0.76 ± 0.12. Larger impulse noise resulted in the lowest performance (DSC = 0.60 ± 0.27). Detailed per-class DSC and HD95 results for the ablation experiments are provided in Supplementary Figure S2 ( Section E ). Discussion Accurate segmentation of ocular anatomy and surgical instruments is essential for enabling intelligent surgical assistance. Anatomic segmentation facilitates tasks such as eye localization, incision planning, and intraocular lens alignment, 1 whereas instrument segmentation supports analysis of tool–tissue interactions and intraoperative skill assessment. Due to the annotation burden in surgical datasets, we built upon the MT semi-supervised framework 33 for multi-class segmentation of cataract surgical images under limited labeled data conditions. These four structures are meaningful because the iris and pupil provide spatial orientation within the eye, the instruments indicate the safety of surgical maneuvers, and the IOL is central to postoperative visual outcomes. Across all experiments, our MT framework consistently improved performance over the fully supervised UNet baseline, particularly when labeled data were limited. With only 100 labeled frames, MT increased the mean DSC from 0.59 ± 0.17 to 0.76 ± 0.12 and reduced HD95 from 99.82 ± 49.16 to 65.78 ± 49.45, indicating improvement in both overlap and boundary accuracy. These improvements were statistically significant across all evaluated label scenarios ( P < 0.05). In addition, as more labeled data became available, the performance gap between MT and the UNet baseline narrowed, yet MT maintained consistent advantages in both DSC and HD95 across all scenarios. For example, with the full labeled dataset, MT achieved a DSC of 0.90 ± 0.06, compared with 0.86 ± 0.08 for the UNet baseline. A per-class analysis shows that MT improved performance across all anatomic classes, with particularly noticeable improvement for smaller or visually complex structures such as the IOL and surgical instruments, which are more prone to overfitting due to limited pixel coverage and high variability. With 300 labeled images, MT improved instrument DSC from 0.66 ± 0.14 to 0.80 ± 0.12 and reduced HD95, with similar improvements observed for the IOL across all label scenarios. These results suggest that the teacher–student consistency mechanism helps the model learn more stable, discriminative features for under-represented structures. To contextualize MT's performance, we also compared it with fully supervised state-of-the-art models reported previously. 6 DeepLabV3+ and UPerNet achieved average DSCs of 0.85 and 0.89; however, differences in training protocols limit direct comparison. For a consistent baseline, we trained SwinUNet on the full labeled dataset, which reached DSC 0.89 ± 0.06 and HD95 17.04 ± 17.71, closely matching MT under full supervision (0.90 ± 0.06 and 17.78 ± 24.93, respectively). However, per-class results show that MT provided a small DSC gain for instrument segmentation (0.85 ± 0.12 vs. 0.83 ± 0.11; t = 3.32, P < 0.05). In addition, unlike SwinUNet, MT achieves this performance without additional architectural complexity, retaining a lightweight design that reuses the same architecture for both student and teacher. Beyond the supervised comparisons, recent SSL approaches 46 , 51 , 52 frequently rely on transformer-based backbones or multi-stage teacher–student pipelines, which increase architectural and computational overhead. In contrast, our results show that MT achieves comparable performance with a far simpler design, indicating that lightweight consistency-based methods remain competitive and effective for cataract surgical image segmentation. We further evaluated the scalability of the MT framework with increasing amounts of unlabeled data. Performance improved progressively as the number of unlabeled frames increased from 5000 to 40,000, indicating that additional unlabeled data continued to contribute positively to segmentation accuracy under the current model configuration. Our ablation studies demonstrated that MT performance is sensitive to both hyperparameters and input noise. Optimal results were achieved with a consistency weight (λ) of 0.1 and EMA decay α of 0.995, consistent with prior SSL segmentation studies. 53 To simulate real-world variability in cataract surgery images, we applied Gaussian noise (representing sensor fluctuations), Gaussian blur (caused by defocus or motion), and impulse noise (resulting from high-frequency distortions). Without noise, the MT model outperformed the supervised UNet baseline (DSC = 0.66 ± 0.15 vs. 0.59 ± 0.17), confirming the value of consistency training alone. Adding noise further improved performance, with Gaussian noise at σ = 15 yielding the highest DSC (0.76 ± 0.12). However, excessive noises, such as 5% impulse noise, led to degraded accuracy (DSC = 0.60 ± 0.27), comparable to the supervised baseline. These findings support earlier studies, 39 , 54 showing that the performance of consistency-based semi-supervised learning is influenced by both the type and intensity of the noise used. External validation demonstrated that the MT framework maintained its performance across the external validation dataset. This suggests that consistency-based training can improve robustness to domain shift, which is common in multicenter ophthalmic datasets. 34 , 55 Such domain shifts may result from differences in physical imaging parameters, 56 including sensor resolution, illumination conditions, and optical characteristics, all of which can alter the visual appearance of surgical scenes. Moreover, variability stemming from data collected across different institutions, imaging protocols, and patient demographics further amplifies this effect. 55 , 57 Additionally, inconsistencies in ground-truth labels, arising from variations in annotation protocols, annotator expertise, and clinical objectives, introduce additional variability that contributes to domain shift in medical image segmentation. 58 Variability in surgical protocols, 59 such as differences in procedural steps or instrument positioning, 60 further contributes to domain shift by influencing the appearance of key anatomic structures. Although the results are promising, several limitations should be acknowledged. First, our approach performs frame-wise segmentation without modeling temporal continuity. This may limit its ability to capture dynamic context in surgical videos. Incorporating temporal models in future work could enable analysis of clinically important motion patterns, such as instrument tip trajectories, iris billowing, and IOL stability, which are valuable for surgical training and performance evaluation. Second, although we used a large unlabeled dataset from Cataract-1K, incorporating additional unlabeled data from diverse clinical settings may improve cross-institutional generalizability. 1 Third, class imbalance remains an issue not only in our dataset but also in many surgical image segmentation tasks. 61 Whereas the MT framework improved segmentation for small structures to some degree, future work can involve custom loss functions or adaptive sampling strategies to address this issue to a certain extent. 61 In addition, as noted earlier, domain shifts may influence model performance, and future analyses should investigate the factors that most strongly drive these shifts. Last, it is worth noting that the proposed framework is designed for offline segmentation and extending it for real-time inference remains an area for future work. In brief, the MT model provides an efficient and scalable semi-supervised method for cataract surgical image segmentation, particularly in cases with limited annotation. Given that the annotation of surgical images is costly and involves clinical expertise, techniques like MT that utilize unlabeled data offer practical value. Its ability to generalize, simplicity, and compatibility with real training conditions support its readiness for use in future smart surgical devices and learning platforms. Supplementary Material Supplement 1 tvst-15-4-5_s001.docx (4.9MB, docx) Acknowledgments Supported by an Unrestricted Grant from Research to Prevent Blindness, generous donation from Cless Family Foundation, NIH P30 EY001792 Core Grant. Code and Data Availability: The Cataract-1K dataset is available via Synapse at https://www.synapse.org/Synapse:syn53404507 , the CaDIS dataset can be accessed through the Grand Challenge platform at https://cataracts-semantic-segmentation2020.grand-challenge.org/Data/ , and the CatInstSeg dataset is available at https://ftp.itec.aau.at/datasets/ovid/InSegCat/ . To facilitate reproducibility, the full source code and utility scripts will be released via our public GitHub repository: https://github.com/mahtabfaraji1/Semi-supervised-segmentation-of-cataract-surgical-images . Trained model weights will be made available upon reasonable request. Disclosure: M. Faraji , None; D. Yi , None; M.J. Heiferman , None; H. Rashidisabet , None References 1. Zhang Y, Pan Y, Ou M, et al.. Domain adaptation for semantic segmentation of cataract surgical images based on masked image consistency. In: Applied Intelligence . New York, NY: Springer Nature; 2025: 243–254. [ Google Scholar ] 2. Müller S, Jain M, Sachdeva B, et al.. Artificial intelligence in cataract surgery: a systematic review. Transl Vis Sci Technol . 2024; 13(4): 20. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 3. Nespolo R, Nahass GR, Faraji M, Yi D, Leiderman YI. Assessing vitreoretinal surgical training experience by leveraging instrument maneuvers and visual attention with deep learning neural networks. Invest Ophthalmol Vis Sci . 2024; 65(7): 900. [ Google Scholar ] 4. Bhattarai B, Subedi R, Gaire RR, Vazquez E, Stoyanov D. Histogram of oriented gradients meet deep learning: a novel multi-task deep network for 2D surgical image semantic segmentation. Med Image Anal . 2023; 85: 102747. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 5. Sachdeva B, Akash N, Ashraf T, et al.. Phase-informed tool segmentation for manual small-incision cataract surgery. a rXiv (Preprint) . Available at: http://arxiv.org/abs/2411.16794 . 6. Ghamsarian N, El-Shabrawi Y, Nasirihaghighi S, et al.. Cataract-1K dataset for deep-learning-assisted analysis of cataract surgery videos. Sci Data . 2024; 11(1): 373. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 7. Li H, Ou M, Li H, et al.. Multi-view test-time adaptation for semantic segmentation in clinical cataract surgery. IEEE Trans Med Imaging . 2025; 44: 2307–2318. [ DOI ] [ PubMed ] [ Google Scholar ] 8. Fox M, Taschwer M, Schoeffmann K. Pixel-based tool segmentation in cataract surgery videos with Mask R-CNN. In: 2020 IEEE 33rd International Symposium on Computer-Based Medical Systems (CBMS) . 2020: 565–568. 9. Pissas T, Ravasio CS, Da Cruz L, Bergeles C. Effective semantic segmentation in cataract surgery: what matters most? In: Medical Image Computing and Computer Assisted Intervention – MICCAI 2021 . 2021: 509–518. 10. Yeh HH, Sen S, Chou JC, Christopher KL, Wang SY. PhacoTrainer: automatic artificial intelligence-generated performance ratings for cataract surgery. Transl Vis Sci Technol . 2025; 14(5): 2. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 11. Ma J, He Y, Li F, Han L, You C, Wang B. Segment anything in medical images. Nat Commun . 2024; 15(1): 654. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 12. Wang S, Li C, Wang R, et al.. Annotation-efficient deep learning for automatic medical image segmentation. Nat Commun . 2021; 12(1): 5915. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 13. Ran L, Li Y, Liang G, Zhang Y. Pseudo labeling methods for semi-supervised semantic segmentation: a review and future perspectives. IEEE Trans Circuits Syst Video Technol . 2025; 35(4): 3054–3080. [ Google Scholar ] 14. Ma S, Du H, An Y, et al.. Deep learning approaches for medical imaging under varying degrees of label availability: a comprehensive survey. arXiv (Preprint) . Available at: https://arxiv.org/abs/2504.11588 . 15. Zhong Y, Tang C, Yang Y, et al.. Weakly-supervised medical image segmentation with gaze annotations. In: Medical Image Computing and Computer Assisted Intervention – MICCAI 2024 . Springer Nature Switzerland; 2024: 530–540. [ Google Scholar ] 16. Kuang Z, Yan Z, Yu L. Weakly supervised learning for multi-class medical image segmentation via feature decomposition. Comput Biol Med . 2024; 171: 108228. [ DOI ] [ PubMed ] [ Google Scholar ] 17. Xiao Z, Su Y, Deng Z, Zhang W. Efficient combination of CNN and transformer for dual-teacher uncertainty-guided semi-supervised medical image segmentation. Comput Methods Programs Biomed . 2022; 226: 107099. [ DOI ] [ PubMed ] [ Google Scholar ] 18. Chen X, Zhou HY, Liu F, Guo J, Wang L, Yu Y. MASS: modality-collaborative semi-supervised segmentation by exploiting cross-modal consistency from unpaired CT and MRI images. Med Image Anal . 2022; 80: 102506. [ DOI ] [ PubMed ] [ Google Scholar ] 19. Lu L, Yin M, Fu L, Yang F. Uncertainty-aware pseudo-label and consistency for semi-supervised medical image segmentation. Biomed Signal Process Control . 2023; 79: 104203. [ Google Scholar ] 20. Sohn K, Berthelot D, Li CL, et al.. FixMatch: simplifying semi-supervised learning with consistency and confidence. arXiv ( Preprint) . Available at: http://arxiv.org/abs/2001.07685 . 21. Faraji M, Rashidisabet H, Nahass GR, Chan RP, Vajaranant TS, Yi D. Trends, challenges, and future directions in deep learning for glaucoma: a systematic review. arXiv (Prepr int) . Available at: http://arxiv.org/abs/2411.05876 . 22. Kugelman J, Alonso-Caneiro D, Read SA, Vincent SJ, Collins MJ. Enhanced OCT chorio-retinal segmentation in low-data settings with semi-supervised GAN augmentation using cross-localisation. Comput Vis Image Underst . 2023; 237: 103852. [ Google Scholar ] 23. Liu S, Hong J, Lu X, et al.. Joint optic disc and cup segmentation using semi-supervised conditional GANs. Comput Biol Med . 2019; 115: 103485. [ DOI ] [ PubMed ] [ Google Scholar ] 24. Wu F, Zhuang X. Minimizing estimated risks on unlabeled data: a new formulation for semi-supervised medical image segmentation. IEEE Trans Pattern Anal Mach Intell . 2023; 45(5): 6021–6036. [ DOI ] [ PubMed ] [ Google Scholar ] 25. Zhou Y, He X, Huang L, et al.. Collaborative learning of semi-supervised segmentation and classification for medical images. In: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 2019: 2074–2083. 26. Faraji M, Rashidisabet H, Heiferman M, Yi D. Evaluating a semi-supervised segmentation method in cataract surgical images. Invest Ophthalmol Vis Sci . 2025; 66(8): 570. [ Google Scholar ] 27. Alapatt D, Mascagni P, Vardazaryan A, et al.. Temporally Constrained Neural Networks (TCNN): a framework for semi-supervised video semantic segmentation. arXiv (Preprint) . Available at: http://arxiv.org/abs/2112.13815 . 28. Zhao Z, Jin Y, Gao X, Dou Q, Heng PA. Learning Motion Flows for Semi-supervised Instrument Segmentation from Robotic Surgical Video. In: Martel AL, Abolmaesumi P, Stoyanov D, et al., eds. Medical Image Computing and Computer Assisted Intervention – MICCAI 2020 . New York, NY: Springer International Publishing; 2020: 679–689. [ Google Scholar ] 29. Ghamsarian N, Nasirihaghighi S, Schoeffmann K, Sznitman R. Feedback-driven pseudo-label reliability assessment: redefining thresholding for semi-supervised semantic segmentation. arXiv (Preprin t) . Available at: http://arxiv.org/abs/2505.07691 . 30. Li D, Hu Y, Shen J, Hao L, Liu J. Semi-supervised surgical video semantic segmentation with cross supervision of inter-frame. In: 2023 IEEE 20th International Symposium on Biomedical Imaging (ISBI) . 2023: 1–5. 31. Chen H, Ma X, Xia T, Jia F. Semi-supervised semantic segmentation of cataract surgical images based on DeepLab v3+. In: Proceedings of the 2021 5th International Conference on Compute and Data Analysis . ICCDA ’21. Association for Computing Machinery; 2021: 112–118. [ Google Scholar ] 32. Grammatikopoulou M, Flouty E, Kadkhodamohammadi A, et al.. CaDIS: cataract dataset for surgical RGB-image segmentation. Med Image Anal . 2021; 71: 102053. [ DOI ] [ PubMed ] [ Google Scholar ] 33. Tarvainen A, Valpola H. Mean teachers are better role models: weight-averaged consistency targets improve semi-supervised deep learning results. a rXiv (Preprint) . Available at: http://arxiv.org/abs/1703.01780 . 34. Ghamsarian N, Tejero JG, Neila PM, et al.. Domain adaptation for medical image segmentation using transformation-invariant self-training. a rXiv (Preprint) . Available at: https://arxiv.org/abs/2307.16660v1 . 35. Chen X, Yuan Y, Zeng G, Wang J. Semi-supervised semantic segmentation with cross pseudo supervision. arXiv (Pre print) . Available at: http://arxiv.org/abs/2106.01226 . 36. Yu L, Wang S, Li X, Fu CW, Heng PA. Uncertainty-aware self-ensembling model for semi-supervised 3D left atrium segmentation. In: Medical Image Computing and Computer Assisted Intervention – MICCAI 2019 . New York, NY: Springer International Publishing; 2019: 605–613. [ Google Scholar ] 37. Li N, Pan Y, Qiu W, Xiong L, Wang Y, Zhang Y. Constantly optimized mean teacher for semi-supervised 3D MRI image segmentation. Med Biol Eng Comput . 2024; 62(7): 2231–2245. [ DOI ] [ PubMed ] [ Google Scholar ] 38. Karri M, Arya AS, Biswas K, et al.. Uncertainty-guided cross attention ensemble mean teacher for semi-supervised medical image segmentation. arXiv ( Preprint) . Available at: http://arxiv.org/abs/2412.15380 . 39. Zhao Z, Yang L, Long S, Pi J, Zhou L, Wang J. Augmentation matters: a simple-yet-effective approach to semi-supervised semantic segmentation. arXiv ( Preprint) . Available at: http://arxiv.org/abs/2212.04976 . 40. Ronneberger O, Fischer P, Brox T. U-Net: convolutional networks for biomedical image segmentation. a rXiv (Preprint) . Available at: http://arxiv.org/abs/1505.04597 . 41. Chen LC, Zhu Y, Papandreou G, Schroff F, Adam H. Encoder-decoder with atrous separable convolution for semantic image segmentation. arXiv (Preprint) . Available at: doi: 10.48550/arXiv.1802.02611. [ DOI ] 42. Xiao T, Liu Y, Zhou B, Jiang Y, Sun J. Unified perceptual parsing for scene understanding. arXiv (Preprint) . Available at: doi: 10.48550/arXiv.1807.10221. [ DOI ] 43. Cao H, Wang Y, Chen J, et al.. Swin-Unet: Unet-like pure transformer for medical image segmentation. arXiv (Preprint) . Available at: doi: 10.48550/arXiv.2105.05537. [ DOI ] 44. Al Hajj H, Lamard M, Conze PH, et al.. CATARACTS: challenge on automatic tool annotation for cataRACT surgery. Med Image Anal . 2019; 52: 24–41. [ DOI ] [ PubMed ] [ Google Scholar ] 45. Jadon S. A survey of loss functions for semantic segmentation. In: 2020 IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB) . 2020: 1–7. 46. Liu Y, Tian Y, Chen Y, Liu F, Belagiannis V, Carneiro G. Perturbed and strict mean teachers for semi-supervised semantic segmentation. arXiv (Preprint) . Available at: doi: 10.48550/arXiv.2111.12903 [ DOI ] 47. Gao F, Hu M, Zhong ME, et al.. Segmentation only uses sparse annotations: unified weakly and semi-supervised learning in medical images. Med Image Anal . 2022; 80: 102515. [ DOI ] [ PubMed ] [ Google Scholar ] 48. Reinke A, Tizabi MD, Sudre CH, et al.. Common limitations of image processing metrics: a picture story. a rXiv (Preprint) . Available at: doi: 10.48550/arXiv.2104.05642. [ DOI ] 49. Hara C, Maruyama K, Wakabayashi T, et al.. Choroidal vessel and stromal volumetric analysis after photodynamic therapy or focal laser for central serous chorioretinopathy. Transl Vis Sci Technol . 2023; 12(11): 26. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 50. Musial G, Queener HM, Adhikari S, et al.. Automatic segmentation of retinal capillaries in adaptive optics scanning laser ophthalmoscope perfusion images using a convolutional neural network. Transl Vis Sci Technol . 2020; 9(2): 43. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 51. Luo X, Hu M, Song T, Wang G, Zhang S. Semi-supervised medical image segmentation via cross teaching between CNN and Transformer. arXiv (Preprint) . Available at: doi: 10.48550/arXiv.2112.04894. [ DOI ] 52. Zhao Z, Wang Z, Wang L, Yu D, Yuan Y, Zhou L. Alternate diverse teaching for semi-supervised medical image segmentation. a rXiv (Preprint) . Available at: doi: 10.48550/arXiv.2311.17325. [ DOI ] 53. Gao N, Zhou S, Wang L, Zheng N. PMT: progressive mean teacher via exploring temporal consistency for semi-supervised medical image segmentation. arXiv (Preprint) . Available at: http://arxiv.org/abs/2409.05122 . 54. Zhao X, Wang W. Semi-supervised medical image segmentation based on deep consistent collaborative learning. J Imaging . 2024; 10(5): 118. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 55. Rashidisabet H, Sethi A, Jindarak P, et al.. Validating the generalizability of ophthalmic artificial intelligence models on real-world clinical data. Transl Vis Sci Technol . 2023; 12(11): 8. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 56. Kilim O, Olar A, Joó T, Palicz T, Pollner P, Csabai I. Physical imaging parameter variation drives domain shift. Sci Rep . 2022; 12(1): 21302. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 57. Rashidisabet H, Chan RVP, Leiderman YI, Vajaranant TS, Yi D. Robust uncertainty-informed glaucoma classification under data shift. Transl Vis Sci Technol . 2025; 14(6): 3. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 58. Nichyporuk B, Cardinell J, Szeto J, et al.. Rethinking generalization: the impact of annotation style on medical image segmentation. arXiv (Preprint) . Available at: doi: 10.48550/arXiv.2210.17398. [ DOI ] 59. Nasirihaghighi S. Data-efficient learning for generalizable surgical video understanding. arXiv (Preprint) . Available at: https://arxiv.org/html/2508.10215v2 . 60. Galvez-Yanjari V, de la Fuente R, Munoz-Gama J, Sepúlveda M. The sequence of steps: a key concept missing in surgical training-a systematic review and recommendations to include it. Int J Environ Res Public Health . 2023; 20(2): 1436. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 61. Pissas T, Ravasio CS, Da Cruz L, Bergeles C. Effective semantic segmentation in cataract surgery: what matters most? In: Medical Image Computing and Computer Assisted Intervention – MICCAI 2021 . New York, NY: Springer International Publishing; 2021: 509–518. [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials Supplement 1 tvst-15-4-5_s001.docx (4.9MB, docx) Articles from Translational Vision Science & Technology are provided here courtesy of Association for Research in Vision and Ophthalmology ACTIONS View on publisher site PDF (6.0 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top

Record · ID 13649 · SHA-256 93b57ae851c57c1f
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.