ConceptioArchiveNCBI PubMed Central
NCBI PubMed Centralopen access

Research on the construction of an AI diagnostic model for plus disease of retinopathy of prematurity based on cross-center fusion datasets.

Zhang X et al. · ncbi_pmc
NCBI PubMed Central · Papers · License: Open Access
Open Source ↗Direct PDF ↓
legalinformatics
legal informatics

Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Front Pediatr . 2026 Mar 24;14:1765353. doi: 10.3389/fped.2026.1765353 Search in PMC Search in PubMed View in NLM Catalog Add to search Research on the construction of an AI diagnostic model for plus disease of retinopathy of prematurity based on cross-center fusion datasets Xiqianru Zhang Xiqianru Zhang 1 The First School of Clinical Medicine, Lanzhou University, Lanzhou, Gansu, China Writing – original draft Find articles by Xiqianru Zhang 1, † , Huichun Liang Huichun Liang 2 School of Medical Informatics and Engineering, Gansu University of Chinese Medicine, Lanzhou, Gansu, China Software, Writing – review & editing, Visualization Find articles by Huichun Liang 2, † , Ruifeng Wang Ruifeng Wang 3 Key Laboratory of Dunhuang Medical and Transformation, Ministry of Education of the People’s Republic of China, Gansu University of Chinese Medicine, Lanzhou, Gansu, China Data curation, Investigation, Writing – review & editing Find articles by Ruifeng Wang 3 , Rouqing Wu Rouqing Wu 1 The First School of Clinical Medicine, Lanzhou University, Lanzhou, Gansu, China Data curation, Writing – review & editing, Investigation Find articles by Rouqing Wu 1 , Xiao Shen Xiao Shen 4 Department of Ophthalmology, The First Hospital of Lanzhou University, Lanzhou, Gansu, China Resources, Project administration, Writing – review & editing Find articles by Xiao Shen 4, * , Yuemei Zhang Yuemei Zhang 1 The First School of Clinical Medicine, Lanzhou University, Lanzhou, Gansu, China 4 Department of Ophthalmology, The First Hospital of Lanzhou University, Lanzhou, Gansu, China Supervision, Methodology, Writing – review & editing, Project administration, Resources Find articles by Yuemei Zhang 1, 4, * Author information Article notes Copyright and License information 1 The First School of Clinical Medicine, Lanzhou University, Lanzhou, Gansu, China 2 School of Medical Informatics and Engineering, Gansu University of Chinese Medicine, Lanzhou, Gansu, China 3 Key Laboratory of Dunhuang Medical and Transformation, Ministry of Education of the People’s Republic of China, Gansu University of Chinese Medicine, Lanzhou, Gansu, China 4 Department of Ophthalmology, The First Hospital of Lanzhou University, Lanzhou, Gansu, China * Correspondence: Yuemei Zhang [email protected] Xiao Shen [email protected] † These authors have contributed equally to this work and share first authorship Roles Xiqianru Zhang : Writing – original draft Huichun Liang : Software, Writing – review & editing, Visualization Ruifeng Wang : Data curation, Investigation, Writing – review & editing Rouqing Wu : Data curation, Writing – review & editing, Investigation Xiao Shen : Resources, Project administration, Writing – review & editing Yuemei Zhang : Supervision, Methodology, Writing – review & editing, Project administration, Resources Received 2025 Dec 11; Revised 2026 Jan 20; Accepted 2026 Feb 23; Collection date 2026. © 2026 Zhang, Liang, Wang, Wu, Shen and Zhang. This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY) . The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms. PMC Copyright notice PMCID: PMC13055548  PMID: 41953179 Abstract Purpose Retinopathy of prematurity (ROP) is a leading cause of blindness in infants. Early and accurate screening is essential. Current deep learning systems can help, yet their accuracy drops when used on different population. We aimed to find the best deep learning model for plus disease in ROP and to test it on a multi-center dataset. Methods We built a cross-center retinal image database by merging public and private sets (FARFUM-RoP, HVDROPDB, LAN-RoP, Preterm infants <34 weeks GA from three tertiary NICUs, 2635 images). Nine types models were compared: ResNet34, ResNet50, DenseNet, Inception, MobileNet, VGG16, VGG19, EfficientNet, and Swin-Transformer. We compared FLOPs, parameter count, accuracy, recall, precision, and F1 score. ResNet50 showed the best balance and was kept. Ten-fold cross-validation was run on FARFUM-RoP alone, LAN-RoP alone, and their combined set. Results Across the three diagnostic tasks, the ResNet50 algorithm attained area-under-the-curve (AUC) values of 0.97, 0.95 and 1.00 (95% CI 0.94–0.99, 0.91–0.98, 0.97–1.00) for Normal, Pre-plus and Plus disease, respectively. When trained on the consolidated multi-centre cohort, the model achieved optimal overall performance, delivering an accuracy of 92.60%, recall of 92.58%, precision of 92.69% and F1-score of 92.60%—all metrics surpassing those obtained with any single-centre training set. Conclusion Compared with single-centre training, the cross-centre fusion strategy significantly enhanced the generalisability of the artificial-intelligence model, yielded superior diagnostic indices, and improved diagnostic accuracy for infants from diverse demographic backgrounds. Keywords: artificial intelligence (AI), convolutional neural network (CNN), cross-center datasets, retinopathy of prematurity (ROP), ROP plus disease Introduction Retinopathy of prematurity (ROP) mainly affects preterm and low birth weight infants and is a major cause of childhood blindness ( 1 ). As more premature babies survive, the incidence of ROP is rising ( 2 ), placing a heavy burden on families and society. Timely screening and prompt treatment can sharply reduce the risk of severe ROP and blindness ( 3 ). Yet there are too few ophthalmologists trained to read ROP images, especially in resource-limited regions ( 4 ). A fast, scalable tool for diagnosis is urgently needed. Recent advances in artificial intelligence, particularly convolutional neural networks (CNN) for medical imaging, have opened new avenues for automated ROP diagnosis ( 5 – 7 ). Most current models, however, are trained and tested on single, homogeneous populations ( 8 – 10 ). Retinal structure and lesion features differ among ethnic groups ( 11 ), so these models lose accuracy when applied to new racial or geographic settings. This performance degradation reflects limited model generalizability across populations and compromises diagnostic reliability ( 12 – 16 ). Fusing data from multiple centers constitutes an effective strategy to enhance such generalizability ( 17 – 20 ). To address this issue, we aimed to improve the generalizability of ROP diagnosis across center. We systematically compared nine representative deep learning architectures, from classic VGGNet and ResNet to recent EfficientNet and Swin-Transformer. We measured FLOPs, parameter count, accuracy, recall, precision, and F1 score to quantify each model's performance. We then used two center distinct datasets, FARFUM-RoP and LAN-RoP, along with their merged set, and applied ten-fold cross-validation to study how data heterogeneity affects generalization. Our goal was to build an AI model for plus disease in ROP that leverages a cross-center dataset to enhance reliability. This work supports fair and robust ROP screening and offers new directions for AI in ophthalmology. Methods This study had three parts: dataset preparation, model building and tuning, and external validation. Figure 1 shows the full workflow. Figure 1. Open in a new tab Research flowchart. (A) Images from FARFUM-RoP and LAN-RoP were cropped, screened, normalized, augmented, and labeled; (B) ResNet50 was selected among nine backbones and trained on single and fusion dataset; (C) externally validated on HVDROPDB and assessed by junior, senior, and expert ophthalmologists. Dataset preparation We used two public datasets, FARFUM-RoP and HVDROPDB, and one private datasets, LAN-RoP. FARFUM-RoP: This public datases from Farabi and Ferdowsi Universities in Mashhad, Iran, and covers Caucasian infants ( 21 ). It includes 1,533 images from 68 preterm infants (birth weight <2,000 g, gestational age <34 weeks). RetCam captured 2–12 images per infant in five fields: posterior pole, superior, inferior, nasal, and temporal. Mydriasis was achieved with three drops of 0.5% phenylephrine and 0.5% tropicamide given every ten minutes. Five senior pediatric ophthalmologists labeled each image as normal, pre-plus, or plus. Although the original paper did not report inter-rater reliability coefficients, the expert-consensus labels provided by the dataset have been widely adopted as reference standards in multiple ROP-AI studies; therefore, we followed its annotation protocol without additional reassessment. LAN-RoP: This private retrospective set covers infants screened at the The First Hospital of Lanzhou University from January 2019 to March 2025. It includes 1,002 images from 52 preterm infants (birth weight <2,000 g, gestational age <34 weeks) were included. No infant had received bevacizumab or laser before imaging. Mydriasis used compound tropicamide. RetCam wide-field systems obtained 2–15 images per eye in the same five fields. Three ophthalmologists (junior, attending, and senior) labeled every image as normal, pre-plus, or plus following the third edition of the International Classification of Retinopathy of Prematurity ( 22 ). They had joint training and removed poor-quality images. We used the mean across-category accuracy of the three ophthalmologists as a qualitative indicator of inter-rater agreement. The hospital Ethics Committee approved the study. Mixed dataset: We merged FARFUM-RoP and LAN-RoP to create a larger, more diverse training set. HVDROPDB: This public multimodal dataset was collected at H. V. Desai Eye Hospital in India from 2020 to 2023 ( 23 ). It contains 3,000 RetCam and Neo images. We used a subset (100 images) for external validation. All images were preprocessed in the same way: resized, denoised, and contrast-enhanced to reduce equipment differences. Model selection and comparison To find the best model, we ran all experiments on identical hardware and software: GPU:RTX 3,090(24GB) * 1, CPU:14 vCPU Intel(R) Xeon(R) Gold 6,330 CPU @ 2.00 GHz, PyTorch 2.3.0, Python 3.12. We trained on FARFUM-RoP with fixed hyperparameters (epochs, batch size, AdamW optimizer, and initial learning rate).Models included classic CNNs (VGGNet, ResNet, DenseNet, Inception, MobileNet), the efficient network EfficientNet, and the vision transformer Swin-Transformer ( 24 – 28 ). Each model started from ImageNet pretrained weights. We measured parameters, FLOPs, accuracy, recall, precision, F1 score, ROC curve, and AUC ( 29 , 30 ). These metrics were derived from true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN). The performance metrics were calculated as follows ( Equations 1 – 4 ): Accuracy = TP + TN TP + FP + TN + FN (1) Precision = TP TP + FP (2) Recall = TP TP + FN (3) F 1 = 2 * Precision * Recall Precision + Recall (4) ROC curve plots the true positive rate (TPR) against the false positive rate (FPR) at every threshold. The area under this curve (AUC) gives a single number that shows overall model strength. An AUC close to 1 means strong classification. Parameters count tells us how large the model is and how much memory it needs. FLOPs count the floating point operations used in one forward pass. It is a hardware-free measure of model speed. Ten-fold cross-validation Based on the comparison, we chose ResNet50 as the final backbone. Its skip connections ease gradient vanishing and let the network go deeper without adding too many parameters. ResNet50 has shown strong results in many medical image tasks. We ran three separate ten-fold cross-validation experiments to test the model. In each experiment we split the dataset into ten equal parts. Nine parts trained the model and one part served as validation. We rotated the validation part ten times so every image was used for both training and testing once. We report the mean of the ten runs. This approach reduces variance and gives a reliable estimate of model performance ( Table 1 ). Table 1. Ten-fold cross-validation experimental design. Model Ten-fold cross-validation Test datasets FARFUM-RoP Model FARFUM-RoP FARFUM-RoP Test + LAN-RoP Test LAN-RoP Model LAN-RoP LAN-RoP Test + FARFUM-RoP Test FARFUM + LAN RoP Model FARFUM-RoP + LAN-RoP LAN-RoP Test + FARFUM-RoP Test Open in a new tab Ten-fold cross-validation was performed on FARFUM-RoP, LAN-RoP, and their fusion dataset, with model performance evaluated independently on the FARFUM-RoP test set and the LAN-RoP test set. Statistical analysis We used SPSS 29.0. Results are shown as mean ± standard deviation. We compared multiple groups with one-way ANOVA. For pairwise comparisons we used the LSD test. A P value below 0.05 was considered significant. This study conforms to the TRIPOD-AI 2024 statement for reporting machine-learning-based diagnostic models ( 31 ). Results Model interpretability To clarify how the model makes decisions, we created heat maps ( Figure 2 ). Each map marks the image regions that drew the model's attention. Warmer colors show higher activation, moving from purple to red. The maps reveal that the network focuses on the same signs clinicians use: dilated and tortuous vessels in the posterior pole. The heat-maps partially overlap with the posterior pole vessel regions, providing some qualitative support for the model's clinical relevance, though quantitative validation remains for future work. Figure 2. Open in a new tab Heat map of the fundus image. The model highlights regions of interest in purple-to-red; deeper red indicates higher activation. The network concentrates on dilated, tortuous vessels in the posterior pole, aligning with clinical signs and confirming interpretability. Model selection and comparison Figure 3 ; Table 2 show the test set results. VGGNet reaches reasonable accuracy but uses far more parameters and FLOPs. MobileNetV3Large and EfficientNetB3 are light yet give slightly lower accuracy. ResNet50 offers the best balance. It keeps high accuracy while keeping both FLOPs and parameter count moderate. We therefore chose ResNet50 as the backbone for the next steps. Figure 3. Open in a new tab Comparison of FLOPs and accuracy rate of different deep learning models. The x -axis shows FLOPs (G), the y -axis shows test accuracy (%). ResNet50 achieves the highest accuracy at ≈4 G FLOPs, outperforming VGG variants and other lightweight networks while maintaining efficiency. Table 2. Comparison of parameters of different deep learning models. Model Accuracy (%) Precision (%) Recall (%) F1-Score (%) Total Parameters FLOPs ResNet34 89.61 89.58 89.61 89.59 21,286,211.00 3.68 GFLOPs ResNet50 93.51 93.49 93.51 93.41 23,514,179.00 4.13 GFLOPs DenseNet121 85.71 85.98 85.71 85.83 6,956,931.00 2.90 GFLOPs VGG16 91.56 91.71 91.56 91.61 134,272,835.00 15.47 GFLOPs VGG19 94.16 94.23 94.16 94.18 139,582,531.00 19.63 GFLOPs InceptionV3 77.92 78.58 77.92 78.15 21,791,715.00 2.85 GFLOPs EfficientNetB3 85.06 85.15 85.06 84.35 10,700,843.00 1.02 GFLOPs SwinTransformer 90.26 90.32 90.26 90.29 58,716,291.00 10.22 GFLOPs MobileNetV3Large 83.77 84.17 83.77 83.93 4,205,875.00 0.23 GFLOPs Open in a new tab VGG19 achieved the highest accuracy, precision, recall and F1 score, but required the most parameters and FLOPs. ResNet50 delivered 93.51% accuracy with only 4.13 GFLOPs, offering the best trade-off and was selected as the optimal model. Ten-fold cross-validation results We kept ResNet50 as the backbone and trained three models: FARFUM-RoP Model, LAN-RoP Model, and FARFUM + LAN-RoP Model. Each model went through ten-fold cross-validation on its own or the cross-dataset test set. Figure 4 ; Table 3 give the mean scores. The FARFUM + LAN-RoP Model reached the highest accuracy, recall, and precision and showed the smallest spread across folds. The single-set models scored lower and varied more, suggesting they learnt dataset-specific patterns and lost generalizability. The ROC curve of the merged model ran closest to the upper left corner and its AUC was near 1. Figure 4. Open in a new tab Comparison of ROC curves of three AI models on different test sets. Three models were trained via 10-fold cross-validation on different datasets and evaluated on the FARFUM-RoP and LAN-RoP test sets. The FARFUM + LAN-RoP fusion model produced superior ROC curves and higher AUCs across all test sets, whereas single-dataset models showed marked performance drops on cross-dataset testing, confirming that fusion training mitigates racial bias and enhances generalization. Table 3. Shows the average results of ten-fold cross-validation on different test sets. Model FARFUM-RoP Test LAN-RoP Test Accuracy(%) F1 score(%) Precision (%) Recall (%) Accuracy(%) F1 score(%) Precision (%) Recall (%) FARFUM-RoP Model 92.08 ± 1.50 92.04 ± 1.53 92.14 ± 1.52 92.08 ± 1.50 36.63 ± 2.80 35.05 ± 3.29 46.73 ± 4.80 36.63 ± 2.80 LAN-RoP Model 42.08 ± 4.51 40.14 ± 5.85 49.69 ± 5.01 42.08 ± 4.51 93.66 ± 1.89 93.65 ± 1.86 93.81 ± 1.85 93.66 ± 1.89 FARFUM + LAN RoP Model 92.60 ± 1.65 92.58 ± 1.64 92.69 ± 1.62 92.60 ± 1.65 94.55 ± 2.39 94.55 ± 2.39 94.59 ± 2.39 94.55 ± 2.39 Open in a new tab Single-dataset models excelled on their own test set but dropped sharply when tested across datasets; the FARFUM + LAN-RoP fusion model sustained high accuracy, F1-score, and stability on both test sets, confirming that fusion training mitigates ethnic bias and enhances generalizability. AI vs. ophthalmologists We compared the best model (FARFUM + LAN RoP Model) with three experienced ophthalmologists on the same test set. Average accuracies for the doctors were 86% for Normal, 72% for Preplus, and 86% for Plus. The AI model reached 95%, 86%, and 95%. Thus, AI model's accuracy surpassed the doctors, especially junior ones, and approached expert level. The model also processed images faster, making large-scale ROP screening feasible ( Figure 5 ). Figure 5. Open in a new tab Comparison between AI models and ophthalmologists. The AI achieved 95%, 86%, 95% for Normal, Pre-plus, Plus, outperforming the average clinician 86%, 72%, 86%, and approaching expert performance. *** P < 0.001. External validation We tested the FARFUM + LAN RoP Model on a Plus-disease subset from the public HVDROPDB dataset. Figure 6 ; Table 4 show that the model separated Normal and Plus cases well but was weaker on Pre-Plus. The overall AUC on this external set was 0.83. Future work will refine the algorithm to improve Pre-Plus detection, and a broader multi-centre cohort will be pursued in future work to further assess generalisability. Figure 6. Open in a new tab External validation results (matrix and ROC curve). The FARFUM + LAN-RoP model achieved an overall AUC of 0.83 on the HVDROPDB subset, with strong discrimination for Normal (92.8%) and Plus (100%) but weaker performance for pre-plus (56.0%). Table 4. Shows the results of external validation dataset. Class Precision (%) Recall (%) F1 score (%) Normal 96.15 89.29 92.59 Pre-Plus 14.29 8.33 10.53 Plus 26.67 100.00 42.11 Weighted Average 83.55 80.00 80.73 Overall Accuracy 80.00 Open in a new tab Discussion We built an AI model for plus disease in ROP using a cross-center dataset and showed that this approach enhances the generalizability of the AI model. The combined dataset outperformed single-center sets, proving that diverse training data helps the model learn universal retinal features instead of population-specific ones. This finding offers a new tool for early ROP detection and has clear clinical value. But Limitations remain, Performance still differed slightly across center, perhaps because retinal structure and lesion patterns vary among groups. The sample size, especially for Han infants, is modest. Although we used data augmentation and transfer learning, overfitting may still occur ( 32 ) and robustness needs improvement ( 33 , 34 ). Future work will enlarge the dataset, with more Han infant images, to boost generalizability. We will test newer architectures and extend tasks to lesion segmentation and quantitative grading of vessel tortuosity and dilation. Multi-center studies will confirm utility in varied clinical settings. We will also explore ways to embed the model in daily workflow to raise efficiency and reduce resource use. In summary, our cross-center AI model improves accuracy and the generalizability. Despite its limits, it lays a solid foundation for future research and clinical use. Funding Statement The author(s) declared that financial support was received for this work and/or its publication. Natural Science Foundation of Gansu Province (Grant No. 24JRRA321). Footnotes Edited by: Phil Fischer , Mayo Clinic, United States Reviewed by: Anantha Krishnan , Amity University Haryana, India Firas Zaier , Institut National de la Santé, Tunisia Data availability statement The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation. Ethics statement The studies involving humans were approved by Ethics Committee of LZU No. l Hospital (approval No. LDYYLL-2025-960). The studies were conducted in accordance with the local legislation and institutional requirements. Written informed consent for participation was not required from the participants or the participants' legal guardians/next of kin in accordance with the national legislation and institutional requirements. Author contributions XZ: Writing – original draft. HL: Software, Writing – review & editing, Visualization. RW: Data curation, Investigation, Writing – review & editing. RW: Data curation, Writing – review & editing, Investigation. XS: Resources, Project administration, Writing – review & editing. YZ: Supervision, Methodology, Writing – review & editing, Project administration, Resources. Conflict of interest The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. Generative AI statement The author(s) declared that generative AI was not used in the creation of this manuscript. Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us. Publisher's note All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher. Supplementary material The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fped.2026.1765353/full#supplementary-material Datasheet1.pdf (874.9KB, pdf) Datasheet2.pdf (128.1KB, pdf) References 1. Coyner AS, Oh MA, Shah PK, Singh P, Ostmo S, Valikodath NG, et al. External validation of a retinopathy of prematurity screening model using artificial intelligence in 3 low- and middle-income populations. JAMA Ophthalmol. (2022) 140(8):791–8. 10.1001/jamaophthalmol.2022.2135 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 2. Quinn GE, Ying GS, Bell EF, Donohue PK, Morrison D, Tomlinson LA, et al. Incidence and early course of retinopathy of prematurity: secondary analysis of the postnatal growth and retinopathy of prematurity (G-ROP) study. JAMA Ophthalmol. (2018) 136(12):1383–9. 10.1001/jamaophthalmol.2018.4290 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 3. Ekinci DY, Ugurlu A, Tasli NG. What is the incidence of retinopathy of prematurity (ROP). In “big” babies?: results of a retrospective multicenter study. Ophthalmic Epidemiol. (2021) 28(2):138–43. 10.1080/09286586.2020.1793372 [ DOI ] [ PubMed ] [ Google Scholar ] 4. Park JH, Hwang JH, Chang YS, Lee MH, Park WS. Survival rate dependent variations in retinopathy of prematurity treatment rates in very low birth weight infants. Sci Rep. (2020) 10(1):19401. 10.1038/s41598-020-76472-w [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 5. Brown JM, Campbell JP, Beers A, Chang K, Ostmo S, Chan RVP, et al. Automated diagnosis of plus disease in retinopathy of prematurity using deep convolutional neural networks. JAMA Ophthalmol. (2018) 136(7):803–10. 10.1001/jamaophthalmol.2018.1934 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 6. Mao J, Luo Y, Liu L, Lao J, Shao Y, Zhang M, et al. Automated diagnosis and quantitative analysis of plus disease in retinopathy of prematurity based on deep convolutional neural networks. Acta Ophthalmol (Copenh). (2020) 98(3):e339–e45. 10.1111/aos.14264 [ DOI ] [ PubMed ] [ Google Scholar ] 7. Maji D, Sekh AA. Automatic grading of retinal blood vessel in deep retinal image diagnosis. J Med Syst. (2020) 44(10):180. 10.1007/s10916-020-01635-1 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 8. Tan Z, Simkin S, Lai C, Dai S. Deep learning algorithm for automated diagnosis of retinopathy of prematurity plus disease. Transl Vis Sci Technol. (2019) 8(6):23. 10.1167/tvst.8.6.23 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 9. Agrawal R, Kulkarni S, Walambe R, Kotecha K. Assistive framework for automatic detection of all the zones in retinopathy of prematurity using deep learning. J Digit Imaging. (2021) 34(4):932–47. 10.1007/s10278-021-00477-8 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 10. Gupta K, Campbell JP, Taylor S, Brown JM, Ostmo S, Chan RVP, et al. A quantitative severity scale for retinopathy of prematurity using deep learning to monitor disease regression after treatment. JAMA Ophthalmol. (2019) 137(9):1029–36. 10.1001/jamaophthalmol.2019.2442 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 11. Jin J, Friess A, Hendricks D, Lehman S, Salvin J, Reid JE, et al. Effect of gestational age at birth, sex, and race on foveal structure in children. Graefe’s Arch Clin Exp Ophthalmol Albrecht von Graefes Archiv fur Klinische und Experimentelle Ophthalmologie. (2021) 259(10):3137–48. 10.1007/s00417-021-05191-3 [ DOI ] [ PubMed ] [ Google Scholar ] 12. Puyol-Antón E, Ruijsink B, Harana M, Piechnik J, Neubauer SK, Petersen S, et al. Fairness in cardiac magnetic resonance imaging: assessing sex and racial bias in deep learning-based segmentation. Front Cardiovasc Med. (2022) 9:859310. 10.3389/fcvm.2022.859310 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 13. Trentz C, Engelbart J, Semprini J, Kahl A, Anyimadu E, Buatti J, et al. Evaluating machine learning model bias and racial disparities in non-small cell lung cancer using SEER registry data. Health Care Manag Sci. (2024) 27(4):631–49. 10.1007/s10729-024-09691-6 [ DOI ] [ PubMed ] [ Google Scholar ] 14. Ganta T, Kia A, Parchure P, Wang MH, Besculides M, Mazumdar M, et al. Fairness in predicting cancer mortality across racial subgroups. JAMA Netw Open. (2024) 7(7):e2421290. 10.1001/jamanetworkopen.2024.21290 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 15. Chen F, Wang L, Hong J, Jiang J, Zhou L. Unmasking bias in artificial intelligence: a systematic review of bias detection and mitigation strategies in electronic health record-based models. J Am Med Inform Assoc JAMIA. (2024) 31(5):1172–83. 10.1093/jamia/ocae060 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 16. Sasseville M, Ouellet S, Rhéaume C, Sahlia M, Couture V, Després P, et al. Bias mitigation in primary health care artificial intelligence models: scoping review. J Med Internet Res. (2025) 27:e60269. 10.2196/60269 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 17. Yu Y, Cai G, Lin R, Wang Z, Chen Y, Tan Y, et al. Multimodal data fusion AI model uncovers tumor microenvironment immunotyping heterogeneity and enhanced risk stratification of breast cancer. MedComm. (2024) 5(12):e70023. 10.1002/mco2.70023 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 18. Zhang T, Ding R, Luong KD, Hsu W. Evaluating an information theoretic approach for selecting multimodal data fusion methods. J Biomed Inform. (2025) 167:104833. 10.1016/j.jbi.2025.104833 [ DOI ] [ PubMed ] [ Google Scholar ] 19. Preto AJ, Chanana S, Ence D, Healy MD, Domingo-Fernández D, West KA. Multi-omics data integration identifies novel biomarkers and patient subgroups in inflammatory bowel disease. J Crohn’s Colitis. (2025) 19(1):jjae197. 10.1093/ecco-jcc/jjae197 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 20. Khan S, Akram BA, Zafar A, Wasim M, Khurshid KS, Pires IM. Locustlens: leveraging environmental data fusion and machine learning for desert locust swarm prediction. PeerJ Comp Sci. (2024) 10:e2420. 10.7717/peerj-cs.2420 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 21. Akbari M, Pourreza HR, Khalili Pour E, Dastjani Farahani A, Bazvand F, Ebrahimiadib N, et al. FARFUM-RoP, a dataset for computer-aided detection of retinopathy of prematurity. Sci Data. (2024) 11(1):1176. 10.1038/s41597-024-03897-7 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 22. Chiang MF, Quinn GE, Fielder AR, Ostmo SR, Paul Chan RV, Berrocal A, et al. International classification of retinopathy of prematurity, third edition. Ophthalmology. (2021) 128(10): e51–68. 10.1016/j.ophtha.2021.05.031 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 23. Agrawal R, Walambe R, Kotecha K, Gaikwad A, Deshpande CM, Kulkarni S. HVDROPDB datasets for research in retinopathy of prematurity. Data Brief. (2024) 52:109839. 10.1016/j.dib.2023.109839 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 24. Aljohani A, Aburasain RY. A hybrid framework for glaucoma detection through federated machine learning and deep learning models. BMC Med Inform Decis Mak. (2024) 24(1):115. 10.1186/s12911-024-02518-y [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 25. Chuntranapaporn S, Choontanom R, Srimanan W. Ocular duction measurement using three convolutional neural network models: a comparative study. Cureus. (2024) 16(11):e73985. 10.7759/cureus.73985 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 26. Singh S, Banoub R, Sanghvi HA, Agarwal A, Chalam KV, Gupta S, et al. An artificial intelligence driven approach for classification of ophthalmic images using convolutional neural network: an experimental study. Curr Med Imaging. (2024) 20:e15734056286918. 10.2174/0115734056286918240419100058 [ DOI ] [ PubMed ] [ Google Scholar ] 27. Li Z, Yang J, Wang X, Zhou S. Establishment and evaluation of intelligent diagnostic model for ophthalmic ultrasound images based on deep learning. Ultrasound Med Biol. (2023) 49(8):1760–7. 10.1016/j.ultrasmedbio.2023.03.022 [ DOI ] [ PubMed ] [ Google Scholar ] 28. Goh JHL, Ang E, Srinivasan S, Lei X, Loh J, Quek TC, et al. Comparative analysis of vision transformers and conventional convolutional neural networks in detecting referable diabetic retinopathy. Ophthalmol Sci. (2024) 4(6):100552. 10.1016/j.xops.2024.100552 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 29. Kohn MA, Newman TB. The walking man approach to interpreting the receiver operating characteristic curve and area under the receiver operating characteristic curve. J Clin Epidemiol. (2023) 162:182–6. 10.1016/j.jclinepi.2023.07.020 [ DOI ] [ PubMed ] [ Google Scholar ] 30. Ozenne B, Subtil F, Maucort-Boulch D. The precision–recall curve overcame the optimism of the receiver operating characteristic curve in rare diseases. J Clin Epidemiol. (2015) 68(8):855–9. 10.1016/j.jclinepi.2015.02.010 [ DOI ] [ PubMed ] [ Google Scholar ] 31. Collins GS, Moons KGM, Dhiman P, Riley RD, Beam AL, Van Calster B, et al. TRIPOD + AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. Br Med J. (2024) 385:e078378. 10.1136/bmj-2023-078378 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 32. Agrawal V, Jagtap J, Patil S, Kotecha K. Performance analysis of hybrid deep learning framework using a vision transformer and convolutional neural network for handwritten digit recognition. MethodsX. (2024) 12:102554. 10.1016/j.mex.2024.102554 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 33. Shaban M, Ogur Z, Mahmoud A, Switala A, Shalaby A, Abu Khalifeh H, et al. A convolutional neural network for the screening and staging of diabetic retinopathy. PLoS One. (2020) 15(6):e0233514. 10.1371/journal.pone.0233514 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 34. Thomas A, Harikrishnan PM, Ramachandran R, Ramachandran S, Manoj R, Palanisamy P, et al. A novel multiscale and multipath convolutional neural network based age-related macular degeneration detection using OCT images. Comput Methods Programs Biomed. (2021) 209:106294. 10.1016/j.cmpb.2021.106294 [ DOI ] [ PubMed ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials Datasheet1.pdf (874.9KB, pdf) Datasheet2.pdf (128.1KB, pdf) Data Availability Statement The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation. Articles from Frontiers in Pediatrics are provided here courtesy of Frontiers Media SA ACTIONS View on publisher site PDF (1.7 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top

Record · ID 978 · SHA-256 a4912256bd23afe8
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.