ConceptioArchiveNCBI PubMed Central
NCBI PubMed Centralopen access

Convolutional Neural Networks in Chronic Wound Segmentation and Tissue Classification Using Real-World Images.

Huttunen E et al. · ncbi_pmc
NCBI PubMed Central · Papers · License: Open Access
Open Source ↗Direct PDF ↓
machine learning systems

Convolutional Neural Networks in Chronic Wound Segmentation and Tissue Classification Using Real‐World Images - PMC Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Int Wound J . 2026 Apr 16;23(4):e70912. doi: 10.1111/iwj.70912 Search in PMC Search in PubMed View in NLM Catalog Add to search Convolutional Neural Networks in Chronic Wound Segmentation and Tissue Classification Using Real‐World Images Ellen Huttunen Ellen Huttunen 1 Faculty of Medicine and Health Technology, Tampere University, Tampere, Finland Find articles by Ellen Huttunen 1 , Teija Kimpimäki Teija Kimpimäki 1 Faculty of Medicine and Health Technology, Tampere University, Tampere, Finland 2 Department of Dermatology, Tampere University Hospital, Wellbeing Services County of Pirkanmaa, Tampere, Finland Find articles by Teija Kimpimäki 1, 2 , Jenni E Salenius Jenni E Salenius 1 Faculty of Medicine and Health Technology, Tampere University, Tampere, Finland Find articles by Jenni E Salenius 1 , Ilkka Pölönen Ilkka Pölönen 3 Faculty of Information Technology, University of Jyväskylä, Jyväskylä, Finland Find articles by Ilkka Pölönen 3 , Thomas Yambasu Thomas Yambasu 4 Department of Dermatology, Wellbeing Services County of Central Finland, Hospital Nova of Central Finland, Jyväskylä, Finland Find articles by Thomas Yambasu 4 , Maria Huttunen Maria Huttunen 4 Department of Dermatology, Wellbeing Services County of Central Finland, Hospital Nova of Central Finland, Jyväskylä, Finland Find articles by Maria Huttunen 4 , Teea Salmi Teea Salmi 1 Faculty of Medicine and Health Technology, Tampere University, Tampere, Finland 2 Department of Dermatology, Tampere University Hospital, Wellbeing Services County of Pirkanmaa, Tampere, Finland Find articles by Teea Salmi 1, 2, ✉ Author information Article notes Copyright and License information 1 Faculty of Medicine and Health Technology, Tampere University, Tampere, Finland 2 Department of Dermatology, Tampere University Hospital, Wellbeing Services County of Pirkanmaa, Tampere, Finland 3 Faculty of Information Technology, University of Jyväskylä, Jyväskylä, Finland 4 Department of Dermatology, Wellbeing Services County of Central Finland, Hospital Nova of Central Finland, Jyväskylä, Finland * Correspondence: Teea Salmi ( [email protected] ) ✉ Corresponding author. Revised 2026 Mar 31; Received 2025 Sep 3; Accepted 2026 Apr 2; Collection date 2026 Apr. © 2026 The Author(s). International Wound Journal published by Medicalhelplines.com Inc and John Wiley & Sons Ltd. This is an open access article under the terms of the http://creativecommons.org/licenses/by-nc-nd/4.0/ License, which permits use and distribution in any medium, provided the original work is properly cited, the use is non‐commercial and no modifications or adaptations are made. PMC Copyright notice PMCID: PMC13084538  PMID: 41988844 ABSTRACT Chronic wounds cause a significant burden to affected patients and to society. Effective and objective diagnostic and monitoring methods are needed in wound care, and artificial intelligence offers one promising alternative. In this study, real‐world wound images were used to train a convolutional neural network to automatically segment wound area and wound tissues on an image. The study included altogether 362 images of venous, arterial, vasculitis and pyoderma gangrenosum wounds. The model was based on a convolutional neural network architecture U‐Net, and fully supervised learning was utilised during the training phase. Wound area reached a Dice Similarity Coefficient (DSC) of 0.927 and Intersection over Union (IoU) of 0.868 using an augmented dataset with pretraining. Fibrinous exudate and granulation performed fairly well with DSC 0.750 and 0.696, and with IoU 0.659 and 0.601, respectively. Necrosis present in only 56 images achieved lower performance with DSC 0.503 and IoU 0.502. In conclusion, this study suggested that it is possible to train a neural network to perform well with images taken for purely clinical purposes. Besides wound area, several wound structures can be identified, but wound structure identification performance is dependent on the number of images featuring the structure. Keywords: artificial intelligence, computer, leg ulcer, neural networks, supervised machine learning, wound healing Summary Clinical images taken in real‐world conditions can be utilised with CNN to provide wound area segmentation and tissue classification. The combination of augmented data, optimised input size and appropriate epoch count achieved the best segmentation performance among the models tested. Wound area and the most common structures in chronic wounds, like fibrinous exudation and granulation, can be detected with CNN quite reliably. Traditional augmentation methods did not improve the CNN performance in wound segmentation as expected; the best way to boost performance was by adding more images. 1. Introduction Chronic wounds pose a major health concern globally, affecting up to 3% of the population [ 1 ] and causing significant direct and indirect health‐care costs. For example, almost a decade ago, venous leg ulcer treatment was estimated to cost some £7600 annually per patient in the UK [ 2 ], and the cost has likely risen since then. From the patient perspective, chronic wounds are associated with multiple comorbidities [ 3 ], poorer quality of life [ 4 , 5 ] and even increased mortality [ 6 ]. Identification of wound aetiology is of the utmost importance in the aspect of effective treatment. Furthermore, reliable monitoring methods are vital during treatment to evaluate the response to the care. For both diagnostics and follow‐up of wounds, wound area and wound structure assessment are crucial. However, in clinical work wound evaluation is currently rather subjective and susceptible to error as during wound treatment professionals often change and each may have different expertise, assessment criteria or interpretation of the same wound structure. Further, wound area is generally measured with a ruler and by determining the length and perpendicular width, which is quite a rough estimate of wound surface area and may not always reflect the evolution of the wound. Wound images are quite routinely taken during wound care but are seldom used for other than inspecting wound development by simple visual evaluation, which is prone to misinterpretations. More thorough analysis of all gathered information is warranted to improve the treatment outcome of chronic wounds. It would be optimal to develop a new, user‐independent, and ideally also less time‐consuming method for measuring the wound area and identifying wound structures. Such a method would have significant clinical value by reducing subjectivity in wound assessment and decreasing the reliance on specialised expertise. Artificial intelligence (AI) has gained significant recognition in medicine, with dermatology being particularly well‐suited to AI innovations due to its reliance on visual assessment. In dermatology, AI has most often been used in the classification of skin lesions to neoplastic and non‐neoplastic, but other applications have also been studied [ 7 ]. In wound care, most AI studies have focused on wound area measurement and the identification of specific features, such as infection or ischaemia [ 8 , 9 ]. While some studies have attempted to identify wound structures in images [ 10 , 11 , 12 ], many have relied on image bank data rather than clinical photographs. Further, most chronic wound studies concerning AI have centred on pressure ulcers and diabetic foot ulcers, with only a few studies researching vascular and atypical wounds [ 8 , 13 , 14 ]. Neural networks are advanced machine learning applications that mimic the brain's neuronal patterns. They enhance workflow efficiency and optimise computing power. Machine learning could potentially offer an entirely new approach to observing and documenting wound healing. In principle, wound photographs contain analysable data that could be classified and processed to assess and potentially predict healing more reliably than conventional clinical evaluation, as deep‐learning neural networks may be capable of identifying image features relevant to chronic wound healing. In wound care, neural networks could thus possibly be exploited in wound diagnostics and monitoring, and in the long run might even facilitate improved quality of care. However, when developing a tool intended for clinical use, it is essential that the material employed in testing reflects the material encountered in actual clinical practice. In this study, a convolutional neural network (CNN) was chosen for its ability to identify abstract objects in images, such as wounds. The primary aim of this study was to develop a model for wound area segmentation and the secondary aim for tissue classification using real‐world images of vascular as well as vasculitis and pyoderma gangrenosum (PG) wounds, which are one of the most common atypical wounds [ 15 ]. 2. Materials and Methods 2.1. Study Material and Protocol The images of chronic wounds used in this study were obtained from the electronic patient record systems of Hospital Nova, Central Finland, Jyväskylä, during 2020 and 2021. Hospital Nova is a consulting hospital that mostly treats patients with complex and hard‐to‐heal wounds. The images were collected from patients treated at the Departments of Dermatology or Surgery between 2018 and 2020, that is, the study period, with International Classification of Diseases version 10 (ICD‐10) codes consistent with arterial (I70.2), venous (I83.0), vasculitis (L95.8) or PG wounds (L88). Only patients having clinical images, most often taken by experienced nurses working with wound patients daily, during clinic visits for wound documentation purposes were included in the initial search. None of the images were originally intended to be used for research; they were all taken solely for clinical purposes and hence imaging conditions, such as lightning, were not fully standardised. All images and patient records were manually reviewed, and only wounds located between knee and ankle were included, except for PG ulcers, where wounds from all skin areas were included. PG wound images were close‐ups and anatomical landmarks were not present in order to avoid confusion. Wounds previously treated with surgery, such as skin grafting, were excluded. All images were taken before wound debridement. Image imperfections, such as rulers, were allowed. Image quality was evaluated separately for each image by assessing the exposure, focus, and resolution of the image, and only images of adequate quality for annotation process were included. Images had been taken during routine visits by multiple nurses using different cameras, and as we had no access to the metadata of the images due to the image archiving system erasing the data before saving, we could not confirm which cameras were used. If the patient had multiple simultaneous or sequential wounds that met the inclusion criteria during the study period, all wounds and images of these were included in the study, as were follow‐up images of the same wounds. The data collected consisted of 62 patients, of whom 41 were female. Patients' median age was 71.5 (range 32–96) years. These patients had 72 wounds included in the study and altogether 362 wound images. Most of the wounds were venous (53%), followed by vasculitis (18%), arterial (15%) and PG wounds (8%). Some wounds (6%) were considered to have a mixed aetiology of both venous and arterial origin. None of the images had been previously used in any AI study, and the data set was completely novel for AI training. Wound area was annotated first, and different wound structures were annotated in this area. Wound structures identified during the annotation process were fibrinous exudation, granulation, necrosis, bone/tendon (combined class) and unhealthy granulation. These structures were chosen, as they are the most common structures in chronic wounds. The annotation of the images was primarily done by a medical student (E.H.) and was initially checked by a resident dermatologist (J.E.S.), both thoroughly trained to identify wound structures. The annotation was ultimately checked by one to three dermatologists with special competence in wound care and each having > 20 years of clinical experience in wound care (T.S., T.K., M.H.), and annotation was corrected if needed. The study protocol was approved by the Medical Director of the Central Finland Central Hospital. Approval from the ethics committee was not required since the study was a registry study. The study followed the guidelines of the Declaration of Helsinki as revised in 2024. 2.2. Technical Part The CNN architecture employed in this study was U‐Net, which has yielded promising results in earlier AI studies in dermatology [ 16 , 17 ]. U‐Net was first published in 2015 [ 18 ], and the version used in this study was a variation with ResNet34‐encoder pretrained with ImageNet‐database. In this study, we used supervised learning. Training was performed using a single NVIDIA RTX3060 GPU with a data loader amount set to two. Model performance was evaluated after every epoch and model status was saved after the three best performing epochs. Full details of the software environment and hardware configuration are provided in the Technical Appendix A1 . Most of the images were large (average dimension 5152 × 3456 pixels). Only 18 images were medium sized (1024 × 1024 pixels or smaller). After annotation, original high‐resolution images were resized to model input dimensions. We experimented with 256 × 256 and 512 × 512 pixels. Images were resized using bilinear interpolation; ground truth masks used nearest‐neighbour interpolation to preserve label integrity. The orientation of the images varied: 27 images had greater height than width, five were squares, and 329 had greater width than height. The images were randomised into a training set ( n = 289) and a validation set ( n = 73). The size of the dataset was limited, so it was decided not to include a separate test set while acknowledging the impact this might have on the results. Data augmentation was used to enhance model performance, enabling improved functionality under diverse conditions and enhancing the model's generalisation capabilities. Data augmentation was implemented using imgaug‐library. Augmentation techniques used in this study included salt‐and‐pepper noise, colour space conversion, Gaussian blur, Gamma contrast and JPEG compression [ 19 ]. Each technique was activated with a 20% probability in random images. Because of the class imbalances, the images containing necrosis were augmented by splitting images in half and combining the halves randomly to achieve more images for training. The augmentation strategies increased training set sizes substantially; detailed image counts for each class before and after augmentation are provided in the Technical Appendix (Table A5 ). For statistical analysis, Dice Similarity Coefficient (DSC) and Intersection over Union (IoU), along with specificity and sensitivity with 95% confidence intervals, were used. DSC is a statistical test used for comparing the AI outcome against the original, human‐made mask. IoU measures the intersection between AI outcome and the original mask (ground truth), 1.00 being a perfect match. 3. Results Wound area was present in all 362 images, fibrinous exudate in 274, granulation in 191 and necrosis in 56 images. Hypergranulation (five images), bone/tendon (two images) and unhealthy granulation (seven images) were annotated but ultimately excluded from the analysis because of the small number of images in which these structures were presented. In wound area segmentation, the CNN was able to detect the wound area very well in most images (Figure 1A ), but in some segmented the wound border slightly more roughly than the human‐annotated ground truth (Figure 1B–D ). Since no masks were annotated to the images nor were the images cropped to show only the wound area, the CNN analysed the whole image looking for wound structures, and in a few cases, the neural network mistakenly identified wound structures outside the wound and even skin (Figure 1E ). FIGURE 1. Open in a new tab Selected examples (A–E) of original wound images and their corresponding ground truth of wound area and output segmentation using Convolutional Neural Networks. Of the parameters investigated, wound area achieved the highest sensitivity (0.814) and specificity (0.989) (Table 1 ). All wound structures included in the final analysis achieved high specificity, fibrinous exudate having the lowest with 0.972. The results varied more in sensitivity, where granulation achieved the lowest sensitivity of 0.566. Due to the data augmentation performed, it was not possible to calculate the sensitivity of the necrosis present in only 56 images in the original dataset. To improve the model's performance on this class, extensive data augmentation was applied by combining and altering images to generate additional samples. Although ground truth masks were available for the augmented dataset, the small size and homogeneity of the dataset posed significant challenges in calculating sensitivity, where enough positive samples are required to accurately measure the proportion of true positives out of all positives. Due to the limited number of distinct necrosis examples, sensitivity measurements were unstable and less reliable; sensitivity could not be counted even after augmentation. In contrast, specificity focusing on true negatives, which were abundant and well‐represented in both the original and augmented datasets, could still be calculated accurately. TABLE 1. Result metrics for investigated parameters. Wound structure Architecture Epoch count Sensitivity (95% CI) Specificity (95% CI) DSC IoU Wound area Albunet mod. ResNet34 30 0.814 ± 0.0234 0.989 ± 0.000238 0.927 0.868 Fibrinous exudate U‐net vgg11 15 0.773 ± 0.0321 0.972 ± 0.0159 0.750 0.659 Granulation Albunet mod. ResNet34 30 0.566 ± 0.0348 0.988 ± 0.000216 0.696 0.601 Necrosis U‐net vgg11 20 x 1.000 ± 0.000141 0.503 0.502 Open in a new tab Note: x = sensitivity could not be counted. Abbreviations: CI, confidence interval; DSC, dice similarity coefficient; IoU, intersect over union. In further analysis, wound area had the best performance of classes analysed based on DSC and IoU values (0.927 and IoU of 0.868 respectively), and necrosis the lowest performance (DSC 0.503 and IoU 0.502) (Table 1 ). Due to necrosis being an underrepresented class, the images were heavily augmented during the training of the CNN to increase the amount of data available by combining images featuring necrosis. After this augmentation DSC increased to 0.799 and IoU to 0.722. Several models with different input sizes and epoch counts were used to analyse wound area (Table 2 ). Non‐augmented performance is reported as baseline comparison demonstrating the substantial improvement achieved through augmentation. Given the limited dataset size, augmentation was essential for clinically useful segmentation performance. While LinkNet345 on ResNet34 with 512 × 512 input and 35 epochs was the top performer on the non‐augmented dataset, the AlbuNet model, modified with ResNet34 and trained with a 256 × 256 input size for 30 epochs, demonstrated the best performance overall, achieving a DSC of 0.927 and an IoU of 0.868. This result was consistent with the overall pattern across the evaluated models, indicating that the combination of augmented data, optimised input size and appropriate epoch count achieved the best segmentation performance among the models tested. The coherence across models confirms the robustness of the augmentation strategy and the reliability of the performance metrics, reducing the likelihood of isolated or model‐specific effects. TABLE 2. Comparison of wound area segmentation performance using different models. Model Input size Epoch count Mean DSC Mean IoU U‐net vgg11 on ImageNet NA 256 × 256 20 0.834 0.747 Unet vgg16 on ImageNet NA 512 × 512 35 0.673 0.594 LinkNet345 on Resnet34 on ImageNet NA 512 × 512 35 0.877 0.799 AlbuNet mod. Resnet34 on ImageNet, augmented dataset 256 × 256 30 0.927 0.868 Open in a new tab Abbreviations: DSC, dice similarity coefficient; IoU, intersect over union; NA, dataset not preprocessed. 4. Discussion Our data implies that wound images taken purely for clinical purposes can be successfully used to train CNN to identify wound area and different wound structures. If further studies comparing CNN ability to clinicians provide proof that CNN is superior in analysing wound images, digital wound assessment and precise classification of wound structures could potentially significantly enhance clinical accuracy in diagnosis and treatment, and reduce inaccuracies caused by professionals' differing subjective interpretations. Earlier AI studies related to wounds have mostly investigated wound area segmentation, but a few have also investigated other wound structures, particularly granulation, necrosis and fibrinous exudation (slough) [ 20 , 21 , 22 , 23 , 24 , 25 ]. In some studies, the total number of images is not specified [ 25 , 26 ], making it difficult to evaluate and compare the methods and results with ours. There are also multiple studies in which the datasets consist partly or wholly of the same images [ 9 , 12 , 25 ]. Most of the studies where the model segmented tissues utilised support vector machines. Only a few previous studies have utilised CNN models in automatic tissue segmentation [ 10 , 11 , 12 , 27 ], but in this study a CNN based approach was chosen due to its ability to lead to better performance with larger datasets [ 18 ]. A CNN‐based model presented in a study published in 2018 reached promising results [ 10 ], but the dataset consisted of only 22 images featuring stage III‐IV pressure ulcers. A larger dataset with images of 193 pressure ulcers was used in another study [ 12 ] and achieved a better performance on necrosis (DSC 0.966) than that was presented in the present study. These authors, however, did not specify how many images featured necrosis in their dataset and hence the difference between model results compared to our study could be explained by the amount of data. One study with images of venous wounds [ 27 ] did not report DSC, but the overall sensitivity (0.733) and specificity (0.946) had good performance on the 5‐image test set. Their training strategy differed greatly from the one used in this study, as they divided images into subimages containing one class only, whereas the model presented in this study used the whole image throughout the training phase. A more recent study [ 11 ] proposed a framework utilising a generative adversarial network to expand the dataset, and achieved the highest DSC of 0.744 with EfficientNet and precision of 0.760 with YOLOv5. While the model is CNN based, the architecture differs greatly from those of other studies. Further, in one study comprising venous and PG wounds CNN performance was compared to dermatologists, and CNN achieved higher sensitivity and dermatologists higher specificity (0.970 vs. 0.727, 0.889 vs. 0.833, respectfully) [ 14 ]. Finally, based on available evidence at the time, the use of AI in wound assessment has been discussed in a systematic review published in 2022 [ 8 ]. The authors highlight the potential of AI‐driven tools to enhance diagnostic accuracy, streamline clinical workflows and improve patient care, but note existing challenges such as limited datasets and variability in system performance to be limiting factors in implementing AI to wound care. In this study, the performance of the neural network was clearly linked to the amount of image material available for each structure. The best results were achieved with the tissues that were most often present in image data, such as wound area and fibrinous exudate. Necrosis present in 56 images remained borderline, and we could not analyse bone/tendon, hypergranulation or unhealthy granulation due to the small number of images in which they occurred. Traditional augmentation methods, such as salt‐and‐pepper noise and Gaussian blur, did not enhance the performance of the neural network as expected. In this study, the most effective way to improve the performance of the neural network was to obtain more images featuring new wounds or augmenting the original images to obtain more images. In general, datasets in clinical healthcare overall tend to be limited in size compared to the massive datasets of image banks. However, using real‐world data with images taken by different professionals and with varying equipment is closer to the authentic clinical circumstances and is essential when evaluating the generalisability of the model in clinical use. In patient care, wound images are often taken and available but are primarily used for rough and subjective visual assessment of wound structures, and hence currently tend to be underutilised. In our study we utilised data that has never been used in any AI studies. Compared to earlier studies of similar nature, our dataset is medium‐sized. Comparison of augmented and non‐augmented datasets (Table 2 ) showed that the augmented dataset achieved better results than the non‐augmented dataset due to the increased amount of data after preprocessing. In preprocessing it was decided to use fully supervised learning in our study. This leads to better results but is also quite labour‐intensive and requires segmented ground truth. Ground truth, in the context of AI, refers to the accurate and reliable data or information that serves as the standard against which an AI model's performance is measured and evaluated with [ 28 ]. In our study, ground truth is an interpretation made by a healthcare professional of a structure in an image AI had to replicate. Training the models with manually annotated, pixel‐level ground‐truth masks (i.e., fully supervised learning) proved effective for achieving accurate wound segmentation in our dataset. The CNN architectures used in this study were selected based on their strong performance in medical image segmentation tasks and their complementary feature‐learning properties. U‐Net was employed for wound boundary segmentation due to its encoder–decoder structure with skip connections that preserve fine spatial details [ 29 ]. LinkNet/AlbuNet‐34 models, incorporating a ResNet‐34 backbone, were used for tissue classification because their residual connections enable efficient deep feature extraction while maintaining computational efficiency [ 30 ]. In Table 2 , several architectures appear together because they were evaluated independently on the same task—not used as an ensemble unless explicitly stated. Their inclusion allows comparison across model families with different representational capacities. Pretrained ImageNet weights were used for the encoder components to stabilise training on a limited dataset and to leverage generic low‐level features. Our model was trained by focusing on one class at a time to avoid confusion between classes. This may happen because of an unbalanced dataset or visual similarity between classes, both of which were a concern in our study. A significant challenge during training was the imbalance in our dataset, which predominantly featured granulation and fibrinous exudate, reflecting their real‐world prevalence. Such an imbalance can lead to overfitting. To mitigate these issues, it was decided to exclude certain annotated wound structures from the analysis. With an expanded dataset in the future, it may become feasible to analyse additional structures such as hypergranulation, bone‐tendon interfaces, and unhealthy granulation. A notable strength in our study was that we utilised images obtained during routine patient visits, and that CNN performance was investigated in patients with vascular and atypical wounds, where AI performance has previously scarcely been explored. Focusing on one class was an effective strategy in this study, but it must be noted that the general applicability of the model to other classes has not been evaluated and requires further studies. Further, it should also be noted that our dataset contained images of only fair‐skinned patients and CNN performance on coloured skin may vary considerably from the results reported here. We also concede that as our dataset is from a single wound centre treating hard‐to‐heal wounds, it may not fully reflect the entirety of the chronic wound patient population. We also note that the performance of our model was validated using the validation set and no external validation set was used. Performance on external data may differ from the results reported here. The study does not include a dedicated measure of annotation reliability such as Cohen's kappa or ICC. Although annotations underwent a multistep review process, initially created by a trained annotator and subsequently checked and corrected by dermatologists with expertise in wound care, this sequential consensus approach did not provide independent parallel annotations needed to compute formal agreement statistics. This represents a limitation, and future work will incorporate structured multiannotator labelling and formal reliability assessment to fully characterise annotation quality. Randomisation was performed image‐wise rather than patient‐wise. While this maximises data utilisation, it may introduce dependencies between training and validation sets when multiple images from the same patient are present, potentially leading to optimistic performance estimates. Patient‐wise splitting or external validation would provide a more rigorous assessment of generalisability. A common challenge in medical imaging research is the limited availability of large annotated datasets. Medical datasets are often small due to difficulties in data collection, patient privacy considerations, and the substantial time and expertise required for manual annotation, particularly when pixel‐level ground‐truth masks are needed [ 31 ]. These constraints may affect the robustness and generalisability of machine learning models. Several strategies can help mitigate these limitations. Data augmentation techniques can artificially increase the diversity of training samples, while cross‐validation provides a more reliable estimate of model performance in small datasets [ 31 , 32 , 33 ]. Model overfitting can also be reduced by using architectures with fewer parameters, applying regularisation techniques, and limiting model complexity [ 31 , 32 , 33 ]. Another important approach is the establishment of multicenter collaborations, which can increase dataset size and distribute the annotation workload across institutions. In addition, weak supervision, pseudo‐label generation or foundation‐model‐based approaches may help reduce the burden of manual annotation [ 34 ]. Nevertheless, when the number of samples in a class becomes extremely small, meaningful training and evaluation may not be possible. It should be noted that there is no single universal threshold for determining when a dataset is too small, as this depends on several factors, including model complexity, class balance, data variability, and the intended analysis. This was the case for the necrosis class in our dataset, where the number of available samples was insufficient for reliable model development or interpretation. In such situations, results should be interpreted with caution, and further data collection is required before drawing firm conclusions. Given the limited dataset size, k‐fold cross‐validation would provide a more comprehensive estimate of model variability by allowing each sample to serve as both training and validation data across different folds. This approach could help assess performance stability and reduce the risk that results are influenced by a particular train/validation split, especially since the current study used image‐wise randomisation, which may introduce some dependency between sets when multiple images originate from the same patient. Although k‐fold cross‐validation was not implemented in the present study due to computational constraints and the need for a consistent evaluation pipeline across models, it represents an important direction for methodological strengthening in future work. Additionally, external validation using an independent dataset, preferably from a different clinical site or imaging environment, would offer a more rigorous assessment of model generalisability and improve clinical relevance. Although our dataset includes images captured by several nurses using different digital cameras, which introduces natural variability in lighting, perspective and acquisition conditions that likely improves model robustness, the reported Dice and IoU values may still overestimate performance in entirely new clinical environments. Because validation was conducted only on images originating from the same institution, there remains a risk that the metrics partially reflect site‐specific characteristics rather than full real‐world generalisability. In conclusion, our study suggests that a reliable CNN can be trained using real‐world images also featuring wound aetiologies so far only rarely used in AI wound studies. Real‐world data use in CNN training is important to gain an understanding of model performance with images that are not standardised and include the variability that is to be expected in clinical work. In the future it is possible to extend the capability of the model to include other clinically significant classes, but that would require a larger and more balanced dataset. The results of this study are of value given the need for new methods for wound assessment that are less subjective and less dependent on professional expertise. Reliable tissue identification creates the technological base to classify wounds based on their aetiology, and wound area recognition and measurement for improved tracking of wound development. Therefore, automatic image‐based analysis has great future potential in enhancing precision in the diagnosis and follow‐up of chronic wounds. However, while leveraging existing image data can support less experienced clinicians and AI may aid decision‐making under uncertainty, its limitations must be understood and it should never be used as the sole basis for diagnosis. Funding This study was financially supported by the Competitive State Research Financing of the Expert Responsibility Area of Tampere University Hospital (9AA082 and 9AC033) and the Tampere University Hospital Support Foundation (MK326). Ethics Statement The study protocol was approved by the Medical Director of the Central Finland Central Hospital. Approval from the ethics committee was not required since the study was a registry study. The study followed the guidelines of the Declaration of Helsinki as revised in 2024. Conflicts of Interest The authors declare no conflicts of interest. Supporting information Appendix A1. Table A1 . Comparison of segmentation architectures. Table A2 . Training hyperparameters. Table A3 . Models evaluated for wound area segmentation. Table A4 . Models used for tissue classification. Table A5 . Training set image counts after augmentation. IWJ-23-e70912-s001.docx (27.5KB, docx) Acknowledgements Open access publishing facilitated by Tampereen yliopisto ja Tampereen ammattikorkeakoulu, as part of the Wiley ‐ FinELib agreement. Data Availability Statement The data that support the findings of this study are available from the corresponding author upon reasonable request. References 1. Graham I. D., Harrison M. B., Nelson E. A., Lorimer K., and Fisher A., “Prevalence of Lower‐Limb Ulceration: A Systematic Review of Prevalence Studies,” Advances in Skin & Wound Care 16, no. 6 (2003): 305–316. [ DOI ] [ PubMed ] [ Google Scholar ] 2. Guest J. F., Fuller G. W., and Vowden P., “Venous Leg Ulcer Management in Clinical Practice in the UK: Costs and Outcomes,” International Wound Journal 15, no. 1 (2018): 29–37. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 3. Kuikko K., Salmi T., Huhtala H., and Kimpimäki T., “Characteristics of Chronic Ulcer Patients by Gender and Ulcer Aetiology From a Multidisciplinary Wound Centre,” International Wound Journal 21, no. 8 (2024): e70012. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 4. Persoon A., Heinen M. M., van der Vleuten C. J. M., de Rooij M. J., van de Kerkhof P. C. M., and van Achterberg T., “Leg Ulcers: A Review of Their Impact on Daily Life,” Journal of Clinical Nursing 13, no. 3 (2004): 341–354. [ DOI ] [ PubMed ] [ Google Scholar ] 5. Kimpimäki T., Karhu M., Vaalasti A., et al., “Health‐Related Quality of Life of Patients With Hard‐To‐Heal Ulcers Measured With the 15D Instrument: A Prospective Study,” Journal of Wound Care 34, no. 5 (2025): 350–358. [ DOI ] [ PubMed ] [ Google Scholar ] 6. Salenius J., Suntila M., Ahti T., et al., “Long‐Term Mortality Among Patients With Chronic Ulcers,” Acta Dermato‐Venereologica 101, no. 5 (2021): adv00455. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 7. Thomsen K., Iversen L., Titlestad T. L., and Winther O., “Systematic Review of Machine Learning for Diagnosis and Prognosis in Dermatology,” Journal of Dermatological Treatment 31, no. 5 (2020): 496–510. [ DOI ] [ PubMed ] [ Google Scholar ] 8. Anisuzzaman D. M., Wang C., Rostami B., Gopalakrishnan S., Niezgoda J., and Yu Z., “Image‐Based Artificial Intelligence in Wound Assessment: A Systematic Review,” Advances in Wound Care 11, no. 12 (2022): 687–709. [ DOI ] [ PubMed ] [ Google Scholar ] 9. Yap M. H., Hachiuma R., Alavi A., et al., “Deep Learning in Diabetic Foot Ulcers Detection: A Comprehensive Evaluation,” Computers in Biology and Medicine 135 (2021): 104596. [ DOI ] [ PubMed ] [ Google Scholar ] 10. Zahia S., Sierra‐Sosa D., Garcia‐Zapirain B., and Elmaghraby A., “Tissue Classification and Segmentation of Pressure Injuries Using Convolutional Neural Networks,” Computer Methods and Programs in Biomedicine 159 (2018): 51–58. [ DOI ] [ PubMed ] [ Google Scholar ] 11. Sarp S., Kuzlu M., Pipattanasomporn M., and Guler O., “Simultaneous Wound Border Segmentation and Tissue Classification Using a Conditional Generative Adversarial Network,” Journal of Engineering 2021, no. 3 (2021): 125–134. [ Google Scholar ] 12. García‐Zapirain B., Elmogy M., El‐Baz A., and Elmaghraby A. S., “Classification of Pressure Ulcer Tissues With 3D Convolutional Neural Network,” Medical & Biological Engineering & Computing 56, no. 12 (2018): 2245–2258. [ DOI ] [ PubMed ] [ Google Scholar ] 13. Lucas Y., Niri R., Treuillet S., Douzi H., and Castaneda B., “Wound Size Imaging: Ready for Smart Assessment and Monitoring,” Advances in Wound Care 10, no. 11 (2021): 641–661. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 14. Birkner M., Schalk J., Von Den Driesch P., and Schultz E. S., “Computer‐Assisted Differential Diagnosis of Pyoderma Gangrenosum and Venous Ulcers With Deep Neural Networks,” Journal of Clinical Medicine 11, no. 23 (2022): 7103. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 15. Isoherranen K., O'Brien J. J., Barker J., et al., “Atypical Wounds. Best Clinical Practice and Challenges,” Journal of Wound Care 28, no. 6 (2019): S1–S92. [ DOI ] [ PubMed ] [ Google Scholar ] 16. Arshad S., Amjad T., Hussain A., Qureshi I., and Abbas Q., “Dermo‐Seg: ResNet‐UNet Architecture and Hybrid Loss Function for Detection of Differential Patterns to Diagnose Pigmented Skin Lesions,” Diagnostics (Basel, Switzerland) 13, no. 18 (2023): 2924. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 17. Shah S. M. A. H., Rizwan A., Atteia G., and Alabdulhafith M., “CADFU for Dermatologists: A Novel Chronic Wounds & Ulcers Diagnosis System With DHuNeT (Dual‐Phase Hyperactive UNet) and YOLOv8 Algorithm,” Healthcare (Basel) 11, no. 21 (2023): 2840. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 18. Ronneberger O., Fischer P., and Brox T., “U‐Net: Convolutional Networks for Biomedical Image Segmentation,” in Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Springer Verlag, 2015), 234–241. [ Google Scholar ] 19. Shorten C. and Khoshgoftaar T. M., “A Survey on Image Data Augmentation for Deep Learning,” Journal of Big Data 6, no. 1 (2019): 60. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 20. Wang P., Fan E., and Wang P., “Comparative Analysis of Image Classification Algorithms Based on Traditional Machine Learning and Deep Learning,” Pattern Recognition Letters 141 (2021): 61–67. [ Google Scholar ] 21. Wang C., Anisuzzaman D. M., Williamson V., et al., “Fully Automatic Wound Segmentation With Deep Convolutional Neural Networks,” Scientific Reports 10, no. 1 (2020): 21897. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 22. Malihi L., Hüsers J., Richter M. L., et al., “Automatic Wound Type Classification With Convolutional Neural Networks,” Studies in Health Technology and Informatics 295 (2022): 281–284. [ DOI ] [ PubMed ] [ Google Scholar ] 23. Chino D. Y. T., Scabora L. C., Cazzolato M. T., Jorge A. E. S., Traina‐Jr C., and Traina A. J. M., “Segmenting Skin Ulcers and Measuring the Wound Area Using Deep Convolutional Networks,” Computer Methods and Programs in Biomedicine 191 (2020): 105376. [ DOI ] [ PubMed ] [ Google Scholar ] 24. Hüsers J., Hafer G., Heggemann J., et al., “Automatic Classification of Diabetic Foot Ulcer Images—A Transfer‐Learning Approach to Detect Wound Maceration,” Studies in Health Technology and Informatics 289 (2022): 301–304. [ DOI ] [ PubMed ] [ Google Scholar ] 25. Li F., Wang C., Liu X., Peng Y., and Jin S., “A Composite Model of Wound Segmentation Based on Traditional Methods and Deep Neural Networks,” Computational Intelligence and Neuroscience 2018 (2018): 4149103. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 26. Wang L., Pedersen P. C., Agu E., Strong D. M., and Tulu B., “Area Determination of Diabetic Foot Ulcer Images Using a Cascaded Two‐Stage SVM‐Based Classification,” IEEE Transactions on Biomedical Engineering 64, no. 9 (2017): 2098–2109. [ DOI ] [ PubMed ] [ Google Scholar ] 27. Rajathi V., Bhavani R. R., and Wiselin Jiji G., “Varicose Ulcer(C6) Wound Image Tissue Classification Using Multidimensional Convolutional Neural Networks,” Imaging Science Journal 67, no. 7 (2019): 374–384. [ Google Scholar ] 28. Geiger R. S., Cope D., Ip J., et al., ““Garbage in, Garbage Out” Revisited: What Do Machine Learning Application Papers Report About Human‐Labeled Training Data?,” Quantitative Science Studies 2, no. 3 (2021): 795–827. [ Google Scholar ] 29. Alabdulhafith M., Ba Mahel A. S., Samee N. A., et al., “Automated Wound Care by Employing a Reliable U‐Net Architecture Combined With ResNet Feature Encoders for Monitoring Chronic Wounds,” Frontiers in Medicine 11 (2024): 1310137. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 30. Sulaiman A., Anand V., Gupta S., et al., “An Intelligent LinkNet‐34 Model With EfficientNetB7 Encoder for Semantic Segmentation of Brain Tumor,” Scientific Reports 14, no. 1 (2024): 1345. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 31. Piffer S., Ubaldi L., Tangaro S., Retico A., and Talamonti C., “Tackling the Small Data Problem in Medical Image Classification With Artificial Intelligence: A Systematic Review,” Progress in Biomedical Engineering 6 (2024): 32001. [ DOI ] [ PubMed ] [ Google Scholar ] 32. Gröger F., Amruthalingam L., Lionetti S., Navarini A. A., Ille F., and Pouly M., “A Review and Systematic Guide to Counteracting Medical Data Scarcity for AI Applications,” Computer Methods and Programs in Biomedicine Update 8 (2025): 100220. [ Google Scholar ] 33. Salehi A. W., Khan S., Gupta G., et al., “A Study of CNN and Transfer Learning in Medical Imaging: Advantages, Challenges, Future Scope,” Sustainability 15 (2023): 5930. [ Google Scholar ] 34. Zedda L., Loddo A., and Di Ruberto C., “A Foundation Model for Advanced Radiomics and AI ‐Driven Medical Imaging Analysis,” Computers in Biology and Medicine 195 (2025): 110583. [ DOI ] [ PubMed ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials Appendix A1. Table A1 . Comparison of segmentation architectures. Table A2 . Training hyperparameters. Table A3 . Models evaluated for wound area segmentation. Table A4 . Models used for tissue classification. Table A5 . Training set image counts after augmentation. IWJ-23-e70912-s001.docx (27.5KB, docx) Data Availability Statement The data that support the findings of this study are available from the corresponding author upon reasonable request. Articles from International Wound Journal are provided here courtesy of Wiley ACTIONS View on publisher site PDF (663.1 KB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top

Record · ID 28242 · SHA-256 c6cdb0e53b48b4a5
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.