ConceptioArchivearXiv CS
arXiv CSopen access

Predicting Male Fertility Using Machine Learning: A Semen Parameters Based Analysis with the VISEM Dataset

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

Predicting Male Fertility Using Machine Learning: A Semen Parameters Based Analysis with the VISEM Dataset Shahnawaz Qureshi1 , Raja Khurram Shahzad2*, Muhammad Fozan1 , Emal Kawal1 , Syed Aziz Shah3 , Sattam Al-Anazi4 , Syed Muhammad Zeeshan Iqbal5

arXiv:2607.08429v1 [cs.LG] 9 Jul 2026

1

School of Computing, Pak-Austria Fachhochschule: Institute of Applied Sciences and Technology, Haripur, Khyber Pakhtunkhwa, Pakistan. 2 Department of Communication, Quality Management and Information Systems, Mid Sweden University, Östersund Campus, Sweden. 3 Healthcare Sensing Technology, Center for Intelligent Healthcare, Coventry University, United Kingdom. 4 E-Serivces Department, Saudi Standards, Metrology and Quality Organization, Riyadh, Saudi Arabia. 5 Research and Development, Brightware LLC, Riyadh, Saudi Arabia.

*Corresponding author(s). E-mail(s): [email protected]; Contributing authors: [email protected]; [email protected]; [email protected]; [email protected]; [email protected]; [email protected]; Abstract Male infertility is a significant yet often underdiagnosed aspect of reproductive health, with semen analysis serving as the cornerstone of clinical evaluation. To address this problem, this study investigates the use of machine learning algorithms to classify male fertility status based on key semen parameters, i.e., sperm concentration, motility, and morphology, using the VISEM dataset. This dataset includes semen samples from 85 participants, classified into three categories, i.e., Fertile, Sub-Fertile, and Infertile, according to the World Health Organization’s criteria. After pre-processing and feature engineering, the dataset was used to train and assess multiple classification models using the LazyPredict

1

framework. Among the more than 40 algorithms tested, the Nearest Centroid classifier achieved an accuracy of 94.2%, outperforming other models such as Support Vector Machines and Quadratic Discriminant Analysis. The model’s robustness was validated using 5-fold cross-validation and multiclass ROC-AUC analysis. This study illustrates that machine learning models can provide fast, accurate, and objective assessments of semen quality, potentially supporting clinical decision-making in andrology and assisted reproductive technologies. These findings emphasize the growing potential of machine learning to enhance fertility diagnostics and inform patient-specific treatment strategies. Keywords: Male infertility, Semen analysis, Sperm motility, Sperm morphology, Machine learning, Fertility classification, VISEM dataset, Artificial intelligence in reproduction, Reproductive health, Clinical decision support, Reproductive diagnostics

1 Introduction Infertility is defined as the inability to conceive after 12 months of unprotected intercourse and affects approximately 17.5% of the global adult population. Male factors are responsible for nearly half of all infertility cases, occurring either independently (up to 20%) or in combination with female infertility (up to 40%) Vander Borght and Wyns (2018). Semen analysis is the primary clinical tool for assessing male reproductive health. It is used to evaluate critical parameters, such as sperm concentration, motility, and morphology, according to guidelines established by the World Health Organization (WHO) Kumar and Singh (2015), as illustrated in the Figure 1 . However, despite its clinical significance, traditional semen evaluation is performed manually and is subject to observer bias, inter-laboratory variability, and inconsistent interpretation. These factors may result in inaccurate or delayed diagnoses Leslie et al. (2024). As fertility evaluations become increasingly complex, artificial intelligence (AI) and machine learning (ML) are emerging as transformative tools in reproductive medicine. These technologies enable objective analysis of diverse datasets and have demonstrated strong predictive capabilities for classifying male fertility potential. While previous studies have focused on predicting motility or performing image-based analyses, there has been a limited systematic evaluation of the effectiveness of ML models in categorizing fertility status into clinically relevant groups using standardized semen attributes Ottl et al. (2022); Nguyen et al. (2023); Kobayashi et al. (2024); Tiab et al. (2023). This study addresses that gap by applying supervised ML models to the publicly available VISEM dataset1 , which contains semen analysis records from 85 male participants. The primary aim is to classify these samples into three fertility categories, i.e., Fertile, Sub-Fertile, and Infertile, based on WHO thresholds Organisation (1999). Using the LazyPredict framework2 , we evaluated over 40 ML algorithms to identify the most accurate model, assessing performance through cross-validation3 and Receiver

1 2 3

https://datasets.simula.no/visem/ https://pypi.org/project/lazypredict https://en.wikipedia.org/wiki/Cross-validation (statistics)

2

Operating Characteristic - Area Under Curve (ROC-AUC) analysis4 . Our findings suggest that ML techniques, particularly the Nearest Centroid classifier5 , can significantly enhance the accuracy and consistency of fertility assessments. This research has important implications for andrology and assisted reproductive technologies (ART), providing a robust decision-support tool to assist clinicians and couples in fertility planning. The remainder of this paper is organized as follows: Section 2 provides an overview of human fertilization and semen evaluation. Section 3 reviews previous AI/ML-based approaches. Section 4 presents the dataset and methods used in the study. Section 5 presents the experimental results. Section 6 discusses the findings, and Section 7 concludes the paper with conclusions and future directions.

Fig. 1: This figure presents essential semen parameters for successful fertilization according to WHO recommendations. Key thresholds include normal sperm morphology above 4%, sperm concentration above 39 million per ejaculation, and total motility above 40%. It distinguishes between progressive motility (forward movement) and nonprogressive motility (erratic motion, zig-zag motion). The diagram also highlights the main components of sperm, i.e., head, midpiece, and tail, emphasizing their roles in reaching and penetrating the oocyte during fertilization.

2 Background Male reproductive health is influenced by a range of physiological, genetic, and environmental factors, as illustrated in the Figure 1. The human spermatozoon, or sperm, is a specialized motile cell that must travel an average distance of approximately 19 cm in the female reproductive tract to fertilize the oocyte Dcunha et al. (2022). Sperm are generally categorized into progressively motile, non-progressive, and immotile types, with progressive motility considered the strongest predictor of male fertility Brown 4 5

https://en.wikipedia.org/wiki/Receiver operating characteristic https://en.wikipedia.org/wiki/Nearest centroid classifier

3

(1944); Organization (2021). If the total motility is below 40% or if progressive motility is under 32%, both rates are strongly associated with reduced fertility, a condition known as asthenozoospermia Cavarocchi et al. (2025). Cooper et al. (2010); Nayak et al. (2019); McLachlan (2013); Kao et al. (2004). In addition to motility, sperm morphology plays a key role in fertility potential. Abnormal morphology may be linked to genetic defects, such as aneuploidy or mitochondrial dysfunction, and structural deformities in the head or tail Darand et al. (2023); Carlsen et al. (1992). According to WHO criteria, a morphology threshold of 4% normal forms is used to define fertile samples Organisation (1999); Zhou et al. (2021). Sperm concentration is another vital parameter, with a normal range of 15 million sperm per mL or 39 million per ejaculate Bonde et al. (1998). A strong correlation exists between sperm concentration and the likelihood of conception; however, both partners’ ages influence this relationship, particularly the female partner’s, due to the declining quality of oocytes after age 35 Bonde et al. (1998). Advanced fertility treatments, such as in vitro fertilization (IVF), require an evaluation of reproductive parameters for both partners. The success of IVF depends on accurate evaluation of these parameters, yet manual interpretations can introduce variability between clinics Grow et al. (1994); Leslie et al. (2024). These challenges highlight the need for consistent, data-driven tools to support clinical decision-making in reproductive health.

3 Related Work Challenges in Semen Analysis Male infertility remains a significant challenge in reproductive medicine, particularly due to the subjective nature of traditional semen analysis. These evaluations rely on manual microscopy and expert interpretation by andrologists, making the process labor-intensive and vulnerable to inter-observer variability Gbagbo et al. (2024). This subjectivity leads to inconsistent diagnoses, especially across clinics with varying expertise and standardization. To address these limitations, artificial intelligence and machine learning have emerged as promising tools for enhancing diagnostic accuracy and consistency. In particular, the VISEM dataset, which contains annotated videos, images, and semen parameters, has become a benchmark resource for developing intelligent models in reproductive diagnostics.

Machine Learning for Fertility Prediction Most ML-based studies using the VISEM dataset have focused on predicting sperm motility as a continuous variable. For example, Nguyen et al. Nguyen et al. (2023) developed a deep learning model for estimating motility, achieving a Mean Absolute Error (MAE) of 9.322. They emphasized the importance of customized loss functions for clinical relevance. Similarly, Ottl et al. Ottl et al. (2022) introduced motilitAI, a hybrid framework combining support vector regression (SVR) and neural networks, achieving an MAE of 7.31 using engineered movement features. Particularly, Adinugroho et al. Adinugroho and Nakazawa (2023) introduced MotionFlow, a model that combines motion encoding and morphological features. This 4

model achieved an MAE of 6.842% for motility and 4.148% for morphology. However, the focus of these efforts has primarily been on regressing individual parameters rather than on the comprehensive classification of fertility status. Specifically, there is a lack of frameworks that categorize individuals as Fertile, Sub-Fertile, or Infertile based on WHO clinical thresholds. This indicates a gap in the literature regarding interpretable, multi-parameter fertility prediction frameworks.

Advances in Detection and Tracking Accurate detection and tracking are fundamental for Computer-Assisted Sperm Analysis (CASA) systems. Choi et al. Choi et al. (2022) demonstrated that CASA systems rely heavily on robust segmentation algorithms. Thambawita et al. Thambawita et al. (2023) used the VISEM-Tracking Robust dataset, which consists of over 29,000 annotated frames, to facilitate the training of real-time detection systems using YOLOv56 . Similarly, Dobrovolny et al. Dobrovolny et al. (2023) applied YOLOv5 to the VISEM dataset, achieving a mean average precision of 72.15%. Valiuškaitė et al. Valiuškaitė et al. (2023) developed a Rregion-based Convolutional Neural Network (R-CNN) that achieved a detection accuracy of 91.77%, showing a strong correlation with sperm vitality (r = 0.969). Saadat et al. Saadat et al. (2023) further demonstrated that UNet++ with ResNet34 performs well in segmenting sperm heads, although challenges remain in distinguishing sperm from adjacent non-sperm cells.

Feature Engineering and Model Architectures Several studies have highlighted the importance of engineered features and advanced model architectures for improving sperm classification. Ottl et al. Ottl et al. (2021) compared the performance of Support Vector Regression, Convolutional Neural Networks (CNN), and Recurrent Neural Networks (RNN) for tracking sperm motion using feature aggregation techniques such as Crocker-Grier vectorsCrocker and Grier (1996) and bag-of-words methods. More recently, Tiab et al. Tiab et al. (2023) evaluated MobileNet, YOLOv5s, and DeepSort across the VISEM and another dataset, achieving a classification accuracy of 99.2% and a Multi-Object Tracking Precision of 99%. This research confirms the effectiveness of transfer learning approaches and the benefits of dataset diversity.

Gaps in Literature and Our Contribution While the use of artificial intelligence (AI) and machine learning in andrology is increasing, many existing studies often focus on predicting isolated parameters such as motility or morphology, rather than clinically meaningful fertility categorization. Moreover, most of these models optimize for predictive accuracy without accounting for interpretability or integration with clinical decision-making workflows. These limitations restrict their real-world applicability, particularly in assisted reproductive technologies (ART) and fertility counseling. In this study, we address these existing gaps by implementing and comparing over 40 supervised ML classifiers using the LazyPredict framework Musthyala et al. 6

https://github.com/ultralytics/yolov5

5

(2024) on the VISEM dataset. We classify samples into Fertile, Sub-Fertile, and Infertile groups based on WHO thresholds for concentration, morphology, and progressive motility. Our evaluation includes metrics such as precision, recall, F1-score, and AUC to identify robust, interpretable models that can support clinical decision-making. This work contributes to the literature by shifting the focus from isolated parameter predictions to a comprehensive classification of fertility-relevant parameters for clinical settings.

Fig. 2: Workflow for Fertility Classification Using Semen Parameters and Machine Learning. This flowchart presents a method for classifying male fertility status using semen parameters. The LazyPredict framework evaluated and ranked 40 machine learning models based on cross-validated accuracy on the VISEM dataset.

4 Methodology This study used the VISEM dataset, comprising 85 semen samples, to develop machine learning models for classifying fertility based on parameters defined by the World Health Organization: sperm concentration, morphology, and progressive motility. After preprocessing the data and labeling the samples, we benchmarked 40 classifiers using LazyPredict. The top-performing models were Nearest Centroid, Support Vector Machine (SVM), and Quadratic Discriminant Analysis (QDA), which achieved the highest accuracies. To evaluate model performance, we employed 5fold cross-validation and performed ROC-AUC analysis. Moreover, correlation studies revealed strong interdependencies among the seminal parameters that are important for predicting fertility. The overall workflow is illustrated in Figure 2. 6

Dataset Description The VISEM dataset Haugen et al. (2019) was developed by the Simula Research Laboratory in Norway. The dataset was created to facilitate research on the relationship between semen quality and male fertility. It contains semen analysis data from 85 healthy male participants aged 18 or older. All samples were evaluated by trained embryologists using a 40x magnification microscope with a heated stage maintained at 37°C to simulate physiological conditions. The dataset comprises various quantitative semen parameters, including sperm concentration, total sperm count, sperm motility, sperm morphology, ejaculate volume, and sperm viability Haugen et al. (2019); Thambawita et al. (2021).

Data Preprocessing and Feature Engineering As shown in Figure 2, preprocessing was performed to remove unnecessary spaces and missing values, ensuring consistency across all samples. A unique identifier (ID) was assigned to each entry to maintain data integrity. The following key features were extracted, normalized, and engineered for training purposes:

• ID: A unique identifier assigned to each semen sample. • Sperm concentration: The number of sperm per milliliter, which reflects testicular function. • Total sperm count: The total number of sperm in the ejaculate, calculated by multiplying sperm concentration by ejaculate volume. • Normal spermatozoa: The percentage of morphologically normal sperm, determined using the Teratozoospermia Index (TZI): TZI =

h+m+t ab

where h = head defects, m = mid-piece defects, t = tail defects, and ab = total abnormal sperm Carlsen et al. (1992). • Progressive motility: The percentage of sperm that are moving actively in a straight line or in large circles Organization (2021). • Non-progressive motility: The percentage of sperm exhibiting erratic or lowenergy movement. • Immotile sperm: The percentage of sperm showing no movement, which indicates low vitality.

Fertility Classification Criteria To categorize male fertility potential, we used classification thresholds based on the WHO’s 6th edition manual for semen evaluation Organization (2021). The dataset was labeled into three fertility classes, i.e, Fertile, Sub-Fertile, and Infertile, based on progressive motility (≥ 40%, 20–39%, and < 20%), morphology (≥ 4%, 2–3%, and < 2%), and sperm concentration (≥ 15, 10–14, and < 10 million/mL) Cooper et al. (2010), as illustrated in the table 1. Non-progressive motility was excluded, as studies suggest it contributes minimally to live birth outcomes Zhou et al. (2021). 7

Table 1: Semen Characteristics and Fertility Classifications for Model Development. Condition

Progressive Motility (%)

Fertile or Fast Sub-Fertile or Average

≥ 40% > 20% and < 39% < 20%

Infertile or Slow

Normal Morphology (%) ≥ 4% 2%-3%

Sperm Concentration (million/mL)

2%

< 10 million/mL

≥ 15 million/mL 10-14 million/mL

Machine Learning Models’ Evaluation The performance of the models was evaluated using accuracy, F1-score, and balanced accuracy, all calculated through 5-fold cross-validation Raschka (2018). To test for statistical significance, we performed a paired t-test with a significance threshold of p < 0.05. To address sample dependence across the folds, a correction factor was utilized Dietterich (1998). Models that achieved over 90% accuracy demonstrated a significant improvement compared to the ZeroR baseline, which is a naı̈ve classifier that predicts only the majority class.

LazyPredict Benchmarking For our machine learning baseline, we used the LazyPredict Musthyala et al. (2024) package to evaluate the performance of several well-known machine learning algorithms on the semen analysis data. LazyPredict automates the training and testing of over 40 models and ranks them according to performance metrics such as accuracy, F1-score, and balanced accuracy. The models that achieved the highest performance in this experiment are Nearest Centroid, Support Vector Machine, Quadratic Discriminant Analysis, Gaussian Naive Bayes, and Random Forest. LazyPredict streamlines the modeling process by automatically feeding data into various models without the need for hyperparameter tuning. It offers performance metrics for easy comparison. Model performance is evaluated using 5-fold cross-validation, and models are ranked according to their average cross-validation accuracy. The framework uses a simple calculation to effectively evaluate and compare the performance of each model, as illustrated in Equation 1 : Performance Score =

True Positives + True Negatives Total Sample

(1)

Nearest Centroid Model The Nearest Centroid classifier assigns samples to a class based on their distance to the centroid of each class in the feature space. The centroid is calculated by averaging the feature vectors of all samples within that class, as shown in Equation 2 Garcı́a 8

et al. (2018). n

1X Xi (2) n i=1 Where Xi refers to the feature vectors of each sample, while n denotes the total number of samples in the class. Centroid =

Support Vector Machine A Support Vector Machine creates a hyperplane in a high-dimensional space to effectively separate different classes. It maximizes the margin between these classes by solving the related optimization problem, as shown in Equation 3: ! n X 1 2 ∥w∥ + C max(0, 1 − yi (w · Xi + b)) (3) min 2 i=1 In this context, w represents the weight vector, C is the regularization parameter, and b represents the bias. Support Vector Machines are particularly effective when samples are expected to be closely clustered in the feature space Cortes and Vapnik (1995). This concept is particularly relevant in studies of sperm morphology and motility, where scores can vary significantly.

Quadratic Discriminant Analysis Quadratic Discriminant Analysis (QDA) was used to address the classification problem, assuming that different classes follow a normal distribution and that the decision boundary between them is quadratic in shape Ghojogh and Crowley (2019). The posterior probability for each class is calculated using Bayes’ theorem as shown in Equation 4:

P (x | y = k )P (y = k ) P (y = k | x) = P (4) l P (x | y = l)P (y = l) This probabilistic model assumes features are normally distributed and allows for nonlinear class boundaries Ghojogh and Crowley (2019). Thus, QDA is especially useful for assessing sperm parameters, as these parameters are assumed to follow a normal distribution.

Gaussian Naive Bayes The Gaussian Naive Bayes (GaussianNB) classifier uses Bayes’ Theorem to estimate the lthe probability of a class based on the measured features. GaussianNB assumes that the input predictors, such as sperm concentration, morphology, and motility, follow a normal distribution. Consequently, the model makes predictions on test samples based on this assumption. The probability is calculated using the normal probability density function, as shown in Equation 5. 1

P (Xi | C ) = √

2πσ 2 9

· e−

(Xi −µ)2 2σ 2

(5)

Here, Xi represents the actual measured feature (for example, sperm concentration in millions per milliliter), µ is the mean of the features for class C , and σ is the variance of the distribution of those features for that class. The GaussianNB classifier is known for its robustness, interpretability, and efficiency Zhang (2004).

Random forest Classifier The Random Forest Classifier is an adaptive learning technique that employs multiple decision trees to minimize the chances of errors in decision-making Breiman (2001). This method is particularly useful in systems where various factors must be considered to reach a conclusion, such as in semen analysis. In a Random Forest, the final prediction is made by averaging the outputs from all the decision trees in the forest, as shown in Equation 6.   m X ŷ = arg max  Tj (X ) (6) C

j=1

The term Tj (X ) refers to the prediction made by the jth tree in the forest, and m represents the total number of trees.

Cross-Validation and Statistical Testing To minimize overfitting and evaluate model performance on unseen data, all models were assessed using 5-fold cross-validation. This technique is beneficial because it exposes models to various subsets of data, helping prevent overfitting. Moreover, the Naive Bayes model was hyperparameter-tuned using GridSearchCV to identify the optimal parameters. The primary focus during this tuning process was the Var smoothing parameter, which significantly enhanced the model’s performance on the semen analysis dataset.

ROC-AUC Analysis In addition to basic performance measures that focus on classification accuracy, Receiver Operating Characteristic curves (ROC) were used to evaluate the model’s performance. Since the dataset involves a multi-class problem, a one-vs-rest approach was used, treating each class as positive and the others as negative. The performance of the level one models was satisfactory, with no misclassifications. Consequently, the Area Under the Curve (AUC) was calculated for each class to assess the model’s ability to differentiate between classes. The resulting Area Under the Curve scores were:

• Fast Class: AUC = 0.95 • Average Class: AUC = 1.00 • Slow Class: AUC = 0.97 The resulting AUC scores indicate that the classifier effectively distinguishes between different sperm motility classes. 10

Feature Visualization To investigate the relationships among progressive motility, morphology, and sperm concentration, we generated density plots for each feature. These plots helped us evaluate how effectively we could differentiate between various fertility classes and seminal characteristics. They supported the classification boundaries we selected Zhou et al. (2021), providing deeper insights into the factors that influence classification.

5 Results This study examined semen samples from the VISEM dataset, which included 85 male participants aged 18 years and older. After preprocessing and labeling according to WHO standards, various machine learning classifiers were applied to predict fertility status.

Dataset Statistics and Class Distribution The VISEM dataset comprises 85 samples and 21 attributes that capture essential sperm characteristics, including concentration, morphology, quality, motility, and deoxyribonucleic acid (DNA) fragmentation. Key columns in the dataset include sperm concentration (measured in ×106 per ml), the percentage of normal spermatozoa, morphology, and progressive motility percentages. The label column categorizes each sample into one of three classes, i.e., Fast, Average, or Slow, based on these parameters as shown in Table 2. The dataset revealed that the sperm concentration (measured

Table 2: Fertility Class Distribution in the VISEM Dataset. Class distribution Fast or Fertile Average or Sub-Fertile Slow or Infertile

No.of Samples 46 31 08

in millions per milliliter, 106 /ml) had a mean value of 81.01 ± 65.34, with a range from 3 to 350. Progressive motility had a mean of 42.67% ± 20.18%, with a minimum of 0% and a maximum of 76%. Correlation analysis between parameters such as total sperm concentration, motility, and various morphological attributes indicated moderate to strong positive correlations. These correlations suggest multicollinearity, which can negatively affect model performance (Please see the Figure 3).

Feature Visualization and Interpretation Figures 3, 4, and 5 illustrate the distributions of progressive motility, morphology, and sperm concentration across fertility classes. The Fast class consistently showed higher median values in all three metrics. Specifically: 11

Fig. 3: Progressive Motility Distribution Across Fertility Classes. This boxplot depicts the distribution of progressive motility percentages across three fertility classes: Fast, Average, and Slow. The Fast group exhibits the highest median progressive motility at 60%, ranging from 4% to 70%. The Average group shows a median of 30% with a broader range (0–50%), while the Slow group has the lowest median at approximately 20%, with most values between 10% and 30%. The visual representation clearly shows class separability, reinforcing the association between higher progressive motility and a superior fertility classification.

Fig. 4: Normal Sperm Morphology Across Fertility Classes. This boxplot displays the distribution of normal spermatozoa morphology percentages across three motility-based fertility classes: Fast, Average, and Slow. The Fast class has the highest median values, mostly ranging from 3% to 6%, with an outlier at 8.9%. The Average class is centered between 1% and 4%, while the Slow class shows even lower percentages. This suggests a potential link between higher motility and improved sperm morphology, underscoring their combined importance in fertility assessment.

• Progressive Motility (see Figure 3): The Fast group had a median motility near 60%, while the Average and Slow groups had median values of approximately 30% and 20%, respectively. • Sperm Morphology (see Figure 4): The Fast group showed the highest median percentage of normal forms, with values reaching up to 8.9%. • Sperm Concentration (see Figure 5): The Fast group had median concentrations ranging between 50 and 150 million/mL, which were significantly higher than those in the Average and Slow categories. 12

These visualizations support the hypothesis that higher sperm quality, indicated by concentration, morphology, and motility, correlates with improved fertility classification.

Fig. 5: Sperm Concentration Across Fertility Classes. This boxplot shows sperm concentration (millions per milliliter) across three fertility classes based on motility, i.e., Fast, Average, and Slow. The Fast group has a median concentration of about 50-150 million/mL, with some outliers exceeding 300 million/mL. The Average group ranges from 25 to 75 million/mL. The Slow group is around 25 million/mL with minimal variability. This visualization highlights the relationship between sperm concentration and motility-based fertility classification.

Feature Correlation Analysis Pair plots (see Figure 6) revealed clear separability among fertility classes based on sperm concentration, morphology, and motility. The diagonal kernel density plots illustrated distinct distributions: Fast samples (blue) exhibited higher values across all three dimensions, while Slow samples (green) were clustered at the lower end. Average samples (orange) fell in between. Figure 7 presents a heatmap illustrating the correlations among these three core features. Key findings include:

• Sperm concentration and progressive motility: correlation = 0.50 (moderate positive) • Morphology and motility: correlation = 0.35 (weaker positive) • Concentration and morphology: correlation = –0.11 (slight negative) These results indicate that concentration and motility are more closely linked than morphology, although all three factors significantly contribute to fertility prediction.

Model Performance and Accuracy Using LazyPredict, we evaluated more than 40 machine learning classifiers. The best performers, i.e., Nearest Centroid, SVM, and QDA, were further analyzed in detail (see Figure 8). The Nearest Centroid method has a modest ability to separate classes and indicate reliably by SVM and QDA, both around 94%. Gaussian Naive Bayes and Random Forest also demonstrated good performance, with Gaussian Naive Bayes reaching 13

Fig. 6: Correlation Between Sperm Parameters Across Fertility Classes. This scatter plot matrix illustrates the relationships among sperm concentration (millions/mL), normal morphology (%), and progressive motility (%) across three fertility classes, i.e., Fast (blue), Average (orange), and Slow (green). The diagonal features kernel density plots that reveal distinct distributions for each parameter. Fast samples exhibit higher concentrations and motility, while Slow samples cluster at lower values, highlighting interdependencies among these fertility parameters.

an accuracy of 91%. The ROC-AUC values supported these findings. Specifically, the Nearest Centroid model achieved:

• AUC = 0.95 for Fertile class • AUC = 1.00 for Sub-Fertile class • AUC = 0.97 for Infertile class These high scores demonstrate strong ability to separate classes and indicate reliable classification.

Confusion Matrix and ROC Curve Analysis The confusion matrix for the Nearest Centroid classifier (as shown in Figure 9) demonstrated excellent class-specific performance:

• Perfect classification was achieved in the Fast and Slow groups. • One sample in the Average group was misclassified as Slow. Despite these minimal misclassifications, the model achieved high accuracy in distinguishing between classes. The misclassifications, especially between the Average and Slow groups, may arise from overlapping morphological or motility traits in borderline samples. This suggests opportunities to improve the model by expanding the training set or including additional biological features. The multiclass ROC curve (see Figure 10) further validated the model’s discriminative power. All curves significantly surpassed the diagonal baseline, and the AUC values confirmed the model’s strong performance across all fertility categories. 14

Fig. 7: Correlation Heatmap of Key Sperm Parameters. This heatmap shows the correlation matrix for sperm concentration (millions/mL), normal morphology (percentage), and progressive motility (percentage). There is a moderate positive correlation of 0.50 between concentration and motility, a weak positive correlation of 0.35 between morphology and motility, and a slight negative correlation of -0.11 between concentration and morphology. Color intensity indicates the strength of the correlation, from red (high positive) to blue (low or negative), illustrating the relationships among these factors in predicting fertility.

Fig. 8: Accuracy of Machine Learning Models for Fertility Prediction. This bar chart compares the performance of machine learning classifiers in predicting fertility status from seminal parameters. The top models ,i.e., Nearest Centroid, Support Vector Machine, and Quadratic Discriminant Analysis achieved over 90% accuracy, while the ZeroR Classifier served as a lower baseline. Rankings are based on average cross-validation accuracy, demonstrating their effectiveness for real-world applications.

15

6 Discussion 6.1 Interpretation of Key Findings Our study reaffirms the crucial role of sperm concentration, morphology, and progressive motility in evaluating male fertility, consistent with existing literature. Previous research Westerman (2020) has shown that higher sperm concentration and improved morphology are positively correlated with enhanced motility, underscoring their importance in clinical fertility evaluations. Our correlation analysis provides empirical support for this relationship. Specifically, we found that samples with higher sperm concentration and superior morphological characteristics exhibited significantly better progressive motility. This observation aligns with Agarwal et al. (2022), who emphasized the strong relationship between these parameters and fertility outcomes. Furthermore, our results are consistent with Hook and Fisher (2020), who argued that sperm morphology directly influences motility and, consequently, fertility success rates. The correlation matrix and visualizations (see Figures 5, 4, and 6) also revealed an interdependence between sperm traits. Progressive motility showed a moderate to strong positive correlation with both morphology and concentration, reinforcing the hypothesis that high-quality sperm are characterized by an alignment of multiple parameters rather than isolated features.

Fig. 9: Confusion Matrix for Nearest Centroid Classifier. This confusion matrix shows the Nearest Centroid model’s performance across three fertility classes, i.e., Fast, Average, and Slow. The model accurately classified all Fast and Slow samples but misclassified one Average sample as Slow. The color intensity indicates the number of predictions per cell. With an overall accuracy of 94.2%, the matrix highlights the model’s strong ability to differentiate among fertility categories, with minimal errors.

16

6.2 Comparative Performance of Classifiers In evaluating predictive performance, the Nearest Centroid classifier achieved the highest cross-validated accuracy at 94.2%, outperforming other models tested using LazyPredict. This finding supports the conclusions of Hicks et al. (2019), who highlighted the effectiveness of centroid-based models in biological classification tasks, especially when data distributions are both compact and distinct. Although it was slightly less accurate, the Gaussian Naive Bayes model still achieved 91% accuracy. This result aligns with Wood et al. (2019), which showed that Naive Bayes performs well in medical and biological domains where input features are only moderately correlated. Given its computational simplicity and interpretability, it remains a valuable option, particularly in clinical decision-support contexts with limited resources. It is worth noting that all of the top models exhibited strong discriminative power, as indicated by high ROC-AUC scores (see Figure 10). The Nearest Centroid model achieved AUC values of 0.95, 1.00, and 0.97 for the Fertile, Sub-Fertile, and Infertile classes, respectively, exceeding the performance thresholds typically expected in fertility diagnostics.

6.3 Limitations and Areas for Improvement While the models demonstrated high overall accuracy, some limitations were identified. The confusion matrix (as shown in Figure 9) revealed minor misclassifications, particularly between the Average and Slow categories. This overlap may be attributed to morphological and concentration similarities in borderline cases, an observation also noted in You et al. (2021). Although the Nearest Centroid model performed well overall, its sensitivity for false negatives in these adjacent categories needs improvement. As highlighted by Idowu et al. (2015), minimizing false negatives is critical in clinical contexts. Misclassifying an infertile patient as sub-fertile may lead to inappropriate treatment recommendations. To address this challenge, it may be beneficial to enrich the dataset with a broader range of borderline cases, fine-tune feature selection, or integrate additional biomarkers, such as sperm DNA fragmentation or hormonal profiles. Moreover, despite the assumption of class separability, fertility data are often influenced by complex, nonlinear interactions and latent variables that are not captured by conventional semen analysis. Future research work could explore the application of ensemble methods or neural networks capable of modeling such nonlinearities, as well as semi-supervised learning techniques to leverage unlabelled clinical data. Finally, our findings confirm that the Area Under the Curve remains a reliable metric for assessing model performance in medical contexts, as highlighted by Christodoulou et al. (2019). The high AUC scores reported in this study validate the model’s ability to accurately differentiate between fertility categories, reinforcing the potential for machine learning to enhance sperm quality evaluation and inform clinical decision-making. 17

Fig. 10: ROC Curves for Fertility Classification Across Three Classes. This figure shows multiclass ROC curves for the Nearest Centroid model. True Positive Rates and False Positive Rates are displayed for each fertility class, i.e., Fertile (AUC = 0.95), Sub-Fertile (AUC = 1.00), and Infertile (AUC = 0.97). The diagonal dashed line indicates the random classification baseline, while the curves highlight the model’s strong predictive ability and excellent class separability.

7 Conclusion This study aimed to address a significant challenge in reproductive medicine, i.e., the subjective and often inconsistent evaluation of male fertility through traditional semen analysis. By leveraging machine learning models to classify sperm motility based on various seminal parameters, we demonstrated that automated, data-driven methods can provide more objective, reproducible, and clinically meaningful insights into male reproductive potential. Our findings showed that models such as Nearest Centroid and Support Vector Machine accurately classified fertility status based on sperm concentration, morphology, and progressive motility, achieving cross-validated accuracies that exceeded 90%. Visual analyses and correlation matrices further supported these results, revealing strong associations between sperm morphology, motility, and concentration. ROCAUC evaluations further confirmed the high discriminative power of these models, emphasizing their potential for real-world clinical application. These findings collectively support the integration of machine learning frameworks into clinical fertility assessment workflows. This integration can help reduce human error, improve diagnostic precision, and provide evidence-based guidance to both patients and clinicians. Furthermore, expanding these models with additional biological markers, such as hormonal profiles and molecular-genetic indicators, could enhance prediction accuracy and allow for more personalized fertility treatment planning. More broadly, this research highlights the transformative role of AI in reproductive health, not only in diagnosing male infertility but also in shaping the future of precision 18

medicine. In short, empowering clinicians with intelligent, data-driven tools may redefine how fertility is understood, diagnosed, and managed, ushering in a new era of reproductive care focused on predictive accuracy and patient-centered outcomes.

References Adinugroho S, Nakazawa A (2023) Deep learning-based sperm motility and morphology estimation on stacked color-coded motionflow. Computers in Biology and Medicine n/a:107381. https://doi.org/https://doi.org/10.1016/j.cmpb.2023.107381 Agarwal A, Sharma R, Gupta S, et al (2022) Sperm morphology assessment in the era of intracytoplasmic sperm injection: reliable results require focus on standardization, quality control, and training. The World Journal of Men’s Health 40(3):347 Bonde JPE, Ernst E, Jensen TK, et al (1998) Relation between semen quality and fertility: a population-based study of 430 first-pregnancy planners. The Lancet 352(9135):1172–1177 Breiman L (2001) Random forests. Machine learning 45(1):5–32 Brown RL (1944) Rate of transport of spermia in human uterus and tubes. American Journal of Obstetrics and Gynecology 47(3):407–411 Carlsen E, Giwercman A, Keiding N, et al (1992) Evidence for decreasing quality of semen during past 50 years. British medical journal 305(6854):609–613 Cavarocchi E, Drouault M, Ribeiro JC, et al (2025) Human asthenozoospermia: Update on genetic causes, patient management, and clinical strategies. Andrology 13(5):1044–1064 Choi Jw, Alkhoury L, Urbano LF, et al (2022) An assessment tool for computerassisted semen analysis (casa) algorithms. Scientific Reports 12(1):16830. https://doi.org/10.1038/s41598-022-20943-9, URL https://doi.org/10.1038/ s41598-022-20943-9 Christodoulou E, Ma J, Collins GS, et al (2019) A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models. Journal of clinical epidemiology 110:12–22 Cooper TG, Noonan E, Von Eckardstein S, et al (2010) World health organization reference values for human semen characteristics. Human reproduction update 16(3):231–245 Cortes C, Vapnik V (1995) Support-vector networks. Machine learning 20:273–297 Crocker JC, Grier DG (1996) Methods of digital video microscopy for colloidal studies. Journal of Colloid and Interface Science 179(1):298–310. https://doi.org/https:// 19

doi.org/10.1006/jcis.1996.0217 Darand M, Salimi Z, Ghorbani M, et al (2023) Obesity is associated with quality of sperm parameters in men with infertility: a cross-sectional study. Reproductive Health 20(1):134. https://doi.org/10.1186/s12978-023-01664-2, URL https:// doi.org/10.1186/s12978-023-01664-2 Dcunha R, Hussein RS, Ananda H, et al (2022) Current insights and latest updates in sperm motility and associated applications in assisted reproduction. Reproductive sciences pp 1–19 Dietterich TG (1998) Approximate statistical tests for comparing supervised classification learning algorithms. Neural computation 10(7):1895–1923 Dobrovolny M, Benes J, Langer J, et al (2023) Study on sperm-cell detection using yolov5 architecture with labeled dataset. Faculty of Informatics and Management, University of Hradec Kralove Unpublished Garcı́a V, Sánchez JS, Marqués A, et al (2018) A regression model based on the nearest centroid neighborhood. Pattern Analysis and Applications 21(4):941–951 Gbagbo FY, Ameyaw EK, Yaya S (2024) Artificial intelligence and sexual reproductive health and rights: a technological leap towards achieving sustainable development goal target 3.7. Reproductive Health 21(1):196. https://doi.org/10.1186/ s12978-024-01924-9, URL https://doi.org/10.1186/s12978-024-01924-9 Ghojogh B, Crowley M (2019) Linear and quadratic discriminant analysis: Tutorial. arXiv preprint arXiv:190602590 Grow DR, Oehninger S, Seltman HJ, et al (1994) Sperm morphology as diagnosed by strict criteria: probing the impact of teratozoospermia on fertilization rate and pregnancy outcome in a large in vitro fertilization population. Fertility and sterility 62(3):559–567 Haugen TB, Hicks SA, Andersen JM, et al (2019) Visem: A multimodal video dataset of human spermatozoa. In: Proceedings of the 10th ACM Multimedia Systems Conference. ACM, pp 261–266 Hicks SA, Andersen JM, Witczak O, et al (2019) Machine learning-based analysis of sperm videos and participant data for male fertility prediction. Scientific reports 9(1):16770 Hook KA, Fisher HS (2020) Methodological considerations for examining the relationship between sperm morphology and motility. Molecular reproduction and development 87(6):633–649

20

Idowu PA, Sarumi S, Balogun JA (2015) A prediction model for the likelihood of infertility in women. In: 9TH International Conference on Information and Communications Technology(ICT) Applications, Ilorin, Kwara, pp 78–88 Kao SH, Chao HT, Liu HW, et al (2004) Sperm mitochondrial dna depletion in men with asthenospermia. Fertility and sterility 82(1):66–73 Kobayashi H, Uetani M, Yamabe F, et al (2024) A new model for determining risk of male infertility from serum hormone levels, without semen analysis. Scientific Reports 14(1):17079 Kumar N, Singh AK (2015) Trends of male factor infertility, an important cause of infertility: A review of literature. Journal of human reproductive sciences 8(4):191– 196 Leslie SW, Soon-Sutton TL, Khan MA (2024) Male infertility. In: StatPearls [internet]. StatPearls Publishing McLachlan RI (2013) Approach to the patient with oligozoospermia. The Journal of Clinical Endocrinology & Metabolism 98(3):873–880 Musthyala R, Narayanan A, Nistala A, et al (2024) An ai framework for predicting the winner of the grammys. In: 9th International Conference on Big Data Analytics (ICBDA), IEEE, pp 20–25 Nayak J, Jena SR, Samanta L (2019) Oxidative stress and sperm dysfunction: An insight into dynamics of semen proteome. In: Oxidants, Antioxidants and Impact of the Oxidative Status in Male Reproduction. Elsevier, p 261–275 Nguyen VD, Ngo TTH, Duong LM, et al (2023) Assessing sperm motility using deep learning on the visem dataset. In: Conference/Journal Name if available, unpublished Organisation WH (1999) WHO laboratory manual for the examination of human semen and sperm-cervical mucus interaction. Cambridge university press Organization WH (2021) WHO laboratory manual for the examination and processing of human semen, 6th edn. World Health Organization Ottl S, Amiriparian S, Gerczuk M, et al (2021) A machine learning framework for automatic prediction of human semen motility. arXiv preprint arXiv:2109.08049. https://doi.org/https://doi.org/10.48550/arXiv.2109.08049 Ottl S, Amiriparian S, Gerczuk M, et al (2022) motilitai: A machine learning framework for automatic prediction of human sperm motility. iscience 25(8) Raschka S (2018) Model Evaluation, Model Selection, and Algorithm Selection in Machine Learning. arXiv preprint arXiv:1811.12808 21

Saadat H, Sepehri MM, Borna MR, et al (2023) A modified u-net to detect real sperms in videos of human sperm cell. Unpublished Journal/Conference Name if available Thambawita V, Hicks SA, Halvorsen P, et al (2021) Deep learning-based video analysis of human spermatozoa for automated motility assessment. Scientific Reports 11(1):1–14 Thambawita V, Hicks SA, Storås AM, et al (2023) The visem-tracking dataset for sperm motility analysis. Scientific Data 10. https://doi.org/https://doi.org/10. 1038/s41597-023-02051-3 Tiab I, Belaid A, Yahiaoui S, et al (2023) Deep characterization for sperm analysis: Leveraging a novel dataset to enhance motility assessments. IEEE DOI or URL if available Valiuškaitė V, Raudonis V, Maskeliūnas R, et al (2023) Deep learning based evaluation of spermatozoid motility for artificial insemination. Sensors 21:3422. https://doi. org/https://doi.org/10.3390/s21103422 Vander Borght M, Wyns C (2018) Fertility and infertility: Definition and epidemiology. Clinical biochemistry 62:2–10 Westerman R (2020) Biomarkers for demographic research: sperm counts and other male infertility biomarkers. Biodemography and Social Biology 65(1):73–87 Wood A, Shpilrain V, Najarian K, et al (2019) Private naive bayes classification of personal biomedical data: application in cancer data analysis. Computers in biology and medicine 105:144–150 You JB, McCallum C, Wang Y, et al (2021) Machine learning for sperm selection. Nature Reviews Urology 18(7):387–403 Zhang H (2004) The optimality of naive bayes. AA 1(2):3 Zhou WJ, Huang C, Jiang SH, et al (2021) Influence of sperm morphology on pregnancy outcome and offspring in in vitro fertilization and intracytoplasmic sperm injection: a matched case-control study. Asian Journal of Andrology 23(4):421–428

22

Record · ID 353071 · SHA-256 1fadd9008ddd2d13
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.