arXiv:2606.24047v1 [cs.AI] 23 Jun 2026
Ensemble Feature Selection and Harris Hawks Optimization for Explainable Mental Health Risk Prediction in Female Sex Workers Ahnaf Atef Choudhury
Md. Parvej Hoque Palash
Department of Information Sciences and Technology George Mason University USA [email protected]
Department of Computer Science and Engineering Jahangirnagar University Bangladesh [email protected]
Ramkrishna Saha
Shahriar Siddique Ayon Department of Computer Science AIUB Bangladesh [email protected]
Abdullah Al Mamun
Department of Computer Science Department of Computer Science and Engineering The University of Texas at Dallas Dhaka University of Engineering and Technology USA Bangladesh [email protected] [email protected]
Abstract—One of the significant mental health issues affecting female sex workers (FSWs) is mental disorders, especially depression. Exposure to violence, stigma, and economic hardship further increases their psychological risk. Current machine learning (ML) models are typically ineffective at capturing the high-dimensional and complex risk patterns that exist in this marginalized group. This paper suggests a hybrid predictive model that merges an ensemble feature selection strategy using ANOVA and mutual information and Harris Hawks optimization-tuned logistic regression and represents a new application of swarm intelligence to predict mental health in vulnerable groups. The explainable AI (XAI) methods can be used to understand the factors of trauma associated with model predictions. When applied to a group of 3,005 FSWs, it can be seen that the proposed model is more effective than traditional classifiers, with an accuracy of 95.78%, an F1 score of 95.77%, and an AUC of 0.96, and identifying posttraumatic stress, client-related violence, and occupational factors as major contributors to depression. This work bridges the gaps between conventional and ML approaches to develop an XAI tool that enables vulnerable groups to receive early assistance, evidence-based targeted psychosocial care, and health planning. Index Terms—Depression prediction, Mental disorder, Swarm intelligence, Stress disorder, Explainable artificial intelligence
I. I NTRODUCTION Mental health issues are still a big problem for public health around the world in the 21st century. Prediction of mental health status among female sex workers (FSWs) is still a crucial public health issue given the confluence of structural violence, stigma, economic vulnerability, and occupational risks [1]. According to the World Health Organization (WHO), nearly one in seven people worldwide, about 1.1 billion individuals, were living with a mental disorder in 2021, with depression and anxiety being the most prevalent conditions [2]. Depression affects about 4% of the global population, including 5.7% of adults, with higher prevalence among women than men, and together with anxiety disorders leads to an estimated annual loss of nearly US$1 trillion in productivity [3]. These disorders are a leading cause of long-
term disability worldwide and are expected to rise significantly by 2040 [4]. Vulnerable groups exposed to overlapping structural, social, and occupational stressors face higher rates of depression, post-traumatic stress disorder (PTSD), anxiety, and suicidality. Among these groups, FSWs are highly marginalized and face severe risks, including violence from partners and clients, stigma, social exclusion, HIV exposure, and economic and housing insecurity [5]. These factors intensify psychological distress and limit access to care, with studies reporting depression rates up to 52.7% and PTSD rates of 53.6% among FSWs, far higher than in the general population [6]. Traditional methods has identified these associations, whereas machine learning (ML) has demonstrated significant potential in modeling intricate behavioral and health outcomes, including treatment-seeking behavior, risky sexual behavior, and HIV risk stratification [7]. Existing techniques frequently depend on traditional statistics or simplistic classifiers that inadequately address high-dimensional, nonlinear, and imbalanced datasets, as well as intricate interactions between exposure to violence, socioeconomic variables, and health indicators [8], [9]. Its insufficient hybrid feature selection and sophisticated optimization further deteriorate the levels of accuracy, stability, and interpretability [10], [11]. These constraints highlight the importance of more complex and reliable predictive models to capture the complex mental health risk profile of FSWs and enable early intervention, psychosocial support, and planning of health services. To address these gaps, this study proposes a hybrid model, which integrates ensemble feature selection with Harris Hawks Optimization (HHO)-tuned Logistic Regression (LR). The algorithm is tried on a community-based dataset, which is already cleaned up. Explainable AI (XAI) LIME also provides understandable, instance-level explanations of model predictions, and, therefore, the decisions made are easier to
understand and trust. Major Initiatives of this work: • An ensemble feature selection method that involved ANOVA and mutual information to better identify feature relevance as compared to other methods. • HHO is the first mental health prediction model to be applied to marginalized populations with high accuracy and HFO-optimized LR, making it better than the baseline models. • Implement XAI techniques to demonstrate important traumatic and occupational variables to predict depression, enhancing transparency and clinical applicability. The rest of the paper is structured in the following way: Section II provides the review of the relevant work, Section III describes the proposed methodology, Section IV provides and discusses the results of the experiment, and the last section V provides a conclusion of the study and the future directions of the research. II. L ITERATURE R EVIEW Interdependent risks, including violence, HIV exposure, and occupational risks, should be modeled to make accurate predictions of FSWs’ mental health. Existing stressors such as housing and food insecurity further exacerbate the risk of depression and PTSD [12]. Depression and anxiety are also strongly linked to intimate partner violence (IPV) and violence committed by clients, with emotional IPV having the strongest effect [13]. Muhlen et al. assert that stigma and social exclusion contribute to enhancing mental health risks further, which explains why psychosocial interventions are highly necessary [14]. Lowering violence and improving psychosocial support are important steps toward helping FSWs with their mental health issues. Structural vulnerabilities such as violence and economic instability profoundly affect the mental health and HIV risk of FSWs. Jewkes et al. [15] through a multi stage, community based cross sectional survey of 3,005 FSWs, reported alarmingly poor mental health outcomes, with 52.7% experiencing depression and 53.6% meeting criteria for PTSD. Similarly, Machisa et al. [16], using structural equation modeling on a sample of 1,292 participants, found high prevalence rates of binge drinking (50%), depressive symptoms (43%), PTSD symptoms (9%), and suicidal ideation (21%). Studies consistently show that FSWs face high rates of depression, PTSD, and suicidality, which can be effectively modeled using advanced ML techniques. Zhang et al. [17] evaluated seven ML algorithms—including logistic regression (LR), support vector machines (SVM), random forests (RFC), XGBoost, k-nearest neighbors (KNN), naïve Bayes (NB), and neural networks (NN)—reporting acceptable performance (ROC > 0.51) in predicting risky sexual behaviors, with an overall accuracy of 78% but limited sensitivity (11%). Nethi et al. [18] found that the light gradient boosting machine achieved the strongest predictive power for identifying candidates suitable for HIV pre-exposure prophylaxis (PrEP), with an area under the curve (AUC) of 0.88. In a similar study, it was showed that decision trees (DT), SVM, and RFC were better at handling SMOTE-processed data and that the accuracy and precision of decision-making processes were 0.871 and
0.960, respectively [19]. Qualitative studies emphasise stigma, violence, and inadequate psychosocial support as significant factors contributing to poor mental health among FSWs [20]. These structural and psychosocial factors can enhance ML models for more accurate mental health prediction in this population. The mental health prediction among FSWs depends on such methods as chi-square, ANOVA, mutual information (MI), Boruta, and tree-based models to model intricate risk patterns. Chi-squared tests were applied to examine relationships between categorical variables, such as work conditions and mental health outcomes [10]. The RFC model is the most effective among the four ML models used by Fauste et al. [21] with a 75% accuracy of comorbidity of mental diseases and an 88.8% accuracy of factors that lead to mental health vulnerability. Integrating ML and XAI enhances the accuracy and interpretability of mental health assessments, making it possible to intervene and support them more effectively. According to Saha et al. [22], the dynamic ensemble selection methods, including KNORA, have an accuracy of 81%, which clarifies the importance of using multiple models to enhance predictive accuracy. Masudur et al. [23] used the SMOTE to overcome the class imbalance and used SHapley Additive exPlanations (SHAP) to explain model results. SHAP successfully described feature contributions, and the RFC model performed the best (accuracy = 0.66, recall = 0.69). According to these studies, integrating ML and XAI enhances predictability and interpretability in practical applications related to mental health. Existing research relies on limited models and struggles with complex risk factors that depend on each other. Its performance is moderate, and it poorly generalizes on imbalanced data. Most studies also lack integrated feature selection, optimization, and XAI, which lowers both accuracy and interpretability. This study addresses these gaps with a unified hybrid ML and XAI framework. III. M ETHODOLOGY This section is a clear and well-organized presentation of the overall methodological framework of the study. It begins with the data collection, preprocessing, feature selection, and the development of optimized models, and XAI is used to extract important positive and negative features. Fig. 1 outlines the suggested framework to predict mental health status among FSWs on the basis of several perspectives. A. Data Collection, Cleaning, and Preprocessing The statistics in this article are based on a giant communitybased cross-sectional survey of 3,005 adult FSWs carried out in all nine provinces of South Africa [24]. The research was aimed at learning important details about their lives, such as HIV status, mental health problems, such as depression and PTSD, violence experience, and work-related issues. The data were gathered in locations where sex worker support programs already exist, and hence it is easier to access the participants by the virtue of having trusted networks. Accurate responses were collected using structured questionnaires and recorded in real time using REDCap. FSWs were also engaged in the entire process, during the survey design and in data collection,
Data Preprocessing Feature Selection & Comparison Mental Health Risks and Challenges for Female Sex Workers
Mental Health Data Data Splitting
Train Set
Validation Set
Test Set
Classification Model Selection
Best Model Optimized with Harris Hawks Optimizer
Explainable AI Use
Evaluation and Result Analysis
Fig. 1: Proposed Methodological Framework for Depression Prediction among Female Sex Workers.
which made the information more applicable, trustworthy, and realistic. The dataset initially had 20 columns with 3,005 rows and seven missing columns. In the numerical columns, Years worked as sex worker, Age of first sex, Number of clients in past day, and Earning potential per client, the missing values were replaced with the mean. Missing values in categorical columns were completely removed. Label encoding of all categorical variables transformed them into numerical form to analyze them using MLS. The column of enrolment date was divided into day, month, and year columns. Superfluous or redundant columns were also eliminated in order to simplify the dataset. The processed final dataset will have 2,911 rows and 19 feature columns. B. Hybrid Ensemble Feature Selection Strategy After preparing the data, we using feature selection to determine the most significant variables. We tested a variety of approaches such as chi-squared, ANOVA, MI, Boruta, treebased algorithms, and ensemble. Among them, there is the ensemble ANOVA and MI technique, which is a combination of statistical analysis of variance and information gain to sturdily detect the most significant features. The ranking of features according to the importance of features using the ensemble feature selection method is shown in figure 2.
1.0 0.5
ote Loca nti tio al n pe rc in lie pa nt st 12 Co mo nd nth om s us ei Pro nt Yea vin h r ce ep Int sw im a ork st ate mo ed pa n as rtn sex th er Nu wo ph mb ysi rke Ch ca l/se er of ilhoo r da clie xu al bu Po ab nts i lice n p se us e ast ph in ysi pa da st ca l/se 12 y m xu al Ag onth ab eo s us f ei n p first sex ast Fir 1 st 2m sex ua onth le s xp eri en ce HIV Fir st Ou st rap atus tdo id or HIV ve rsu tes s In en t do r or ba ol_da sed y sex w en ork rol _m on th
0.0
gp
C. Model Selection and Hyperparameter Tuning We tested several ML and DL models for depression prediction, tuning their hyperparameters for maximum performance. Random Forest Classifier (RFC) is an ensemble of 100 decision trees with max depth 10 and minimum 2 samples per leaf, designed to reduce overfitting. k-Nearest Neighbors (kNN) classifies instances based on the majority label of the 5 nearest neighbors using Euclidean distance. Support Vector Classifier (SVC) separates classes using an RBF kernel with C = 1.0 and γ = 0.1 to find the optimal hyperplane. Light Gradient Boosting Machine (LGBM) is a fast, gradient-boosted tree model using 200 estimators, learning rate of 0.05, and maximum depth of 7. Artificial Neural Network (ANN) with three dense layers (64, 32, 16 neurons), ReLU activation, and Adam optimizer (learning rate 0.001) was used to capture nonlinear relationships. Logistic Regression (LR) models the probability of a binary outcome using a logistic function, configured with L2 regularization, regularization strength C = 1.0, and the ’lbfgs’ solver. Among all models, LR performed best, and its performance was further enhanced using advanced swarm intelligence optimization techniques, including Particle Swarm, Ant Colony, Genetic Algorithm, and Harris Hawks Optimization (HHO), with HHO-based LR achieving the highest overall performance. HHO-based LR enhances standard LR by integrating HHO to efficiently search the parameter space and identify optimal model weights. The model estimates the probability of the positive class as:
Cli
en tp
hy
sic
al/
sex
ua
la
rni n Ea
bu se
PT
SD
ou tco
me
Importance Score
Top Features Driving Depression Prediction (ANOVA + Information Gain) 1.5
Figure 2 presents the ranking of features influencing depression prediction among sex workers, as determined by the ensemble ANOVA and MI method. PTSD outcome was the most important (1.558), followed by Location (1.195) and Earning potential per client (0.710). Other key features included Client physical/sexual abuse (0.688), Province (0.682), and Condom use in the past month (0.347). Other features such as years worked as a sex worker (0.218), childhood abuse (0.206), and number of clients in the past day (0.195) also played a role; the rest of the features had lower scores (0.184-0.019), which means that they had less impact on the depression prediction. According to the results of the ensemble feature selection, the least significant features were filtered off, resulting in an optimized set of 11 features and 2,911 observations. Location and Province capture regional structural factors (e.g., violence, access to care, and economic disadvantage) not directly observed but relevant in South Africa. They are used only for contextual risk stratification and referral guidance, not as causal individual predictors. The target variable, Depression, consists of 1,530 positive cases and 1,381 negative cases, corresponding to labels 1 and 0, respectively. The data was divided into 80% training and 20% testing; 20% of the training was used as a validation.
Features
Fig. 2: Key Features Driving Depression Prediction Identified by Ensemble Feature Selection.
P (y = 1|X) =
1 1 + e−(Xw+b)
(1)
Where w and b are the weights and bias. The HHO algorithm was configured with a population size of 30, maximum iterations of 50, and an escape energy parameter E0 = 2, which
together maximized predictive performance and effectively handled class imbalance.
TABLE I: Baseline Model Performance Comparison for Depression Prediction
Model RFC k-NN SVC LGBM ANN LR HHO-LR
D. Evaluation Metrics for Classification
Accuracy (%) 87.85 89.22 87.36 90.42 91.35 92.28 93.06
Precision Recall (%) (%) 87.62 87.94 89.05 89.40 87.15 87.50 90.18 90.55 91.10 91.55 92.05 92.50 92.82 93.20
F1-score (%) 87.78 89.22 87.32 90.36 91.32 92.27 93.01
AUC
92.82 94.17 93.22 92.62 94.32 93.47 94.55 95.77
0.93 0.95 0.94 0.93 0.95 0.94 0.95 0.96
0.88 0.89 0.87 0.91 0.92 0.93 0.94
In binary classification, the accuracy is the general percentage of correct predictions. Precision is the ratio between the number of predicted positives that are actually positive, whereas recall (or sensitivity) represents the number of true positive cases detected by the model. The F1-score gives an B. Performance After Feature Selection equilibrium between precision and recall, and the AUC gives After feature selection (Table II), performance improved the capacity of the model to differentiate the two classes, in across all methods, with the proposed ANOVA + IG achieving which the higher the values, the better the discrimination is. the best results. HHO-based LR reached the highest performance with 95.78% accuracy, 95.60% precision, 95.95% recall, 95.77% F1-score, and 0.96 AUC, followed by LR E. Explainable AI LIME (94.54%, 0.95 AUC). In comparison, Boruta and ensemble XAI assists in making ML models understandable, indicat- (MI + RFE) methods showed slightly lower performance, with ing why they make particular predictions. We applied LIME HHO-based LR achieving up to 94.31% accuracy and 0.95 in the study, which clarifies individual predictions by pointing AUC. to the contribution of each feature, which offers clear insights TABLE II: Performance Comparison After Feature Selection on the local level that complement global approaches such as Feature Model Accuracy Precision Recall F1 score AUC SHAP and SHAPASH [25]. The LIME explanatory model is Selection (%) (%) (%) (%) LR 93.14 92.95 93.35 93.15 0.94 represented as follows: Boruta
ĝ (x) = argmingϵG L (f, g, πx´) + Ω (g)
(2)
Ensemble (MI+RFE) Proposed
ANN HHO-LR LR XGBoost HHO-LR ANN LR HHO-LR
92.81 94.16 93.22 92.61 94.31 93.46 94.54 95.78
92.60 93.95 93.05 92.40 94.10 93.25 94.35 95.60
93.05 94.40 93.40 92.85 94.55 93.70 94.75 95.95
(ANOVA In LIME, the complex model f is approximated locally by + IG) an interpretable model ĝ(x) ∈ G, which minimizes the loss function L(f, g, πx′ ) to achieve a balance between fidelity to HHO improved recall by 1.20% and F1-score by 1.22% the original model and interpretability. over grid-search LR; while the absolute gain is modest, this corresponds to approximately ∼ 36 additional correct depression detections in the cohort, justifying its use in high-stakes IV. R ESULTS AND D ISCUSSION A NALYSIS screening settings where the cost of false negatives outweighs All models were developed and tested using the free version the minimal computational overhead. Figure 3 compares violence exposure between depressed and of Google Colab, providing a simple and efficient platform for experimentation. Models were tuned using 5 fold cross- non-depressed groups, showing consistently higher prevalence validation to optimize performance on the training data. We among the depressed cohort. Client abuse shows the largest compared model performance before and after applying feature gap (68.6% vs. 44.5%, ∆ = +24.0%), followed by intimate selection and correlation analysis to better understand their partner abuse (52.7% vs. 39.5%, ∆ = +13.2%). Childhood effects. To enhance interpretability, LIME was used to identify abuse remains highest overall (92.6% vs. 83.6%, ∆ = +8.9%), features that positively or negatively influenced predictions. while police abuse is lower but still elevated (17.1% vs. The overall process and key findings are outlined step by step 10.6%, ∆ = +6.6%). On the whole, these findings suggest that depression is closely related to a greater exposure to below. various types of violence, especially those related to clients and partners. A. Baseline Model Performance Table I, the baseline model comparison shows that HHO- C. Explaining Predictions Using LIME
based LR achieved the best overall performance with 93.06% accuracy, 92.82% precision, 93.20% recall, 93.01% F1-score, and 0.94 AUC. Among standard models, LR performed strongest with 92.28% accuracy and 0.93 AUC, followed by ANN (91.35%, 0.92 AUC) and LGBM (90.42%, 0.91 AUC). k-NN showed moderate performance (89.22%, 0.89 AUC), while RFC (87.85%, 0.88 AUC) and SVC (87.36%, 0.87 AUC) achieved comparatively lower results. This highlights the effectiveness of the HHO-based optimization in improving model performance.
We used LIME tabular and feature importance plots to highlight how individual features influenced the model’s predictions. To assess generalizability beyond a single case, LIME explanations were aggregated across all 582 test instances. PTSD outcome and client abuse consistently emerged as the top two predictors in 94.3% and 87.1% of cases, respectively, and their importance closely aligned with the ANOVA+MI ranking (Spearman ρ = 0.91, p < 0.001), indicating strong population-level stability. The Figure 4 presents a tabular explanation of a binary prediction with LIME that gives more
Fig. 5: LIME Feature Importance for Correct Depression Prediction.
Fig. 3: Violence Exposure Prevalence Among Depressed and Non-Depressed Groups.
preference to class 1 (0.57 vs. 0.43), which implies that the model predicts higher risk. PTSD outcome (+0.27) is the strongest contributor, followed by recent client abuse (+0.13) and years in sex work (+0.06). Other factors, including earning potential, client number, location, and province, show smaller effects (+0.02–0.03), while the remaining features contribute minimally. Overall, the prediction is primarily driven by trauma-related and occupational exposure factors.
factors such as the factors of PTSD and violence exposure. Cross validation assists the internal validity but external generalization is untested. Few differences to the previous work (Zhang et al. [17]; Ndikumana et al. [21]) exist, which is due to the complexity of the datasets and the tasks instead of the superiority of the models. The hybrid model is an ensemble of ensemble feature selection, HHO-optimized LR, and XAI, which enhances both the accuracy and interpretability of mental health predictions. It has a light weight design, so that it can be used on devices without cloud support in real time and on low cost. The model can be integrated into any REDCap-based system to support CHWs in making interpretable depression risk screenings and providing referral support. V. C ONCLUSION AND F UTURE W ORK
Fig. 4: LIME Explanation of Key Features Driving Depression Prediction.
Figure 5 presents a LIME feature importance plot for a correct prediction (actual = 1, predicted = 1). The top contributors driving the prediction toward class 1 are PTSD outcome ≤ 0.00 (+0.27), client abuse in the past 12 months (+0.13), and years worked as a sex worker (+0.06). Moderate positive effects come from earning potential (+0.03), number of clients (+0.02), location (+0.02), and province (+0.02), while intimate partner abuse (+0.01) and features like age of first sex and condom use ( 0.00) have minimal impact. Overall, trauma-related and occupational factors dominate, resulting in a confident classification into the higher-risk class. Past research on mental health prediction for FSWs has shown moderate performance, which is limited by suboptimal feature selection, limited optimization, inability to deal with high dimensional interactions, lack of XAI, and generalization issues [15], [18], [22], [23]. Our model, on the other hand, has high performance and is mostly led by clinically important
The current study suggested a novel hybrid approach for mental health prediction for FSWs based on the combination of ensemble feature selection and HHO-tuned LR. The model achieved better performance than baseline models and other optimization-based models in mental health datasets through the incorporation of the above techniques: ANOVA, MI and swarm intelligence. Further, LIME increased the interpretability of the findings: they identified important trauma-related and socioeconomic factors that could influence depression in FSWs, which would allow for more explainable and data-based mental health interventions. Future work will extend the proposed framework through mixed-methods validation, fairness-aware evaluation, multicountry generalization, and real-world deployment. Qualitative follow-up interviews with FSWs will confirm the congruence of the risk factors identified by LIME with the lived experience of FSWs and help to identify other structural determinants for improved modeling. Future studies will test the framework with cohorts from East Africa and South Asia to address the geographic and selection biases of the current dataset from South Africa, and will also include fairness metrics across different subgroups defined by demographic characteristics and exposure to violence. The modular HHO-LR framework can be adapted to other mental health settings, but there is a need for rigorous external validation and benchmarking before claims to broader generalizability can be made. Future studies will also explore the creation of an offline-first mHealth system with encrypted on-device data storage and privacy protection, along with testing in the field to assess usability, cultural acceptability
of XAI outputs, and long-term predictive stability in real-world [16] M. T. Machisa, E. Chirwa, P. Mahlangu, N. Nunze, Y. Sikweyiya, E. Dartnall, M. Pillay, and R. Jewkes, “Suicidal thoughts, depression, public health environments. DATA AVAILABILITY The dataset used in this study is available on Mendeley at the following link: https://data.mendeley.com/datasets/hfr552s47v/ 1 R EFERENCES [1] G. Kaya, O. Kalinowski, F. Kroehn-Liedtke, A. Lotysh, H. Mihaylova, L. Zerbe, W. Rössler, and M. Schouler-Ocak, “The impact of selfstigmatization on the mental health of female sex workers (fsws),” Frontiers in Public Health, vol. 13, p. 1679876, Nov 2025. Impact Factor: 3.4, Q1. [2] World Health Organization, “Mental disorders,” 2025. Accessed: 202604-16. [3] World Health Organization, “Depressive disorder (depression),” 2025. Accessed: 2026-04-16. [4] Z. Zhang, X. Chen, S. Wu, X. Chen, X. Wang, C. Liu, N. Zeng, Y. Liu, T. Huo, X. Liu, et al., “Global, regional and national burden of anxiety and depression disorders from 1990 to 2021, and forecasts up to 2040,” Journal of Affective Disorders, p. 120299, 2025. [5] N. T. Tutlam, S. Kizito, P. Nabunya, M. Naseh, I. Nabbosa, I. Kwesiga, P. Namatovu, O. S. Bahar, N. Nakasujja, and F. M. Ssewamala, “Social determinants of mental health outcomes among refugee adolescents and youth living with hiv in refugee settlements in uganda: A cross-sectional analysis,” AIDS and Behavior, vol. 29, pp. 3432–3443, Nov 2025. Impact Factor: 2.4, Q2. Epub 2025 Jun 16. [6] National Alliance on Mental Illness (NAMI), “Mental health by the numbers,” 2025. Accessed: 2026-04-16. [7] A.-A. Kebede Kassaw, T. Melese Yilma, Y. Sebastian, A. Yeneneh Birhanu, M. Sharew Melaku, and S. Surur Jemal, “Spatial distribution and machine learning prediction of sexually transmitted infections and associated factors among sexually active men and women in ethiopia, evidence from edhs 2016,” BMC Infectious Diseases, vol. 23, no. 1, p. 49, 2023. [8] M. Abubakkar, K. S. Sharif, I. Ahmad, D. M. Tabila, F. A. Alsaud, and S. Debnath, “Explainable suicide risk prediction with deepfusion: A hybrid intelligence approach,” in 2025 4th International Conference on Electronics Representation and Algorithm (ICERA), pp. 455–460, IEEE, 2025. [9] E. R. Bernal-Monroy, E. D. Castañeda-Monroy, R. R. Rentería-Ramos, S. E. Campaña-Bastidas, J. Barrera, T. M. Palacios-Yampuezan, O. L. González Gustin, C. F. Tobar-Torres, and Z. R. Ceballos-Villada, “Detection of victimization patterns and risk of gender violence through machine learning algorithms,” in Informatics, vol. 12, p. 21, MDPI, 2025. [10] F. Kroehn-Liedtke, O. Kalinowski, G. Kaya, A. Lotysh, H. Mihaylova, K. Sipos, A. Strunk, L. Zerbe, W. Rössler, and M. Schouler-Ocak, “A quantitative study on female sex workers’ mental health in germany,” Frontiers in Public Health, vol. Volume 13 - 2025, 2025. [11] S. S. Ayon, A. Al Mamun, M. E. Hossain, W. Alamro, Y. M. Allawi, N. N. I. Prova, M. S. U. Miah, S. M. Sultan, and A. Abadleh, “Explainable ai framework for improved thalassemia mental health classification and feature selection,” PLoS One, vol. 21, no. 1, p. e0341168, 2026. [12] C. Tomko, R. J. Musci, M. R. Kaufman, C. R. Underwood, M. R. Decker, and S. G. Sherman, “Mental health and hiv risk differs by cooccurring structural vulnerabilities among women who sell sex,” AIDS Care, vol. 35, no. 2, pp. 205–214, 2023. [13] M. Leis, M. McDermott, A. Koziarz, L. Szadkowski, A. Kariri, T. S. Beattie, R. Kaul, and J. Kimani, “Intimate partner and client-perpetrated violence are associated with reduced hiv pre-exposure prophylaxis (prep) uptake, depression and generalized anxiety in a cross-sectional study of female sex workers from nairobi, kenya,” Journal of the international AIDS society, vol. 24, p. e25711, 2021. [14] A. Mühlen, J. Rudy, A. Böckmann, and D. Deimel, “Psychische gesundheit von sexarbeiter* innen in europa: ein scoping-review,” Das Gesundheitswesen, vol. 85, no. 06, pp. 561–567, 2023. [15] R. Jewkes, M. Milovanovic, K. Otwombe, E. Chirwa, K. Hlongwane, N. Hill, V. Mbowane, M. Matuludi, K. Hopkins, G. Gray, and J. Coetzee, “Intersections of sex work, mental ill-health, ipv and other violence experienced by female sex workers: Findings from a cross-sectional community-centric national study in south africa,” International Journal of Environmental Research and Public Health, vol. 18, no. 22, 2021.
post-traumatic stress, and harmful alcohol use associated with intimate partner violence and rape exposures among female students in south africa,” International Journal of Environmental Research and Public Health, vol. 19, no. 13, 2022. [17] F. Zhang, S. Zhu, S. Chen, Z. Hao, Y. Fang, H. Zou, Y. Cai, B. Cao, K. Zhang, H. Cao, Y. Chen, T. Hu, and Z. Wang, “Application of machine learning for risky sexual behavior interventions among factory workers in china,” Frontiers in Public Health, vol. 11, p. 1092018, 2023. [18] A. K. M. S. Nethi, M. Karam, Albert George M. S., K. S. Alvarez, A. E. Luque, A. E. Nijhawan, E. Adhikari, and H. L. King, “Using machine learning to identify patients at risk of acquiring hiv in an urban health system,” JAIDS Journal of Acquired Immune Deficiency Syndromes, vol. 97, pp. 40–47, Sep 2024. [19] J. He, J. Li, S. Jiang, W. Cheng, J. Jiang, Y. Xu, J. Yang, X. Zhou, C. Chai, and C. Wu, “Application of machine learning algorithms in predicting hiv infection among men who have sex with men: Model development and validation,” Frontiers in Public Health, vol. 10, p. 967681, 2022. [20] L. Morgan, H. R. Welborn, G. Feist-Paz, et al., “Mental ill health experiences of female sex workers and their perceived risk factors: A systematic review of qualitative studies,” Nov 2023. Preprint, Version 1. [21] F. Ndikumana, J. Izabayo, J. Kalisa, M. Nemerimana, E. C. Nyabyenda, S. H. Muzungu, I. Komezusenge, M. Uwase, S. Ndagijimana, C. Twizere, and V. Sezibera, “Machine learning-based predictive modelling of mental health in rwandan youth,” Scientific Reports, vol. 15, p. 16032, May 2025. Q1, Impact Factor: 3.9. [22] Y. Saha and H. S, “Dynamic ensemble selection for mental health prediction : A path towards explainable, scalable and high-impact ai solutions,” in 2025 International Conference on Intelligent Computing and Knowledge Extraction (ICICKE), pp. 1–8, 2025. [23] M. R. Kanchon, J. Sani, T. Ahmed, et al., “Interpretable machine learning for predicting early mental health care-seeking among reproductive-age women in bangladesh using bdhs 2022 data,” Feb 2026. Preprint, Version 1. [24] M. Milovanovic, “Cross sectional study of female sex workers in south africa,” 2021. [25] S. Siddique Ayon, M. Ebrahim Hossain, M. S. Ullah Miah, M. M. Rahman, and M. Mahmud, “Explainable ai in feature selection: Improving classification performance on imbalanced datasets,” in International conference on neural information processing, pp. 303–318, Springer, 2024.