Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Sci Rep . 2026 Mar 1;16:11495. doi: 10.1038/s41598-026-41953-x Search in PMC Search in PubMed View in NLM Catalog Add to search Explainable machine learning prediction of tracheostomy after craniotomy for supratentorial intracerebral hemorrhage Feiyu Qiao Feiyu Qiao 1 Department of Neurosurgery, Weifang People’s Hospital, Shandong Second Medical University, No. 151 Guangwen Road, Weifang, Shandong China Find articles by Feiyu Qiao 1 , Xin Xue Xin Xue 2 Department of Neurosurgery, Weifang Hospital of Traditional Chinese Medicine, Weifang, China Find articles by Xin Xue 2 , Huanhuan Yu Huanhuan Yu 2 Department of Neurosurgery, Weifang Hospital of Traditional Chinese Medicine, Weifang, China Find articles by Huanhuan Yu 2 , Yanrui Cai Yanrui Cai 1 Department of Neurosurgery, Weifang People’s Hospital, Shandong Second Medical University, No. 151 Guangwen Road, Weifang, Shandong China Find articles by Yanrui Cai 1 , Di Tian Di Tian 1 Department of Neurosurgery, Weifang People’s Hospital, Shandong Second Medical University, No. 151 Guangwen Road, Weifang, Shandong China Find articles by Di Tian 1 , Yuting Wang Yuting Wang 1 Department of Neurosurgery, Weifang People’s Hospital, Shandong Second Medical University, No. 151 Guangwen Road, Weifang, Shandong China Find articles by Yuting Wang 1 , Qi Liu Qi Liu 1 Department of Neurosurgery, Weifang People’s Hospital, Shandong Second Medical University, No. 151 Guangwen Road, Weifang, Shandong China Find articles by Qi Liu 1 , Quancai Li Quancai Li 1 Department of Neurosurgery, Weifang People’s Hospital, Shandong Second Medical University, No. 151 Guangwen Road, Weifang, Shandong China Find articles by Quancai Li 1, ✉ Author information Article notes Copyright and License information 1 Department of Neurosurgery, Weifang People’s Hospital, Shandong Second Medical University, No. 151 Guangwen Road, Weifang, Shandong China 2 Department of Neurosurgery, Weifang Hospital of Traditional Chinese Medicine, Weifang, China ✉ Corresponding author. Received 2025 Oct 21; Accepted 2026 Feb 23; Collection date 2026. © The Author(s) 2026 Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/ . PMC Copyright notice PMCID: PMC13057019 PMID: 41765996 Abstract Accurate prediction of tracheostomy after craniotomy for supratentorial intracerebral hemorrhage (sICH) remains challenging. This study aimed to develop, externally validate, and interpret a machine learning model for individualized risk prediction. A retrospective multicenter cohort was constructed, including 738 patients from Weifang People’s Hospital and 186 from Weifang Hospital of Traditional Chinese Medicine who underwent craniotomy between January 2017 and December 2024. Predictor variables were screened using least absolute shrinkage and selection operator (LASSO) and multivariate logistic regression. Logistic regression, random forest, and extreme gradient boosting (XGBoost) models were trained with repeated 10-fold cross-validation and assessed for discrimination, calibration, and clinical utility. Five key predictors were identified: Glasgow Coma Scale, age, hematoma volume, operative time, and serum bicarbonate. In external validation, XGBoost demonstrated the most balanced and robust performance, with an AUROC of 0.86 and a Brier score of 0.15, and showed superior net benefit on decision curve analysis. SHapley Additive exPlanations confirmed clinical plausibility, and a web-based dynamic nomogram was developed for individualized prediction. This explainable XGBoost model provides reliable and interpretable estimation of postoperative tracheostomy risk, facilitating evidence-based perioperative decision-making and resource allocation in neurocritical care. Supplementary Information The online version contains supplementary material available at 10.1038/s41598-026-41953-x. Keywords: Supratentorial intracerebral hemorrhage, Tracheostomy, Craniotomy, Explainable machine learning, Risk prediction Subject terms: Diseases, Medical research, Neurology, Neuroscience, Risk factors Introduction Stroke patients often experience impaired consciousness, respiratory compromise, and dysphagia, leading to prolonged mechanical ventilation and the need for tracheostomy 1 – 3 . Tracheostomy facilitates airway management and ventilator weaning, with reported rates of 15–45% among stroke patients, substantially higher than in the general ICU population 4 – 6 . Among these, patients with supratentorial intracerebral hemorrhage (sICH) who undergo craniotomy represent a particularly high-risk subgroup 7 – 9 , and conventional craniotomy remains widely performed in China 10 , 11 . However, tracheostomy predictive tools specifically designed for sICH patients undergoing craniotomy are lacking. Existing prediction tools for tracheostomy in stroke rely mainly on traditional regression-based methods 12 – 15 and often target mixed stroke cohorts 12 , subarachnoid hemorrhage 14 , 15 , or non-surgical patients 13 . These limitations highlight the need for accurate, interpretable, and clinically applicable models to enable early risk prediction of tracheostomy in this specific population. Machine learning (ML) can capture complex, non-linear relationships among clinical variables and has demonstrated strong predictive capability across diverse medical applications 16 , 17 . Nevertheless, the “black-box” nature of many ML algorithms hinders clinical adoption. SHapley Additive exPlanations (SHAP), a game-theoretic interpretability method, can quantify each predictor’s contribution and enhance transparency at both cohort and individual levels 18 , 19 . Accordingly, we developed and externally validated explainable ML models to predict postoperative tracheostomy after craniotomy for sICH. We compared multiple algorithms to identify the best-performing model through a comprehensive evaluation of predictive metrics. Model interpretation was achieved using SHAP, and the final model was deployed as an interactive web-based nomogram to facilitate clinical application. Methods This study followed the TRIPOD statement for reporting prediction model development and validation 20 , 21 . The overall workflow is illustrated in Fig. 1 . Fig. 1. Open in a new tab Overall workflow of the study. Study design and population This retrospective dual-center cohort study included adult sICH patients who underwent craniotomy between January 2017 and December 2024. Data were collected from Weifang People’s Hospital (development cohort) and Weifang Hospital of Traditional Chinese Medicine (external validation cohort). Inclusion criteria were: (1) age ≥ 18 years; (2) sICH confirmed by non-contrast computed tomography (CT); and (3) craniotomy for hematoma evacuation within 24 h of symptom onset. Exclusion criteria were: (1) infratentorial hemorrhage; (2) secondary causes such as tumor, vascular malformation, or trauma; and (3) prior tracheostomy. This retrospective study was conducted in accordance with the principles of the Declaration of Helsinki. It was approved by the Institutional Review Board of Weifang People’s Hospital (Approval No. KYLL20250106-9) and the Ethics Committee of Weifang Hospital of Traditional Chinese Medicine (Approval No. 2025-WFSZYY-016). The requirement for informed consent was waived by both committees (the Institutional Review Board of Weifang People’s Hospital and the Ethics Committee of Weifang Hospital of Traditional Chinese Medicine) owing to the retrospective design and use of anonymized data. Craniotomy and tracheostomy indications At both centers, craniotomy treatment followed institutional protocols, broadly consistent with Chinese national guidelines. Indications for craniotomy in sICH included hematoma volume > 30 mL, midline shift > 1 cm, Glasgow Coma Scale (GCS) ≤ 8, signs of impending herniation, and failure of medical management. Tracheostomy was also performed according to institutional protocols, generally consistent with accepted standards of critical care. Indications included prolonged respiratory failure requiring ventilation for > 7 days, the need for airway protection in patients at high risk of aspiration or with functional/mechanical obstruction, persistent need for suctioning of tracheal secretions, and the presence of dysphagia. Data collection and variables Demographic, clinical, radiological, intraoperative, and laboratory variables were extracted from the electronic medical record systems of both centers. Two trained investigators independently collected the data, with cross-checking to ensure accuracy and inter-center consistency. Demographic variables included age and sex. Clinical variables at admission were Glasgow Coma Scale (GCS), systolic blood pressure (SBP), and diastolic blood pressure (DBP). Radiological variables included hematoma volume (HVol), hematoma location (HLoc), and midline shift (MLS). HVol was estimated using the ABC/2 method (A, maximum diameter on the largest hematoma slice; B, perpendicular diameter; C, number of slices with hemorrhage multiplied by slice thickness). HLoc was categorized as lobar or deep, with the latter defined as hemorrhage involving the basal ganglia (caudate, putamen, globus pallidus) or thalamus. MLS was defined as the distance of septum pellucidum deviation from the midline, measured in millimeters. Intraoperative variables included operative time (OT) and intraoperative blood loss (IBL). Laboratory parameters within 24 h of admission included white blood cell count (WBC), hemoglobin (Hb), platelet count (PLT), lymphocytes, monocytes, neutrophils, serum bicarbonate (HCO₃⁻), and serum glucose (Glu). Derived inflammatory indices comprised the neutrophil-to-lymphocyte ratio (NLR), systemic inflammatory response index (SIRI), and platelet-to-lymphocyte ratio (PLR). The primary outcome was postoperative tracheostomy, defined as any tracheostomy performed during hospitalization following craniotomy. Continuous variables were analyzed in their native units. Binary variables were dummy coded (e.g., sex: female = 0, male = 1; tracheostomy: no = 0, yes = 1). All variables were harmonized across centers using standardized definitions and coding. Missing data were evaluated for all candidate variables. If the proportion of missing values exceeded 5% for any variable, multiple imputation was planned. Otherwise, complete case analysis was applied. Feature selection Feature selection was performed in the development cohort to prevent data leakage. A two-step approach combined penalized regression with multivariable modeling. First, least absolute shrinkage and selection operator (LASSO) regression identified the most informative predictors. Variables with non-zero coefficients were then entered into a multivariable logistic regression to confirm independent associations ( p < 0.05). This sequential strategy ensured a clinically interpretable and parsimonious predictor set, minimizing overfitting and enhancing model generalizability. Model development and hyperparameter optimization Three candidate machine learning models were developed: logistic regression (LR), random forest (RF), and extreme gradient boosting (XGBoost) 22 – 24 . Model training and tuning were conducted in the development cohort using repeated 10-fold cross-validation (100 repetitions). To prevent information leakage, all preprocessing steps, feature selection procedures, and hyperparameter tuning were performed exclusively within the training folds of the cross-validation framework. The validation folds were not involved in any stage of preprocessing, feature selection, or model fitting. For LR, elastic-net regularization was applied with grid search over penalty (λ) and mixing (α) parameters. RF tuning with grid search covered mtry and min_n. XGBoost tuning included mtry, number of trees, min_n, tree_depth, learn_rate, loss_reduction, and sample_size, optimized with the tune_race_anova() function across 1,000 candidate configurations. Performance was evaluated on the resampled training sets using area under the receiver operating characteristic curve (AUROC), area under the precision-recall curve (AUPRC), accuracy, sensitivity, specificity, precision, F1-score, and Brier score. The hyperparameter configuration that achieved the highest mean AUROC across resamples was selected for each model and carried forward for external validation. Model validation External validation was performed using the independent cohort from Weifang Hospital of Traditional Chinese Medicine. Each model, trained with the optimal hyperparameters identified in the development cohort, was directly applied to the external dataset without additional recalibration, providing an unbiased estimate of out-of-sample performance. Model performance was comprehensively evaluated across three domains: discrimination, calibration, and clinical utility. Discrimination was quantified using the area under the receiver operating characteristic curve (AUROC), the area under the precision–recall curve (AUPRC), and threshold-dependent metrics including accuracy, sensitivity, specificity, precision, and F1-score. Calibration was assessed using the Brier score, the Hosmer–Lemeshow goodness of fit test, and visual inspection of calibration plots comparing predicted and observed event rates across risk deciles. Clinical utility was evaluated using decision curve analysis (DCA), which estimates the net clinical benefit of each model across a range of threshold probabilities relative to default strategies of treating all or treating none. This multifaceted validation strategy enabled a balanced and rigorous assessment of model performance, capturing not only statistical discrimination but also the reliability of probability estimates and the potential clinical value of each predictive model. Model interpretation Model interpretability was explored using SHAP, which attributes a contribution value to each predictor, reflecting its impact on the predicted probability of tracheostomy. SHAP analysis was performed on the best-performing model to quantify predictor importance and the direction of their effects at the global level. Patient-level SHAP values were also generated to produce individualized explanation plots, illustrating how specific variables influenced tracheostomy risk in individual cases. Web-based deployment To facilitate clinical translation, the final model was deployed as an interactive web-based nomogram using the Shiny framework in R. The tool allows clinicians to input patient-specific variables and obtain individualized probabilities of postoperative tracheostomy. The interface also generates SHAP-based waterfall plots, providing case-specific visual explanations of how each predictor contributed to the estimated probability. The application incorporates the best-performing externally validated algorithm, retains all predictors in their original clinical units, and is accessible via a secure URL across devices. Sample size considerations No a priori sample size calculation was conducted because of the retrospective design. The required sample size for a binary outcome can be approximated as , where φ is the expected outcome ratio (φ = 0.43), δ is the set margin of error (δ = 0.05) 25 . According to this formula, the minimum sample size required for model development would be approximately 494 participants. Furthermore, as suggested by Collins et al. for external validation of prognostic models, at least 100 events are recommended for external validation 26 . Both cohorts (738 for development, 186 for validation) met these thresholds. Statistical analysis Categorical variables were expressed as counts (%) and continuous variables as mean ± standard deviation (SD) or median (interquartile range, IQR), as appropriate. Group differences were tested using the Mann–Whitney U test or Pearson’s χ² test. A two-sided p < 0.05 was considered statistically significant. All analyses were performed in R. Data preprocessing used the tidyverse package suite. Machine learning procedures—including feature selection, model training, resampling, hyperparameter tuning, and evaluation—were implemented with tidymodels. Data visualization was performed using ggplot2, and SHAP analyses were conducted with kernelshap and shapviz. The web-based nomogram was developed with Shiny. Details of all R packages and versions are provided in Table S1 . Results Patient characteristics A total of 738 patients were included in the development cohort (Weifang People’s Hospital) and 186 in the external validation cohort (Weifang Hospital of Traditional Chinese Medicine). In the development cohort, missing values were observed in 27 patients, corresponding to low per-variable missingness (maximum 2.0% for serum glucose and 1.6% for serum bicarbonate); therefore, a complete-case analysis was applied, and the reported sample size represents the final analytic cohort. No missing data were observed in the external validation cohort (Table S2 ). Overall, tracheostomy was performed in 43% of patients in the development cohort and 49% in the validation cohort, respectively. Baseline characteristics of patients in the development and validation cohorts are summarized in Table 1 , and detailed in Table S3 . Compared with the development cohort, patients in the validation cohort had a higher proportion of deep hematomas, lower hemoglobin levels, and shorter operative times, while other demographic and clinical features were broadly comparable. Table 1. Comparison of baseline characteristics between the development (WFPH) and validation (WHTCM) cohorts. Development N = 738 1 Validation N = 186 1 p -value 2 Tracheostomy status 0.13 Yes 317 (43%) 92 (49%) No 421 (57%) 94 (51%) Demographics Age, years 59 (51, 68) 56 (46, 69) 0.2 Clinical characteristics GCS 8.0 (5.0, 11.0) 7.0 (5.0, 9.0) 0.009 SBP, mmHg 180 (159, 200) 181 (168, 210) 0.018 DBP, mmHg 101 (92, 112) 103 (94, 114) 0.052 Radiological findings HVol, ml 60 (50, 80) 60 (50, 80) 0.2 HLoc (Deep) 562 (76%) 170 (91%) < 0.001 MLS, mm 7.2 (5.0, 10.0) 6.0 (5.0, 8.5) 0.001 Surgical characteristics OT, hours 4.00 (3.00, 4.50) 3.20 (2.50, 4.00) < 0.001 IBL, ml 500 (300, 600) 400 (300, 450) < 0.001 Laboratory parameters WBC 11.6 (9.1, 15.1) 11.0 (9.1, 14.5) 0.11 HCO 3 − , mmol/L 23.3 (21.1, 25.5) 25.6 (22.6, 27.3) < 0.001 Open in a new tab ¹ Data are presented as median (Q1, Q3) for continuous variables and number (percentage) for categorical variables. ² p-values were calculated using the Wilcoxon rank-sum test for continuous variables and Pearson’s Chi-squared test for categorical variables. GCS, Glasgow Coma Scale; SBP, systolic blood pressure; DBP, diastolic blood pressure; HVol, hematoma volume; HLoc, hematoma location; MLS, midline shift; OT, operative time; IBL, intraoperative blood loss; WBC, white blood cell count; HCO₃⁻, serum bicarbonate. Within the development cohort (Table S4 ), patients who underwent tracheostomy were older, had lower GCS scores at admission, larger hematoma volumes, greater midline shift, longer operative times, and lower serum bicarbonate levels compared with those without tracheostomy ( P < 0.05 for all). Similar patterns were observed in the validation cohort (Table S5 ), where tracheostomy patients were older, presented with lower GCS, larger hematomas, more pronounced midline shift, and longer operative times ( P < 0.05 for all). Feature selection Using the LASSO regression (Fig. 2 A) followed by multivariable logistic regression analysis (Table 2 ; Fig. 2 B), five predictors were retained in the development cohort: GCS, age, HVol, OT, and HCO₃⁻. These predictors were statistically significant in the multivariable model ( p < 0.05 for all) and were considered clinically plausible. The chosen features were integrated into three machine learning classifiers—LR, RF and XGBoost—to develop predictive models in subsequent model development and validation. Fig. 2. Open in a new tab Feature selection and hyperparameter tuning of machine learning models. ( A ) LASSO regression path across a sequence of penalty values (λ). ( B ) Forest plot of multivariate logistic regression identifying independent predictors of postoperative tracheostomy; blue bars denote significant variables ( p < 0.05). ( C ) Cross-validated area under the receiver operating characteristic curve (AUROC) of LR across penalty–mixture combinations. ( D ) AUROC surface of RF across mtry and min_n. ( E ) AUROC optimization of XGB tuned via a performance-based racing strategy. Table 2. Multivariate logistic regression analysis. Predictor β SE OR p -value Age 0.024 0.008 1.024(1.008–1.041) 0.005 ** GCS −0.439 0.043 0.644(0.561–0.728) < 0.001 *** HVol 0.018 0.006 1.018(1.007–1.030) 0.002 ** MLS 0.032 0.028 1.032(0.977–1.088) 0.262 OT 0.450 0.103 1.568(1.366–1.770) < 0.001 *** HCO 3 − −0.070 0.032 0.932(0.870–0.994) 0.026 * Open in a new tab GCS, Glasgow Coma Scale; HVol, hematoma volume; MLS, midline shift; OT, operative time; HCO₃⁻, serum bicarbonate. Hyperparameter tuning All three candidate models underwent systematic hyperparameter optimization in the development cohort using repeated 10-fold cross-validation. The best-performing configurations are illustrated in Figs. 2 C–E, with detailed parameter grids provided in Tables S6 – S8 . For logistic regression, the optimal penalty and mixture values were 0.10 and 0.00, respectively, corresponding to a purely LASSO-penalized model. For the random forest, the best performance was achieved with an mtry of 1 and a minimum node size of 6. For XGBoost, the optimal configuration included mtry = 5, number of trees = 830, min_n = 18, tree depth = 3, learning rate = 1.13 (on a negative log₁₀ scale, corresponding to an actual value of 0.074), loss reduction = 3.28, and sample size = 0.98. These tuned hyperparameters defined the final models, which were subsequently carried forward for external validation and interpretability analyses. Development and validation of prediction models In the development cohort, LR, RF, and XGBoost demonstrated comparable discrimination, with AUROC values ranging from 0.87 to 0.89 and AUPRC values from 0.90 to 0.91 (Fig. 3 E, Table S9 ). RF showed slightly higher AUROC and lower Brier score than the other models, while accuracy and F1-score were also marginally superior. Calibration performance in the internal training cohort is shown in Figure S1 . In the validation cohort, a similar pattern was observed: LR achieved the highest AUROC (0.88), RF attained the highest accuracy (0.81) and sensitivity (0.93), whereas XGBoost performed intermediately across most conventional metrics (Figs. 4 G-H, Table S10 ). Fig. 3. Open in a new tab Feature importance and cross-validated performance of machine learning models in the development cohort. ( A–C ) Permutation-based feature importance for LR, RF, and XGB models. ( D–I ) Violin plots showing the distribution of performance metrics across 100 iterations of repeated 10-fold cross-validation. Fig. 4. Open in a new tab External validation and comparative performance of machine learning models in the validation cohort. ( A–C ) Normalized confusion matrices for LR, RF, and XGB showing classification accuracy. ( D–F ) Calibration plots showing agreement between predicted and observed probabilities. ( G ) ROC curves with corresponding AUROC values indicating discriminative ability. ( H ) Precision–recall curves showing model performance in imbalanced data. ( I ) Decision curve analysis (DCA) demonstrating net clinical benefit across threshold probabilities. Abbreviations: TT, tracheostomy; NT, no tracheostomy. Despite these broadly comparable results, calibration and decision-analytic performance distinguished the models. The Hosmer–Lemeshow test revealed significant miscalibration for LR ( p < 0.001) and RF ( p < 0.05), whereas XGBoost showed no evidence of lack of fit ( p = 0.11) (Figs. 4 D-F). Calibration plots confirmed better alignment between predicted and observed risks for XGBoost (Figs. 4 D-F). Moreover, DCA demonstrated that XGBoost provided the greatest net clinical benefit, particularly at threshold probabilities > 0.5 (Fig. 4 I), which is clinically relevant for identifying high-risk patients likely to require tracheostomy. Taken together, although XGBoost did not outperform LR or RF on discrimination metrics alone, it provided the most favorable balance of calibration and clinical utility. Accordingly, XGBoost was selected as the final model for interpretation and deployment. Model interpretation and visualization Global SHAP summary plots for LR, RF, and XGBoost demonstrated consistent feature importance rankings, with GCS and HVol among the most influential predictors across models, followed by age, OT, and HCO₃⁻ (Figs. 5 A–C). Fig. 5. Open in a new tab SHAP-based interpretation and visualization of feature effects. ( A–C ) SHAP summary plots showing the distribution, magnitude, and direction of feature contributions for LR, RF, and XGB models. ( D ) Bar plot of mean absolute SHAP values from the XGB model ranking overall feature importance. ( E–I ) SHAP dependence plots derived from the XGB model showing the marginal effect of each feature on predicted tracheostomy probability. Abbreviations: OT, operative time; MLS, midline shift; HVol, hematoma volume; HCO₃⁻, serum bicarbonate; GCS, Glasgow Coma Scale. SHAP dependence plots indicated that lower GCS, larger HVol, longer OT, and lower HCO₃⁻ were associated with higher predicted tracheostomy risk, whereas increasing age showed a gradual rise in predicted probability (Figs. 5 E–I). At the individual level, SHAP waterfall plots from the external validation cohort illustrated how patient-specific feature values contributed to predicted risk (Fig. S2). In a representative high-risk patient (Fig. S2A), low GCS, large hematoma volume, advanced age, prolonged operative time, and low HCO₃⁻ collectively increased the predicted probability of tracheostomy. In contrast, a representative low-risk patient (Fig. S2B) exhibited non-extreme values of these predictors, corresponding to a substantially lower predicted probability. Web application deployment To support bedside use, we developed an interactive web-based nomogram based on the final XGBoost model. The tool accepts patient-specific inputs for the five selected predictors and returns an individualized probability of postoperative tracheostomy (Fig. 6 ). For transparency, the application also generates a SHAP-based visualization that illustrates how each variable contributes to the predicted probability. The interface was designed for clinical usability, featuring clear labels and original measurement units. The web application is freely accessible at https://feiyuqiao.shinyapps.io/sich-cr-trach-nomogram/ , and the source code is available upon request for academic use. Fig. 6. Open in a new tab Web-based dynamic nomogram for individualized prediction of postoperative tracheostomy in sICH patients undergoing craniotomy. The nomogram, derived from the XGB model, allows users to input patient-specific variables and obtain individualized predicted probabilities, accompanied by SHAP-based visualization of feature contributions. Discussion In this dual-center study, XGBoost proved to be the most clinically useful model for predicting postoperative tracheostomy in sICH. Although the three algorithms showed comparable discrimination, XGBoost achieved superior calibration and net clinical benefit, indicating more accurate probability estimation and better generalizability. A notable finding was the strong concordance in feature importance across algorithms and cohorts (Figs. 3 A–C and 5 A–C), underscoring the stability and reproducibility of the selected predictors. These findings suggest that the identified variables capture core aspects of postoperative airway vulnerability rather than model-specific artifacts. The final model incorporated five interpretable predictors—GCS, age, HVol, OT, and HCO₃⁻—each reflecting a distinct dimension of perioperative risk. Consistently, GCS emerged as the strongest determinant across models and cohorts (Figs. 3 A–C and 5 A–C), with lower scores indicating higher tracheostomy risk (Fig. 5 E). Clinically, reduced GCS reflects impaired consciousness, diminished respiratory drive, and suppression of cough and swallowing reflexes, all of which increase aspiration risk and necessitate prolonged airway protection. This finding aligns with a large meta-analysis identifying low GCS as the most robust predictor of extubation failure across heterogeneous ICU populations 27 , as well as multiple studies in neurocritical cohorts reporting that higher GCS was associated with successful extubation 28 , 29 . Furthermore, both the TRACH score and the SETscore—two established tracheostomy prediction tools in stroke populations—explicitly incorporate GCS as a central component 12 , 13 , underscoring its prognostic value. Advanced age was also independently associated with an increased likelihood of tracheostomy, consistent with prior evidence linking older age to extubation failure—a closely related endpoint. A meta-analysis confirmed age as an independent predictor of extubation failure 27 and consistent with original studies showing higher failure rates among elderly patients, particularly those aged > 65 years 28 , 30 – 32 . Collectively, these findings suggest that aging-related declines in respiratory reserve and airway reflexes contribute to impaired weaning and prolonged ventilatory support, thereby increasing tracheostomy risk. Larger hematoma volume was also significantly associated with an increased likelihood of tracheostomy, indicating that hematoma burden may prolong the need for airway protection. This finding is consistent with prior evidence in deep-seated ICH, where hematoma volume > 30 mL independently predicted prolonged mechanical ventilation 33 . Additionally, longer operative time emerged as a significant predictor of tracheostomy risk, likely reflecting greater surgical complexity and extended anesthesia exposure. These factors can delay neurological recovery and predispose to pulmonary complications, as reported in neurosurgical series linking procedures exceeding 6 h to higher rates of extubation failure or reintubation 34 , and in general surgery meta-analyses showing stepwise increases in postoperative pneumonia with longer operative durations 35 . Finally, lower serum bicarbonate (HCO₃⁻) levels were significantly associated with increased tracheostomy risk, suggesting that systemic acid–base imbalance may impair airway protection or weaning capacity. Several ICU studies have identified low bicarbonate or acidemia as independent predictors of extubation failure or prolonged ventilation 36 – 39 . The I-TRACH model likewise incorporated HCO₃⁻ <20 mmol/L as a key predictor of tracheostomy 37 . Although this association has not been consistently observed in stroke-specific cohorts 40 , our findings highlight HCO₃⁻ as a potential metabolic marker of airway vulnerability in postoperative sICH. Several prior models have aimed to predict tracheostomy in stroke-related neurocritical populations, although most were developed in mixed or non-surgical cohorts, thereby limiting their applicability to postoperative sICH. The TRACH score was developed in supratentorial ICH but excluded surgical cases 13 , restricting its postoperative utility. The RAISE score 14 and an aneurysmal SAH nomogram 15 were derived in distinct disease contexts with different pathophysiology and management. The SETscore, although conceptually similar, required 15 variables and was designed for a mixed cerebrovascular cohort 12 , limiting feasibility for bedside use. Its authors also highlighted the need for simplified tools in homogeneous populations, which directly motivated the present work. This discussion also connects to the ongoing debate on tracheostomy timing 2 , 3 , 41 – 45 . Randomized trials such as TracMan, SETPOINT, and SETPOINT2 have failed to demonstrate consistent benefit of early versus late tracheostomy 41 , 44 , 45 , a discrepancy that is likely attributable to two major factors: heterogeneous study populations and variable clinical criteria for tracheostomy indication. By focusing on postoperative sICH, our model provides a standardized, individualized estimate of tracheostomy risk, which may serve as a useful reference for more standardized and evidence-based patient selection in future studies. This study addresses a previously underserved population—patients undergoing craniotomy for sICH—by providing an interpretable, externally validated model with direct clinical applicability. SHAP-based visualizations enhanced transparency, and the web-based dynamic nomogram enables real-time, individualized risk estimation across devices. Together, these features facilitate integration into perioperative workflows and support data-driven decision-making in neurocritical care. Several limitations of this study should be acknowledged. First, the retrospective design precludes causal inference and carries the risk of residual confounding. In addition, predictors were limited to admission and operative variables and did not incorporate dynamic perioperative or longitudinal factors. Moreover, relevant unmeasured confounders, such as pre-existing comorbidities, surgeon experience, and variability in postoperative airway management and critical care practices, were not explicitly captured and may have influenced tracheostomy decision-making. Although external validation was performed, the validation cohort was relatively small and geographically close to the development cohort, which may limit generalizability to broader populations. Importantly, however, notable case-mix differences were observed between the two cohorts, including a higher proportion of deep hematomas and shorter operative times in the validation cohort. Despite these differences, the model—particularly XGBoost—maintained stable discrimination, good calibration, and favorable clinical utility, suggesting a degree of robustness across related but non-identical clinical settings. Nevertheless, further prospective validation in larger, geographically diverse, and multi-regional cohorts is warranted to confirm generalizability and to assess performance under different surgical practices and perioperative care pathways. Conclusion We developed and externally validated an interpretable machine learning model for predicting tracheostomy after craniotomy in sICH. Among candidate algorithms, XGBoost demonstrated the most balanced performance and has been implemented as an open-access web application with SHAP-based explanations. This tool may assist in airway planning, improve communication, and optimize resource allocation. With further prospective validation, it has the potential to enhance perioperative management and support individualized, evidence-based decisions in neurocritical care. Supplementary Information Below is the link to the electronic supplementary material. Supplementary Material 1 (396.5KB, pdf) Acknowledgements The authors sincerely thank the Department of Neurosurgery teams at Weifang People’s Hospital and Weifang Hospital of Traditional Chinese Medicine for their valuable assistance in data management and case verification. Author contributions F.Q. and Q.Li conceived and designed the study. F.Q. and X.X. collected and curated the clinical data. F.Q. performed data preprocessing, feature selection, model development, visualization, and web deployment. F.Q. drafted the manuscript. H.Y., Y.C., D.T., Y.W., and Q.Liu contributed to manuscript revision. Q.Li supervised the project, interpreted the findings, and critically revised the manuscript. All authors reviewed and approved the final version of the manuscript. Funding This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors. Data availability The datasets analyzed during the current study are not publicly available due to institutional data-use regulations and patient confidentiality policies. De-identified data and analytic code are available from the corresponding author upon reasonable request and approval by the participating institutions. Declarations Competing interests The authors declare no competing interests. Footnotes Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. References 1. Bösel, J. Use and Timing of Tracheostomy After Severe Stroke. Stroke 48 , 2638–2643 (2017). [ DOI ] [ PubMed ] [ Google Scholar ] 2. Lais, G. & Piquilloud, L. Tracheostomy: Update on why, when and how. Curr. Opin. Crit. Care 31 , 101–107 (2025). [ DOI ] [ PubMed ] [ Google Scholar ] 3. Premraj, L. et al. Tracheostomy timing and outcome in critically ill patients with stroke: A meta-analysis and meta-regression. Crit. Care 27 , 132 (2023). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 4. Kurtz, P. et al. How does care differ for neurological patients admitted to a neurocritical care unit versus a general ICU?. Neurocrit. Care 15 , 477–480 (2011). [ DOI ] [ PubMed ] [ Google Scholar ] 5. Pelosi, P. et al. Management and outcome of mechanically ventilated neurologic patients. Crit. Care Med. 39 , 1482–1492 (2011). [ DOI ] [ PubMed ] [ Google Scholar ] 6. Steidl, C. et al. Tracheostomy, extubation, reintubation: Airway management decisions in intubated stroke patients. Cerebrovasc. Dis. 44 , 1–9 (2017). [ DOI ] [ PubMed ] [ Google Scholar ] 7. Abulhasan, Y. B., Teitelbaum, J., Al-Ramadhani, K., Morrison, K. T. & Angle, M. R. Functional outcomes and mortality in patients with intracerebral hemorrhage after intensive medical and surgical support. Neurology 100 , e1985–e1995 (2023). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 8. Hoffman, H., Jalal, M. S. & Chin, L. S. Prediction of mortality after evacuation of supratentorial intracerebral hemorrhage using NSQIP data. J. Clin. Neurosci. 77 , 148–156 (2020). [ DOI ] [ PubMed ] [ Google Scholar ] 9. Ho, U.-C., Hsieh, C.-J., Lu, H.-Y., Huang, A.-H. & Kuo, L.-T. Predictors of extubation failure and prolonged mechanical ventilation among patients with intracerebral hemorrhage after surgery. Respir. Res. 25 , 19 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 10. Hu, X. et al. Surgical outcomes from haematoma evacuation for intracerebral haemorrhage in the INTERACT3 study. Lancet Reg. Health - Western Pac. 62 , 101669 (2025). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 11. Yu, Z. et al. Chinese multidisciplinary guideline for management of hypertensive intracerebral hemorrhage. Chin. Med. J. (Engl.) 135 , 2269–2271 (2022). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 12. Schönenberger, S., Al-Suwaidan, F., Kieser, M., Uhlmann, L. & Bösel, J. The SETscore to predict tracheostomy need in cerebrovascular neurocritical care patients. Neurocrit. Care 25 , 94–104 (2016). [ DOI ] [ PubMed ] [ Google Scholar ] 13. Szeder, V., Ortega-Gutierrez, S., Ziai, W. & Torbey, M. T. The TRACH score: Clinical and radiological predictors of tracheostomy in supratentorial spontaneous intracerebral hemorrhage. Neurocrit. Care 13 , 40–46 (2010). [ DOI ] [ PubMed ] [ Google Scholar ] 14. Rass, V. et al. Factors associated with prolonged mechanical ventilation in patients with subarachnoid hemorrhage—the RAISE score*. Crit. Care Med. 50 , 103 (2022). [ DOI ] [ PubMed ] [ Google Scholar ] 15. Chen, X.-Y. et al. A nomogram for predicting the need of postoperative tracheostomy in patients with aneurysmal subarachnoid hemorrhage. Front. Neurol. 12 , 711468 (2021). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 16. Zhang, Z. et al. Prediction of microvascular obstruction from angio-based microvascular resistance and available clinical data in percutaneous coronary intervention: An explainable machine learning model. Sci. Rep. 15 , 3045 (2025). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 17. Ballı, M., Dogan, A. E., Senol, S. H. & Eser, H. Y. Machine learning based identification of suicidal ideation using non-suicidal predictors in a university mental health clinic. Sci. Rep. 15 , 13843 (2025). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 18. Lundberg, S. M. & Lee, S. I. A unified approach to interpreting model predictions. in Advances in Neural Information Processing Systems vol. 30Curran Associates, Inc., (2017). 19. Lundberg, S. M. et al. From local explanations to global understanding with explainable AI for trees. Nat Mach Intell 2 , 56–67 (2020). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 20. Collins, G. S., Reitsma, J. B., Altman, D. G. & Moons, K. G. M. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): The TRIPOD statement. Ann. Intern. Med. 162 , 55–63 (2015). [ DOI ] [ PubMed ] [ Google Scholar ] 21. Collins, G. S. et al. Protocol for development of a reporting guideline (TRIPOD-AI) and risk of bias tool (PROBAST-AI) for diagnostic and prognostic prediction model studies based on artificial intelligence. BMJ Open 11 , e048008 (2021). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 22. Jr, D. W. H., Lemeshow, S. & Sturdivant, R. X. Applied Logistic Regression (Wiley, 2013). 23. Breiman, L. Random forests. Mach. Learn. 45 , 5–32 (2001). [ Google Scholar ] 24. Chen, T., Guestrin, C. & XGBoost: A scalable tree boosting system. in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining 785–794Association for Computing Machinery, New York, NY, USA, (2016). 10.1145/2939672.2939785 25. Riley, R. D. et al. Calculating the sample size required for developing a clinical prediction model. BMJ 368 , m441 (2020). [ DOI ] [ PubMed ] [ Google Scholar ] 26. Collins, G. S., Ogundimu, E. O. & Altman, D. G. Sample size considerations for the external validation of a multivariable prognostic model: A resampling study. Stat. Med. 35 , 214–226 (2016). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 27. Torrini, F. et al. Prediction of extubation outcome in critically ill patients: A systematic review and meta-analysis. Crit. Care 25 , 391 (2021). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 28. Asehnoune, K. et al. Extubation success prediction in a multicentric cohort of patients with severe brain injury. Anesthesiology 127 , 338–346 (2017). [ DOI ] [ PubMed ] [ Google Scholar ] 29. Namen, A. M. et al. Predictors of successful extubation in neurosurgical patients. Am. J. Respir. Crit. Care Med. 163 , 658–664 (2001). [ DOI ] [ PubMed ] [ Google Scholar ] 30. Thille, A. W. et al. Risk factors for and prediction by caregivers of extubation failure in ICU patients: A prospective study. Crit. Care Med. 43 , 613–620 (2015). [ DOI ] [ PubMed ] [ Google Scholar ] 31. El Solh, A. A., Bhat, A., Gunen, H. & Berbary, E. Extubation failure in the elderly. Respir. Med. 98 , 661–668 (2004). [ DOI ] [ PubMed ] [ Google Scholar ] 32. Lai, C.-C. et al. Establishing predictors for successfully planned endotracheal extubation. Medicine 95 , e4852 (2016). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 33. Lehmann, F. et al. Prolonged mechanical ventilation in patients with deep-seated intracerebral hemorrhage: Risk factors and clinical implications. J. Clin. Med. 10 , 1015 (2021). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 34. Cai, Y.-H., Wang, H.-T. & Zhou, J.-X. Perioperative predictors of extubation failure and the effect on clinical outcome after infratentorial craniotomy. Med. Sci. Monit. 22 , 2431–2438 (2016). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 35. Cheng, H. et al. Prolonged operative duration is associated with complications: A systematic review and meta-analysis. J. Surg. Res. 229 , 134–144 (2018). [ DOI ] [ PubMed ] [ Google Scholar ] 36. Boniatti, V. M. C. et al. The modified integrative weaning index as a predictor of extubation failure. Respir. Care 59 , 1042–1047 (2014). [ DOI ] [ PubMed ] [ Google Scholar ] 37. Clark, P. A., Inocencio, R. C. & Lettieri, C. J. I-TRACH: Validating a tool for predicting prolonged mechanical ventilation. J. Intensive Care Med. 33 , 567–573 (2018). [ DOI ] [ PubMed ] [ Google Scholar ] 38. Al-Ali, A. H. et al. Independent risk factors of failed extubation among adult critically ill patients: A prospective observational study from Saudi Arabia. Saudi J. Med. Med. Sci. 12 , 216–222 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 39. Chang, Y.-C. et al. Ventilator dependence risk score for the prediction of prolonged mechanical ventilation in patients who survive sepsis/septic shock with respiratory failure. Sci. Rep. 8 , 5650 (2018). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 40. Maier, I. L. et al. Predictive factors for the need of tracheostomy in patients with large vessel occlusion stroke being treated with mechanical thrombectomy. Front. Neurol. 12 , 728624 (2021). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 41. Young, D., Harrison, D. A., Cuthbertson, B. H., Rowan, K. & TracMan Collaborators, F. T. Effect of early vs late tracheostomy placement on survival in patients receiving mechanical ventilation: The TracMan randomized trial. JAMA 309 , 2121 (2013). [ DOI ] [ PubMed ] [ Google Scholar ] 42. Catalino, M. P. et al. Early versus late tracheostomy after decompressive craniectomy for stroke. J. Intensive Care 6 , 1 (2018). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 43. Chen, W. et al. Timing and outcomes of tracheostomy in patients with hemorrhagic stroke. World Neurosurg. 131 , e606–e613 (2019). [ DOI ] [ PubMed ] [ Google Scholar ] 44. Bösel, J. et al. Stroke-related early tracheostomy versus prolonged orotracheal intubation in neurocritical care trial (SETPOINT). Stroke 44 , 21–28 (2013). [ DOI ] [ PubMed ] [ Google Scholar ] 45. Bösel, J. et al. Effect of early vs standard approach to tracheostomy on functional outcome at 6 months among patients with severe stroke receiving mechanical ventilation: The SETPOINT2 randomized clinical trial. JAMA 327 , 1899–1909 (2022). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials Supplementary Material 1 (396.5KB, pdf) Data Availability Statement The datasets analyzed during the current study are not publicly available due to institutional data-use regulations and patient confidentiality policies. De-identified data and analytic code are available from the corresponding author upon reasonable request and approval by the participating institutions. Articles from Scientific Reports are provided here courtesy of Nature Publishing Group ACTIONS View on publisher site PDF (4.5 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top