ConceptioArchiveNCBI PubMed Central
NCBI PubMed Centralopen access

Machine-learning prediction and risk stratification of 12-month cognitive decline in Alzheimer's disease using routine clinical and MRI data.

Geng Y et al. · ncbi_pmc
NCBI PubMed Central · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cognitive psychology

Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Sci Rep . 2026 Mar 12;16:12227. doi: 10.1038/s41598-026-43321-1 Search in PMC Search in PubMed View in NLM Catalog Add to search Machine-learning prediction and risk stratification of 12-month cognitive decline in Alzheimer’s disease using routine clinical and MRI data Yinghui Geng Yinghui Geng 1 School of Nursing, Jinzhou Medical University, Jinzhou, Liaoning China Find articles by Yinghui Geng 1 , Huijun Zhang Huijun Zhang 2 Jinzhou Medical University, Jinzhou, Liaoning China Find articles by Huijun Zhang 2, ✉ Author information Article notes Copyright and License information 1 School of Nursing, Jinzhou Medical University, Jinzhou, Liaoning China 2 Jinzhou Medical University, Jinzhou, Liaoning China ✉ Corresponding author. Received 2025 Nov 7; Accepted 2026 Mar 3; Collection date 2026. © The Author(s) 2026 Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/ . PMC Copyright notice PMCID: PMC13076664  PMID: 41820490 Abstract Early identification of patients with Alzheimer’s disease (AD) who will experience near-term cognitive decline can support trial enrichment and risk-stratified follow-up. Using the Alzheimer’s Disease Neuroimaging Initiative (ADNI), we developed two prognostic models for 12-month Mini-Mental State Examination (MMSE) decrease (≥ 3 points): (i) a clinical logistic-regression model and (ii) a random-forest model combining clinical variables with MRI-derived volumetric measures. In 306 participants with baseline AD and complete 12-month MMSE (mean age 74.8 years; baseline MMSE 23.1), 131 (42.8%) declined. Five-fold stratified cross-validation with within-fold preprocessing and imputation was used for internal validation. The clinical model achieved an area under the ROC curve (AUC) of 0.755, while the random-forest model achieved an AUC of 0.773 and provided higher net benefit across threshold probabilities of 0.20–0.80 in decision-curve analysis. Risk stratification using pre-specified cut-offs (< 0.25, 0.25–0.50, ≥ 0.50) yielded monotonic observed decline rates (13.2%, 35.3%, 67.2%). These findings suggest that a transparent two-model framework based on ADNI data provides moderate prognostic accuracy and clinically interpretable three-tier risk stratification; however, external validation and local recalibration are required before clinical implementation. Supplementary Information The online version contains supplementary material available at 10.1038/s41598-026-43321-1. Keywords: Alzheimer’s disease, MMSE, machine learning, random forest, structural MRI, risk stratification Subject terms: Diseases, Medical research, Neurology, Neuroscience Introduction Alzheimer’s disease (AD) remains the most common cause of dementia in older adults. According to the Alzheimer’s Association 2024 report, an estimated 6.9 million U.S. adults aged ≥ 65 years were living with AD dementia in 2024. Without effective preventive or disease-modifying strategies, this number may reach 13–14 million by 2060, indicating a sustained rise in clinical and economic burden on memory clinics and long-term care systems 1 . At the same time, recent high-level reviews emphasise that AD is biologically and clinically heterogeneous—patients follow different atrophy patterns, carry different biomarker profiles, and progress at different speeds—which makes it difficult to identify, at baseline, those who will deteriorate rapidly enough to justify intensified monitoring or trial enrolment 2 . Longitudinal cohort analyses have shown that even within already diagnosed AD populations, some individuals experience a clinically meaningful decline over 12–18 months (for example, MMSE decrease ≥ 3 points), whereas others remain comparatively stable; such heterogeneity reduces statistical power in AD trials and complicates real-world follow-up planning 3 , 4 . A pragmatic solution is to develop prediction models that can flag, at baseline, those patients who are likely to decline over the next year, so that they can be prioritised for disease-modifying or non-pharmacological interventions, or simply be scheduled for closer cognitive/functional assessments 3 , 4 . The Alzheimer’s Disease Neuroimaging Initiative (ADNI) provides precisely the kind of multimodal, prospectively collected data that support such models. Several 2023–2024 studies using ADNI or ADNI-like cohorts have demonstrated that combining structural MRI, cognitive scales and other clinical features in machine-learning (ML) frameworks (gradient boosting, random forests, or deep multimodal networks) improves prediction of AD progression or conversion compared with single-modality models 3 – 5 , 8 . However, most of these studies focused on MCI-to-AD conversion or multi-year trajectories, and only a minority reported short-term (12-month) cognitive decline in patients who already had AD, or translated their model outputs into clinically interpretable risk strata. Moreover, many ML papers reported AUC only, without complementary metrics such as Brier score, calibration slope/intercept, or—most importantly—decision-curve analysis (DCA), even though current reporting guidance considers DCA essential for judging whether a new model offers net benefit over “treat all” or “treat none.” 9 , 10 . Methodological updates such as the 2024 TRIPOD + AI statement and the BMJ series on evaluation of clinical prediction models now explicitly recommend two things that are directly relevant to ADNI-based machine-learning work: (i) any complex or AI-based model should be compared with a simple, clinically plausible baseline model built from routinely available variables (e.g. age, sex, baseline MMSE, ADAS-Cog); and (ii) model performance should be presented together with decision-curve analysis (DCA) across the range of thresholds that are realistic for clinical decision-making 6 , 7 , 9 , 10 . Following this framework, our study—“Machine-learning-based prediction and risk stratification of 12-month cognitive decline in Alzheimer’s disease: an ADNI analysis with model comparison and decision-curve evaluation”—used ADNI participants with complete 12-month MMSE to: (1) build a transparent clinical baseline comparator; (2) develop and tune an RF model on the full candidate-predictor set; (3) compare discrimination, calibration and Brier score under five-fold cross-validation; and (4) apply DCA to identify probability thresholds (~ 0.20–0.80) with the highest net benefit and to derive a three-tier risk-stratification scheme for memory-clinic use. Methods Data source and study design We conducted a retrospective prognostic modelling study using the Alzheimer’s Disease Neuroimaging Initiative (ADNI) ADNIMERGE dataset. The baseline visit was defined as the first ADNI visit at which participants fulfilled the AD diagnosis. All candidate predictors were obtained at baseline, and the outcome was assessed 12 months later. The analysis followed recent guidance for reporting machine-learning prediction models (TRIPOD-AI) and for evaluating clinical prediction models. The ADNI study was approved by the Institutional Review Board of the University of Southern California (USC) and by the institutional review boards of all participating sites. All participants provided written informed consent. As this study used de-identified secondary data, no additional ethical approval was required. Participants From ADNIMERGE we first identified all individuals with a baseline diagnosis of AD and baseline cognitive, demographic and MRI-derived variables available ( n = 411; Table 1 ). Model development and validation were then restricted to participants who also had 12-month Mini-Mental State Examination (MMSE) data available ( n = 306); this analytic sample was used for all modelling steps and for Tables 2 , 3 and 4 ; Figs. 1 , 2 and 3 . Table 1. Baseline characteristics of participants with Alzheimer’s disease in the ADNI cohort ( n = 411). Variable n Mean SD Median Min Max Missing Age, years 411 74.75 7.94 75.3 55.1 90.9 0 MMSE at baseline 411 23.15 2.19 23.0 16.0 30.0 0 ADAS-Cog 13 at baseline 402 29.96 8.0 29.33 12.67 54.67 9 CDR-SB at baseline 411 4.43 1.69 4.5 1.0 10.0 0 Hippocampal volume, mm³ 338 5770.49 1025.37 5655.15 2991.0 9572.0 73 Intracranial volume, mm³ 402 1528337.54 224488.96 1494575.0 1071900.0 3315210.0 9 MMSE at 12 months 306 20.91 4.46 22.0 4.0 29.0 105 12-month MMSE change 306 -2.34 3.87 -2.0 -18.0 5.0 105 Open in a new tab Note: Continuous variables are shown as mean (SD); for skewed distributions, median (min, max) is also presented. “Missing” indicates the number of participants without data for the corresponding variable. MMSE at 12 months and MMSE change were available for 306/411 participants and analyses in Tables 2 , 3 and 4 were restricted to this subgroup. AD Alzheimer’s disease, ADAS-Cog 13 Alzheimer’s Disease Assessment Scale–Cognitive Subscale (13-item), CDR-SB Clinical Dementia Rating–Sum of Boxes, ICV intracranial volume, MMSE Mini-Mental State Examination. Table 2. Risk stratification of 12-month cognitive decline based on random-forest–predicted probabilities (ADNI participants with 12-month MMSE, n = 306). Risk stratum N Events, n Event rate, % Mean MMSE change, points Low (< 0.25) 68 9 13.2 0.04 Intermediate (0.25–0.50) 119 42 35.3 -1.62 High (≥ 0.50) 119 80 67.2 -4.42 Open in a new tab Note: Cognitive decline was defined as a decrease of ≥ 3 points in MMSE from baseline to 12 months. Predicted probabilities were obtained as out-of-fold calibrated probabilities from five-fold cross-validation. Risk groups were prespecified as low (< 0.25), intermediate (0.25–0.50), and high (≥ 0.50). Event rate = events / N. Mean MMSE change is shown for descriptive purposes only. MMSE Mini-Mental State Examination. Table 3. Threshold-specific performance of the random forest (RF) model for 12-month cognitive decline (ADNI participants with 12-month MMSE, n = 306). Threshold type Threshold Prevalence, % Sensitivity, % Specificity, % PPV, % NPV, % Accuracy, % F1, % Balanced accuracy, % Youden’s J Best Youden’s J 0.40 42.8 78.6 65.1 62.8 80.3 70.9 69.8 71.9 0.438 Preset 0.25 0.25 42.8 93.1 33.7 51.3 86.8 59.2 66.1 63.4 0.268 Preset 0.5 0.50 42.8 61.1 77.7 67.2 72.7 70.6 64.0 69.4 0.388 Open in a new tab Note: Cognitive decline was defined as an MMSE decrease of ≥ 3 points from baseline to 12 months. “Best Youden’s J” denotes the probability threshold that maximized (sensitivity + specificity − 1). PPV and NPV were influenced by the observed prevalence of cognitive decline in this sample (42.8%). Accuracy = (TP + TN) / total; F1 = 2 × (precision × recall) / (precision + recall); balanced accuracy = (sensitivity + specificity) / 2. FN false negative, FP false positive, MMSE Mini-Mental State Examination, NPV negative predictive value, PPV positive predictive value, RF random forest; TN true negative, TP true positive. Table 4. Comparative performance of baseline and machine-learning models for predicting 12-month cognitive decline (MMSE decrease ≥ 3 points) in ADNI participants ( n = 306). Model Predictors included AUC Brier score Calibration slope Calibration intercept Net benefit (0.20–0.80) Model 1 (baseline clinical) Age, sex, baseline MMSE, ADAS-Cog 13 0.755 0.199 0.140 0.474 Low, mainly at ≤ 0.30 Model 2 (RF, main model) All candidate predictors, tuned random forest 0.773 0.192 0.237 0.511 Highest across 0.20–0.80 Open in a new tab Note: Both models used the same analytic sample (ADNI participants with complete 12-month MMSE, n = 306) and five-fold cross-validation with out-of-fold predictions. Calibration slope and intercept were estimated from cross-validated predictions; a calibration slope < 1 suggests overfitting, while the intercept reflects calibration-in-the-large. The RF model included MRI-derived volumes (including ventricular volume and ICV), which were additionally explored in sensitivity analyses using ventricles/ICV (Supplementary Table S5). Bootstrap 95% confidence intervals for discrimination and calibration metrics are provided in Supplementary Table S6. ADNI Alzheimer’s Disease Neuroimaging Initiative, MMSE Mini-Mental State Examination, RF random forest. Fig. 1. Open in a new tab Receiver operating characteristic (ROC) curve of the random forest (RF) model for predicting 12-month cognitive decline in patients with Alzheimer’s disease (AD). Fig. 2. Open in a new tab Calibration plot of the RF model for predicting 12-month cognitive decline in patients with AD (MMSE decrease ≥ 3 points). Fig. 3. Open in a new tab Decision-curve analysis (DCA) of the RF model for 12-month cognitive decline in patients with AD. Outcome definition The primary outcome was 12-month cognitive decline, defined a priori as a decrease of ≥ 3 points in MMSE from baseline to the 12-month visit. This threshold has been used in AD clinical studies to indicate a clinically meaningful decline and was consistent with the observed event rate in this ADNI sample (42.8%). Candidate predictors Candidate predictors were selected to reflect variables routinely collected in ADNI memory-clinic settings and those shown to be important in previous ADNI machine-learning studies: Demographics/genetics: age, sex, years of education, APOE ε4 carrier (≥ 1 allele; derived from APOE4 allele count in ADNIMERGE); Baseline cognition and clinical severity: baseline MMSE, ADAS-Cog 13, ADAS-Cog 11, ADAS-Q4, CDR-Sum of Boxes (CDR-SB), and Functional Activities Questionnaire (FAQ); ADAS-Cog 11 and ADAS-Cog 13 are expected to be highly correlated; we retained both to capture overlapping but non-identical scoring ranges used in routine assessments, noting that tree-based models can accommodate correlated predictors, and we interpret feature importance with redundancy in mind. MRI-derived structural measures: hippocampal volume, ventricular volume, whole-brain volume, entorhinal cortex volume, fusiform volume, middle temporal gyrus volume, and intracranial volume (ICV) (see Supplementary Methods for details); Other care-related or potentially modifiable factors that were repeatedly selected in the RF model were explored in sensitivity analyses (see Supplementary Methods). Data preprocessing All variables were inspected for missingness; missingness percentages for each candidate predictor in the analytic sample are summarized in Supplementary Table S4. Within each cross-validation training fold, continuous variables with missing values were imputed using the median and categorical variables were imputed using the most frequent category (single imputation). This pragmatic approach may introduce bias and underestimate uncertainty, particularly for MRI-derived variables with higher missingness (Supplementary Table S4); alternative strategies (e.g., multiple imputation) may yield different estimates and should be evaluated in future work. We did not formally test missingness mechanisms (e.g., MCAR/MAR/MNAR) or perform multiple imputation; given the modest sample size and the goal of a reproducible modelling pipeline, we used simple within-fold imputation and report missingness transparently. Continuous predictors were then standardised to zero mean and unit variance. Categorical predictors (e.g. sex, APOE4) were one-hot encoded. All preprocessing steps were performed within each cross-validation fold to avoid data leakage. MRI-derived volumetric measures were taken as provided in ADNIMERGE (centrally processed with standardized pipelines); we did not apply additional site/scanner harmonization methods (e.g., ComBat). Model specification and comparators To make the incremental value of machine learning explicit, we compared two prespecified models, both trained and evaluated on the same analytic sample ( n = 306): Model 1 (baseline clinical model): logistic regression including age, sex, baseline MMSE and ADAS-Cog 13; presented in Table 4 . Model 2 (machine-learning model): tuned RF using all candidate predictors (clinical variables and MRI-derived brain volumes). This model was used to generate the ROC, calibration and decision-curve plots (Figs. 1 , 2 and 3 ) and is presented in Table 4 . Key tuned random-forest hyperparameters are reported in Supplementary Table S9 (n_estimators = 200, max_features=sqrt, max_depth=None; full list in Table S9). Sensitivity analysis: because ventricular volume can be influenced by inter-individual head size, we repeated the main analyses replacing absolute ventricular volume and ICV with the ventricles/ICV ratio. Findings were consistent; details are provided in the Supplementary Methods and Supplementary Table S5. Internal validation and performance metrics We used five-fold stratified cross-validation. Given the moderate event rate (42.8%), we used outcome-stratified folds and did not apply additional imbalance techniques (e.g., SMOTE or class-weighting). In each fold, models were trained on 80% of the data and tested on the remaining 20%; all out-of-fold predictions were concatenated to obtain overall out-of-fold probabilities. Model performance was summarized as: (i) area under the ROC curve (AUC) for discrimination (Fig. 1 ); (ii) Brier score for overall accuracy of predicted probabilities (Table 4 ); (iii) calibration slope and intercept estimated from out-of-fold probabilities by fitting a logistic calibration model regressing the observed outcome on the logit of the out-of-fold predicted probabilities (Fig. 2 ; Table 4 ); and (iv) threshold-specific sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV) and Youden’s J at probability thresholds of 0.25, 0.40 and 0.50 (Table 3 ). Within each cross-validation training fold, we fitted an isotonic regression calibrator on the training predictions and applied it to the validation fold to obtain out-of-fold calibrated probabilities. Decision-curve analysis To assess potential clinical utility, we performed decision-curve analysis (DCA) 9 , 10 on the out-of-fold calibrated probabilities of both models. Net benefit was calculated across probability thresholds from 0.20 to 0.80, a range relevant to risk-stratified follow-up in memory-clinic settings. Thresholds below 0.20 or above 0.80 were not emphasised because they imply very low or very high risk tolerance, where net benefit becomes unstable and decisions are less plausible in routine memory-clinic workflows. Risk stratification and clinical use case Based on decision-curve analysis (DCA) and the distribution of predicted probabilities from the random forest model, we defined three risk strata low (< 0.25), intermediate (0.25–0.50) and high (≥ 0.50). For each stratum we calculated the number of patients, the observed 12-month decline rate and the mean MMSE change (Table 2 ). We further present an illustrative clinical application scenario: low-risk patients could be considered for routine annual assessment; intermediate-risk patients for closer monitoring (e.g., 3–6 months) and optimisation of modifiable factors; and high-risk patients for closer follow-up and further clinical assessment, subject to local resources and external validation. In the RF model, observed decline rates increased monotonically across strata: 13.2% (9/68) in low-risk, 35.3% (42/119) in intermediate-risk, and 67.2% (80/119) in high-risk groups. Median predicted risks were 0.161, 0.370, and 0.644, respectively (IQRs 0.126–0.209; 0.305–0.429; 0.553–0.709). Software All data management and modelling were performed in Python 3.13 (pandas 2.x, scikit-learn 1.4) and R 4.3 (for calibration and decision-curve plotting). Cross-validation and model tuning were scripted to ensure reproducibility. A large language model tool (ChatGPT, OpenAI) was used to assist with language editing only; all scientific content, analyses, and interpretations were performed and verified by the authors. Model interpretability was assessed using SHAP (SHapley Additive exPlanations). SHAP values were computed from a final random forest model refitted on the full analytic dataset using the tuned hyperparameters. These analyses were conducted post hoc and were not used in model training, feature selection, or performance evaluation. Results Participants and baseline characteristics A total of 411 ADNI participants met the AD diagnosis and had baseline cognitive, demographic and MRI-derived variables recorded (mean age 74.8 ± 7.9 years; baseline MMSE 23.15 ± 2.19). Of these, 306/411 (74.5%) also completed the 12-month MMSE assessment, and all analyses, tables (Tables 2 , 3 and 4 ) and figures (Figs. 1 , 2 and 3 ) were based on this analytic sample ( n = 306). Additional baseline categorical characteristics are summarised in Supplementary Table S2. Primary outcome At 12 months, 131/306 participants met the predefined endpoint of MMSE decrease ≥ 3 points, giving an event rate of 42.8%. This event rate was used in the threshold-based operating-characteristics analysis (Table 3 ) and in the decision-curve analysis. Model performance in cross-validation Both prespecified models were trained and evaluated using five-fold stratified cross-validation on the same 306 participants. The baseline clinical model (age, sex, baseline MMSE, ADAS-Cog 13) achieved an AUC of 0.755 and a Brier score of 0.199. The RF model achieved an AUC of 0.773 based on out-of-fold predictions from five-fold stratified cross-validation with a Brier score of 0.192. The absolute gain in discrimination over the baseline clinical model was modest, so decision-curve analysis is emphasized as the primary evidence for added clinical value. Calibration slopes remained below 1 for both models (Table 4 ), indicating some overfitting and underscoring the need for local recalibration when transported. Decision-curve analysis showed that the RF model provided higher net benefit than the “treat-all” and “treat-none” strategies across threshold probabilities of 0.20–0.80 (Fig. 3 ); ROC and calibration plots are shown in Figs. 1 and 2 . A sensitivity analysis adjusting ventricular volume for intracranial volume (ventricles/ICV) produced consistent results (Supplementary Table S5). To provide a clinically interpretable summary, we examined the out-of-fold predictions at the probability threshold that maximised Youden’s J (≈ 0.40). At this cut-off the confusion matrix was TP 103, FP 61, TN 114, FN 28, corresponding to sensitivity 78.6% and specificity 65.1%, in line with Table 3 . This cut-off classified 164/306 participants (53.6%) as test-positive (predicted probability ≥ 0.40), capturing 103/131 decliners (78.6%) while yielding 61/175 false positives; PPV was 62.8% and NPV was 80.3%. At a threshold probability of 0.40 in DCA, the RF model achieved a net benefit of 0.204, equivalent to 20.4 net true-positive decisions per 100 patients versus treating none and a net reduction of 23.5 unnecessary interventions per 100 patients compared with treating all. The corresponding confusion matrix is provided in Supplementary Table S1 . Threshold-based discrimination Table 3 reports operating characteristics at prespecified thresholds of 0.25 and 0.50 and at the best Youden’s J (0.40). At 0.25, sensitivity was very high (93.1%) but specificity was low (33.7%), which is suitable for rule-out or early recall. At 0.50, specificity increased to 77.7% while sensitivity dropped to 61.1%, which is suitable for rule-in or intensified follow-up. The best Youden’s J was 0.438 at 0.40, with balanced accuracy 71.9% and overall accuracy 70.9%. Additional threshold-specific metrics are provided in Supplementary Table S3, and a clinician-oriented interpretation of key cutoffs is summarized in Supplementary Table S7. Risk stratification When RF-predicted probabilities were grouped into the three prespecified strata, 68 (22.2%) patients were low risk (< 0.25), 119 (38.9%) were intermediate risk (0.25–0.50), and 119 (38.9%) were high risk (≥ 0.50). Observed 12-month cognitive-decline rates increased monotonically from 13.2% (low) to 35.3% (intermediate) and 67.2% (high); mean MMSE change also followed this gradient (≈ 0, − 1.6, − 4.4 points, respectively). This confirms that RF-generated probabilities can be translated into clinically interpretable tiers (Table 2 ). The distribution of participants across these strata is shown in Supplementary Figure S3. Observed decline rates by stratum are shown in Supplementary Figure S4. Clinical utility (decision-curve analysis) Decision-curve analysis based on out-of-fold calibrated probabilities showed that the RF model provided net benefit across a wide range of threshold probabilities (approximately 0.20–0.80) compared with the “treat-all” and “treat-none” strategies (Fig. 3 ). In practical memory-clinic ranges (≈ 0.20–0.50), using the RF model could identify more true decliners without generating excessive false positives. Net benefit at representative decision thresholds is reported in Supplementary Table S8. The model was developed with 5-fold cross-validation and evaluated on out-of-fold predictions among participants with complete 12-month MMSE data ( n = 306) from the ADNI cohort. Cognitive decline was defined as an MMSE decrease ≥ 3 points. The diagonal line represents a non-informative classifier. Predicted probabilities were grouped into deciles, and the observed event rates within each decile were plotted against the corresponding predicted risks ( n = 306). The dashed 45° line indicates perfect calibration, and the solid line shows the model-estimated calibration. The RF model (solid line) was compared with the “treat-all” and “treat-none” strategies using out-of-fold calibrated probabilities. Within the clinically relevant threshold range of 0.20–0.80, the RF model provided higher net benefit, supporting its use for risk-stratified follow-up. Discussion Summary of key findings In this ADNI-based prognostic study, we developed a tuned random-forest (RF) model to predict clinically meaningful 12-month cognitive decline (MMSE decrease ≥ 3 points) in patients with Alzheimer’s disease and benchmarked it against a transparent baseline clinical logistic-regression model. The RF model achieved moderate discrimination (out-of-fold AUC ~ 0.77) with acceptable Brier score and showed positive net benefit across clinically relevant thresholds on decision-curve analysis 11 . When translated into three prespecified risk strata, 12-month decline rates increased monotonically (~ 13% to ~ 35% to ~ 67%), suggesting potential utility for trial enrichment and risk stratification, pending external validation 12 , 13 . Independent external validation and local recalibration are required before clinical deployment. Relation to recent literature Our findings align with recent ADNI-anchored and multimodal studies showing that combinations of routine cognitive measures and structural MRI can support short-term prognosis in AD and AD-spectrum cohorts, and that longitudinal modelling of ADNI participants enables more precise characterization of individual progression trajectories 12 , 14 – 16 . Although more complex deep/ensemble approaches that integrate PET, CSF, or validated plasma markers such as p-tau217 can further improve discrimination 17 , 18 , our results support the view that MRI+clinical pipelines remain clinically meaningful and far more deployable in settings where advanced biomarkers are not routinely available. The current ADNI programme and its prospective extensions continue to provide the methodological and data infrastructure for such ML-based prognostic studies 16 , 19 . Model performance, calibration and decision-analytic value In line with contemporary recommendations on evaluating ML-based prediction models, we benchmarked complex models against parsimonious clinical baselines and reported discrimination, Brier score, calibration slope/intercept, and DCA 6 , 7 , 20 . The fact that the RF model maintained net benefit over a wide range of risk thresholds underscores that decision-analytic reporting—not accuracy alone—is essential when a model is intended to guide follow-up intensity, biomarker testing, or trial screening 11 . This approach is consistent with recent BMJ guidance on the development–validation continuum for clinical prediction models 20 . Calibration slopes substantially below 1 indicate potential overfitting and overly extreme risk predictions. Therefore, model outputs should be interpreted with caution, and recalibration is essential before any clinical application. Predictors, interpretability and biological plausibility Model explainability using SHAP (Supplementary Figures S1 –S2) highlighted baseline cognitive severity—especially ADAS-Cog components and global MMSE—together with medial-temporal and ventricular volumetrics as dominant risk drivers 21 . These attributions are biologically plausible given ADNI syntheses linking regional neurodegeneration, global atrophy and near-term clinical worsening 14 , and they are also consistent with emerging biomarker evidence: recent plasma p-tau217 work shows that blood-based markers can flag early amyloid/tau pathology in clinically unimpaired individuals 17 , while earlier studies demonstrated that p-tau217 discriminates AD from other neurodegenerative disorders 18 . Presenting patient-level feature-contribution plots can therefore bridge statistical output and clinician–patient communication, in line with broader work on explainable AI for medical imaging and neurodegenerative disease 22 . Consistent with clinical expectations, the most influential predictors included baseline ADAS-Cog 13, age, and MRI-derived volumetric measures (e.g., ICV, ventricles, hippocampal and temporal lobe volumes). Risk stratification: trial enrichment and risk-stratified follow-up Translating continuous probabilities into low-, intermediate- and high-risk tiers produced clearly separated 12-month decline rates (for example, ≈ 13% → ≈35% → ≈67%), supporting two potential use-cases: (i) prognostic enrichment to raise event rates and reduce follow-up burden in disease-modifying or proof-of-concept trials; and (ii) risk-stratified follow-up to prioritize confirmatory biomarkers or closer monitoring for the highest-risk group 12 , 13 . This mirrors recent deep-learning–based patient-stratification work that seeks to optimize clinical dementia trials by oversampling those most likely to decline 13 . Strengths Key strengths of this study include: the use of a harmonized multimodal cohort (ADNI); head-to-head comparison of ML algorithms under stratified cross-validation with probability recalibration; explicit reporting of decision-analytic results; and the integration of explainability to improve clinician trust. These elements are in line with current expectations for transparent, evaluable clinical prediction models and with the movement toward end-to-end clinical-AI pipelines 20 , 23 . Limitations and equity considerations ADNI participants are typically research volunteers with relatively high education levels and limited ethnic diversity; this may restrict transportability to community, hospital or multiethnic cohorts, making independent, locally representative external validation and local recalibration essential 14 , 16 , 19 . ADNI is a multi-site study; residual scanner/site differences can affect volumetric predictors, and we did not apply additional harmonization. The MMSE-based decline definition may be influenced by test–retest variability and ceiling/floor effects; future work should evaluate robustness using alternative or complementary endpoints (e.g., ADAS-Cog or CDR-SB change), repeated assessments, and sensitivity analyses around the decline threshold. We used single imputation (median/mode) within training folds, which may introduce bias and underestimate uncertainty, particularly for MRI-derived predictors with non-trivial missingness (Supplementary Table S4). Our feature set intentionally emphasized routinely available clinical and MRI variables to enhance feasibility, which will likely cap discrimination compared with models that incorporate PET, CSF or plasma p-tau217 17 , 18 . We also did not include white matter hyperintensity burden or a broader set of subcortical volumes, which may contribute to cognitive decline and may interact with grey matter atrophy; these features should be explored in future studies when consistently available and during external validation. Moreover, ML models for AD progression can exhibit performance and calibration differences across subgroups defined by race/ethnicity, education or language, so subgroup-specific auditing, reporting and recalibration (e.g., intercept/slope adjustment, Platt scaling or isotonic regression) should be planned 24 – 26 . These issues become particularly salient when such models are translated to health-care systems that rely on culture-fair cognitive tools or cross-culturally adapted neuropsychological assessment 25 , 26 . Practical deployment considerations For real-world adoption, we recommend: (a) prespecifying calibration-monitoring and model-updating triggers to handle dataset shift; (b) embedding clinician-facing explanations (top features and directionality) in the interface; (c) mapping each risk tier to clear standard operating procedures (follow-up interval, confirmatory biomarker strategy, trial referral); and (d) conducting prospective or pragmatic impact evaluations within an end-to-end clinical-AI framework such as SALIENT to quantify effects on workflow, patient-centred outcomes and costs 22 , 23 . Clinical implications With appropriate external validation and local recalibration, a 12-month decline predictor can target patients for prioritized biomarker testing or intensified monitoring, support shared decision-making using calibrated probabilities plus the dominant contributors for a given patient, and assist trial teams in enriching for likely decliners within typical trial windows 12 , 13 , 17 . Where MRI or advanced biomarkers are limited, simplified cognitive-centric workflows informed by the same modelling principles may still offer actionable stratification, provided local data are used for recalibration 16 , 19 . Future directions Future work should prioritize: (1) multi-region, multiethnic external validation in both clinical and community cohorts; (2) incremental improvement via validated blood biomarkers (for example, p-tau217) and PET where feasible; (3) integration of longitudinal digital phenotypes from wearable devices and other remote-monitoring technologies, as demonstrated in recent ML studies of MCI trials 27 ; (4) pragmatic or stepped-wedge impact trials of model-guided care pathways; and (5) embedding continuous fairness monitoring and governance within implementation frameworks to ensure safe, equitable translation 23 – 26 . Conclusions A transparent ML workflow using routinely collected cognitive and structural MRI variables provided moderate discrimination and positive decision-analytic value for predicting 12-month MMSE decline in AD. Translating calibrated probabilities into three risk tiers offers a practical route for trial enrichment and risk-stratified care in settings without routine PET/CSF. Calibration required and remained imperfect in internal validation, and external/cross-cultural validation with fairness auditing and prospective impact evaluation are needed before routine clinical use. Supplementary Information Below is the link to the electronic supplementary material. Supplementary Material 1 (780.9KB, pdf) Acknowledgements Data collection and sharing for the ADNI project were funded by the Alzheimer’s Disease Neuroimaging Initiative (ADNI) (National Institutes of Health, U.S.) and by the ADNI private sector contributors. The investigators within ADNI contributed to the design and implementation of ADNI and/or provided data but did not participate in the analysis or the writing of this manuscript. We thank the ADNI participants and study staff for their time and commitment.We used a large language model–based tool (ChatGPT, OpenAI) to assist with language editing of the manuscript. All content was reviewed and approved by the authors. Abbreviations AD Alzheimer’s disease ADNI Alzheimer’s Disease Neuroimaging Initiative ADAS-Cog 13 Alzheimer’s Disease Assessment Scale – Cognitive Subscale (13-item) AUC Area under the receiver operating characteristic curve CDR-SB Clinical Dementia Rating – Sum of Boxes CSF Cerebrospinal fluid DCA Decision-curve analysis ICV Intracranial volume MCI Mild cognitive impairment ML Machine learning MMSE Mini-Mental State Examination MRI Magnetic resonance imaging NPV Negative predictive value OOF Out-of-fold PET Positron emission tomography PPV Positive predictive value RF Random forest ROC Receiver operating characteristic Author contributions Y.G. conceived the study, designed the analysis plan, performed data curation and statistical/machine-learning analyses, prepared the figures and drafted the manuscript. H.Z. contributed to study conceptualisation and methodology, provided clinical expertise and supervision, and critically revised the manuscript for important intellectual content. Both authors read and approved the final version of the manuscript. Data availability The datasets analysed in this study were obtained from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) database ( https://adni.loni.usc.edu ) upon application and data use agreement. The ADNIMERGE dataset used in this analysis is available to qualified researchers from the ADNI repository. Analysis scripts and derived outputs are available from the corresponding author on reasonable request. Code availability All analysis scripts required to reproduce the reported results are publicly available in the following repository: https://github.com/gengyinghui/ADNI-12M-Cognitive-Decline-ML . Declarations Competing interests The authors declare no competing interests. Ethics approval and consent to participate This study is a secondary analysis of de-identified data from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) database. The ADNI study was approved by the Institutional Review Board of the University of Southern California (USC) and by the institutional review boards of all participating sites. All procedures were conducted in accordance with the Declaration of Helsinki, Good Clinical Practice guidelines, and U.S. 21 CFR Part 50 and Part 56. Written informed consent was obtained from all participants at each site. As this study used de-identified secondary data with no direct interaction with participants, no additional ethical approval was required for the present study. Further details on ADNI ethics approvals are available at https://adni.loni.usc.edu . Footnotes Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. References 1. Alzheimer’s, A. Alzheimer’s disease facts and figures. Alzheimers Dement. 20 (5), 3708–3821. 10.1002/alz.13809 (2024). [ DOI ] [ PMC free article ] [ PubMed ] 2. Zhang, J. et al. Recent advances in Alzheimer’s disease: Mechanisms, clinical trials and new drug development strategies. Signal. Transduct. Target. Ther. 9 (1), 211. 10.1038/s41392-024-01911-3 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 3. Zhang, S. et al. Alzheimer’s Disease Neuroimaging Initiative and the Australian Imaging Biomarkers and Lifestyle Study of Aging. Machine learning on longitudinal multi-modal data enables the understanding and prognosis of Alzheimer’s disease progression. iScience 27 (7), 110263. 10.1016/j.isci.2024.110263 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 4. Wang, Y. et al. Alzheimer’s Disease Neuroimaging Initiative. Predicting long-term progression of Alzheimer’s disease using a multimodal deep learning model incorporating interaction effects. J. Transl Med. 22 (1), 265. 10.1186/s12967-024-05025-w (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 5. Lee, M. W. et al. A multimodal machine learning model for predicting dementia conversion in Alzheimer’s disease. Sci. Rep. 14 (1), 12276. 10.1038/s41598-024-60134-2 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 6. Collins, G. S. et al. TRIPOD + AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 385 , e078378. Erratum in: BMJ 385 , q902. 10.1136/bmj-2023-078378. 10.1136/bmj.q902 (2024). [ DOI ] [ PMC free article ] [ PubMed ] 7. Collins, G. S. et al. Evaluation of clinical prediction models (part 1): from development to external validation. BMJ 384 , e074819. 10.1136/bmj-2023-074819 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 8. Wang, C. et al. A multimodal deep learning approach for the prediction of cognitive decline and its effectiveness in clinical trials for Alzheimer’s disease. Transl Psychiatry . 14 (1), 105. 10.1038/s41398-024-02819-w (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 9. Zhao, L. et al. Understanding decision curve analysis in clinical prediction model research. Postgrad. Med. J. 100 (1185), 512–515. 10.1093/postmj/qgae027 (2024). [ DOI ] [ PubMed ] [ Google Scholar ] 10. Vickers, A. J., Van Calster, B., Wynants, L. & Steyerberg, E. W. Decision curve analysis: confidence intervals and hypothesis testing for net benefit. Diagn Progn Res. 7:11. 10.1186/s41512-023-00148-y (2023). Correction in: Diagn Progn Res. 9 (1), 8. 10.1186/s41512-025-00188-6 (2025). [ DOI ] [ PMC free article ] [ PubMed ] 11. Piovani, D., Sokou, R., Tsantes, A. G., Vitello, A. S. & Bonovas, S. Optimizing Clinical Decision Making with Decision Curve Analysis: Insights for Clinical Investigators. Healthc. (Basel) . 11 (16), 2244. 10.3390/healthcare11162244 (2023). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 12. Devanarayan, V. et al. Alzheimer’s Disease Neuroimaging Initiative (ADNI). Predicting clinical progression trajectories of early Alzheimer’s disease patients. Alzheimers Dement. 20 (3), 1725–1738. 10.1002/alz.13565 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 13. Birkenbihl, C., de Jong, J., Yalchyk, I. & Fröhlich, H. Deep learning-based patient stratification for prognostic enrichment of clinical dementia trials. Brain Commun. 6 (6), fcae445. 10.1093/braincomms/fcae445 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 14. Veitch, D. P. et al. Alzheimer’s Disease Neuroimaging Initiative. The Alzheimer’s Disease Neuroimaging Initiative in the era of Alzheimer’s disease treatment: A review of ADNI studies from 2021 to 2022. Alzheimers Dement. 20 (1), 652–694. 10.1002/alz.13449 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 15. Maheux, E. et al. Forecasting individual progression trajectories in Alzheimer’s disease. Nat. Commun. 14 (1), 761. 10.1038/s41467-022-35712-5 (2023). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 16. Weiner, M. W. et al. Alzheimer’s Disease Neuroimaging Initiative. Overview of Alzheimer’s Disease Neuroimaging Initiative and future clinical trials. Alzheimers Dement. 21 (1), e14321. 10.1002/alz.14321 (2025). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 17. Janelidze, S. et al. Plasma Phosphorylated Tau 217 and Aβ42/40 to Predict Early Brain Aβ Accumulation in People Without Cognitive Impairment. JAMA Neurol. 81 (9), 947–957. 10.1001/jamaneurol.2024.2619 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 18. Palmqvist, S. et al. Discriminative Accuracy of Plasma Phospho-tau217 for Alzheimer Disease vs Other Neurodegenerative Disorders. JAMA 324 (8), 772–781. 10.1001/jama.2020.12134 (2020). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 19. Toga, A. W., Neu, S., Sheehan, S. T. & Crawford, K. Alzheimer’s Disease Neuroimaging Initiative. The informatics of ADNI. Alzheimers Dement. 20 (10), 7320–7330. 10.1002/alz.14099 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 20. Riley, R. D. et al. Evaluation of clinical prediction models (part 2): how to undertake an external validation study. BMJ 384 , e074820. 10.1136/bmj-2023-074820 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 21. Lundberg, S. M. & Lee, S-I. A unified approach to interpreting model predictions. Adv. Neural Inf. Process. Syst. (NeurIPS) . 30 , 4765–4774 (2017). [ Google Scholar ] 22. Muhammad, D. & Bendechache, M. Unveiling the black box: A systematic review of Explainable Artificial Intelligence in medical image analysis. Comput. Struct. Biotechnol. J. 24 , 542–560. 10.1016/j.csbj.2024.08.005 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 23. van der Vegt, A. H. et al. Implementation frameworks for end-to-end clinical AI: derivation of the SALIENT framework. J. Am. Med. Inf. Assoc. 30 (9), 1503–1515. 10.1093/jamia/ocad088 (2023). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 24. Yuan, C., Linn, K. A. & Hubbard, R. A. Algorithmic Fairness of Machine Learning Models for Alzheimer Disease Progression. JAMA Netw. Open. 6 (11), e2342203. 10.1001/jamanetworkopen.2023.42203 (2023). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 25. Chithiramohan, T. et al. Culture-Fair Cognitive Screening Tools for Assessment of Cognitive Impairment: A Systematic Review. J. Alzheimers Dis. Rep. 8 (1), 289–306. 10.3233/ADR-230194 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 26. Gkintoni, E. & Nikolaou, G. The Cross-Cultural Validation of Neuropsychological Assessments and Their Clinical Applications in Cognitive Behavioral Therapy: A Scoping Analysis. Int. J. Environ. Res. Public. Health . 21 (8), 1110. 10.3390/ijerph21081110 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 27. Rykov, Y. G. et al. Predicting cognitive scores from wearable-based digital physiological features using machine learning: data from a clinical trial in mild cognitive impairment. BMC Med. 22 (1), 36. 10.1186/s12916-024-03252-y (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials Supplementary Material 1 (780.9KB, pdf) Data Availability Statement The datasets analysed in this study were obtained from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) database ( https://adni.loni.usc.edu ) upon application and data use agreement. The ADNIMERGE dataset used in this analysis is available to qualified researchers from the ADNI repository. Analysis scripts and derived outputs are available from the corresponding author on reasonable request. All analysis scripts required to reproduce the reported results are publicly available in the following repository: https://github.com/gengyinghui/ADNI-12M-Cognitive-Decline-ML . Articles from Scientific Reports are provided here courtesy of Nature Publishing Group ACTIONS View on publisher site PDF (1.5 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top

Record · ID 13486 · SHA-256 911e50b6b737b70f
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.