ConceptioArchiveNCBI PubMed Central
NCBI PubMed Centralopen access

Multimodal MRI-based two-stage artificial intelligence framework for renal fibrosis classification in chronic kidney disease.

Li X et al. · ncbi_pmc
NCBI PubMed Central · Papers · License: Open Access
Open Source ↗Direct PDF ↓
machine learning systems

Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice BMC Nephrol . 2026 Mar 4;27:229. doi: 10.1186/s12882-026-04869-2 Search in PMC Search in PubMed View in NLM Catalog Add to search Multimodal MRI-based two-stage artificial intelligence framework for renal fibrosis classification in chronic kidney disease Xiaojing Li Xiaojing Li 1 Department of Radiology, The Second Affiliated Hospital of Soochow University, 1055 Sanxiang Road, Suzhou, Jiangsu 215004 China Find articles by Xiaojing Li 1, # , Yirui Li Yirui Li 1 Department of Radiology, The Second Affiliated Hospital of Soochow University, 1055 Sanxiang Road, Suzhou, Jiangsu 215004 China Find articles by Yirui Li 1, # , Qing Ma Qing Ma 1 Department of Radiology, The Second Affiliated Hospital of Soochow University, 1055 Sanxiang Road, Suzhou, Jiangsu 215004 China Find articles by Qing Ma 1, # , Yilin Xu Yilin Xu 2 Department of Nephrology, The Second Affiliated Hospital of Soochow University, Suzhou, 215004 China Find articles by Yilin Xu 2 , Ye Zhu Ye Zhu 2 Department of Nephrology, The Second Affiliated Hospital of Soochow University, Suzhou, 215004 China Find articles by Ye Zhu 2 , Jing Zhang Jing Zhang 3 MR Research Collaboration Team, Siemens Healthineers Ltd. 399 Haiyang West Road, Shanghai, 200126 China Find articles by Jing Zhang 3 , Junkang Shen Junkang Shen 1 Department of Radiology, The Second Affiliated Hospital of Soochow University, 1055 Sanxiang Road, Suzhou, Jiangsu 215004 China Find articles by Junkang Shen 1 , Wu Cai Wu Cai 1 Department of Radiology, The Second Affiliated Hospital of Soochow University, 1055 Sanxiang Road, Suzhou, Jiangsu 215004 China Find articles by Wu Cai 1 , Chaogang Wei Chaogang Wei 1 Department of Radiology, The Second Affiliated Hospital of Soochow University, 1055 Sanxiang Road, Suzhou, Jiangsu 215004 China Find articles by Chaogang Wei 1, ✉ , Zhen Jiang Zhen Jiang 1 Department of Radiology, The Second Affiliated Hospital of Soochow University, 1055 Sanxiang Road, Suzhou, Jiangsu 215004 China Find articles by Zhen Jiang 1, ✉ Author information Article notes Copyright and License information 1 Department of Radiology, The Second Affiliated Hospital of Soochow University, 1055 Sanxiang Road, Suzhou, Jiangsu 215004 China 2 Department of Nephrology, The Second Affiliated Hospital of Soochow University, Suzhou, 215004 China 3 MR Research Collaboration Team, Siemens Healthineers Ltd. 399 Haiyang West Road, Shanghai, 200126 China ✉ Corresponding author. # Contributed equally. Received 2025 Aug 22; Accepted 2026 Feb 24; Collection date 2026. © The Author(s) 2026 Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/ . PMC Copyright notice PMCID: PMC13067705  PMID: 41782094 Abstract Background Renal fibrosis (RF) is a key pathological hallmark and prognostic indicator of chronic kidney disease (CKD). Accurate evaluation of RF is critical for risk stratification and therapeutic decision-making, yet the current assessment relies mainly on renal biopsy, which has several limitations. This study aimed to develop a two-stage artificial intelligence framework that integrates multimodal MRI (native T 1 mapping, ADC, and T 2 * mapping) and clinical indicators for the noninvasive assessment of RF in patients with CKD. Methods This prospective study included 152 patients with biopsy-proven CKD (RF 1: no RF, 34 patients; RF 2: mild RF, 69 patients; and RF 3: moderate to severe RF, 49 patients). The dataset was randomly partitioned into training and test cohorts at a 2:1 ratio. A two-stage model combining MobileNetV2-SE-based deep learning features with clinical indicators was developed for RF classification. Two binary tasks were performed: RF presence (RF 1 vs . RF 2 and RF 3) and severity (RF 2 vs . RF 3). Nested cross-validation was applied for model development and hyperparameter tuning, and bootstrapping was used to assess performance robustness. Model performance was evaluated with the area under the curve (AUC), calibration curves, decision curve analysis (DCA), and SHapley Additive exPlanations (SHAP) visualization. Results Compared with single-modality models, the multimodal deep learning model (DL-combine), which is based exclusively on native T 1 mapping, ADC, and T 2 * mapping, demonstrated favourable and stable performance (test AUC: 0.930). Among the 14 classifiers, XGBoost performed numerically better in terms of RF presence (mean AUCs: 0.986, 0.887; accuracy: 0.947, 0.829), whereas ExtraTree performed better in terms of RF severity assessment (mean AUCs: 0.935, 0.883; accuracy: 0.886, 0.848). Calibration curves and DCA confirmed robust predictive reliability and clinical utility. SHAP analysis highlighted the relative contributions of the DL-sign and eGFR. Conclusion This two-stage multimodal MRI-based framework provides accurate and interpretable assessment of RF across different stages of CKD, supporting noninvasive risk stratification and complementary clinical decision-making. Supplementary Information The online version contains supplementary material available at 10.1186/s12882-026-04869-2. Keywords: Multimodal magnetic resonance imaging, Deep learning, Renal fibrosis, Machine learning, Chronic kidney disease Background Chronic kidney disease (CKD) has emerged as a global health care problem, with an estimated 700 million people living with any stage of CKD [ 1 , 2 ]. Renal fibrosis (RF), a pathological hallmark of CKD progression, is often irreversible and closely associated with renal prognosis [ 3 ]. While renal biopsy remains the gold standard for the diagnosis of RF, its invasiveness limits its clinical utility [ 4 ]. Functional magnetic resonance imaging (MRI) techniques provide noninvasive assessments of CKD and RF [ 5 , 6 ]. Diffusion-weighted imaging (DWI) measures the mobility of water molecules through apparent diffusion coefficient (ADC) values, with more severe RF correlating with restricted diffusion and lower ADC [ 7 ]. T 1 mapping, particularly non-contrast native T 1 mapping, detects increasing cortical T 1 values and decreasing differences in the corticomedullary region as fibrosis progresses [ 8 , 9 ]. Similarly, T 2 * mapping reflects renal tissue hypoxia through R 2 * values, showing elevated levels in cortical and medullary regions with worsening kidney damage [ 10 ]. Conventional medical imaging is associated with operator bias and limited feature extraction. Radiomics improves objectivity by quantifying subtle features but still relies on manual engineering [ 11 ]. Deep learning (DL) overcomes this by automatically extracting complex patterns from raw images [ 12 ] and has shown great promise in tumour diagnosis and treatment [ 13 ]. However, its application in CKD is still in the early stages. Recent studies suggest that DL models based on single-modality MRI can identify CKD patients [ 14 ]. However, these models overlook the complementary pathological insights provided by sequences such as T 1 mapping, ADC, and T 2 * mapping, which reflect distinct fibrotic mechanisms. [ 15 ]. Multimodal MRI fusion could address these limitations, yet no studies have explored combining multimodal MRI with DL for renal fibrosis assessment in patients with CKD. In this study, we propose a two-stage, clinically oriented framework for noninvasive RF assessment in patients with CKD, with the aims of evaluating the utility of multimodal MRI-based deep learning for capturing fibrosis-related imaging information, examining the added value of multimodal integration over single-sequence models, and exploring whether a deep learning-derived imaging signature can provide a stable and interpretable representation for subsequent machine learning-based fibrosis assessment. Multimodal MRI-based deep learning models were developed using native T 1 mapping, ADC, and T 2 * mapping to capture complementary structural and microstructural features, followed by integration of the derived DL signs with routine clinical indicators. This integrative two-stage framework is intended to support noninvasive fibrosis assessment and risk stratification across different stages of CKD, particularly in clinical settings where biopsy is constrained by eligibility or feasibility concerns. Methods Subjects This single-centre, prospective study was approved by the ethics committee of our institution (JD-LK-2022–060-01). All patients underwent MRI voluntarily and signed informed consent. A prospective analysis of 216 CKD patients between September 2021 and November 2024 revealed that these patients were scheduled for renal biopsy and agreed to undergo renal MRI. The inclusion criteria included (1) meeting the clinical diagnostic criteria for CKD [ 16 ], and (2) performing multimodal MRI of the kidney within 3 days prior to scheduled renal biopsy. However, 64 patients were excluded because (1) 3 patients had claustrophobia during the MRI (2) 26 patients could not complete all the sequence scans and the MRI images were incomplete (3) 17 patients with severe image artifacts and poor image quality could not perform the post-processing, and (4) 18 patients ultimately did not consent to renal biopsy. Finally, a total of 152 CKD patients were enrolled. The inclusion and exclusion criteria are shown in Fig. 1 . Fig. 1. Open in a new tab Flowchart for the inclusion and exclusion of patients Multimodal MRI examination and preprocessing MRI of both kidneys was performed using a Siemens Prisma 3.0T magnetic resonance scanner with an 18-channel body coil. Multimodal MRI sequences included native T 1 mapping, DWI and T 2 * mapping imaging. MRI images of both kidneys were acquired in the coronal plane. The image acquisition details are summarized in Supplementary Table 1 . ROIs of the right renal cortex were manually delineated by a radiologist with 7 years of genitourinary experience on coronal native T 1 mapping images. The right renal cortex was selected to ensure anatomical consistency with the biopsy site. T 2 * and ADC volumes were aligned to the T 1 mapping space using an advanced normalization tools (ANTs)-based multistep registration strategy. Additional preprocessing included denoising, spatial resampling, and field-of-view cropping. A rigid-affine registration served as the primary alignment approach, and a lightweight symmetric normalization (SyN) refinement was selectively applied when minor residual mismatches remained after affine alignment. All the registered images underwent manual quality control to ensure accurate multimodal correspondence. [ 17 ]. Clinical data Clinical data such as age, sex, height, weight, body mass index (BMI), blood pressure, blood glucose, serum creatinine (Scr), blood urea nitrogen (BUN) and 24-hour urinary protein (24h-UPRO) were collected. The estimated glomerular filtration rate (eGFR) was calculated according to the Chronic Kidney Disease Epidemiology Collaboration (CKD-EPI) formula [ 18 ]. Data splitting and augmentation The dataset was randomly split into training and test sets at a 2:1 ratio. In each round, the model was trained on two folds and evaluated on the remaining fold, which served as the test data for that specific iteration. The performance metrics were calculated for each round and subsequently averaged across the three rounds. To address class imbalance, negative samples in the training set were duplicated, and weighted loss was applied. To augment the training dataset, random flipping and rotation were used. Performance-sample size analysis To evaluate whether the available sample size was sufficient to support stable model performance, a performance-sample size curve analysis was conducted. We incrementally increased the training sample size for two research objectives: “RF presence” and “RF severity,” and calculated the area under the receiver operating characteristic curve (AUC) on a fixed, independent test set. Deep learning modelling MobileNetV2-SE, a lightweight convolutional network with a squeeze-and-excitation (SE) module, was used to enhance feature extraction efficiency [ 19 ]. Multimodal input channels were concatenated and passed through a multihead self-attention module, followed by a multilayer perceptron. Binary cross-entropy loss was optimized using the Adam optimizer (mini-batch size = 8, learning rate = 0.001), with batch normalization, dropout, and L2 regularization to mitigate overfitting. A softmax classifier generated probabilistic outputs. Category probabilities were further integrated into a composite DL-sign score using the following formula: [\mathrm{DL\text{-}sign}=\sum_{i = 0}^{ n -1} v_i\cdot p_i] The composite DL-sign score is a scalar value. It is mathematically defined as the expected value of the categorical outcome and is calculated by summing the products of each category’s assigned value (v i ) and its corresponding predicted probability (p i ).(n: The total number of categories. Indexing ranges from (0) to ( n -1). i: The category index (the i-th category). p i : The predicted probability of a sample belonging to the i-th category (from the model’s final softmax output). Value range: ([0,1]) Constraint: ) To ensure the reliability of the composite DL-sign score, probability calibration was applied to the model outputs prior to DL-sign construction. Temperature scaling was performed on the logits of the model output on the test set (learning only a single temperature parameter T, estimated by minimizing the NLL/cross-entropy objective function), and was used to generate calibrated multiclass probabilities for calculating the DL-sign in the testing phase. The test set was not involved in fitting any calibration parameters to avoid information leakage. To evaluate the calibration effect and the stability of the DL-sign, a supplementary analysis was conducted: DL-signs derived twice from the same batch of samples were compared, and their consistency was quantified using Spearman’s rank correlation coefficient. An ablation analysis was performed by systematically removing or replacing selected architectural components, including the SE module and transformer-based fusion, while keeping all other experimental conditions unchanged. MM-MNetV2-MLP was consisted of a MobileNetV2 backbone (without SE), simple fusion, and an MLP head; MM-MNetV2SE-MLP incorporated an SE module while keeping all the other components unchanged; and MM-MNetV2-Trans replaced the simple fusion with a transformer fusion module but without using the SE module. MobileNetV2-SE represents the complete proposed framework and was used as the primary model for subsequent analyses, including DL-sign derivation. Macro-averaged performance metrics were calculated by averaging the one-vs-rest performance of each class. Machine learning modelling Logistic regression revealed significant clinical indicators ( p < 0.05), which along with the DL-sign, formed a model to assess RF presence and severity. Fourteen machine learning models were used, including logistic regression (LR), naive Bayes (NB), SVM variants, decision tree (DT), random forest, ExtraTree, XGBoost, AdaBoost, multi-layer perceptron (MLP), gradient boosting machines (GBM), and light GBM. The full workflow is summarized in Fig. 2 . Fig. 2. Open in a new tab The full workflow. MobileNetV2-SE-based DL models extract DL-sign features from multimodal MRI images (native T 1 mapping, ADC and T 2 * mapping). The DL-sign combined with selected clinical indicators (eGFR, Scr, BUN) were input into 14 machine learning classifiers for evaluation of the presence and severity of RF in patients with CKD Model stability was evaluated using three-fold cross-validation. The entire dataset was randomly partitioned into three folds of comparable size and distribution. In each round, two folds were used to train the model, and the remaining single fold served exclusively as the held-out set for testing within that round. This process was repeated three times so that each fold was used as the test set exactly once. The performance metrics obtained from the three rounds were averaged to produce the final aggregated estimate. Multimodal MRI samples with different RF stages are shown in Figs. 3 A–D. Fig. 3. Open in a new tab Multimodal MRI images (native T 1 mapping, ADC and T 2 * mapping) of different degrees of renal fibrosis (RF) in CKD patients, and using the shapley additive explanations (SHAP) to visualize the interaction contribution of features in the optimal model. ( A ) RF 1 patients (no fibrosis); ( B ) RF 2 patients(mild fibrosis); ( C ) RF 3 Patients(moderate fibrosis); ( D ) RF 3 patients (severe fibrosis); ( E, F ) SHAP summary plots show that DL-sign contributes the highest SHAP value to XGBoost and ExtraTree models, followed by eGFR To further assess the robustness of the main findings and to reduce potential model selection bias, we additionally performed nested cross-validation (nested CV) for representative models corresponding to the two primary clinical tasks (RF presence and RF severity). Specifically, a three-fold outer cross-validation loop was used exclusively for unbiased performance estimation. Within each outer training fold, a three-fold inner cross-validation loop was applied for hyperparameter tuning. Model optimization was strictly confined to the inner loop, whereas the outer loop served solely for independent performance evaluation, thereby reducing the risk of information leakage and optimism bias. Model interpretability was enhanced using SHAP (SHapley Additive exPlanations), highlighting how the DL-sign interacts with clinical biomarkers (eGFR, Scr, BUN) [ 20 ]. Detailed pseudo code is provided via a link in the Supplementary Materials. Renal histopathology An ultrasound-guided renal biopsy was conducted within 3 days after the MRI by an experienced nephrologist. Patients were positioned prone with a sandbag under the abdomen to minimize kidney movement. Typically, the lower pole of the right kidney is the preferred puncture site. [ 21 ]. Following standard histopathological procedures, the degree of fibrosis in kidney biopsy specimens was assessed by Masson’s staining. According to the degree of fibrosis, the RF group was divided into a no RF group (referred as “RF 1”, no fibrosis), a mild RF group (“RF 2”, fibrosis proportion ≤ 25%), and a moderate to severe RF group (“RF 3”, fibrosis proportion > 25%) [ 22 ]. Statistical analysis Statistical analyses were conducted with SPSS 25.0 (IBM Corp.) and MedCalc 15.2.2 (MedCalc Software Ltd.). The MobileNetV2-SE network was developed in Python 3.8 using PyTorch 2.0. Continuous variables with a normal distribution were analysed using ANOVA and are reported as the mean ± standard deviation. Non-normal variables are presented as median (interquartile range) and were analysed with the Kruskal-Wallis test. Categorical variables were compared using the chi-squared test. Model discrimination performance was evaluated primarily using the AUC. A bootstrap-based approach was applied to estimate AUC differences and corresponding confidence intervals in the independent test set. Model calibration and clinical utility were assessed using calibration curves and decision curve analysis (DCA), respectively. Threshold-dependent performance metrics, including sensitivity, specificity, accuracy, positive predictive value (PPV), negative predictive value (NPV), and F1 score, were also reported to provide a comprehensive evaluation beyond the AUC alone. All tests were two-tailed with a significance threshold of p < 0.05. To evaluate sample size sufficiency, a performance-sample size curve analysis was performed. Results Patient baseline characteristics According to the degree of fibrosis, the 152 patients with CKD included in this study were divided into the RF 1 group (no fibrosis, n = 34), RF 2 group (mild fibrosis, n = 69), and RF 3 group (moderate to severe fibrosis, n = 49), as shown in Table 1 . Table 1. Basic characteristics RF 1 (n = 34) RF 2 (n = 69) RF 3 (n = 49) Statistics P Age (years) 44 ± 13 49 ± 15 47 ± 13 1.430 0.243 Gender (male, %) 14(41.2%) 33(47.8%) 30(61.2%) 3.633 0.163 Height (m) 1.64 ± 0.09 1.61(1.59, 1.70) 1.67 ± 0.08 4.286 0.117 BMI (kg/m 2 ) 24.97 ± 3.18 24.26 ± 3.69 24.53 ± 3.34 0.475 0.623 Blood pressure (hypertension, %) 21(61.8%) 40(58.0%) 32(65.3%) 0.655 0.721 Blood glucose (mmol/L) 4.71(4.40, 5.06) 4.72(4.46, 5.31) 5.00(4.47, 5.47) 3.881 0.144 eGFR (mL/min/1.73 m 2 ) 109.39 ± 19.96 90.20 ± 28.26 59.23 ± 22.19 44.662 <0.001* Scr (μmol/L) 61(48, 72.5) 74(56, 96) 110(100, 139) 58.265 <0.001* BUN (mmol/L) 4.80(3.50, 5.95) 5.30(4.10, 6.70) 6.90(6.00,9.40) 33.864 <0.001* 24 h-UPRO (g/24 h) 2.92(1.37, 5.84) 2.81(1.51, 5.24) 2.00(0.96,4.02) 3.006 0.222 Open in a new tab RF, renal fibrosis; RF 1, no RF (0% fibrosis); RF 2, mild RF (≤25% fibrosis); RF 3, moderate to severe RF (>25% fibrosis); CKD, chronic kidney disease; BMI, body mass index; eGFR, estimated glomerular filtration rate; Scr, serum creatinine; BUN, blood urea nitrogen; 24 h-UPRO, 24-hour urinary protein; Hypertension was defined as systolic/diastolic blood pressure ≥ 140/90 mmHg; values are mean with standard deviation or median with lower and upper quartile in parentheses or number with percentage in parentheses; *statistically significant Patient age, sex, height, weight, BMI, blood pressure, blood glucose, and 24 h-UPRO did not significantly differ among the three RF groups ( p > 0.05). The eGFR, Scr and BUN values of CKD patients were significantly different among the three RF groups ( p < 0.05). With the worsening of fibrosis, the eGFR levels gradually decreased, while the SCr and BUN levels gradually increased. Ablation analysis The contributions of individual components were evaluated through ablation experiments (Table 2 ). Table 2. Ablation analysis of model variants for RF classification (macro-averaged metrics on the test set) Model Fusion Val AUC (macro) Val ACC (macro) Val F1 (macro) Notes MobileNetV2-SE Full fusion 0.880 0.843 0.763 Best overall, used for downstream signature MM-MNetV2-MLP Concat 0.674 0.660 0.597 No SE, no Transformer MM-MNetV2-Trans Transformer 0.593 0.608 0.555 Transformer fusion only (no SE) MM-MNetV2SE-MLP Concat 0.637 0.647 0.501 SE only (no Transformer) Open in a new tab RF, renal fibrosis; SE: Squeeze-and-Excitation; MLP: Multi-Layer Perceptron; MM-MNetV2-Trans, MobileNetV2 (without SE) + Transformer fusion; MM-MNetV2SE-MLP, MobileNetV2 + SE + simple fusion + MLP head (no Transformer); MM-MNetV2-MLP, MobileNetV2 (without SE) + simple fusion + MLP head (no Transformer); MobileNetV2-SE (Full model), complete model used as the main baseline; AUC, area under the curve; ACC, accuracy Compared with the MM-MNetV2-MLP model, incorporating the SE module (MM-MNetV2SE-MLP) did not improve the generalization performance and resulted in reduced macro AUC (0.637 vs. 0.674), macro accuracy (0.647 vs. 0.660), and macro F1 score (0.501 vs. 0.597). Similarly, replacing the simple fusion with transformer fusion (MM-MNetV2-Trans) led to a further decrease in the macro AUC (0.593 vs. 0.674), macro accuracy (0.608 vs. 0.660), and macro F1 score (0.555 vs. 0.597). Overall, the complete framework (MobileNetV2-SE) achieved the best performance among all ablation settings. Removal or modification of individual components consistently degraded performance; therefore, MobileNetV2-SE was selected to derive the DL-sign for subsequent machine learning models. Detailed performance metrics are provided in Supplementary Table 6 . Optimization of multimodal MRI-based deep learning models The native T 1 mapping, ADC, and T 2 * mapping images were input into the MobileNetV2-SE architecture to construct four DL models (DL-native T 1 mapping, DL-ADC, DL-T 2 * mapping and DL-combine), which can be used to discriminate RF 1, RF 2, and RF 3 in CKD patients. Comparisons of the RF diagnostic performance of these MRI-based DL models are shown in the training and test cohorts (Table 3 ). Table 3. Comparisons of diagnostic performance of deep learning models based on multimodal MRI sequences for renal fibrosis AUC (95%CI) SEN SPE ACC PPV NPV F1-score DL-combine RF 1 Training 0.934 (0.875,0.977) 0.826 0.910 0.891 0.731 0.947 0.776 Test 0.930 (0.837,0.988) 0.818 0.900 0.882 0.692 0.947 0.750 RF 2 Training 0.883 (0.822,0.948) 0.932 0.789 0.851 0.774 0.938 0.845 Test 0.827 (0.700,0.910) 0.720 0.846 0.784 0.818 0.759 0.766 RF 3 Training 0.925 (0.879,0.966) 0.765 0.925 0.871 0.839 0.886 0.800 Test 0.882 (0.768,0.963) 0.800 0.889 0.863 0.750 0.914 0.774 DL-native T 1 mapping RF 1 Training 0.902 (0.840,0.950) 0.913 0.756 0.792 0.525 0.967 0.667 Test 0.777 (0.643,0.907) 0.909 0.650 0.706 0.417 0.963 0.571 RF 2 Training 0.766 (0.670,0.848) 0.864 0.579 0.703 0.613 0.846 0.717 Test 0.588 (0.449,0.745) 0.760 0.462 0.608 0.576 0.667 0.655 RF 3 Training 0.873 (0.798,0.938) 0.882 0.716 0.772 0.612 0.923 0.723 Test 0.657 (0.502,0.820) 0.733 0.639 0.667 0.458 0.852 0.564 DL-ADC RF 1 Training 0.941 (0.897,0.974) 0.783 0.936 0.901 0.783 0.936 0.783 Test 0.727 (0.547,0.862) 0.727 0.675 0.686 0.381 0.900 0.500 RF 2 Training 0.572 (0.439,0.697) 0.909 0.421 0.634 0.548 0.857 0.684 Test 0.512 (0.362,0.685) 0.960 0.154 0.549 0.522 0.800 0.676 RF 3 Training 0.847 (0.743,0.918) 1.000 0.701 0.802 0.630 1.000 0.773 Test 0.541 (0.408,0.741) 0.667 0.556 0.588 0.385 0.800 0.488 DL-T 2 * mapping RF 1 Training 0.871 (0.785,0.927) 0.913 0.718 0.762 0.488 0.966 0.636 Test 0.723 (0.574,0.855) 0.909 0.500 0.588 0.333 0.952 0.488 RF 2 Training 0.741 (0.630,0.842) 0.682 0.737 0.713 0.667 0.750 0.674 Test 0.597 (0.455,0.738) 0.640 0.577 0.608 0.593 0.625 0.615 RF 3 Training 0.803 (0.703,0.887) 0.765 0.821 0.802 0.684 0.873 0.722 Test 0.648 (0.495,0.815) 0.600 0.778 0.725 0.529 0.824 0.562 Open in a new tab RF, renal fibrosis; DL-combine, deep learning model based on the combination of native T 1 mapping, ADC and T 2 * mapping images; DL-native T 1 mapping, deep learning model based on native T 1 mapping images; DL-ADC, deep learning model based on ADC images; DL- T 2 * mapping, deep learning model based on T 2 * mapping images; RF 1, no RF (0% fibrosis); RF 2, mild RF (≤25% fibrosis); RF 3, moderate to severe RF (>25% fibrosis); AUC, area under the curve; SEN, sensitivity; SPE, specificity; ACC, accuracy; PPV, positive predictive value; NPV, negative predictive value; CI, confidence interval When the three RF groups were compared, the AUCs (0.930, 0.827, and 0.882), accuracy (0.882, 0.784, and 0.863) and F1 score (0.750, 0.766, and 0.774) were greater in the DL-combine group than in the other groups in the test cohort. Additionally, the overall performance of DL-combine was consistently favourable in the training cohort. Thus, compared with the three single models, the combined DL model demonstrated improved discriminative ability for differentiating RF in patients with CKD. The weighted probability of the DL-combine model (DL-sign) was obtained. Spearman’s rank correlation analysis demonstrated a strong monotonic association between the DL-sign values before and after probability calibration (ρ=0.963, p = 5.3 × 10 -8 7 ), indicating the high stability of the DL-sign rankings across the calibration procedures (Supplementary Figure 1 ). RF evaluation of models integrating deep learning with clinical indicators The performance-sample size curve analysis demonstrated that, at the current sample size, the generalization performance of both models essentially converged (Supplementary Figures 2 and 3 ). In terms of identifying the presence of RF (RF 1 vs. RF 2 and RF 3), compared with the other machine learning algorithms, the XGBoost model demonstrated a numerically favourable overall diagnostic performance, with a mean AUC and a mean ACC of 0.986 and 0.947,respectively, in the training cohort and 0.887 and 0.829,respectively, in the test cohort (Table 4 , Supplementary Table 2 ). Table 4. Cross-validation results of the optimal machine learning classification model (XGBoost) for identifying the presence of renal fibrosis (RF 1 vs. RF 2 and RF 3) Fold number Training cohort Test cohort AUC ACC AUC ACC Fold 1 0.980 0.931 0.911 0.843 Fold 2 0.991 0.950 0.895 0.765 Fold 3 0.987 0.961 0.855 0.880 Mean 0.986 0.947 0.887 0.829 Open in a new tab AUC, area under the curve; ACC, accuracy The confusion matrix revealed that most of the no RF and RF cases were correctly classified, with only a small number of false-negative and false positive predictions (Figs. 4 A–C) Fig. 4. Open in a new tab Confusion matrices for renal fibrosis classification. ( A-C ) XGBoost model for distinguishing non-fibrosis from fibrosis; ( D-F ) ExtraTree model for distinguishing mild fibrosis from moderate to severe fibrosis For identifying RF severity (RF 2 vs. RF 3), the Extratree model showed better overall diagnostic performance, with a mean AUC and a mean ACC of 0.935 and 0.886, respectively, in the training cohort and 0.883 and 0.848, respectively, in the test cohort (Table 5 , Supplementary Table 3 ). Table 5. Cross-validation results of the optimal machine learning classification model (ExtraTree) for identifying the severity of renal fibrosis (RF 2 vs. RF 3) Fold number Training cohort Test cohort AUC ACC AUC ACC Fold 1 0.941 0.897 0.852 0.800 Fold 2 0.926 0.861 0.918 0.872 Fold 3 0.939 0.899 0.880 0.872 Mean 0.935 0.886 0.883 0.848 Open in a new tab ACC, accuracy; AUC, area under the curve Within the nested CV framework, the XGBoost model showed relatively stable performance across different outer test folds for RF presence classification. The AUC values in the outer testing sets were distributed mainly between 0.85 and 0.95, with consistent accuracy and F1-scores across folds (Table 6 ). Table 6. Nested cross-validation results of the optimal machine learning classification model (XGBoost) for identifying the presence of renal fibrosis (RF 1 vs. RF 2 and RF 3) Source Model ACC AUC SEN SPE NPV PPV F1 inner-train-fold0 XGBoost 0.896 0.958 0.887 0.929 0.684 0.979 0.931 inner-test-fold0 XGBoost 0.765 0.856 0.680 1 0.529 1 0.810 outer-test-fold1 XGBoost 0.922 0.951 0.950 0.818 0.818 0.950 0.950 inner-train-fold1 XGBoost 0.851 0.910 0.808 1 0.600 1 0.894 inner-test-fold1 XGBoost 0.882 0.894 0.923 0.750 0.750 0.923 0.923 outer-test-fold0 XGBoost 0.902 0.918 0.950 0.727 0.800 0.927 0.938 inner-train-fold2 XGBoost 0.779 0.912 0.706 1 0.531 1 0.828 inner-test-fold2 XGBoost 0.909 0.929 0.926 0.833 0.714 0.962 0.943 outer-test-fold0 XGBoost 0.902 0.924 0.925 0.818 0.750 0.949 0.937 inner-train-fold0 XGBoost 0.851 0.952 0.811 1 0.583 1 0.896 inner-test-fold0 XGBoost 0.912 0.926 0.929 0.833 0.714 0.963 0.945 outer-test-fold0 XGBoost 0.784 0.899 0.730 0.929 0.565 0.964 0.831 inner-train-fold1 XGBoost 0.762 0.925 0.714 1 0.407 1 0.833 inner-test-fold1 XGBoost 0.912 0.953 0.92 0.889 0.800 0.958 0.939 outer-test-fold1 XGBoost 0.725 0.867 0.622 1 0.500 1 0.767 inner-train-fold2 XGBoost 0.912 0.942 0.925 0.867 0.765 0.961 0.942 inner-test-fold2 XGBoost 0.606 0.811 0.536 1 0.278 1 0.698 outer-test-fold1 XGBoost 0.824 0.864 0.838 0.786 0.647 0.912 0.873 inner-train-fold0 XGBoost 0.912 0.938 0.925 0.867 0.765 0.961 0.942 inner-test-fold0 XGBoost 0.735 0.824 0.625 1 0.527 1 0.769 outer-test-fold2 XGBoost 0.720 0.854 0.659 1 0.391 1 0.794 inner-train-fold1 XGBoost 0.897 0.964 0.863 1 0.708 1 0.926 inner-test-fold1 XGBoost 0.735 0.815 0.731 0.750 0.462 0.905 0.809 outer-test-fold2 XGBoost 0.860 0.923 0.854 0.889 0.571 0.972 0.909 inner-train-fold2 XGBoost 0.779 0.887 0.740 0.889 0.552 0.949 0.831 inner-test-fold2 XGBoost 0.971 0.989 0.963 1 0.875 1 0.981 outer-test-fold2 XGBoost 0.800 0.940 0.756 1 0.474 1 0.861 Open in a new tab AUC, area under the curve; ACC, accuracy; SEN, sensitivity; SPE, specificity; NPV, negative predictive value; PPV, positive predictive value For the task of renal fibrosis severity stratification, the ExtraTrees model demonstrated stable classification performance across repeated conventional cross-validation. Across different test folds, the model achieved a moderate-to-high AUC and balanced sensitivity and specificity (Table 7 ). Table 7. Nested cross-validation results of the optimal machine learning classification model (ExtraTree) for identifying the severity of renal fibrosis (RF 2 vs. RF 3) Source Model ACC AUC Sensitivity Specificity NPV PPV F1 inner-train-fold0 ExtraTree 0.865 0.912 0.889 0.853 0.935 0.762 0.821 inner-test-fold0 ExtraTree 0.885 0.950 1 0.813 1 0.769 0.870 outer-test-fold1 ExtraTree 0.800 0.886 0.952 0.632 0.923 0.741 0.833 inner-train-fold1 ExtraTree 0.942 0.981 0.889 0.971 0.943 0.941 0.914 inner-test-fold1 ExtraTree 0.846 0.853 0.8 0.875 0.875 0.800 0.800 outer-test-fold0 ExtraTree 0.775 0.858 0.810 0.737 0.778 0.773 0.791 inner-train-fold2 ExtraTree 0.942 0.958 0.850 1 0.914 1 0.919 inner-test-fold2 ExtraTree 0.692 0.781 1 0.556 1 0.5 0.667 outer-test-fold0 ExtraTree 0.800 0.853 0.857 0.737 0.824 0.783 0.818 inner-train-fold0 ExtraTree 0.904 0.944 0.958 0.857 0.960 0.852 0.902 inner-test-fold0 ExtraTree 0.815 0.843 0.692 0.929 0.765 0.900 0.783 outer-test-fold0 ExtraTree 0.872 0.951 1 0.815 1 0.706 0.828 inner-train-fold1 ExtraTree 0.868 0.924 0.783 0.933 0.848 0.900 0.837 inner-test-fold1 ExtraTree 0.769 0.821 0.857 0.667 0.800 0.750 0.800 outer-test-fold1 ExtraTree 0.821 0.923 0.917 0.778 0.954 0.647 0.759 inner-train-fold2 ExtraTree 0.887 0.945 0.889 0.885 0.885 0.889 0.889 inner-test-fold2 ExtraTree 0.885 0.872 0.900 0.875 0.933 0.818 0.857 outer-test-fold1 ExtraTree 0.897 0.923 0.917 0.889 0.96 0.786 0.846 inner-train-fold0 ExtraTree 0.942 0.980 0.947 0.940 0.969 0.900 0.923 inner-test-fold0 ExtraTree 0.778 0.838 0.643 0.923 0.706 0.900 0.750 outer-test-fold2 ExtraTree 0.821 0.856 0.8125 0.826 0.864 0.765 0.788 inner-train-fold1 ExtraTree 0.868 0.95 0.913 0.833 0.926 0.808 0.857 inner-test-fold1 ExtraTree 0.808 0.875 0.9 0.750 0.923 0.692 0.783 outer-test-fold2 ExtraTree 0.846 0.871 0.8125 0.870 0.870 0.813 0.813 inner-train-fold2 ExtraTree 0.868 0.933 0.917 0.828 0.923 0.815 0.863 inner-test-fold2 ExtraTree 0.923 0.974 1 0.882 1 0.818 0.900 outer-test-fold2 ExtraTree 0.821 0.889 0.813 0.826 0.864 0.765 0.788 Open in a new tab AUC, area under the curve; ACC, accuracy; SEN, sensitivity; SPE, specificity; NPV, negative predictive value;PPV, positive predictive value The nested CV results yielded performance estimates that were generally consistent with those obtained from conventional cross-validation, supporting the robustness of the main findings under a more rigorous evaluation framework. The corresponding confusion matrix indicated that both mild and moderate to severe fibrosis cases were predominantly correctly identified, with limited misclassifications (Figs. 4 D–F ). As shown in the SHAP summary plot, the DL-sign emerged as the most influential feature, contributing the highest relative SHAP values to the XGBoost and ExtraTree models, followed by the eGFR (Figs. 3 E, 3 F). ROC curves revealed that the XGBoost and ExtraTree models demonstrated strong diagnostic performance for RF evaluation, with AUC values greater than 0.85 for the test cohort (Fig. 5 ). Fig. 5. Open in a new tab Receiver operating characteristic curves of the optimal machine learning model for assessment of renal fibrosis (RF) in training and test cohorts. ( A-C ) XGBoost model for identifying the presence of RF (RF 1 vs. RF 2 and RF 3); ( D-F ) ExtraTree model for identifying the severity of RF (RF 2 vs. RF 3) The calibration curves revealed that both models were generally close to the ideal line in the training and test cohorts but with some biases in the local range (Fig. 6 ). Fig. 6. Open in a new tab Calibration curves of the optimal machine learning classification model for assessment of renal fibrosis (RF) in training and test cohorts. ( A-C ) XGBoost model for identifying the presence of RF (RF 1 vs. RF 2 and RF 3); ( D-F ) ExtraTree model for identifying the severity of RF (RF 2 vs. RF 3) DCA indicated that both models had net benefits in both the training and test cohorts (Fig. 7 ). Fig. 7. Open in a new tab Decision curves of the optimal machine learning classification model for assessment of renal fibrosis (RF) in training and test cohorts. ( A-C ) XGBoost model for identifying the presence of RF (RF 1 vs. RF 2 and RF 3); ( D-F ) ExtraTree model for identifying severity of RF (RF 2 vs. RF 3) Bootstrap resampling on the test set revealed that, for most model comparisons, the estimated AUC differences were small, with 95% confidence intervals overlapping zero, whereas a limited number of classifiers clearly exhibited lower discriminative performance. Detailed bootstrap results are provided in Supplementary Tables 4 – 5 . On the basis of the overall performance observed across these analyses, XGBoost and ExtraTrees were used for the subsequent evaluation of RF presence and RF severity, respectively. Discussion This study introduces a novel two-stage deep learning approach using multimodal MRI (native T 1 mapping, ADC, and T 2 * mapping) for differentiating RF in patients with CKD, offering advantages over previous single-modality methods. The DL-combine model had the highest AUCs in both the training and test groups, outperforming the single-modality models in terms of the diagnosis of RF. To our knowledge, multimodal MRI-based DL frameworks for RF assessment in CKD have not been extensively investigated, as previous research focused primarily on MRI texture-based ML methods [ 23 , 24 ]. Our previous studies demonstrated that native T 1 mapping-based radiomic models accurately assess renal function and fibrosis in patients with CKD [ 17 ]. However, single-modality imaging only partially captures fibrosis features and can be affected by confounding factors such as inflammation and oedema, limiting its specificity. In this study, native T 1 mapping was used to assess fibrosis and extracellular matrix changes, the ADC showed microstructural disruption from restricted water diffusion, and T 2 * mapping indicated hypoxia-induced iron deposition. Together, these methods provide a comprehensive view of fibrotic mechanisms. We employed the MobileNetV2-SE architecture, which combines the lightweight MobileNetV2 network with a squeeze-and-excitation (SE) attention mechanism, owing to its favourable balance between computational efficiency and representational capacity. The SE module is designed to recalibrate channelwise feature responses and has been shown to be effective at enhancing discriminative features in various medical imaging tasks [ 19 ]. Lightweight and transfer learning-based convolutional architectures have demonstrated robust performance across diverse medical image classification problems, including mammographic breast cancer identification and chest radiograph analysis [ 25 , 26 ]. These studies highlight the practical advantages of resource-efficient CNN frameworks in real-world clinical settings. However, our ablation analysis indicates that, under the current data setting, simply introducing SE modules or increasing architectural complexity does not necessarily translate into improved generalization performance. Instead, the observed performance gains are primarily attributable to the overall design of the proposed framework, including multimodal feature integration and the two-stage modelling strategy. In contrast to segmentation-focused models such as U-Net and its variants [ 27 , 28 ], which are computationally intensive and less suited for real-time applications, the proposed framework adopts a classification-oriented and resource-efficient design. Various registration strategies have been developed and applied in multimodal medical imaging to facilitate image alignment, including both intensity-based and feature-based approaches [ 29 ]. In the present study, an intensity-based registration framework was adopted to ensure consistent voxel-wise correspondence across MRI sequences. The use of depthwise separable convolutions enables effective feature extraction with reduced computational cost. Moreover, compared with heavier architectures such as ResNet and DenseNet [ 30 ], this lightweight design is better suited for small-sample medical imaging scenarios, offering a practical balance between efficiency and robustness. Overall, these characteristics support the applicability of the proposed framework as a resource-efficient tool for renal fibrosis staging. We further combined the DL-sign with clinical indicators (eGFR, Scr, BUN) to construct machine learning models to identify both the presence and severity of RF. XGBoost emerged as the optimal model for detecting the presence of RF, whereas ExtraTree excelled for evaluating the severity of RF. During model comparison, bootstrap-based AUC analyses indicated that the differences in discrimination between the selected optimal models and alternative classifiers were generally modest, with overlapping confidence intervals in most comparisons. These observations suggest that several candidate models achieved comparable AUC performance when similar feature representations were used. Importantly, model selection was not based on the AUC alone but was informed by a comprehensive assessment integrating multiple threshold-dependent metrics, including the F1-score, sensitivity, specificity, and predictive values, as well as model stability across cross-validation folds. From an integrated evaluation perspective, the selected optimal models consistently ranked among the top-performing classifiers across multiple performance metrics and exhibited limited fold-to-fold variability, supporting their robustness and reliability in practical RF assessment. This multi-dimensional evaluation strategy helps mitigate potential bias associated with reliance on a single performance metric. Similar multimodal MRI-based approaches have demonstrated clinical feasibility in CKD studies. A single-centre retrospective study combined a T2WI-based model with ADC and R 2 * values to assess renal function in patients with diabetic nephropathy [ 24 ]. Hua et al. employed an SVM model integrating T 1 mapping and DWI to effectively evaluate CKD and RF [ 15 ]. In our study, the inclusion of clinical biomarkers in the multimodal MRI-based DL framework improved both diagnostic accuracy and clinical interpretability. While eGFR reflects glomerular filtration, Scr and BUN represent nitrogen metabolism. Their combination with DL-derived structural features offers a comprehensive perspective on both functional decline and tissue remodelling. Notably, DL-sign features from native T 1 mapping, ADC, and T 2 * mapping consistently ranked highest in SHAP analyses, indicating the pivotal role of multimodal MRI-derived features in RF stratification and highlighting the complementary value of clinical markers. The consistent dominance of the DL-sign in both the XGBoost and ExtraTree models further supports its robustness and diagnostic relevance across machine learning frameworks. This proposed two-stage framework has meaningful potential for clinical adoption. By integrating quantitative multiparametric MRI with DL-derived features, the approach could support more standardized and objective evaluation of RF within routine kidney MRI workflows. In practice, such a system may assist radiologists in improving the consistency of fibrosis assessment and could facilitate earlier identification of patients at risk of progressive CKD. However, several practical considerations may influence large-scale deployment. The availability of scanners remains heterogeneous across institutions, and the cost and accessibility of multiparametric MRI sequences may limit their use in low-resource settings. In addition, variations in acquisition protocols across centres may introduce challenges to model transferability, underscoring the need for harmonization strategies. In the future, federated learning represents a promising direction for enhancing multicenter generalizability while avoiding raw data sharing [ 31 ]. Further development of more efficient and streamlined inference pipelines may also enable real-time implementation in clinical systems. These future improvements have the potential to expand the applicability of the proposed framework in broader clinical environments. This study had several limitations. First, although 152 biopsy-proven CKD patients were included, cases with advanced fibrosis remained relatively rare. In addition, the use of relatively broad fibrosis grading thresholds resulted in the combination of moderate and severe fibrosis into a single RF3 category, which may introduce intragroup heterogeneity. Second, the dataset was derived from a single centre, which may limit generalizability. Although a multicentre collaboration for CKD imaging and biopsy data collection has been initiated by our team, the external datasets are still in the early acquisition stage and were not yet suitable for model validation in the present study. Third, the moderate sample size constrained further analysis of the associations between DL-derived features and different CKD aetiologies, larger multicentre cohorts will be needed to identify aetiology-specific fibrosis signatures. Finally, we did not compare the proposed lightweight MobileNetV2-SE-based feature extractor with deeper architectures such as ResNet, EfficientNet, or the Swin Transformer. Given the current dataset size, deeper networks pose a high risk of overfitting, however, future studies with expanded multicentre datasets will incorporate such architectures to further evaluate model robustness and benchmark model performance comprehensively. In conclusion, this study demonstrates that a multimodal MRI-based deep learning framework can effectively capture fibrosis-related imaging information and enable noninvasive assessment of RF in patients with CKD. By integrating complementary information from native T 1 mapping, ADC, and T 2 * mapping, the proposed approach provides added value over single-sequence models for evaluating both the presence and severity of RF. Furthermore, the derived DL-sign serves as a stable and continuous imaging representation that can be readily combined with routine clinical indicators. Using this two-stage strategy, XGBoost was identified as the optimal classifier for detecting the presence of renal fibrosis, whereas ExtraTree showed superior performance for assessing fibrosis severity. Collectively, these findings highlight the potential clinical utility of combining multimodal MRI with clinical data to support noninvasive RF evaluation and personalized management in patients with CKD. Electronic supplementary material Below is the link to the electronic supplementary material. Supplementary material 1 (746.4KB, docx) Acknowledgements We thank AJE Editing Service for editing this manuscript. Abbreviation CKD Chronic kidney disease RF Renal fibrosis MRI Magnetic resonance imaging DWI Diffusion-weighted imaging ADC Apparent diffusion coefficient ANTs Advanced normalization tools SyN Symmetric normalization DL Deep learning BMI Body mass index Scr Serum creatinine BUN Blood urea nitrogen 24 h-UPRO 24-hour urinary protein eGFR Estimated glomerular filtration rate CKD-EPI Chronic Kidney Disease Epidemiology Collaboration SE Squeeze-and-Excitation ROI Region of interest LR Logistic regression NB Naive bayes SVM Support Vector Machine DT Decision Tree ExtraTree Extremely Randomized Trees XGBoost eXtreme Gradient Boosting MLP Multi-Layer Perceptron GBM Gradient Boosting Machines AUC Area under the curve ROC Receiver operating characteristic DCA Decision curve analysis Author contributions Xiaojing Li: Conceptualization, Formal analysis, Writing-Original Draft. Yirui Li: Visualization, Formal analysis, Writing- Original Draft. Qing Ma: Investigation, Data Curation. Yilin Xu: Resources. Ye Zhu: Investigation. Jing Zhang: Software. Junkang Shen: Validation. Wu Cai: Supervision. Zhen Jiang, Chaogang Wei: Project administration, Writing-Review & Editing. Funding This study was financially supported by the Suzhou Medical College-QiLu Medical Research Program of Soochow University (24QL200214); the Suzhou Science and Technology Healthcare Innovation Project (SYW2025047); the Suzhou Science and Education Strong Health Project (MSXM2025013); the National Natural Science Foundation of China (81801754); the Project of State Key Laboratory of Radiation Medicine and Protection, Soochow University (GZK12025016). Declarations Data availability The datasets generated and analysed during the current study are not publicly available due to ethical obligations to protect patient confidentiality and the dataset’s integral role in an ongoing longitudinal research program, but are available from the corresponding author on reasonable request. Ethics approval and consent to participate This study was performed in line with the principles of the Declaration of Helsinki. Approval was granted by the Ethics Committee of The Second Affiliated Hospital of Soochow University. (Approval number: JD-LK-2022–060-01).Written informed consent was obtained from all individual participants included in the study. Consent for publication All patients signed informed consent regarding publishing their data and photographs. Competing interests The authors declare no competing interests. Footnotes Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Xiaojing Li, Yirui Li and Qing Ma contributed equally to this work as co-first authors. Contributor Information Chaogang Wei, Email: [email protected]. Zhen Jiang, Email: [email protected]. References 1. Stewart S, Kalra PA, Blakeman T, Kontopantelis E, Cranmer-Gordon H, Sinha S. Chronic kidney disease: detect, diagnose, disclose-a UK primary care perspective of barriers and enablers to effective kidney care. BMC Med. 2024;22:331. 10.1186/s12916-024-03555-0. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 2. Francis A, Harhay MN, Ong A, et al. Chronic kidney disease and the global public health agenda: an international consensus. Nat Rev Nephrol. 2024;20:473–85. 10.1038/s41581-024-00820-6. [ DOI ] [ PubMed ] [ Google Scholar ] 3. Panizo S, Martínez-Arias L, Alonso-Montes C, et al. Fibrosis in chronic kidney disease: pathogenesis and consequences. Int J Mol Sci. 2021;22:408. 10.3390/ijms22010408. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 4. Schnuelle P. Renal biopsy for diagnosis in kidney disease: indication, technique, and safety. J Clin Med. 2023;12:6424. 10.3390/jcm12196424. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 5. Jiang B, Liu F, Fu H, Mao J. Advances in imaging techniques to assess kidney fibrosis. Ren Fail. 2023;45:2171887. 10.1080/0886022X.2023.2171887. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 6. Li J, An C, Kang L, Mitch WE, Wang Y. Recent Advances in magnetic resonance imaging assessment of renal fibrosis. Adv Chronic Kidney Dis. 2017;24:150–53. 10.1053/j.ackd.2017.03.005. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 7. Berchtold L, Crowe LA, Combescure C, et al. Diffusion-magnetic resonance imaging predicts decline of kidney function in chronic kidney disease and in patients with a kidney allograft. Kidney Int. 2022;101:804–13. 10.1016/j.kint.2021.12.014. [ DOI ] [ PubMed ] [ Google Scholar ] 8. Berchtold L, Friedli I, Crowe LA, et al. Validation of the corticomedullary difference in magnetic resonance imaging-derived apparent diffusion coefficient for kidney fibrosis detection: a cross-sectional study. Nephrol Dial Transpl. 2020;35:937–45. 10.1093/ndt/gfy389. [ DOI ] [ PubMed ] [ Google Scholar ] 9. Wei CG, Zeng Y, Zhang R, et al. Native T1 mapping for non-invasive quantitative evaluation of renal function and renal fibrosis in patients with chronic kidney disease. Quant Imag Med Surg. 2023;13:5058–71. 10.21037/qims-22-1304. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 10. Inoue T, Kozawa E, Okada H, et al. Noninvasive evaluation of kidney hypoxia and fibrosis using magnetic resonance imaging. J Am Soc Nephrol. 2011;22:1429–34. 10.1681/ASN.2010111143. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 11. Guiot J, Vaidyanathan A, Deprez L, et al. A review in radiomics: making personalized medicine a reality via routine imaging. Med Res Rev. 2022;42:426–40. 10.1002/med.21846. [ DOI ] [ PubMed ] [ Google Scholar ] 12. Zhang M, Ye Z, Yuan E, et al. Imaging-based deep learning in kidney diseases: recent progress and future prospects. Insights Imag. 2024;15:50. 10.1186/s13244-024-01636-5. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 13. Jiang X, Hu Z, Wang S, Zhang Y. Deep learning for medical image-based cancer diagnosis. Cancers (Basel). 2023;15:3608. 10.3390/cancers15143608. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 14. Aslam I, Aamir F, Kassai M, et al. Validation of automatically measured T1 map cortico-medullary difference (ΔT1) for eGFR and fibrosis assessment in allograft kidneys. PLoS One. 2023;18:e0277277. 10.1371/journal.pone.0277277. [ DOI ] [ PMC free article ] [ PubMed ] 15. Hua C, Qiu L, Zhou L, et al. Value of multiparametric magnetic resonance imaging for evaluating chronic kidney disease and renal fibrosis. Eur Radiol. 2023;33:5211–21. 10.1007/s00330-023-09674-1. [ DOI ] [ PubMed ] [ Google Scholar ] 16. Stevens PE, Levin A. Kidney disease: improving global outcomes chronic kidney disease guideline development work group members (2013) evaluation and management of chronic kidney disease: synopsis of the kidney disease: improving global outcomes 2012 clinical practice guideline. Ann Intern Med. 158:825–30. 10.7326/0003-4819-158-11-201306040-00007. [ DOI ] [ PubMed ] 17. Wei C, Jin Z, Ma Q, et al. Native T1 mapping-based radiomics diagnosis of kidney function and renal fibrosis in chronic kidney disease. iScience. 2024;27:110493. 10.1016/j.isci.2024.110493. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 18. Levey AS, Stevens LA, Schmid CH, et al. A new equation to estimate glomerular filtration rate. Ann Intern Med. 2009;150:604–12. 10.7326/0003-4819-150-9-200905050-00006. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 19. Zhu Q, Zhuang H, Zhao M, Xu S, Meng R. A study on expression recognition based on improved mobilenetV2 network. Sci Rep. 2024;14:8121. 10.1038/s41598-024-58736-x. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 20. Chen Z, Wang Y, Ying M, Su Z. Interpretable machine learning model integrating clinical and elastosonographic features to detect renal fibrosis in Asian patients with chronic kidney disease. J Nephrol. 2024;37:1027–39. 10.1007/s40620-023-01878-4. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 21. Xu J, Wu X, Xu Y, et al. Acute kidney disease increases the risk of post-kidney biopsy bleeding complications. Kidney Blood Press Res. 2020;45:873–82. 10.1159/000509443. [ DOI ] [ PubMed ] [ Google Scholar ] 22. Srivastava A, Palsson R, Kaze AD, et al. The prognostic value of histopathologic lesions in native kidney biopsy specimens: results from the Boston kidney biopsy cohort study. J Am Soc Nephrol. 2018;29:2213–24. 10.1681/ASN.2017121260. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 23. Mo X, Chen W, Chen S, et al. MRI texture-based machine learning models for the evaluation of renal function on different segmentations: a proof-of-concept study. Insights Imag. 2023;14:28. 10.1186/s13244-023-01370-4. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 24. Chen W, Zhang L, Cai G, et al. Machine learning-based multimodal MRI texture analysis for assessing renal function and fibrosis in diabetic nephropathy: a retrospective study. Front Endocrinol (Lausanne). 2023;14:1050078. 10.3389/fendo.2023.1050078. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 25. Patel RK, Choudhary A, Kumari N, Lamkuche HS. Pneumonia screening from radiology images using homomorphic transformation filter-based FAWT and customized VGG-16. Int J Imag Syst Technol. 2025;35:e70093. 10.1002/ima.70093. 26. Patel RK, Kashyap M. The study of various registration methods based on maximal stable extremal region and machine learning. Comput Methods Biomech Biomed Eng: Imag Visual. 2023;11(6):2508–15. 10.1080/21681163.2023.2243351. [ Google Scholar ] 27. Siddique N, Sidike P, Elkin C, Devabhaktuni V. U-net and its variants for medical image segmentation: a review of theory and applications. IEEE Access. 2021;9:82031–57. 10.1109/ACCESS.2021.3081920. [ Google Scholar ] 28. Liu J, Yildirim O, Akin O, Tian Y. AI-Driven robust kidney and renal mass segmentation and classification on 3D CT images. Bioeng (Basel). 2023;10:116. 10.3390/bioengineering10010116. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 29. Deshpande S, Chouhan SS, Patel RK, Vishwakarma H. Transfer learning with ResNet50 for enhanced mammographic breast cancer identification. 2024 5th International Conference on Circuits, Control, Communication and Computing (I4C). IEEE; 2024. 30. Sharma N, Gupta S, Gupta D, et al. UMobileNetV2 model for semantic segmentation of gastrointestinal tract in MRI scans. PLoS One. 2024;19(5):e0302880. 10.1371/journal.pone.0302880. [ DOI ] [ PMC free article ] [ PubMed ] 31. Sheller MJ, Edwards B, Reina GA, Martin J, Pati S, Kotrotsou A, et al. Federated learning in medicine: facilitating multi-institutional collaborations without sharing patient data. Sci Rep. 2020;10(1):12598. 10.1038/s41598-020-69250-1. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials Supplementary material 1 (746.4KB, docx) Data Availability Statement The datasets generated and analysed during the current study are not publicly available due to ethical obligations to protect patient confidentiality and the dataset’s integral role in an ongoing longitudinal research program, but are available from the corresponding author on reasonable request. Articles from BMC Nephrology are provided here courtesy of BMC ACTIONS View on publisher site PDF (9.7 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top

Record · ID 748 · SHA-256 1357ca55666525b7
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.