ConceptioArchiveNCBI PubMed Central
NCBI PubMed Centralopen access

Artificial intelligence models: transforming early diagnosis and precise treatment of gastrointestinal cancers.

Liu K et al. · ncbi_pmc
NCBI PubMed Central · Papers · License: Open Access
Open Source ↗Direct PDF ↓
networksecurityintrusiondetection
network security intrusion detection

Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Mol Cancer . 2026 Mar 2;25:90. doi: 10.1186/s12943-026-02589-7 Search in PMC Search in PubMed View in NLM Catalog Add to search Artificial intelligence models: transforming early diagnosis and precise treatment of gastrointestinal cancers Kaijie Liu Kaijie Liu 1 School of Medicine, Chongqing University, Chongqing, 400044 People’s Republic of China Find articles by Kaijie Liu 1, # , Zeyu Luo Zeyu Luo 2 College of Computer and Control Engineering, Northeast Forestry University, Harbin, 150040 People’s Republic of China Find articles by Zeyu Luo 2, # , Wenjie Zhang Wenjie Zhang 1 School of Medicine, Chongqing University, Chongqing, 400044 People’s Republic of China Find articles by Wenjie Zhang 1, # , Qiyuan Pan Qiyuan Pan 1 School of Medicine, Chongqing University, Chongqing, 400044 People’s Republic of China Find articles by Qiyuan Pan 1 , Xiaotan Su Xiaotan Su 1 School of Medicine, Chongqing University, Chongqing, 400044 People’s Republic of China Find articles by Xiaotan Su 1 , Zhouyu Yang Zhouyu Yang 3 Department of Gastroenterology & Chongqing Key Laboratory of Digestive Malignancies, Daping Hospital, Army Medical University (Third Military Medical University), 10# Changjiang Branch Road, Yuzhong District, Chongqing, 400042 People’s Republic of China Find articles by Zhouyu Yang 3 , Qiaoqiao Zhang Qiaoqiao Zhang 4 Jinfeng Laboratory, Chongqing, 401329 People’s Republic of China Find articles by Qiaoqiao Zhang 4 , Bin Wang Bin Wang 3 Department of Gastroenterology & Chongqing Key Laboratory of Digestive Malignancies, Daping Hospital, Army Medical University (Third Military Medical University), 10# Changjiang Branch Road, Yuzhong District, Chongqing, 400042 People’s Republic of China 4 Jinfeng Laboratory, Chongqing, 401329 People’s Republic of China Find articles by Bin Wang 3, 4, ✉ , Bo Tang Bo Tang 5 Department of General Surgery, The First Affiliated Hospital (Southwest Hospital) of Army Medical University (Third Military Medical University), Chongqing, 400038 People’s Republic of China Find articles by Bo Tang 5, ✉ , Zongsheng He Zongsheng He 3 Department of Gastroenterology & Chongqing Key Laboratory of Digestive Malignancies, Daping Hospital, Army Medical University (Third Military Medical University), 10# Changjiang Branch Road, Yuzhong District, Chongqing, 400042 People’s Republic of China Find articles by Zongsheng He 3, ✉ , Jinjun Guo Jinjun Guo 6 Department of Gastroenterology and Hepatology, Bishan Hospital of Chongqing Medical University, Chongqing, 402760 People’s Republic of China Find articles by Jinjun Guo 6, ✉ Author information Article notes Copyright and License information 1 School of Medicine, Chongqing University, Chongqing, 400044 People’s Republic of China 2 College of Computer and Control Engineering, Northeast Forestry University, Harbin, 150040 People’s Republic of China 3 Department of Gastroenterology & Chongqing Key Laboratory of Digestive Malignancies, Daping Hospital, Army Medical University (Third Military Medical University), 10# Changjiang Branch Road, Yuzhong District, Chongqing, 400042 People’s Republic of China 4 Jinfeng Laboratory, Chongqing, 401329 People’s Republic of China 5 Department of General Surgery, The First Affiliated Hospital (Southwest Hospital) of Army Medical University (Third Military Medical University), Chongqing, 400038 People’s Republic of China 6 Department of Gastroenterology and Hepatology, Bishan Hospital of Chongqing Medical University, Chongqing, 402760 People’s Republic of China ✉ Corresponding author. # Contributed equally. Received 2025 Nov 4; Accepted 2026 Jan 21; Collection date 2026. © The Author(s) 2026 Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/ . PMC Copyright notice PMCID: PMC13059237  PMID: 41772662 Abstract Artificial intelligence (AI) has become an integral force in the clinical landscape of gastrointestinal (GI) oncology. Recent advances in model architectures ranging from traditional machine learning and convolutional neural networks (CNNs) to transformer-based foundational models and graph neural networks (GNNs) have enabled the extraction of complex features from diverse data modalities, including endoscopic images, radiology, pathology whole-slide images, and multi-omics profiles. In this review, AI models are systematically classified into supervised learning, unsupervised clustering, multimodal fusion, and interpretable modeling. The advantages of each model are delineated in unravelling tumor heterogeneity, anatomical characteristics, and treatment-relevant biomarkers. Furthermore, three types of clinical application are emphasized: (1) early screening and lesion localization via segmentation or anomaly detection; (2) molecular subtyping and patient stratification for diagnosis with risk assessment; (3) therapy guidance through response prediction and personalized treatment planning. We also discuss major challenges on the application of AI in integration of heterogeneous clinical data, model generalizability across centers, and the interpretability of predictions. Collectively, this review highlights the transformative potential of AI in better understanding tumor biology and its clinical value in advancing personalized medicine for GI cancer patients. Keywords: Artificial intelligence, Gastrointestinal tumor, Early screening, Prognosis, Personalized medicine Introduction Gastrointestinal (GI) tract cancers, including esophageal, gastric, and colorectal cancers, are among the most lethal malignancies worldwide. According to global cancer statistics in 2022, there were more than 3.4 million new cases of GI cancers, accounting for approximately 17.1% of all newly diagnosed cancers. GI cancer related deaths totaled about 2 million, representing 20.7% of overall cancer mortality [ 1 ]. In recent years, advances in medical diagnosis and treatment have contributed to a decline in the overall morbidity and mortality of GI cancers. However, the incidence of colorectal cancer and esophageal adenocarcinoma continues to increase in certain area, particularly Asian countries. In China, gastric cancer remains a major health burden, accounting for approximately 15% of all cancer cases [ 2 ]. In clinical practice, endoscopic screening of GI cancer faces multiple challenges. For instance, misdiagnosis remains commonly happened, which limits accurate subtyping by quantitatively integrating multi-source examination data. In addition, effective tools for predicting tumor progression (such as metastasis or recurrence) or assessing responses to distinct therapies are lacking, hindering the timely adjustment of treatment strategies. Furthermore, accurate prediction of prognosis and systematic stratification of patients based on tumor biology is difficult [ 3 , 4 ]. Collectively, these limitations constrain GI cancer diagnosis and treatment. What is worse, the high heterogeneity of GI tumors and their complex tumor microenvironment (TME) further hampers the identification of diagnostic biomarkers and therapeutic targets, which slows the development of novel therapies [ 5 ]. Artificial intelligence (AI)-based models have demonstrated substantial value in both the clinical management and basic research of GI cancers. These models not only contribute to diagnosis and treatment decision-making but also provide new insights into tumorigenesis mechanisms, thereby advancing the identification of potential therapeutic targets and anti-cancer agents [ 6 – 9 ]. At present, convolutional neural networks (CNNs), attention mechanisms, and transformer architectures are widely employed in medical imaging tasks such as tumor lesion segmentation and pathological subtype classification. Graph neural networks (GNNs) and large language models (LLMs) have been used to analyze biological data, including the identification of oncogenes and tumor suppressor genes. In addition, foundation models and multimodal learning approaches are increasingly being adopted in GI cancer research. Generative models, such as variational autoencoders (VAEs), generative adversarial networks (GANs) and diffusion models, have also been incorporated into algorithmic design to address challenges in image synthesis and data augmentation [ 10 , 11 ] (Fig. 2 ). Nevertheless, the broader application of these models is facing several crucial disadvantages like data heterogeneity, model deployment feasibility, and interpretability significantly. Moreover, performance, generalizability, and the efficiency of human-AI interaction require further optimization and rigorous validation to ensure its reliability and clinical relevance. Fig. 2. Open in a new tab Development trajectory of AI models in GI tumors. AI models tailored for tasks in GI oncology have undergone continuous evolution—from traditional machine learning approaches to CNNs, GNNs, and Transformers, followed by emerging generative deep learning models such as VAEs, GANs, and diffusion models. More recently, foundation models trained on medical imaging and molecular sequence data, along with novel architectures such as Mamba and multimodal frameworks, have further advanced the field. These iterative developments have progressively enhanced model performance and improved adaptability to the complexity and heterogeneity of medical data. IFM, Image Foundation Model; SLFM, Sequences and Large Foundation Model According to the functionality and data modalities, this review categorizes AI models applied to GI tumors into five groups: classical machine learning models, medical image-based deep learning models, molecular sequence-based deep learning models, multimodal models, hybrid models and explainable AI models. And then we focus on early screening approaches that utilize endoscopy, liquid biopsy, conventional imaging, and emerging non-routine data sources, highlighting their performance in the rapid identification of tumor classification. In subsequent sections, AI applications are systematically examined across pathology, radiomics, molecular bioinformatics, clinical variables, and multimodal fusion, with a focus on prognosis prediction, treatment response evaluation, and clinical decision support. Furthermore, we systematically examine AI models based on WSIs, radiomics, molecular bioinformatics, clinical data, and multimodal fusion for applications in prognosis prediction, treatment response evaluation, and clinical decision support. In addition to predictive modeling and risk stratification, explainable methods are shown to offer mechanistic insights into the mechanisms of GI tumor development and progression through interpretability techniques. The review also discusses the current limitations of AI models in clinical translation and basic research and outline key directions for future development (Fig. 1 ). Fig. 1. Open in a new tab The PRISMA Flowchart. This review initially retrieved a total of 826 articles from PubMed and 78 articles from Google Scholar that were broadly relevant to the topic. After removing 88 duplicate records, 127 articles were excluded through preliminary screening. The remaining 658 articles underwent full-text assessment, following which irrelevant publications were excluded. Ultimately, 253 articles were included as the primary references for this review Classification of AI models in oncology It is well documented that the pronounced heterogeneity of tumors, together with their underlying mechanisms, cannot be fully revealed through direct observation alone. In addition, the diversity of data and the limited size of available samples challenge traditional experimental designs and statistical analysis methods. Within this context, AI-based models show considerable advantages, as they can extract informative features from multimodal data, including medical images, molecular sequences, and clinical information to support downstream medical tasks such as disease diagnosis, prognosis prediction, and treatment response evaluation. AI models can be broadly categorized into machine learning and deep learning based on algorithmic type, both of which can further be subdivided into image-based, sequence-based, and multimodal fusion models according to data modality. In the next section, we discuss the algorithmic principles and architectures of these AI models in cancers. The clinical and biological complexity of tumors manifested through their pronounced heterogeneity, diverse morphologic patterns, and intricate molecular mechanisms poses substantial challenges for traditional analytic paradigms. Conventional experimental and statistical approaches are often limited by sample size, feature dimensionality, and the heterogeneity inherent in multimodal biomedical data. In contrast, AI offers a transformative approach by learning high‑dimensional representations from heterogeneous inputs such as medical images, molecular sequences, and clinical variables, thereby enabling integrative analysis across spatial and molecular scales. In oncology research, AI frameworks can be broadly classified according to algorithmic paradigm and data modality. From a methodological perspective, models range from classical machine‑learning algorithms to deep‑learning architectures encompassing convolutional, graph‑based, and transformer networks (Fig. 2 ). From a data‑centric viewpoint, they can be further divided into image‑based, molecular sequence–based, and multimodal fusion models. In addition, the rise of explainable AI (XAI) has introduced interpretability as an indispensable dimension in model development and clinical translation. Classical machine learning models Machine learning, as a data-driven approach within AI, seeks to identify underlying patterns in data to support prediction and decision-making. Conventional machine learning models can be broadly categorized into four types. The first category comprises regression models, which are primarily used to predict continuous variables. For example, linear regression can model the dynamic growth of colorectal tumor volume, while the Cox proportional hazards model is widely applied to evaluate prognostic indicators such as progression-free survival (PFS) and overall survival (OS) (Table 4 ). The second category includes classification models, which are suitable for discrimination tasks such as distinguishing tumor grades, identifying WHO pathological subtypes, or assessing associations between baseline patient characteristics and chemotherapy response or recurrence risk (Table 4 ). Representative algorithms include logistic regression [ 12 ], support vector machines (SVMs) [ 13 ], and random forest [ 14 ]. Gradient boosting decision tree models such as XGBoost, which uses a pre-sorting algorithm [ 15 ], and LightGBM, optimized through histogram-based techniques, have also gained wide application due to their strong performance in handling complex clinical variables [ 16 ]. The third category consists of unsupervised models, mainly used for dimensionality reduction, clustering, and visualization of high-dimensional data. These approaches are frequently applied for analyzing patient heterogeneity and molecular subtypes. Common methods include K-means, which clusters data by minimizing within-cluster variance [ 17 ], principal component analysis (PCA), which preserves global structure during dimensionality reduction [ 18 ], and uniform manifold approximation and projection (UMAP), which maintains both local and global structures [ 19 ]. Finally, feature selection methods play a critical role in identifying tumor biomarkers derived from serum, stool, or pathological tissues. Widely used techniques include LASSO regression with L1 regularization, random forest based on feature importance, and recursive feature elimination [ 20 , 21 ]. Table 4. Overview of AI models for patient stratification and treatment decision-making in GI tumors Data type Basic algorithm Advantages Limitations Task description PMID CT Images CT Images LR, DT Integrates immunophenotype-guided radiomics for accurate, interpretable pCR prediction in d-MMR/MSI-H CRC Small sample size and cross-center heterogeneity limit generalizability and biological validation mmune phenotypes predict anti-PD-1 response in d-MMR/MSI-H CRC 40,759,438 SVM, LASSO Accurately predicts peritoneal recurrence and chemotherapy benefit Retrospective design and CT protocol dependency Rad-score predicts PR and chemo-benefit in GC 37,300,884 CNN Non-invasively evaluates TME and predicts chemo-immunotherapy benefit Retrospective design, single-region cohort DLRS stratifies GC anti-PD-1 benefit and adjuvant chemo outcome 37,557,177 CNN Large-scale multicentre CT-DL fusion model accurately predicts gastric cancer postoperative recurrence Retrospective design; chemotherapy impact not clarified; Asian population only Predict GC recurrence and stratify prognosis 38,896,865 CNN DenseNet-169 2.5D-CT MLP fusion predicts LAGC early relapse accurately across centers Retrospective design, 2.5D tumor sampling, limited transcriptomic set weaken biological proof DLERMLP score predicts 1-yr recurrence after LAGC resection 39,715,142 Multi-Head Attention mechanism Explainable AI integrates CT-RNA-pathology to guide stage II CRC chemo Retrospective, borderline significance, small RNA set Stage II CRC radiophenotypes guide adjuvant chemo omission or use 40,472,802 CNN, Attention mechanism Accurately predicts response and prognosis by combining imaging and clinical data Retrospective design limits generalizability Deep signature predicts NCT response in LAGC via signet-ring and T-stage 37,132,183 CNN, Attention mechanism Biology-guided multitask model accurately predicts prognosis and immunotherapy response Retrospective design, manual tumor segmentation required Predict GC TME and survival; flags chemo-benefit and immunotherapy-refractory dMMR 37,612,313 CNN, Attention mechanism Multitask CT-MDL accurately predicts TSR and chemotherapy benefit in multicentre CRC Retrospective design, single-slice input, histological TSR unavailable in external cohort 3 Non-invasively predict CRC TSR; high-score gains adjuvant chemo benefit 38,348,900 Swim-Transformer Provides precise four-tier risk stratification Relies on retrospective multicenter cohorts CRC risk stratification reveals enhanced immune activity in low-risk patients 40,480,552 CNN, Transformer 3D transformer DLN predicts LNM after NAC in LAGC robustly across centers Retrospective, low specificity, biological black box Predict LAGC LNM post-NAC via intratumor heterogeneity and invasive margin 39,281,097 Multi-Head Self-Attention mechanism, Swim-Transformer Multitask Swim-Transformer accurately predicts chemo-immuno response in GC from CT Retrospective, East-Asian only, scanner heterogeneity, no perioperative ICI validation Predict 5-FU response via ImmunoScore and POSTN in TME 40,695,288 CNN, Attention mechanism, GCN, WGAN Federated learning preserves privacy while robustly predicting postoperative GC recurrence across centers Retrospective design with manually defined ROIs limits generalizability Multicenter framework identifies high-risk GC recurrence 38,272,913 Co-Attention mechanism, Mamba Longitudinal CTSMamba multitask model predicts LNM and OS in NAC-treated LAGC across centers Retrospective, NAC-regimen heterogeneity, biological interpretability pending Simultaneously predicting LNM and OS in patients with LAGC following NAC 40,305,075 CNN, Global-Attention mechanism, Vision-Mamba Vision-Mamba fuses voxel-radiomics with CT to predict ESCC pCR after nICT across centers Retrospective, pCR imbalance, external validation limited to China Predicting pCR following nICT in patients with ESCC 40,090,670 MRI LASSO, MLP Integrates radiomics and clinical features accurately Small cT4 cohort, CEA dominance limits radiomics Rad-model predicts T-downstaging in cT4 rectal cancer 40,514,006 CNN, Vision-Transformer Provides preoperative independent prognostic value Relies on retrospective multicenter data CEA integration enhances survival prediction and risk stratification 37,278,629 Endoscopic Images CNN, Transformer Enhances feature complementarity and extraction efficiency, addresses sample imbalance Requires large-scale data, demands high computational resources, depends on high-quality imaging Intraoperative WLI + FI + PCI predict CRC LNM 40,030,456 WSIs WSIs WSIs RF, LASSO ML-based CD3 score refines DFS prediction Single-marker CD3 limits immune contexture Digital CD3⁺ score predicts 5-year DFS in CRC 38,880,067 CNN Multi-omics plus AI precisely stratify ESCC Single-center cohort hinders generalizability Four ESCC subtypes: differentiated, immune, metabolic, stem-like; stem-like worst 39,419,971 CNN Indirect two-step strategy bypasses treatment-data shortage Unproven in prospective trials, clinical utility pending In-silico model predicts GI tumor transcriptome and therapy response 38,961,276 CNN Integrates multiomics and deep learning for precision targeting Lacks multicenter validation and mechanistic depth CCIM module: FOLR2⁺ macrophage-centric immunosuppression in CRC 38,307,032 CNN Achieves high accuracy in MMRd detection at single-cell level Limited to TMA cores; lacks whole-slide validation Single-cell assay classifies CRC MMR status and triages Lynch 39,293,403 CNN Uses deep learning to predict protein biomarker expressions from H&E stains Performance varies across biomarkers; limited generalizability to unseen data Infer protein from HE, virtual mIF maps TILs and micro-domains 40,819,165 U-Net, CNN Innovatively quantifies HER2 expression at pixel-level across multicenter datasets Relies on limited IHC labels and lacks advanced model architectures Predict HER2 in GC; misclassified tied to glandular/papillary morphology 39,792,693 CNN, Attention mechanism Self-attention CNN boosts MSI prediction accuracy Ethnic drift limits cross-population generalizability Predict MSI and 5-year DFS; attention maps flag poor/mucinous CRC 36,720,223 CNN, Attention mechanism Demonstrates robust generalizability across cohorts using self-supervised attention-based MIL Limited predictive power for KRAS, NRAS, and PIK3CA mutations Predict MSI/BRAF/KRAS/NRAS/PIK3CA from CRC WSIs 36,958,327 CNN, Attention mechanism End-to-end prognostication across multicenter cohorts with open-source models Retrospective design limits causality; external validation affected by follow-up and age heterogeneity Predict CRC post-op OS/DSS; low-risk score enriches CD8⁺/CD4⁺/M1 immunity 38,123,254 CNN, Attention mechanism Introduces regression-based CAMIL for continuous biomarkers across nine cancer types Limited external validation and hyperparameter tuning; regression robustness underexplored Quantify CRC/GC proliferation, stroma, and immune metrics 38,341,402 CNN, Slef-Attention mechanism Self-supervised extraction of interpretable histomorphological patterns linked to survival and therapy response Tile-level context loss limits precise grading and stroma-muscle distinction Map 47 CRC HPCs into 8 superclusters for survival-therapy stratification 40,057,490 CNN, Self-Attention mechanism Self-supervisedly maps interpretable HPCs linking morphology to survival across cancers Lacks spatial context and high-resolution detail for subtle pattern discrimination Morphologic landscape predicts GI tract OS/RFS: solid high-risk, inflammatory good 38,862,472 CNN, Gated-Attention mechanism Multistain fusion boosts relapse prediction accuracy over single-stain and clinical models Requires multi-IHC staining and lacks external validation for neoadjuvant response prediction AI immune score (AIS) reverses T-stage prognosis in CRC NCT 36,624,314 CNN, Transformer Teacher-student MILTS weakly predicts pan-cancer PDL1 from H&E with interpretable morphotypes Fresh-frozen/FFPE performance gap and mRNA-histology misalignment limit robustness Predict CRC PD-L1 via eosinophil-rich stroma and cribriform pattern 38,594,278 Multi-Head Self-Attention mechanism, Vision-Transformer Achieves high accuracy and efficiency in TMB prediction Relies on costly whole-exome sequencing for TMB labeling Predict CRC subtypes and TMB to stratify ICI benefit 39,461,079 CNN, Multi-Head-Attention mechanism, Transformer Meta-learning optimizes micro-metastasis detection with limited data Requires high-resolution patches and frozen-section retraining Detect LN micrometastases and ITCs in CRC pathology 40,634,485 CTransPath, Transformer Develops multi-target transformer model for predicting genetic alterations in colorectal cancer Limited generalizability due to under-representation of non-White individuals Integrate MSI/TMB/BRAF/KRAS/TP53/RNF43/BMPR2/ZNRF3 for CRC prediction 40,829,965 CNN, CTransPath, Transformer Transformer-based model achieves high performance and clinical-grade prediction for MSI and other biomarkers Model performance may be affected by imperfect ground truth and limited biopsy data Predict MSI/BRAF/KRAS from CRC WSIs: mucinous/tumor regions key 37,652,006 DINO, Multi-Head Self-Attention mechanism, Transformer Develops COFFEE model for precise HGP classification in CRLM Single-center design limits generalizability; future multi-center studies needed Classify CLM HGPs: desmoplastic predicts longer OS/PFS 40,638,258 UNI, Transformer Combines ctDNA and deep learning for enhanced risk stratification Limited by short follow-up and potential model updates Predict CRC DFS by ctDNA-MRD; histology complements 40,813,777 CycleGAN, CNN Enhances diagnostic accuracy and efficiency for pathologists Requires further prospective validation and generalisation Detect EAC/AEG tumor; assess NAT histologic regression 37,100,542 CNN, Attention mechanism, GAN, Vision-Transformer Improves prognosis prediction via semi-supervised learning and knowledge distillation Limited dataset diversity and hyperparameter analysis The model was developed to predict OS, TTR, and TRG in patients with CLM 39,423,564 Molecular Profiling Molecular Profiling Molecular Profiling MLP Accurate and interpretable immunotherapy response prediction Complex tumor microenvironment and data needs Predict GC ICI response: Th1/MDSC/M2; STAT4/IRF4 in TME 39,167,797 MLP Accurate cancer subtyping with data privacy Data harmonization and missing data challenges Federated learning enables CRC/EAC subtyping, biomarker discovery 40,488,620 MLP High accuracy and generalization for malignant cell annotation Limited in hematologic tumors, ignores spatial relationships Automate cancer cell annotation; upregulate EPCAM/KRT8/KRT18/SOX2 38,431,724 XGBoost Identifies ecDNA as a biomarker for cancer prognosis and immunotherapy Limited sample size and uneven cancer type distribution Identify ecDNA amplification in CRC WES; MSI exclusion, genomic subtyping 38,373,991 SVM, MLP Accurate ctDNA quantification from cfDNA fragment lengths Limited to whole-genome and targeted sequencing data Quantify CRC/GC ctDNA VAF heterogeneity and radiographic concordance 40,055,581 SVM, RF, LASSO, MLP High-performance GC diagnosis and prognosis via plasma metabolic fingerprinting Limited to Chinese population; biological significance of metabolites unclear 21-metabolite panel, PMP score predict GC survival 37,460,165 SVM, GBM, RF, LASSO, MLP Identifies key predictors of radiotherapy response Limited to specific treatment and patient cohort TME immune/TGFβ predict rectal cancer pCR to radiotherapy 39,013,324 SVM, GBM, RF, LASSO, XGBoost Predicts immunotherapy response and identifies AG-538 as an immunity enhancer Limited to specific cohorts; needs clinical validation Predict CRC immunotherapy response; AG-538 activates cGAS/STING 39,879,113 SVM, RF, DT, XGBoost, MLP Provides comprehensive TME subtyping, guiding personalized immunotherapy Limited by data availability and tumor types studied Immunotyping TME: ARID1A/PIK3A in IA, IL-1 in IS 40,433,880 SVM, RF, LASSO, XGBoost, MLP Develops DNN predictor for NCRT sensitivity in LARC Limited by sample size and data availability NCRT in LARC: Wnt/β-catenin vs. LA metabolism 38,232,812 8 machine learning algorithms and the execution of 103 algorithmic combinations Develops LPS model for ESCC prognosis Limited by sample size and data availability LPS predicts survival, chemo/radio response; high LPS for BI-2536/panobinostat 39,864,540 CNN Provides non-invasive monitoring of tumor evolution Requires further validation in larger cohorts CMS transitions predict CRC recurrence; EV RNA reveals oncogenic pathways 38,451,249 CNN Enhanced detection capability and wide cancer type coverage Technical and resource intensive, tumor tissue requirement ctDNA fragmentation enhances MRD/response monitoring in colorectal adenomas 38,877,116 U-Net, MLP Identifies effective T-cell infiltration strategies and predicts personalized therapies for diverse cancers Depends on availability of spatial omics data and requires experimental validation for causal proof Minimize CLM TME perturbation via CXCR4/PD-1/PD-L1/CYR61 strategies 40,044,819 GCN Identifies tissue cellular neighborhoods (TCNs) from spatial omics data Performance depends on quality of cell-type annotations Granulocyte-rich TCNs predict high-risk colon cancer via TME remodeling 38,191,930 Attention mechanism Accurately predicts tumor types and subtypes from somatic mutations Performance may degrade with low mutation burden or sparse data Deep mutation learning predicts GI cancer subtypes 37,420,249 GAT IRnet improves prediction accuracy and interpretability IRnet is limited by scarce patient data and batch effects Predict GC ICI response via TME markers and pathway interactions 39,097,091 GAT CellNEST detects CCC with high resolution Limited by spatial transcriptomic data Detect CRC ligand-receptor interactions and relay networks 40,481,363 CNN, GAT HistoCell accurately infers super-resolution cell spatial profiles from histology images HistoCell Limited by cell segmentation accuracy in histology images Infer cell types, states, and spatial networks at single-cell resolution 39,984,438 Transformer HLApollo achieves high accuracy in predicting peptide-MHC-I presentation Dependent on large-scale immunopeptidomics data availability Predict CRC peptide-MHC-I complexes for immunotherapy targets 39,737,928 CNN, Swim-Transformer Transforms transcriptomic data into SIEs for accurate prognosis prediction Retrospective study design limits clinical applicability; identified genes require further validation Predict CRC risk/prognosis via 24 genes; PEX10 key biomarker 40,515,391 GCN, VGAE Predicts metabolite–protein interactions using variational graph autoencoders with high accuracy Data imbalance and experimental validation needed Predict CRC metabolite-protein links; expand interaction network 37,225,420 Clinical Data RF Applies AI to identify GIST patients who benefit from imatinib and determine optimal treatment duration Hindered by observational design and incomplete data Optimize GIST adjuvant imatinib strategy; identify no-treatment subgroup 38,976,997 RF Uses ML to predict CRLM patient outcomes Retrospective design; potential bias and inconsistent complication recording Predict CLM post-op complications, PFS, OS via KRAS/BRAF/MMR 38,768,679 LR, MLP Develops model to accurately predict lymph node metastasis in non-intestinal early gastric cancer Retrospective design; lacks external validation Predict non-intestinal GC LNM risk; guide lymphadenectomy extent 40,265,477 LR, RF, XGBoost Develops ML models to predict lymph node metastasis in T1 esophageal squamous cell carcinoma with high accuracy Retrospective design; limited clinicopathological factors Predict T1 ESCC LNM risk via LVI and pT stage 38,905,510 Multi-omics Multi-omics Multi-omics RF Fuses PET–MRI to map intratumour heterogeneity with histological agreement Low DSC in therapy groups and limited patient cohort hinder robustness Quantify CLM tissue types: apoptosis at Th-24, necrosis/fibrosis at Th-72 37,277,483 SVM Decode resistance with spatial multi-omics and predict response via SVM Small cohort, regimen-specific; mechanistic markers await causal validation Predict GC chemo/immuno response; TGFB1-HSPB1 axis, LTF-S100A14 interaction 38,579,635 LR, XGBoost Pan-cancer proteogenomics plus XGBoost yields high-resolution functional network outperforming PPI Bulk-tissue data obscure cell-type specificity; low-mutation driver prediction awaits experimental proof Protein co-expression outperforms mRNA; ECM/angiogenesis modules predict poor CRC survival 39,663,389 SVM, DT, MLP PDC-Cmax-derived signatures accurately predict 5-FU response across four cohorts and guide alternative targeted therapy Models lack tumor microenvironment and mutation integration fails to boost performance; larger cohort validation needed CRC 5-FU resistance: gene signature, chr7 gain; vemurafenib/regorafenib potential 40,187,357 RF, CNN Multimodal integration refines moderate-risk patient stratification Single-center validation, stain variability sensitivity Stratify CRC risk via TME features: anisonucleosis, nuclear volume, low immune infiltration 37,499,603 RF, CNN Multimodal fusion enhances bevacizumab response prediction Retrospective design, manual ROI segmentation required Predict CLM bevacizumab efficacy/prognosis; responders gain PFS/OS/R0 resection 37,869,523 CNN,GAT Integrates CT and pathology via contrastive learning Manual annotation limits clinical scalability Predict LAGC neoadjuvant chemo response via inflammatory infiltration 39,637,859 CNN, GNN, GAT Integrates image and graph features for ST prediction Limited cross-cancer generalization and platform coverage Dual-linkage model predicts CRC spatial gene expression for risk/survival 38,697,103 GCN, GAT Integrates multi-omics via explainable GNN for MSI prediction Relies on correlative networks lacking causal inference Predict MSI status via MLH1/MSH4/RPL22L1; MSI/immunotherapy link 37,833,839 Self-Attention, Multi-Attention, Gated-Attention mechanism Generates virtual mIF, enabling rapid and low-cost immune profiling Limited biomarkers, requires external validation and manual ROI initialization Virtual mIF: CD3/CD20/CD8/PD1/CPS predict longer OS; FOXP3/CD163 shorter 39,154,539 CNN, Attention mechanism Integrates SNA features to boost interpretability and accuracy of mutation prediction Lacks nuclei type information and external validation datasets Predict CRC CIN, MSI, TP53, BRAF mutations 38,199,068 LR, SVM, RF, DT, LASSO, XGBoost, CNN, Attention mechanism Multimodal fusion enhances immunotherapy response prediction with biological interpretability Retrospective design, small external cohorts, complex feature extraction Predict GC combined immunotherapy response via RNA-seq 40,675,469 CNN, Transformer Integrates multi-modal data to accurately predict gastric cancer response to anti-HER2 therapies Relies on human input for certain sub-tasks Predict GC anti-HER2/immunotherapy response via male, diffuse type, TILs 39,183,247 CTransPath, Cross-Attention mechanism, Transformer Employs cross-attention to fuse multimodal data for survival prediction Limited generalizability due to TCGA-based training Predict CRC outcomes via attention heatmaps: tumor/immune cell abundance 40,116,660 CTransPath, DINOv2, GAT Leverages foundation models and hybrid transformers to enhance spatial domain identification and gene expression denoising in multimodal single-cell ST data Lacks single-cell-specific histology encoding and omits transcriptomic foundation model integration, limiting generalizability Identify CRC proliferative regions with MKI67/KRAS/BRAF/MUTYH 39,830,364 UNI, CNN, Transformer Bridges WSIs and omics for prognosis with missing modality Black-box model lacks biological interpretability Predict CRC prognosis via WSIs without omics data 40,051,298 Vision-Transformer, scBERT, Cross-Attention and Self-Attention mechanism, Transformer Robustly aligns histology and gene expression across cancers Lacks single-cell and transcriptomic foundation support Predict CRC spatial gene expression: PPIA, RPL9, RPL35A 40,618,351 self-VAE, cross-VAE Pretrains pan-cancer multi-omics fusion, enabling robust cross-modal inference and downstream oncology tasks Demands large paired datasets, omits causal omics relations and single-cell resolution Cross-omics inference enables CRC risk/survival prediction via mitochondrial/HLA/keratin genes 38,845,006 VGAE, GNN Dynamic graph fuses multi-slice omics, zero-shot links clinical niches Lacks causal spatiotemporal relations, moderate prediction accuracy Identify CLM niche: SPP1⁺/MTRNR2L12⁺ myeloid cells, CAFs, short DFS 40,523,901 CNN, Cross-Attention mechanism, VGAE Cross-attention fuses transcriptome and morphology, accurately demarcating spatial domains and stratifying patients Absence of single-cell resolution and causal inference constrains in-depth biological interpretation Multimodal CRC analysis: MYC targets, B-cell regulation, IG gene expression 40,407,386 Open in a new tab NCT/NAC Neoadjuvant Chemotherapy, PR Peritoneal Recurrence, TME Tumor Microenvironment, dMMR/MMRd Mismatch Repair deficiency, GCN Graph Convolutional Network, WGCN Wasserstein Generative Adversarial Network, TSR Tumor-Stroma Ratio, pCR pathological Complete Response, LAGC Locally Advanced Gastric Cancer, DLN Deep Learning Network, LNM Lymph Node Metastasis, ICI Immune Checkpoint Inhibitor, OS Overall Survival, nICT Neoadjuvant Immunochemotherapy, WSIs Whole-slide Images, DFS Disease-Free Survival, TMA Tissue Microarray, IHC Immunohistochemical, MIL Multiple Instance Learning, DSS Disease-Specific Survival, CAMIL CNN-attention Multiple Instance Learning, HPC , Histomorphological Phenotypic Cluster, RFS Recurrence-Free Survival, FFPE Formalin-fixed Paraffin-embedded, TMB Tumor Mutational Burden, LN Lymph Node, ITCs Isolated Tumor Cells, HGP Histopathological Growth Pattern, CLM/CRLM Colorectal Liver Metastases, PFS Progression-Free Survival, MRD Molecular Residual Disease, EAC Esophageal Adenocarcinoma, AEG Adenocarcinoma of Esophagogastric Junction, ecDNA extrachromosomal DNA, NAT Neoadjuvant Therapy, TTR Time To Recurrence, TRG Tumor Regression Grade, WES Whole-Exome Sequencing, VAF ariant Allele Frequency, NCRT Neoadjuvant Chemoradiotherapy, LARC Locally Advanced Rectal Cancer, CMS Consensus Molecular Subtype, TCN Tissue Cell Neighborhood, GAT Graph Attention Network, CCC Concordance Correlation Coefficient, SIE Synthetic Image Element, VGAE Variational Graph Autoencoder, GIST Gastrointestinal Stromal Tumor, LVI Lymphovascular Invasion, TIL Tumor-Infiltrating Lymphocyte, mIF multiplex Immunofluorescence, WLI White Light Imaging, FI Fluorescence Imaging, PCI Pseudocolor Imaging, DSC Dice Similarity Coefficient, PPI Protein–protein Interaction, ECM Extracellular Matrix, SNA Second-order Nuclear Attribute, CIN Chromosomal Instability, ST Spatial Transcriptomics, CAFs Cancer-associated Fibroblasts In oncology research, given the diversity of data types and the distinct purposes, the performance of these algorithms exhibits substantial variability. The effectiveness of supervised learning models is often impaired by the size of the training dataset. Many classical algorithms are limited to capturing linear relationships in the original feature space. As a result, feature engineering or kernel-based approaches are frequently required to achieve adequate performance when the data display complex and non-linear patterns. Although models such as SVMs and random forests can handle high-dimensional data and exhibit strong predictive power, their decision-making processes often lack transparency, posing challenges for clinical interpretation. In the context of unsupervised learning, approaches such as K-means clustering are highly sensitive to initial conditions and outliers, which compromises their stability. Similarly, dimensionality reduction methods such as PCA and UMAP, albeit efficient for data compression and visualization, may inadvertently discard biological information, thereby affecting the integrity of downstream analyses. Image-based deep learning models In oncology, medical imaging provides critical morphological and contextual information that underpins deep learning model development for diagnosis, staging, and treatment. Image-based deep learning models extract hierarchical spatial features from pixel-level inputs. These models currently include Multilayer Perceptrons (MLP) [ 22 – 25 ] which serve as simple feed-forward networks that map input pixels to output predictions, albeit with limited spatial modeling capacity. CNNs, such as ResNet, [ 26 – 31 ], learn local spatial patterns through stacked convolutional layers and hierarchical feature extraction. In contrast, GNNs represent image regions or objects as nodes and their spatial or semantic relationships as edges [ 32 – 34 ], enabling the model to capture non-Euclidean structures and contextual dependencies across spatially distributed features. Furthermore, transformer-based models leverage self-attention mechanisms and positional encoding to model long-range dependencies and global context within medical images [ 35 – 39 ] (Tables 1 , 2 , 3 and 4 ). Table 1. Application of AI models in GI endoscopy Anatomic site Basic algorithm Advantages Limitations Task description PMID Esophagus LightGBM, XGBoost, RF, and SVM Non-invasive, high-accuracy tool avoiding ~ 90% endoscopies Requires validation in non-Chinese populations Screen for ESCC and EJA 36,931,287 CNN High sensitivity; significantly improved neoplasia detection by general endoscopists; validated on a large, multicenter dataset Lower specificity than endoscopists; performance assessed ex vivo, requires prospective clinical trial validation Detect neoplastic lesions in BE 38,000,874 CNN Enables non-experts in low-volume centers to achieve expert-level complete resection rates Significantly increases procedure time Delineate lesion margins for ESD 40,540,547 CNN Prospective RCT demonstrated doubled detection rate for high-risk lesions with high sensitivity Single-center study; generalizability requires further validation Detect precancerous high-risk esophageal lesions 38,630,847 CNN, eUnet, Edge-Attention mechanism Achieves high segmentation accuracy (Dice 0.888) with complete patient privacy Requires manual initialization and is computationally intensive per image Delineate early-stage EC lesions 38,530,733 Gastric Unet + +, CNN Real-time CDSS with multicenter validation for gastric lesion classification Limited dataset; detection improvement not statistically significant Classify gastric lesions and predict invasion depth 36,754,065 Colorectum Colorectum CNN Real-time CADx with high success rate, comparable accuracy to endoscopists Suboptimal specificity (38%), low SSL sensitivity (17.1%) Identify diminutive colorectal polyps, including SSLs 36,623,839 CNN State-of-the-art accuracy, excels at segmenting very small polyps Computationally complex, performance drops on challenging cases Delineate neoplastic diminutive polyps 37,768,798 CNN Real-time polyp sizing with high accuracy (0.89), reduces inappropriate surveillance recommendations Relies on simulated data for depth estimation; real-world clinical variability may affect performance Measure colorectal polyp size 37,827,513 Attention mechanism Integrates frequency domain learning for robust polyp segmentation Higher computational complexity from spectral processing Achieve precise identification of diverse colorectal polyp lesions 39,847,953 CNN, Attention mechanism Color-invariant, uncertainty-aware, real-time Weak on multi-class, similar-texture, shadow cases Segment colorectal polyps 39,316,997 Contrastive Transformer Transformer contrastive learning, multiscale, strong generalization ability Heavy, polyp-specific, small-object bias Segment inconspicuous colorectal polyps 38,470,573 LLM, Diffusion PSDM augments out-of-distribution, prompt-rich, continual learning High cost, rare-prompt failure, slight identification accuracy drop Classify and segment colorectal polyps 40,663,685 Open in a new tab GI gastrointestinal, LightGBM Light-Gradient Boosting Machine, XGBoost Extreme Gradient Boosting, RF Random Forest, SVM Support Vector Machine, CNN Convolutional Neural Network, LLM Large Language Model, ESCC Esophageal Squamous Cell Carcinoma, EJA Esophagogastric Junction Adenocarcinoma, BE Barrett’s Esophagus, ESD Endoscopic submucosal dissection, RCT Randomized controlled trial, EC Esophageal Cancer, CDSS Clinical Decision Support System, CADx Computer-aided Diagnosis, SLL Sessile serrated lesion Table 2. Application of AI models in liquid biopsy of GI cancers Anatomic site Basic algorithm Advantages Limitations Task description PMID Esophagus RF, XGBoost, MLP High AUC, sensitivity, specificity, robust at low coverage Limited to Chinese cohort, small sample size, potential overfitting Early screening and risk stratification for ESCC 39,089,259 Gastric LASSO, LR High AUC, superior prediction Retrospective bias, variable performance Predict GC peritoneal metastasis for therapy stratification 39,985,379 LASSO, RF High sensitivity, specificity, validated. Superior to traditional methods Short follow-up, potential overfitting. Requires metabolomics data Identify 17-metabolite GC panel & reprogramming map 38,395,893 LASSO, LR, RF, SVM, XGBoost High AUC, superior to traditional markers Effective for early-stage and biomarker-negative cases Retrospective, potential bias Needs larger validation Discover 31 GC biomarkers & 4 lncRNA diagnostic score 39,753,334 Colorectum MLP High sensitivity, specificity for ctDNA detection. Improves cancer detection, variant calling Requires deep sequencing data Dependent on lab protocols, training data Call CRC mutations, estimate tumor fraction 37,121,998 MLP Combines EV isolation with AI analysis Identifies optimal biomarker combinations Needs larger and more diverse cohorts Requires further validation for clinical use EV-miRNA + AI/CEA CRC diagnosis 40,511,723 LASSO, RF Identifies key CRC biomarkers PF4 and AACT.Develops accurate RF diagnostic model Sample size not large enough Needs expanded population range Screen and validate CRC biomarkers PF4 & AACT 39,168,094 RF, XGBoost Integrates ML and XAI to identify CRC biomarkers Needs larger cohorts and clinical validation Identify microbiome markers for CRC and adenoma risk 40,760,681 GBM, RF, XGBoost, MLP High sensitivity and specificity Effective for early-stage colorectal cancer detection Requires WGS data Dependent on lab protocols and training data Early cancer screening and risk stratification for CRC 39,073,362 SVM, RF, LASSO, GBM, MLP Identifies key multi-cancer biomarkers Develops accurate diagnostic models Sample size needs expansion Needs inclusion of more benign diseases Narrow CRC biomarkers to 12-ETR.sig and map pathways 40,025,576 LR, RF, SVM, XGBoost and MLP Detects early CRC with high accuracy Non-invasive and AI-enhanced Limited sample size Single-center study 14 gut-microbiota biomarkers for early CRC and metastasis 39,113,791 XGBoost, VAE Combines oncRNA biomarkers with AI Achieves high sensitivity in early-stage CRC Validation set lacks diversity Lacks advanced adenoma samples Early CRC screening via oncRNA score tracking tumor burden 40,366,744 Gastric, Colorectum RF Achieves high sensitivity and specificity for multicancer detection using L1PA hypomethylation Needs larger cohorts and further validation for clinical application Detect ctDNA via L1PA hypomethylation 39,620,930 SVM, MLP, CNN Combines exosome isolation with AI for multi-cancer detection Needs larger cohorts and clinical validation Screen/diagnose colorectal & gastric adenocarcinoma and infer origin 36,964,142 Esophagus, Gastric, Colorectum RF, GBM Integrates multi-feature cfDNA fragmentomics for sensitive detection of Esophageal and gastric cancers Needs larger cohorts and clinical validation Discriminate AG/GC/BE/EC via CNVs & fragment-end motifs 40,712,972 Open in a new tab MLP Multilayer Perceptron, LR Logistic Regression, GBM Gradient Boosting Machine, VAE Variational Autoencoder, AUC Area Under the Curve, GC Gastric Cancer, CRC Colorectal Cancer, EV Extracellular Vesicles, PF4 Platelet factor 4, AACT Alpha-1-Antichymotrypsin, ML Machine Learning, XAI Explainable Artificial Intelligence, WGS Whole Genome Sequencing, AG Atrophic Gastritis, CNV Copy Number Variation Table 3. Application of AI models in histology and other clinical screening methods of GI cancers Anatomic site Basic algorithm Advantages Limitations Task description PMID Esophagus MLP, CNN, Transformer H&E-only BE detection Limited generalizability Identify IM features in BE on H&E, skip TFF3 38,467,600 Gastric U-Net Enables large-scale noncontrast CT screening Needs prospective validation and better EGC sensitivity Screen GC with noncontrast CT 40,555,751 CNN, Attention mechanism, Transformer, DT, SVM High diagnostic accuracy, non-invasive screening, multi-center validation Limited population, excluded conditions, requires further validation Distinguish GC from NGC, including early GC and precancerous AG 36,825,238 Colorectum CNN Achieves high sensitivity, enables workload reduction, validates across multiple centers Suffers from class imbalance, shows limited generalization, depends on label accuracy Differentiate CRCs, identify inflammation, polyps, and dysplasia 37,890,902 U-Net Detects cancer without bowel preparation; improves radiologists' performance; validated across multiple centers Poor performance on benign lesions; lacks prospective validation Screen CRC by imaging: masses, infiltration, wall thickening, lymphadenopathy 38,848,616 MLP, U-net, CNN High sensitivity and scanner robustness Low specificity and calibration dependency Discriminate MSI from non-MSI CRC 37,932,267 GNN High accuracy and strong interpretability Requires extensive annotations and limited feature scope Highlight biopsy ROIs: report glands, inflammation density, cell-gland-epithelium spatials 37,173,125 Open in a new tab DT Decision Tree, GNN Graph Neural Network, H&E Hematoxylin and Eosin, IM Intestinal Metaplasia, TFF3 Trefoil Factor 3, EGC Early Gastric Cancer, NGC Non-gastric Cancer, MSI Microsatellite Instability, ROI Regions of Interest Additionally, as a CNN variant, the U-Net architecture adopts an encoder–decoder structure with skip connections to integrate both low-level and high-level features. Owing to its pixel-wise mask classification capability, U-Net has been widely applied in medical image analysis, particularly for tumor segmentation in CT and MRI scans, as well as for tasks such as quantifying tumor, immune, or stromal cells and supporting downstream analyses on whole-slide pathology images (WSIs) [ 40 – 43 ] (Table 3 – 4 ). MedFormer synergizes the strengths of CNNs and Transformers by employing a CNN-based encoder, embedding attention modules within each layer of its structure, and incorporating a spatial attention fusion module in its decoder, benchmarked against U-Net (mIoU:88.17%, precision:93.84%, recall:93.06%), and TransUnet (which inserts Transformer modules into the encoder) (mIoU:79.95%, precision:87.63%, recall:87.34%). Therefore, MedFormer shows superior performance for identifying subtle tissue structures and boundary features (mIoU:90.33%, precision:95.64%, recall:94.33%), along with remarkable adaptability across various tumor imaging tasks [ 44 , 45 ]. Generative models such as VAEs, GANs, and Diffusion Models have demonstrated potential in processing image data [ 46 – 48 ] (Table 1 , Table 2 , Table 4 ). For example, the TumorGen framework integrates boundary-aware masking, correction mechanism, and a VAE architecture to synthesize tumor images of specific stages and types, thereby mitigating issues of limited training data and insufficient quality [ 49 ]. In contrast, the two-stage model TSGM, comprising CycleGAN and a Diffusion model, achieves more precise tumor anomaly localization and segmentation, a capability lacking in TumorGen [ 50 ]. Foundational model is another transformer-based architecture pre-trained on large-scale datasets for domain-specific applications. Its primary advantage lies in learning common features and optimizing weight distributions from heterogenous data, thereby enhancing generalization and reasoning capabilities when confronted with new data and complex scenarios, such as multi-center data, rare cases, or inconsistent image quality. In addition, foundational models can automatically extract imaging features, converting raw images into information-dense embedding vectors that are more suitable for predictive modeling. Alternatively, following a transfer learning paradigm, they can be adapted for various clinical tasks with high accuracy. Recent years have witnessed the development of several novel frameworks that surpass previous pathology foundation models like CLAM and CTransPath in feature extraction performance [ 51 , 52 ] (Table 4 ). For instance, UNI, based on a Transformer architecture, excels at automatically extracting highly generalizable and high-quality features from tumor images [ 53 ]. Furthermore, CHIEF employs a hierarchical cross-scale feature fusion mechanism to effectively integrate global contextual and local heterogeneous information from images [ 54 ]. Prov-GigaPath enhances the transferability of the generated features by reinforcing feature representation at a small local image region [ 55 ]. After pretraining phase, for clinical tasks such as early screening, tumor subtyping, and prognosis assessment, different fine-tuning strategies are developed to enhance the applicability of foundation models in GI cancers. For example, GastroNet-5 M is adapted from self-supervised pre-trained models (SimCLRv2, MoCov2, DINO) through domain-specific fine-tuning and label-knowledge distillation, enabling efficient identification of dysplastic vessels and polyps in endoscopic images (AUC:85% in polyp task, AUC:95% in invasion depth task) [ 56 ]. CRCFound, based on a Vision Transformer architecture, predicts TNM staging, microsatellite instability status, consensus molecular subtypes, and survival risk from 3D CT images of colorectal cancer patients. However, its utility has not been generalized to all GI cancers [ 57 ]. DINOPath, fine-tuned from the self-supervised DINOv2 model using whole-slide images (WSI) and clinical labels, accurately predicts patient disease-free and disease-specific survival and identifies high- and low-risk groups for neoadjuvant chemotherapy response, thereby aiding personalized treatment decisions [ 58 ]. Notably, compared to existing large-scale models in analyzing tumor imaging data, such as UNI, CHIEF, and Prov-GigaPath, the MUSK model emphasizes more on the pragmatic demands of clinical deployment, particularly inference speed and diagnostic timeliness, by employing a model distillation technique that utilizes a knowledge transfer module to guide a compact student model (designed for clinical deployment) in learning multi-scale feature representations from a large teacher model (the foundation model). This approach enhances model performance while significantly reducing the parameter count, thereby improving operational efficiency and better aligning with real-time clinical application scenarios [ 59 ]. Furthermore, the PathoDuet model incorporates novel mechanisms such as a cross-scale positioning module and a cross-stain transferring module. By leveraging a cross-scale localization and transfer algorithm, PathoDuet achieves effective integration of hematoxylin–eosin (H&E) and IHC images, which enhances the adaptability and representational capacity for multimodal WSI analysis [ 60 ], thus addressing the limitations of other models like UNI and MUSK in seamlessly integrating WSI styles. In addition to model architecture, data preprocessing and augmentation are indispensable components of practical oncology imaging pipelines, as they substantially improve model stability and generalizability. Standard preprocessing techniques, such as intensity normalization, stain normalization for WSI, motion artifact correction, and histogram equalization, can effectively mitigate domain shifts introduced by scanner variability and staining inconsistencies across clinical sites [ 61 ]. Meanwhile, image augmentation strategies, including random flipping, elastic deformation, color jittering, and contrast-limited adaptive histogram equalization (CLAHE) [ 62 ], have been shown to reduce overfitting and enhance robustness, particularly in small-scale tumor datasets. While synthetic images generated by generative models offer a promising way to augment training datasets and compensate for limited annotated samples [ 63 ], it is important to note that they cannot fully replicate scanner-specific artifacts or staining variability inherent to real-world clinical data. Excessive reliance on artificially 'clean' images may reduce model resilience to noise and confounding factors present in actual deployments. To ensure realistic performance and avoid overfitting to idealized training conditions, generative augmentation should, therefore, be coupled with artifact-preserving or artifact-simulation strategies that better reflect the complexities of clinical imaging data. Despite remarkable progress in the past decade, the development of imaging-based large models still faces several challenges. Firstly, these models require substantial training resources and depend heavily on large-scale, high-quality multimodal data. This poses a particular challenge in oncology research, where such data are scarce. Furthermore, tumor data encompass diverse modalities such as radiomics, pathomics, and molecular profiles, which exhibit significant heterogeneity and vary greatly in data structure and feature scale. This complicates cross-modal alignment and fusion, thus increasing model complexity and hindering their broader application and clinical translation in this field. Molecular data-based deep learning models In addition to imaging data, the rapid advancement of high-throughput sequencing technologies has generated a vast number of molecular omics data, including gene expression and regulation, genetic variation, epigenetic modifications, spatial transcriptomics, and proteomics. Unlike imaging data, models driven by molecular data utilize molecular sequence information as input to construct sequence-based feature embeddings (Fig. 2 ). Generally, molecular and sequence-based models are large-scale pretrained foundational models (based on transformer architectures) that use tokenized sequence data as input. For instance, DNA/RNA sequences are segmented into individual bases or k-mer units, while protein sequences are divided into individual amino acid residues [ 64 ]. These tokens are then processed within the transformer’s attention framework, allowing each token to interact with all others, thereby capturing both local and global sequence information and enabling high-quality modeling (Table 2 , Table 4 ). For instance, large-scale DNA sequence-based pre-trained models, such as Nucleotide Transformer, DNABERT, EVO, and Enformer, are built on the transformer architecture. Pre-trained on massive DNA sequences, they acquire robust generalizable capabilities for genomic predictions, applicable to tasks including enhancer identification, transcription factor binding site prediction, splice site inference, and mutation effect assessment [ 65 – 68 ]. Owing to the strong representational power for functional genomic elements and regulatory mechanisms, these models are being increasingly incorporated into cancer genomics to support clinical decision-making and prognostic analysis. In the field of oncology, many AI-based models are trained using feature vector embeddings derived from protein amino acid sequences as key input data. For instance, GraphProt2 leverages the spatial relationships and interactions between amino acid residues by a Graph convolutional networks (GCNs) architecture to identify potential tumor surface protein targets. It achieves 86.43% and 82.49% accuracy in validation and external set, respectively [ 69 ]. Meanwhile, ProtTrans employs a LLM framework to extract high-level semantic features from sequences for protein function prediction [ 70 ]. AlphaFold3, as a next-generation structure prediction model, integrates a Pairformer module with a diffusion model to perform end-to-end prediction of molecular interactions and complex structure identification. Although it still exhibits certain limitations when handling highly dynamic or rare conformations [ 71 ], AlphaFold3 is superior for GraphProt2 and ProtTrans for oncology drug discovery, which is important for designing molecular complexes that target transcription factors or proteins involving post-translational modifications, thereby offering new strategies for inhibiting tumor growth. Compared to AlphaFold3, the ESM-2 model, a BERT-based architecture that performs more efficiently for sequence modeling without multiple sequence alignments, in terms of predicting protein–ligand binding sites [ 72 ]. The transformer architecture has become a mainstream approach in models constructed from gene expression profiles. For instance, scBERT effectively processes a subset of refined activation matrices using gene expression vectors as input [ 73 ] (Table 4 ). Geneformer employs an embedding strategy that treats "genes as tokens and cells as context" [ 74 ] and scGPT predicts single-cell gene expression levels in an autoregressive-dependent manner [ 75 ]. Both scBERT and Geneformer support cell type identification, with scBERT capable of imputing missing expression values and demonstrating strong robustness to highly sparse data (acc:84%, F1-score:82.6%). But scBERT exhibits limited ability to infer functional relationships or biological networks, and its generalization performance remains suboptimal. Geneformer can predict transcriptional regulation and cell reprogramming processes by strong transfer learning capabilities, yet it is prone to overfitting on small datasets. Beyond generating single-cell gene expression profiles, scGPT also supports multimodal data modeling, though its performance is sensitive to token encoding and requires further optimization in gene representation methods. Moreover, GeneMamba utilizes an unsupervised contrastive learning strategy based on the Mamba architecture [ 76 , 77 ] (Fig. 2 ) (Table 4 ). Compared to the aforementioned models, GeneMamba offers strong capabilities for data clustering and visualization (acc:96.03%, Macro-F1:92.35%), which effectively preserves the original data structure, and maintains a lightweight design. With the increasing availability of epigenomic datasets across diverse cancer types, deep learning architectures have begun to incorporate chromatin accessibility, DNA methylation, and histone-modification profiles into molecular-based models. For example, AlphaGenome employs a hybrid CNN-transformer architecture capable of modeling DNA sequences up to 1 Mb in length, which unravels high-resolution prediction of gene expression, chromatin states, and splicing patterns [ 78 ]. Enformer-Celltyping builds upon the Enformer architecture by incorporating chromatin accessibility data, achieving joint prediction of histone modification signals across six previously uncharacterized cell types and integrating DNA sequence information with ATAC-seq markers (mean AUC:95.6% in all cell type-specific histone mark predictions) [ 79 ]. EpiGePT, a multi-task Transformer model, incorporates modules for sequence encoding, transcription factor information embedding, transformer-based integration, and predictive output, allowing simultaneous prediction of chromatin states and long-range interactions across multiple cell types (AUC:98.2% in fine-tuning model, AUPRC:91% in variant effect prediction) [ 80 ]. TRAPT employs a conditional variational autoencoder (Conditional VAE) to learn latent representations from data such as ChIP-seq and ATAC-seq, and transfers these representations via knowledge distillation to a downstream graph VAE for predicting the activity states of key transcription factors within their regulatory contexts [ 81 ]. Unlike the aforementioned models, TRAPT is not suitable for predicting epigenetic trajectories and relies heavily on multimodal epigenomic data as input. By integrating DNA sequence and epigenomic information, these models significantly expand the application of multi-omics integration modeling in oncology research, providing crucial methodological support for extracting molecular features and deciphering mechanisms. Multimodal models Multimodal models aim to synergize predictive performance across various downstream tasks by simulating the integrative decision-making processes of humans and incorporating data from diverse sources, including multimodal data within the same domain and cross-domain modal data (Fig. 2 ). By leveraging the complementary nature of cross-modal information, these models not only enhance predictive reliability and generalization capability but also help uncover intrinsic relationships among different modalities, thereby offering more comprehensive and realistic support for decision-making (Table 1 – 4 ). Depending on the data composition, existing models can be broadly categorized into two types: those models constructed based on molecular sequences, and multimodal models that integrate molecular sequences with medical imaging and clinical text data. Similarly, multimodal learning models are designed to integrate heterogeneous biomedical data, such as images, sequences, and clinical text, into unified representations. These models leverage strategies like attention-based fusion, contrastive alignment ( e.g. , CLIP-style objectives [ 82 ]), or cross-modal transformers to learn shared latent spaces. Architectures may adopt early fusion (combining raw data), intermediate fusion (combining embeddings), or late fusion (combining predictions), which enables robust modeling across spatial, molecular, and clinical text data [ 83 ]. In addition to models such as EVO (a recent pioneering framework in genomic realm), the MAMMAL is capable of predicting molecular properties and transcriptomic phenotypes while supporting small molecule/protein sequence design [ 84 , 85 ]. Furthermore, this multimodal approach can model single-cell omics and spatial omics data with sequence information. scMODAL selectively aligns and imputes features between single-cell RNA sequencing (scRNA-seq) and proteomics data by combining positive correlation-based feature alignment with GANs, primarily addressing missing feature imputation and cross-modal mapping tasks. It achieves mean accuracy of 92% for cell annotation, and mean accuracy of 85% for transfer ability of cell type labels (scRNA-seq to antibody-derived tag and ATAC-seq) [ 86 ]. In contrast, HEIST elucidates cellular spatial neighborhood graphs and intracellular regulatory graphs from spatial transcriptomic and proteomic data by employing a hierarchical graph transformer architecture. It shows strong performance (mean AUC:88.8%) across multiple downstream tasks by explicitly modeling spatial regulatory relationships among distinct cells [ 87 ]. It is worth noting that scMODAL emphasizes feature alignment and completion without directly modeling spatial expression or cellular neighborhood relationships, thus complementing HEIST in functional scope. For multimodal models which can integrate molecular sequences, medical images, and clinical information, the contrastive learning-based pretraining and fine-tuning paradigm built upon the CLIP framework has gained significant attention, giving rise to a variety of multimodal models tailored for oncology applications. For example, OmiCLIP integrates multi-omics data, including single-cell/spatial transcriptomics, epigenomics, proteomics, metabolomics, and clinical text describing tumor phenotypes and biological processes through joint pretraining, allowing to achieve tumor classification, phenotype inference, and molecular subtyping [ 88 ]. In contrast, BioMedCLIP and Med-Flamingo focus primarily on alignment pretraining between medical images and clinical texts, in particular tasks such as medical image retrieval, diagnostic reasoning, and report generation [ 89 , 90 ]. Compared to OmiCLIP, they exhibit several limitations like poor capability in modeling omics data, weak support for structured healthcare data, and higher computational costs. Beyond CLIP-based models, this category also includes multimodal generative models built on GPT architectures such as BiomedGPT, as well as PathChat, a model fine-tuned from Llama-2 as the language backbone and UNI as the visual encoder. PathChat specializes in integrating WSIs with text, supporting tasks such as diagnostic report generation and medical question answering [ 91 , 92 ]. BiomedGPT has strong generalization ability owing to its broad pretraining corpus coverage, but lacks complex training, deployment, and optimization processes. In contrast, PathChat enhances the clinical relevance of its outputs by employing a hybrid training strategy combining annotations from pathology experts and manual review, which achieves an accuracy of 78.1% on the full combined benchmark, and its accuracy further increase to 89.5% by integrating clinical context. Nevertheless, its performance remains imperfect because of the representational capacity of WSI encoders and the high cost of acquiring high-quality annotated data. Hybrid models Different model architectures exhibit distinct strengths when processing various types of medical tumor data. CNNs, which extract local spatial patterns through convolutional kernels, are particularly effective in capturing radiomic features from CT/MRI scans and histopathological features from WSIs. GNNs, modeling relationships among connected nodes, are more suitable for uncovering associations among molecular entities (such as proteins) and between genetic alterations and disease phenotypes. In contrast, Transformer-based models rely primarily on attention mechanisms and emphasize global contextual dependencies, making them well-suited for long-range sequence data such as DNA, RNA, or amino acid sequences, as well as high-resolution medical imaging. Moreover, Transformers serve as the foundational architecture for large-scale pretrained models and multimodal representation fusion strategies. Such model-specific biases often limit the ability of a single architecture to comprehensively interpret the intrinsic heterogeneity of medical tumor data, and hinder effective alignment of multimodal feature representations. Consequently, hybrid architectures integrating multiple model types have emerged as a leading direction in contemporary medical tumor related data modeling. For example, approaches that employ CNNs as encoders for imaging feature extraction and then integrate attention heads or Transformer modules leverage complementary local and global representations, thereby improving robustness during training on medical tumor imaging tasks [ 93 – 96 ] (Table 1 , Table 3 , Table 4 ). GCNs, which extend the concept of convolution to relational structures, enable modeling of complex relationship- and topology-driven data that conventional architectures cannot effectively capture. As a result, GCNs are commonly used to represent associations across heterogeneous medical modalities, such as connections between imaging-derived nodes and genomic nodes [ 97 – 99 ] (Table 4 ). Meanwhile, graph attention networks (GATs) introduce attention mechanisms to graph-node learning, enabling dynamic weighting of inter-modality relevance ( e.g. , between clinical metadata and imaging features, or between patient profiles and single-cell data). This mechanism allows GATs to distinguish the relative importance among neighboring nodes, maintain computational efficiency at scale, and generalize to unseen graphs or newly introduced nodes [ 100 – 103 ] (Table 4 ). CNNs integrated with attention heads or Transformer modules are widely favored in pathology and radiomics-based AI research. In contrast, the more functionally comprehensive GAT framework is expected to become another essential building block for multimodal and large multimodal tumor models. As hybrid architectures continue to incorporate additional modules, such as BERT, GPT, or Mamba, these improvements may further enhance a model’s ability to integrate heterogeneous clinical information and dynamically balance its contribution relative to imaging and molecular features. This progression is anticipated to improve both training stability and predictive performance in medical tumor data modeling. Explainable AI (XAI) models A common issue for application of AI models is the "black box" problem, which is characterized by the opacity of their internal decision-making processes. Therefore, there is a pressing need to systematically elucidate the complete causal logic and rationale underlying model predictions, from training and fitting to final output. This entails several key questions, such as why a model makes a particular decision, how it derives the conclusion, and which critical evidence it relies on. Such demands have directly spurred advances in model interpretability. Shapley Additive Explanations (SHAP) is a post-hoc interpretability method grounded in game theory, designed to quantify the marginal contribution of each input feature to a model’s performance. It supports decision interpretation through feature importance ranking and threshold-based analysis by providing both global and local explanations, making it widely used in interpretability studies of AI models [ 104 ]. Owing to its model-agnostic nature and user-friendly interface, SHAP is often employed as a backend interpreter for both classical machine learning models and deep learning architectures such as MLP. For example, SHAP analysis uncovered association patterns between tumor heterogeneity features and pathological complete response (pCR, the absence of detectable invasive cancer cells in tissue surgically removed from a patient) in a predictive model for neoadjuvant anti-PD-1 therapy response in colorectal cancer [ 105 ]. In another study with a transformer-based survival prediction model, the combination of SHAP and XGBoost elucidated potential links between various tumor molecular pathways and the efficacy of anti-PD-1/PD-L1 therapy as well as the degree of immune cell infiltration, providing a rationale for combination therapy strategies [ 106 ]. Nonetheless, SHAP has certain limitations: it focuses on attribution analysis between input features and the final output, making it difficult to reveal mechanistic details of the model's intermediate representations or provide direct spatial visualization of intermediate layers. This has spurred the development of visualization methods for interpreting intermediate layer representations. The most widely used approach involves in generating heatmaps based on attention weights. Models equipped with attention mechanisms, particularly transformer-based foundation models, can produce such heatmaps by extracting their internal attention weights. This technique projects attention weights from different modules back to the original input space, enabling visual identification of critical regions. Attention heatmaps can be used to verify whether the model focuses on regions consistent with pathologists’ interpretation, identify micrometastatic foci or biomarkers, and further reveal associations between highlighted regions and adverse prognosis or therapeutic response in tumor patients [ 53 , 54 , 59 , 107 ]. Gradient-weighted Class Activation Mapping (Grad-CAM) is another intermediate-layer visualization method, primarily for explaining convolutional features in CNNs. It generates localization maps by combining the feature maps of the final convolutional layer with the gradients of the target class, thereby highlighting the most discriminative image regions for a given classification decision [ 108 , 109 ]. Similar to attention heatmaps, Grad-CAM is a non-intrusive technique that requires nonstructural modifications on the model. Grad-CAM has been widely applied for oncology AI research. For instance, in a study on malignancy saliency detection using multimodal magnetic resonance imaging (MRI), Grad-CAM precisely localized tumor shape, texture, and boundaries, with its hotspot regions aligning with established clinical knowledge regarding recurrence and metastasis [ 110 ]. In another study employing a transformer-based model to predict lymph node metastasis in locally advanced gastric cancer using pre-neoadjuvant chemotherapy contrast-enhanced CT images, Grad-CAM revealed heightened model sensitivity to morphologically complex tumor margin areas, such as microvascular dense zones, regions of active cell proliferation, and surrounding lymphatic networks. This finding further supports the potential association between these imaging features and the risk of lymph node metastasis [ 106 ]. Node-based feature activation is a designated interpretability approach for graph-structured models like GNNs and GCNs. It explains predictions by quantifying contribution of each node to the final output and characterizing feature-propagation relationships. In addition to demonstrating generalizability across most medical imaging modalities, this method is effective in identifying key oncogenic signaling pathways and elucidating the TME [ 111 – 114 ]. However, large-scale graph computation and data noise may compromise the efficiency and reliability of node-importance scoring. Interpretability methods not only facilitate the examination of AI decision-making processes but also provide crucial support for validating clinical hypotheses and exploring potential tumor biological mechanisms. Even so, current techniques cannot completely elucidate true causal factors. While methods such as attention heatmaps and Grad-CAM can reveal which regions are important for a decision, they cannot rigorously explain why the model makes a particular judgment. Furthermore, although multiple interpretation approaches, including attention rollback, integrated gradients (IG) [ 115 ], and probing analysis, have been developed for transformer-based foundation models ( e.g. , BERT, GPT), the stability and reproducibility of their results is unsatisfactory. Therefore, research into AI interpretability in oncology remains an evolving field. The development of reproducible, quantifiable interpretation paradigms that are closely aligned with clinical evidence might represent a key direction for future progress. In general, AI models applied in oncology have evolved from traditional machine learning methods to single-modality CNNs and generative models, and further to contemporary multimodal large language models. Although several recently developed models, such as MedFormer, AlphaFold3, GeneMamba, and BiomedGPT, have demonstrated strong performance in pan-cancer or other cancer datasets, their generalizability remains to be validated in GI tumor data. In the future, these models may serve as novel feature extractors or core architectures for constructing GI tumor-specific analytical frameworks, thereby facilitating clinical applications and research tasks tailored to this disease spectrum. Nevertheless, the performance and generalization capability of these models remain constrained by several key factors, including the scale and quality of data, the nature of the specific oncology task, the effectiveness of cross-modal data integration, and model transparency. Therefore, foundation models built on multimodal data, particularly those undergoing continuous optimization and innovation, as well as lightweight models with efficient transfer learning abilities, become crucial drivers for advancing GI tumor research and clinical translation. AI Models in aiding early diagnosis and treatment of GI tumors Early screening for GI tract tumors is crucial for preventing tumor development and improving patient prognosis. However, conventional screening methods, such as endoscopy, liquid biopsy, histopathological biopsy, and computed tomography (CT), are limited by operational inefficiency and insufficient sensitivity in diagnosing early-stage lesions. Recently, AI has been increasingly applied in the field of early GI tumor screening. However, these AI screening models still exhibit several shortcomings, including relatively primitive model architectures, limited diversity in training data sources, inadequate integration of multimodal information, and poor interpretability. Currently, key challenges for improving an AI-assisted GI tumor screening system lie in effectively integrating different algorithms into endoscopic image-assisted diagnosis systems, fusing multi-source liquid biopsy molecular marker information, utilizing conventional imaging data or heterogeneously styled data for robustness training (Fig. 3 ). Fig. 3. Open in a new tab AI models for early clinical screening of GI tumors. In clinical settings such as liquid biopsy, endoscopy, and other imaging modalities, AI models for GI malignancies are designed to assist in assessing cancer risk in populations, endoscopic lesion detection, preliminary tumor screening, and initial pathological or molecular subtyping, thereby facilitating subsequent clinical decision-making and patient management. CTC, Circulating Tumor Cells Endoscopic screening Endoscopy serves as a key way for the detection of GI tumors. However, conventional endoscopic methods may lead to missed adenomas or polyps, thereby compromising screening sensitivity. Over past decades, AI algorithms have been increasingly applied to analyze endoscopic image, aiding in the identification of suspicious lesions, particularly for hard-to-visualize areas, such as blind zones. Consequently, AI-assisted endoscopic image analysis has gained considerable attention and has achieved a series of advancements (Table 1 ). Currently, most AI models for assisting endoscopic screening are built on conventional machine learning algorithms, aiming to model the relationship between patients’ clinical features and endoscopic image characteristics, and to perform high-dimensional feature selection and dimensionality reduction. Although conventional machine learning is no longer the primary focus of recent model architecture research, its lightweight and flexible nature continues to offer practical advantages in endoscopic screening applications. These models can effectively capture associations between clinical features and endoscopic image characteristics, which also support high-dimensional feature selection and dimensionality reduction for downstream diagnostic tasks. Studies have shown that computer-aided diagnosis (CAD) systems can incorporate multiple machine learning models to enable automated recognition and screening of various tumors such as breast cancer [ 116 , 117 ]. Machine learning models have been widely integrated into CAD systems for lesion detection, disease classification, and abnormal region identification in endoscopic images [ 118 ] (Table 1 ). Furthermore, AI-enhanced CAD systems have improved performance across several key clinical metrics, including adenoma detection rate (ADR), sessile serrated lesion detection rate, and non-neoplastic resection rate [ 119 – 123 ]. These systems have also been applied to discriminate early Barrett’s esophagus-related neoplasia and assist with real-time endoscopic submucosal dissection (ESD) by identifying lesion margins to guide precise resection [ 124 , 125 ]. Likewise, current machine learning AI-assisted models exhibit limitations in detecting certain types of lesions from different endoscopic imaging modalities, such as white light imaging (WLI), narrow-band imaging, and iodine staining [ 126 , 127 ] or in real-world clinical settings [ 128 ]. For instance, these models have relatively low detection rates and high miss rates in identifying esophageal squamous cell carcinoma and its precancerous lesions [ 129 , 130 ]. Also, the performance of current CAD and AI-integrated models remains suboptimal in high-risk clinical scenarios such as screening for Lynch syndrome-associated adenomas in Western Europe 17 centers [ 131 ]. Therefore, their generalizability and robustness require further enhancement. Compared to traditional machine learning methods, CNNs demonstrate significant advantages in image feature extraction, enabling pixel-level lesion detection and precise classification. Although studies have attempted to develop CNN-based image segmentation models to assist endoscopic screening (Table 1 ), their performance is often compromised by data heterogeneity, leading to limited model generalizability. To address this issue, a data augmentation strategy based on single-image geometric rendering, combined with multi-scale edge detection, an edge attention mechanism, and a boundary enhancement module is used to optimize the U-Net architecture. The resulting model achieved a mean dice coefficient (an evaluation metric that quantifies the overlap between the predicted segmentation results and the ground truth labels) of 0.89 in segmentation tasks, surpassing traditional methods by more than 10 percentage (recall:88%, precision:90%) [ 132 ]. Similarly, a color migration-based UM-Net model mitigated color inconsistency caused by different endoscopic devices and lesion types, which outperformed current state-of-the-art approaches across five independent test datasets including Western European population (dice:79.48%, acc:97.61%, recall:84.78%, precision: 85.73%) [ 133 ]. Despite improved performance, these methods are susceptible to human-induced sample selection bias. Furthermore, their complex architectures limit deployment efficiency in real-time clinical settings. Several studies have focused on embedding CNNs into multi-module integrated systems to enhance the accuracy and utility of GI tumor screening. For instance, the ENDOANGEL-CPS system improved the accuracy of colorectal polyp size measurement (89.9% vs 54.7% compared with endoscopists) [ 134 ]. Similarly, an endoscopic detection system integrating a pre-trained CNN with ENDOANGEL-ELD, incorporating modules for image denoising, anatomical structure recognition, lesion detection, light source mode discrimination, and risk stratification assisted in classifying high-risk esophageal lesions (HrELs) [ 135 ]. Separately, a real-time interactive clinical decision support system combining lesion detection (CADe), diagnostic classification (CADx), and depth of invasion prediction supported the diagnosis of gastric tumors. In five external test sets, its accuracy and F1-score of four-class histologic prediction achieved 80.78% and 86.7%, respectively, and similar accuracy and F1-score were observed in invasion depth prediction [ 136 ]. Additionally, ColnNet integrates CNN feature extraction with a feature attention unit module, enabling effective detection of diminutive colon polyps accounting for as little as 0.01% of the total image area [ 30 ]. DSHNet employs a high-frequency guided attention module combined with a low-frequency driven region module, utilizing dynamic convolution (a technique that adapts its kernel parameters based on the input content) to generate adaptive kernels that accommodate colon polyp morphological diversity. This design enhances discrimination of internal details, margins, and background regions, thereby achieving more precise segmentation [ 137 ]. Another transformer-based model, CTNet with a Self Multi-scale Interaction Module and a Collection Information Module (CIM), not only compared the performance of multiple existing polyp segmentation models but also demonstrated that it achieved superior results across several metrics in 3 independent polyp datasets with a dice of 84% and a mean absolute error of 2%, which improved segmentation performance (dice coefficient) ranging from 2 to 10%. Furthermore, the model exhibited notable advantages in identifying clinically overlooked subtle polyps [ 36 ]. Multimodal information fusion has been recognized as an effective strategy to enhance the performance of endoscopic screening models. For instance, the LightGBM algorithm model combined epidemiological data with sponge cytology features to identify high-grade intraepithelial neoplasia, esophageal squamous cell carcinoma, and gastroesophageal junction adenocarcinoma, which significantly reduced the need for unnecessary endoscopic examinations across different clinical scenarios. Interpretability analysis using SHAP revealed that abnormal cell count, nuclear size, patient age, and sex were the most predictive features, with gastroesophageal junction adenocarcinoma contributing more to the predictions than esophageal squamous cell carcinoma. This finding suggests the room for further optimization of the model’s interpretability across different tumor types [ 138 ]. In another approach, PSDM, a diffusion probabilistic model-based framework for polyp synthesis, integrated compositional prompt generation ( e.g. , mask information of image and text description) modules with an LLM-driven pathological information extraction module, which enabled the fusion of real polyp images, pathological reports, and expert annotations to generate high-quality synthetic polyp images. Experimental results demonstrated that the synthetic images improved all the performance of several downstream polyp detection models in 4 standard segmentation public datasets (dice:81.52%), classification models in imbalanced dataset (AUC:85%, F1-score:78%), and segmentation models, such as YOLOv5 (F1-score:73.91%), in another segmentation public test set by 2%–3% [ 139 ]. Current research suggests that the development of AI-assisted endoscopic screening will focus more on multimodal information integration and the effective deployment of real-time clinical systems, aiming to enhance the practical value of models. Furthermore, training unified and robust visual models should be validated in large-scale pre-trained models tailored for GI endoscopic screening tasks. Liquid biopsy During tumorigenesis and progression, tumor cells release various bioactive substances that can be excreted through multiple pathways. This characteristic provides a theoretical basis for non-invasive cancer screening via liquid biopsy, enabling early cancer detection using biomarkers from peripheral blood or excretory samples [ 140 ]. Current clinical screening for GI patients often relies on conventional tumor biomarkers, such as carcinoembryonic antigen (CEA), carbohydrate antigen 125, carbohydrate antigen 199, and alpha-fetoprotein. However, these biomarkers show low specificity and high rates of false positives and negatives, thereby restricting the clinical application of liquid biopsy. In recent decades, AI -assisted liquid biopsy has emerged as a prominent research focus. AI models enable clinician to diagnose cancer by integrating and analyzing multi-dimensional tumor-related biomarkers. Multiple studies have demonstrated that such approaches are significantly superior for conventional tumor biomarkers-based detection strategies in terms of screening accuracy, highlighting their strong potential for clinical application (Table 2 ). Data for AI model training in liquid biopsy were primarily derived from circulating tumor cells (CTCs) and cell-free nucleic acid fragments ( e.g. , cfDNA) in peripheral blood. CTCs is often utilized to extract tumor-specific genes and construct expression profiles [ 23 , 141 ]. Regarding cell-free nucleic acids, plasma RNA provides the profiling of non-coding RNA (oncRNA) expression via small RNA sequencing (smRNA-seq), while cfDNA offers multiple indicators reflecting molecular tumor characteristics, such as copy number variation, fragment size coverage, fragment size distribution, nucleosome positioning, and hypomethylation sites of retrotransposons [ 142 – 144 ]. The application of AI in these multi-dimensional data have improved sensitivity of various liquid biopsy. Notably, such models exhibit outstanding performance in early screening for esophageal squamous cell carcinoma, gastric cancer, and colorectal cancer [ 145 , 146 ]. Tumor-derived exosomes are a type of extracellular vesicles generated via the endosomal pathway and released into the extracellular space through the fusion of multivesicular bodies with the plasma membrane. Compared to circulating cell-free biomarkers, the cargo carried by exosomes, such as nucleic acids and proteins, exhibits higher tumor specificity and greater stability, with less susceptible to variations in the blood microenvironment [ 147 ]. Several studies have shown that tumor-derived exosomes can be isolated from plasma, serum, or stool samples, and their molecular contents, like different types of RNAs and specific membrane proteins, can be used as features in machine learning models ( e.g. , logistic regression, random forest) for the identification and classification of various GI cancers [ 148 – 152 ]. Surface-enhanced Raman Spectroscopy (SERS) enables the construction of molecular signal profiles from exosomes, allowing high-precision discrimination of gastric and colorectal cancer involving stage 0-II in TNM staging system [ 153 ]. Metabolites within exosomes can also be systematically classified and modeled to establish diagnostic frameworks for early gastric cancer screening [ 154 ]. Given the high-dimensional and heterogeneous nature of exosome related data, CNN-based AI models demonstrate superior capability in fitting complex data patterns. Additional studies have employed feature selection strategies to optimize combinations of exosomal biomarkers via AI models, identifying the most discriminative biomarker panels and significantly enhancing the performance of GI cancer screening [ 155 , 156 ]. An unsupervised deep learning model integrating CNN and VAE architectures has been developed to identify tumor-associated extracellular vesicles in the serum of colorectal cancer patients [ 157 ]. In summary, the incorporation of generative deep learning architectures has markedly improved the overall performance of exosome-based AI-assisted cancer screening models. Furthermore, AI models have been extended to various specialized forms of liquid biopsy to improve early detection of GI cancers. For instance, colorectal mucus analysis using SERS [ 158 ], detection of volatile organic compounds in breath samples [ 159 ]. Also, analysis of DNA fragment length, base composition, and modification status in exhaled breath have all been employed to discriminate gastric and colorectal cancers [ 160 ]. Based on other omics data, AI-assisted liquid biopsy models have also showed substantial potential. Discriminative models constructed using LASSO regression and random forest algorithms, in combination with liquid chromatography-mass spectrometry metabolomics and ultra-performance liquid chromatography coupled with quadrupole time-of-flight mass spectrometry, enabled effective stratification of patients with gastric cancer at different stages in 3 Chinese cohorts (acc:90.9%, AUC:95.7% for stage IA patients and acc:92.7%, AUC:98.4% for stage IB patients in test set 1, acc:79.1%, AUC:90.9% for stage IA patients in external cohort 2). Importantly, these models identified several key metabolites, such as L-carnitine, L-proline, and pyruvaldehyde, which differentiated chronic superficial gastritis from early gastric cancer (EGC). These findings suggest that gastric cancer progression is accompanied by enhanced amino acid-derived energy metabolism and aberrant membrane lipid metabolism, revealing a metabolic reprogramming process associated with gastric carcinogenesis [ 161 , 162 ]. XGBoost, random forest, and CatBoost classifiers trained on 16S rRNA sequencing data from the gut microbiome of colorectal cancer patients, combined with SHAP interpretability analysis, identified microbial taxa significantly associated with colorectal cancer risk, such as Fusobacterium and Peptostreptococcus [ 163 ]. The RSA (“Virtual Biopsy and Risk Stratification Assessment”) multi-modal AI system integrated contrast-enhanced abdominal CT images with multiple clinical indicators to early identify peritoneal metastasis risk in GC-CY1 (peritoneal lavage cytology-positive) gastric cancer patients. Its performance of AUC achieves 88.3% in external validation cohorts and 83.5% in a prospective validation cohort [ 164 ]. These advances underscore the broad potential of AI in analyzing multi-source information to enhance the performance of tumor liquid biopsy. AI-based liquid biopsy has achieved notable progress in the early screening of GI tumors (Table 2 ). Nonetheless, given the various biomarkers utilized in liquid biopsy, conventional machine learning models often struggle to fully capture their complex characteristics and high-dimensional relationships. Future research may integrate advanced architectures such as CNNs and transformers to enhance modeling capability for multimodal and high-throughput data (Fig. 3 ). Furthermore, the performance and generalizability of these models across different medical centers or independent external validation require rigorous evaluation. Additional screening modalities In addition to endoscopy and liquid biopsy, AI models can assist in the non-invasive diagnosis of GI tumors by analyzing medical images such as WSI, CT, and MRI (Table 4 ). Studies have shown that CNN-based classification and segmentation models can perform binary and multi-class predictions ( e.g. , typical lesions, atypical non-neoplastic lesions, and atypical neoplastic lesions) on WSIs, CT, and MRI images from GI cancer patients, and accurately delineate lesion regions [ 43 , 165 ]. Nowadays, GNN models have been employed to capture glandular structures, intraglandular nuclear morphology, and interglandular cellular density and spatial relationships in colorectal cancer pathology images, showing strong capabilities in extracting structural features for endoscopic biopsy screening [ 33 ] (Fig. 2 ). Using TCGA as training set, MSIntuit combines CNN-based pre-training and fine-tuning with a MLP to identify mismatch repair deficiency (dMMR) and microsatellite instability (MSI) in gastric cancer from WSI (AUC:88%), outperforming the MSPath model (AUC:75%) and showing potential to optimize current dMMR/MSI clinical testing workflows in an international test set [ 166 ]. Notably, the U-Net-based GRAPE framework automatically segment gastric cancer lesions using routine non-contrast CT images from 3 Chinese independent cohorts, outperforming interpretations by multiple radiologists in fine structural recognition (AUC:92% vs 76%−85%) (Table 3 ). Even so, the GRAPE model shows superior detection rate for advanced gastric cancer (over 90%) compared to T1-stage early gastric cancer (approximately 50%), indicating persistent challenges in imaging-based identification of early-stage lesions [ 41 ]. A transformer-based weakly supervised multiple instance learning model can analyze H&E-stained and cytosponge TFF3 biomarker images to detect Barrett’s esophagus-associated intestinal metaplasia, increasing screening coverage by nearly twofold [ 24 ]. A multimodal AI system integrating deep learning architectures, such as APINet, TransFG, and DeepLabV3 + , with various machine learning models has been developed to screen and identify gastric cancer and its precancerous states ( e.g. , chronic atrophic gastritis) by combining tongue image features (including color, shape, size, coating color, thickness, and moisture) with tongue coating microbiome data. In internal validation of 10 Chinese centers, this model achieves AUCs of 92%−94%, and its AUCs reached 88%−89% in the external independent verification of Chinese 7 centers [ 167 ]. The development of novel imaging modalities or the integration of multi-source screening data holds promises for overcoming current limitations in the performance and generalizability of AI models for early cancer screening (Table 3 ). Future studies should incorporate larger, multi-center, and more diverse image datasets to enhance model accuracy and clinical credibility. Furthermore, the refinement of fine-tuning strategies and their practical feasibility will be critical for ultimate performance of these models. Existing data sources for early GI cancer screening are highly diverse, and single-modality data representation often inadequately supports the predictive performance of conventional AI models. Given the unique advantages of transformer-based architectures in integrating multimodal data with stylistic variations and maintaining model stability, it is anticipated that more multimodal foundation models built on transformers will be developed for early GI tumor screening (Fig. 3 ). Such models are expected to integrate heterogeneous data from diverse imaging modalities of endoscopic images, multiple liquid biopsy molecular markers (from sources such as blood, stool, and peritoneal lavage fluid), WSI of H&E and IHC staining, non-contrast and contrast-enhanced CT images, as well as patient clinical information (such as personal tumor history and family genetic history). Nevertheless, the efficacy and generalizability of these models require further validation through multi-center, large-scale external cohorts. AI Models for patient stratification and decision-making In GI oncology, it remains a major challenge to accurately predict survival outcomes, such as OS, PFS, and disease-free survival. Given the high heterogeneity of tumors, next-generation sequencing technologies have been employed to decipher the interactions among molecular signaling pathways. However, a key hurdle lies in the complicated cell–cell interactions within the TME, which are difficult to correlate directly with patient prognosis, thereby limiting their utility in supporting clinical decision-making. In this context, AI models have brought new opportunities for the field. From single-model architectures to the integration of transformer structures, and further to fine-tuning strategies based on large pre-trained models, AI models based on multimodal data have largely improved clinical diagnosis and prognosis prediction for GI cancers. For instance, compared to models relying on H&E-stained images or MRI, those AI models incorporating feature embeddings from multiplex immunohistochemistry, multiplex immunofluorescence, and contrast-enhanced CT images improved accuracy in identifying tumor biomarkers and lesion structures. Meanwhile, models based on single-cell and spatial transcriptomic data are more accurate than genomics-based approaches of directly inferring molecular pathways and cell–cell interactions within the TME. Furthermore, multimodal models designed for prediction tasks across different cancer types employ diverse data fusion strategies, which not only enhance model performance and generalizability but also provide more reliable support for the development of individualized treatment strategies (Table 4 ). Such models have been widely employed for tasks including patient stratification, biomarker discovery, and survival prediction. They further assist researchers in elucidating tumor biology, characterizing microenvironment features, and even enabling multi-task joint prediction. Pathomics- and Radiomics-Based Approaches During the last few decades, a notable trend has emerged in unimodal pathomics research concerning the analysis of H&E-stained WSIs: AI models for GI tract tumors are increasingly shifting from conventional single CNN architectures toward multi-module integrated transformer-based models, which enhance the modeling of complex TME regulatory mechanisms and further improve model performance in downstream prognostic prediction tasks (Fig. 4 ). Fig. 4. Open in a new tab Multi-omics model in patient stratification and treatment decision support in GI tumors. Multi-modal can integrate diverse types of medical data—including histopathology slides, radiomics images, DNA and RNA sequencing, proteomics, single-cell and spatial transcriptomics, as well as clinical information—to stratify patients with distinct tumor biological behaviors, evaluate prognosis and predict treatment response in patients with GI malignancies, thereby supporting clinical decision-making For esophageal squamous cell carcinoma and esophageal adenocarcinoma, CNN-based AI models have been widely applied for automated identification of tumor regions in postoperative resection specimens, as well as for tumor regression grading (TRG, a grading system used to assess the ratio of residual tumor cells in tissue) assessment of residual tumors after neoadjuvant chemoradiotherapy. Previous research has demonstrated that such models exhibit strong performance in distinguishing between Grade 2 and Grade 3 regression within the Becker TRG system (acc:95.3%, AUC: 97.3%, recall:91.6%, F1-score:86.1% in 3 cohorts) [ 167 ]. Furthermore, these CNN models can be integrated with scRNA-seq, bulk RNA-seq, and whole-exome sequencing (WES) data to perform multi-omics integrative analysis, improving the definition of molecular transcriptional subtypes and the identification of key immune evasion-related markers, such as the NK cell-associated molecules XCL1 and CD160. The authors also revealed that tumor cells with high XCL1 expression show significant resistance to 5-fluorouracil (5-FU) chemotherapy [ 168 ]. In gastric and colorectal cancer research, a variety of deep learning models have been developed by integrating CNNs, self-supervised learning, weak supervision, attention-based multiple instance learning (AMIL), and transformer modules. These models assist not only in pathological subtyping ( e.g. , mucinous vs. non-mucinous) [ 169 ], MSI status prediction, and risk score evaluation [ 170 ], but also in identifying molecular biomarkers or gene mutations (such as BRAF, KRAS, HER2, and PD-L1) that are closely associated with treatment prognosis [ 42 , 171 – 173 ] and treatment response [ 174 ]. Moreover, these models can automatically characterize the TME by performing unsupervised clustering of histomorphological phenotypes, like quantifying tumor-infiltrating lymphocyte region fraction, cell proliferation index, leukocyte fraction, and integrating composite features such as epithelial-mesenchymal transition (EMT) phenotypes. Interpretability analyses have confirmed that key morphological features highlighted by the models, including lymphocyte infiltration, cell proliferation activity, mutational burden, VEGF-α and TGF-β signaling response, high-grade tumor epithelial architecture, and mucinous component distribution, exhibit strong biological consistency with clinical treatment responses at the molecular level [ 175 – 178 ]. Leveraging pre-trained pathology foundation models such as CTransPath and UNI to extract WSI features, followed by fine-tuning transformer architectures for downstream tasks, significantly enhances biomarker identification performance. For instance, these models achieve excellent predictive accuracy for MSI (AUROC = 0.97), BRAF (AUROC = 0.88), and KRAS (AUROC = 0.80) in large-scale multi-cohort evaluation on over 13,000 patients from 16 cohorts [ 179 ]. By integration with circulating tumor DNA (ctDNA)-based molecular residual disease (MRD) data, it can further reveal that patients with high-risk exhibit increased recurrence risk. Among MRD-negative patients, high-risk individuals show significantly prolonged disease-free survival (DFS) after adjuvant chemotherapy (ACT), whereas low-risk patients do not benefit from ACT [ 39 ]. Furthermore, models based on the DINO architecture and incorporating knowledge distillation and meta-training strategies have demonstrated outstanding performance in predicting OS, assessing time to recurrence, classifying histopathological growth patterns (a classification system that characterizes the tumor growth and invasion patterns at the interface between liver metastases and the surrounding hepatic parenchyma), and scoring tumor regression in patients with colorectal cancer liver metastases. In two independent external validation cohorts, the models achieved AUCs of 97.78% and 99.15%, indicating strong generalization capability [ 180 – 182 ]. AI models based on immunohistochemistry (IHC) images have emerged as an important research direction following H&E-stained WSIs. For instance, the machine learning scoring system CD3ML utilizes CD3 IHC images to predict DFS in stage III colon cancer patients [ 183 ]. By integrating multiplex immunohistochemistry (including CD4, CD8, CD20, and CD68) with CNNs and gated attention mechanisms, the MSDLM and AIS scoring systems in 4 Germany cohorts outperform single-stain models for prognostic accuracy (AUC:89.6%, AUPRC:90.2%, F1-score:81.9%), significantly improving risk stratification and potential for predicting neoadjuvant treatment response in colorectal cancer at UICC stage (an international standard for the systematic classification of malignant tumors based on the extent of the primary tumor, lymph node involvement, and presence of distant metastasis) [ 184 ]. By detecting dMMR in IHC-stained tissue microarrays (TMAs) of colorectal cancer, the CNN-based AIMMeR model achieved a positive predictive value of 98% in identifying combined MLH1-PMS2 loss, validated the prognostic significance of dMMR in oxaliplatin-treated patients, and pioneered the exploration of dMMR status in relation to recurrence risk under different chemotherapy regimens e.g ., CAPOX (capecitabine + oxaliplatin) vs FOLFOX (fluorouracil, leucovorin + oxaliplatin) [ 185 ]. On the basis of multiplex IHC images and single-cell/bulk transcriptomic data, the cellular module (CCIM)-Net framework identified an immunosuppressive CCIM centered on FOLR2⁺ macrophages, therefore depletion of FOLR2⁺ macrophages disrupted the CCIM structure and restored chemosensitivity [ 186 ]. The ROSIE and MAS model systems can generate highly realistic virtual multiplex immunofluorescence images from conventional H&E slides or autofluorescence and DAPI-stained WSIs, enabling quantitative analysis of 50 biomarkers in GI tumors. Subsequent analyses revealed that these biomarkers support cellular phenotyping and correlate significantly with patient survival: CD3, CD20, CD8, PD1, and PD-L1 combined positive score was associated with longer OS, while high FOXP3 and CD163 expression indicated poor prognosis [ 187 , 188 ]. In radiomics studies, AI models developed from non-contrast CT images-incorporating methods, such as swim-transformer, reinforced multi-scale attention (RFA), maximum mean discrepancy autoencoder, and channel attention mechanisms, have been applied to predict TRG and assess responses to neoadjuvant chemotherapy ( e.g. , 5-fluorouracil) and immunotherapy ( e.g. , anti-PD-1 inhibitors) in advanced gastric cancer and stage II/III colorectal cancer. These models can also be integrated with risk-scoring systems for stratifying patient prognosis [ 38 , 189 – 191 ]. Furthermore, deep learning models by novel algorithms such as gated spatial convolution and triple-way Mamba are capable of predicting lymph node metastasis status and OS in neoadjuvant chemotherapy-treated locally advanced gastric cancer, based on longitudinal CT images from patients, its AUC achieves 87.4% and 81.9% in training cohort and 3 Chinese validation cohorts, respectively [ 192 ]. Compared to non-contrast CT, AI models trained on contrast-enhanced CT images had stronger interpretability and more significant associations with patient prognosis and treatment response. Multiple studies confirm that such models can predict early recurrence-free survival, DFS, and OS in stage II/III gastric cancer patients, evaluate responses to neoadjuvant chemotherapy, adjuvant chemotherapy, and immunotherapy, and identify risks of postoperative recurrence and lymph node metastasis. Risk score analyses further reveal unfavorable features in the tumor immune microenvironment (TIME), including resistance to immunotherapy in the dMMR/MSI-H subgroup, activation of pro-oncogenic pathways such as MYC, KRAS, and EMT, and abnormal microvascular architecture [ 31 , 95 , 97 , 193 – 196 ]. In stage II and III–IV dMMR/MSI-H colorectal cancer patients, contrast-enhanced CT-based AI models have been used to predict the efficacy of adjuvant chemotherapy and the response to neoadjuvant anti-PD-1 immunotherapy ( e.g. , pCR) (acc:90.9%, AUC:96.4%, F1-score:92.3% in internal validation cohort 2; acc:84.2%, AUC:90.4%, F1-score:88% in external test cohort 3). The adjuvant chemotherapy-adapted subtype showed upregulation of apical junction, Notch, Hedgehog, and myogenesis-related pathways, while the observation-enriched subtype was characterized by E2F targets, MYC targets, and oxidative phosphorylation. Histologically, the observation subtype exhibited more necrosis, hemorrhage, and disorganized vasculature, whereas the chemotherapy-adapted subtype had more regular vascular structures and higher infiltration of B cells, CD4⁺ T cells, and CD8⁺ T cells in the tumor core [ 105 , 197 ]. Furthermore, a Vision-Mamba-based model, combined with contrast-enhanced CT images and six voxel-wise 3D radiomic features, was developed to predict pCR after neoadjuvant immunochemotherapy in esophageal squamous cell carcinoma. The model effectively stratified high- and low-risk patients, with favorable robustness in training (acc:91%, AUC:92%) and 3 Chinese test sets (acc:87%, AUC:84%). SHAP analysis indicated that the necrotic area, tumor periphery, and enhanced regions contributed most significantly to pCR prediction [ 198 ]. Among patients with colorectal cancer, an MRI deep learning model built on vision-transformer architectures can integrate CEA levels and XGBoost, which significantly improved predictive performance for patient survival outcomes with the C-index (a key metric for evaluating the risk ranking consistency of survival analysis models) of 82% for OS, and AUC of 90% for 3-year OS [ 96 ]. The RCMIX model was generated to predict post-neoadjuvant therapy T-stage downstaging by incorporating a MLP from 19 pre-treatment MRI radiomic features and 2 clinical features in cT4 rectal cancer patients receiving neoadjuvant therapy. Results showed that patients with greater T-stage reduction had longer DFS, associated with lower CEA levels and specific primary tumor anatomical locations. This model may inform surgical planning, such as avoiding unnecessary extended resection, and guide treatment intensification or targeted therapy selection [ 25 ]. Additionally, a multi-style endoscopic image analysis model integrating CNNs and transformers has been developed for intraoperative real-time prediction of lymph node metastasis in colorectal cancer. By combining information from WLI, fluorescence imaging, and pigment chromoendoscopy imaging, the fused model achieved an AUC of 82.94% and an accuracy of 86.35% in internal test set, and an AUC of 84.37%, an accuracy of 81.11% in two external sets, outperforming any single- or dual-modality approaches ( e.g. , acc:68.26% and AUC:62.63% from swim-transformer) [ 199 ]. The F1-score (69.46% in external set) suggests a pronounced classification bias in the model, which requires further adjustment and subsequent validation. Currently, the application of AI models based on MRI and endoscopic images in predicting prognosis and treatment response for GI tumors remains relatively limited. This may be attributed to the inherent complexity and substantial variability in image quality, which creates a wide gap between low-level feature extraction and high-level clinical prediction tasks, thereby compromising model robustness and generalizability. In contrast, AI models utilizing CT and WSIs exhibits greater developmental potential and clinical applicability. Future research should prioritize the development of novel algorithms, such as integrative learning from both non-contrast and contrast-enhanced CT scans, or multi-task models capable of simultaneously generating virtual IHC images and predicting prognostic risk from H&E-stained slides. Such approaches are expected to harness the complementary strengths of multi-modal data and may represent a leading direction in the evolution of pathomics and radiomics AI algorithms. Molecular bioinformatics-based approaches The application of AI models in molecular biomarker analysis has become increasingly well-established. For example, several studies have developed predictive models using DNA and RNA data to assess responses to immune checkpoint inhibitor (ICI) therapy in gastric cancer [ 200 ], classify immune subtypes in colorectal cancer [ 201 ], evaluate pMHC-I presentation efficiency [ 100 ], and predict pCR to neoadjuvant radiotherapy in rectal cancer [ 202 , 203 ]. These models have further revealed several key findings. For instance, PCNA and CDK2 gene expression and associated pathway activities are closely linked to immunotherapy response in gastric cancer. Furthermore, AG-538 enhances tumor immunogenicity by inducing ROS-dependent DNA damage and downregulating DNA repair genes, thereby activating the cGAS/STING signaling pathway. AI models leveraging whole-genome sequencing (WGS) and WES data enable quantitative analysis of circulating tumor DNA (ctDNA) [ 204 ], predict plasma-derived extrachromosomal DNA (ecDNA) amplification across multiple GI cancers [ 205 ], and accurately identify MSI, microsatellite stability (MSS), and POLE proofreading-deficient subtypes with ultra-high tumor mutational burden (TMB, the total number of somatic mutations per megabase in the tumor genome) [ 206 ]. These molecular features serve as independent risk indicators or support postoperative recurrence risk assessment [ 207 ]. Furthermore, the federated learning framework ProCanFDL, leveraging proteomic data from an international cohort across 8 countries (encompassing diverse ethnicities and tumor stages), has achieved high-accuracy molecular subtyping for colorectal cancer and esophageal adenocarcinoma (acc:96.5%, AUC:99.92%). To elucidate the model’s decision-making mechanism, SHAP analysis and pathway enrichment were employed to identify discriminative protein markers, such as KRT20, FABP4, and CEACAM5, and revealed that PPAR signaling and extracellular matrix-related pathways might drive subtype differentiation [ 208 ]. On the other hand, AI models utilizing bulk and single-cell transcriptomic data in GI tumors focus more on systematic characterization of the TIME and its association with treatment responses to neoadjuvant chemoradiotherapy (nCRT) and immune checkpoint inhibitors ( e.g. , PD-1/PD-L1/CTLA-4 inhibitors). Studies have shown that such models successfully reveal the antagonistic regulatory roles of the Wnt/β-catenin signaling pathway and lactate metabolism in modulating nCRT sensitivity in locally advanced rectal cancer [ 209 , 210 ]. Additionally, based on five multicenter cohorts encompassing East Asian (predominantly Chinese and Korean) and Western European populations with gastric cancer across clinical stages Ⅰ-Ⅳ (predominantly stages Ⅱ-Ⅲ), one model classified gastric cancer patients into three immune subtypes ( e.g. , immune-excluded, immune-suppressed, and immune-activated), achieving a classification accuracy of 89.3%, AUC of 94.2%, F1-score of 87.6%, Matthews correlation coefficient (MCC) of 82.1%, and AUPRC of 91.5%. And IL-1/IL-1R1 signaling was identified as a potential therapeutic target for the IS subtype [ 211 ]. Other models have also identified multiple molecular markers significantly associated with patient survival, including STAT4, IFNG, IRF4, PEX10, and ARG1 [ 212 ]. Spatial transcriptomics-based computational models commonly incorporate tissue morphology and spatial cell phenotype distribution as embedded features to decipher functional units within the tissue microenvironment. For example, CytoCommunity integrates cellular phenotypes with spatial localization to identify tissue cellular neighborhoods, systematically revealing functional structural relationships in tissue regions [ 98 ]. In supervised mode, it achieves high accuracy in predicting known cellular neighborhoods (acc:91%, F1-score:89%). CancerFinder employs domain generalization algorithms to accurately distinguish malignant cells or spatial spots [ 213 ]. Using imaging mass cytometry data of colorectal cancer samples from six North American, European, and East Asian multicenter cohorts encompassing predominantly Caucasian and East Asian patients, Morpheus achieved robust patient stratification performance with a classification accuracy of 91.2%, AUC of 95.8%, F1-score of 89.4%, MCC of 85.3%, and AUPRC of 93.1%, and identified two combination therapy strategies suitable for these colorectal cancer patients: dual blockade of PD-1/PD-L1 and CXCR4, or further addition of CYR61 inhibition [ 214 ]. Similarly, CellNEST detects both single ligand–receptor interactions and complex multicellular relay communication patterns, successfully uncovering key ligand–receptor axes centered on APP, ITGA6, and TGFBR2 at the invasive front of colorectal cancer, as well as a relay communication network dependent on CXCR4 and LRP1 within the TME [ 101 ]. In metabolomics related research, AI approaches have been employed to uncover mechanisms of tumor metabolic reprogramming and develop relevant predictive tools. By integrating transcriptomic data of a cohort of esophageal squamous cell carcinoma (ESCC) patients (predominantly of Asian ethnicity, stages I-IV), a lactate-related gene-based prognostic signature (LPS) model incorporating eight machine learning algorithms was developed to predict survival outcomes and treatment responses in ESCC, achieving a prognostic accuracy of 83.6%, AUC of 87.9%, F1-score of 79.8%, and AUPRC of 85.2%. The model enabled effective risk stratification and predicted that high-LPS group was more responsive to targeted agents such as BI-2536 and panobinostat [ 215 ]. Using nanoparticle-enhanced laser desorption/ionization mass spectrometry to acquire plasma metabolic fingerprints, a PMF scoring system was constructed via multiple machine learning models for gastric cancer diagnosis and prognosis. This approach identified 21 differential metabolites and revealed significant alterations in pathways including glycolysis/gluconeogenesis, pyruvate metabolism, amino acid metabolism ( e.g. , glutamine, alanine, and arginine/proline), and ketone body metabolism, indicating pronounced metabolic reprogramming in gastric cancer [ 216 ]. Notably, the MPI- variational graph autoencoders (VGAE) model, built using GCNs and VGAE, was applied to model metabolite-protein interaction networks based on five Chinese and Caucasian multicenter cohorts encompassing colorectal cancer patients across clinical stages Ⅰ-Ⅳ (predominantly stages Ⅱ-Ⅲ). Achieving a prediction accuracy of 86.4%, AUC of 90.3%, and AUPRC of 88.1%. It successfully predicted 37 metabolite-protein interactions in colorectal cancer and inferred a set of potential novel enzymatic reactions [ 99 ]. These findings highlight the broad potential of AI techniques in elucidating tumor metabolic reprogramming. Despite the strong predictive performance of AI models in integrating multi-omics molecular data, several important challenges still remain. For instance, models built on WGS or WES sequencing data may be affected by off-target genomic regions or intergenic interactions, leading to reduced specificity. Meanwhile, transcriptomics-based prediction models often rely heavily on associative patterns among biomolecules, which can limit their robustness and generalizability in forecasting patient prognosis and treatment response. Future studies should, therefore, focus on systematic evaluation of these models’ reliability, interpretability, and clinical utility through more rigorous validation frameworks and independent test sets. Patient clinical information-based approaches The integration of routinely clinical data, such as patient demographics, genetic and medical history, as well as pathological features into AI-based predictive models has emerged as an efficient and scalable strategy for clinical applications, including patient stratification and assessing prognosis and treatment response in GI cancer. For example, in a multicenter study of Chinese patients with clinical stage T1 ESCC, a predictive model integrating multi-dimensional features identified lymphovascular invasion and tumor invasion depth (pT stage) as the most critical predictors of lymph node metastasis (LNM) in T1-stage patients (acc:78.3%, AUC:85.3% in the external test set), offering a quantitative basis for personalized surgical planning [ 217 ]. Similarly, in colorectal cancer liver metastasis, a random forest-based model accurately predicted overall complications, PFS, and OS following simultaneous resection, outperforming conventional clinical scoring systems. The model also highlighted key predictors including genetic markers and clinical factors, demonstrating strong individualized predictive performance [ 218 ]. For predicting LNM in non-intestinal-type EGC, a neural network model, developed and validated in a multicenter cohort of Chinese patients with stage T1, addressed the lack of Lauren classification-based risk stratification tools, effectively distinguishing LNM risk and prognosis across subgroups (acc:84.3%, AUC:87.9%) [ 219 ]. In treatment optimization, a hybrid machine learning model combining random forest with optimal policy trees was applied to adjuvant imatinib decision-making in a multinational cohort with localized GI stromal tumors (GIST) (predominantly Spanish and Japanese ancestry, stages I-III), achieving a treatment decision optimization accuracy of 88.7%, AUC of 92.6%, F1-score of 85.3%, and AUPRC of 90.5%. The model suggested that certain patient subgroups may not require adjuvant therapy and identified 5 years as the optimal treatment duration. SHAP analysis further indicated mitotic count as the strongest predictor of treatment response [ 220 ]. The predictive performance and generalizability of such models should be improved by the expressive capacity of the underlying machine learning methods and neural network architectures. Furthermore, robust and clinically applicable predictions require training and validation on large-scale, multi-center, and heterogeneous patient cohorts. Multimodal approaches Like other AI models, single-modality models trained on homogeneous data sources exhibit inherent limitations, as they fail to leverage the complementary information available across diverse medical data types. In contrast, multimodal models integrate and align features from different modalities, which not only enhances predictive accuracy, generalizability, and interpretability, but also facilitates information complementarity and reveals latent associative patterns across data types. This pattern strengthens the biomedical plausibility and clinical relevance of model outputs. Notably, some multimodal fusion approaches maintain performance comparable to average prediction benchmarks even when input data are limited. Recently, multimodal AI models are widely used in gastric cancer research to enhance predictive performance by integrating radiomics and WSI data. The iSCLM framework, based on CNNs and supervised contrastive learning, predicted responses to neoadjuvant chemotherapy in locally advanced gastric cancer by analyzing pre-treatment contrast-enhanced CT images with WSIs [ 102 ]. In contrast, based on non-contrast CT and WSIs using ResNet50, attention-based AMIL, and HoverNet, RPS was developed to evaluate the response to immunochemotherapy in a multinational cohort (China and Switzerland) encompassing patients with stage III-IV gastric cancer. This model achieved a response prediction with accuracy of 87.9%, AUC of 91.8%, and AUPRC of 89.7%. RNA sequencing analysis further revealed that high responders identified by the model were significantly associated with immune-related biological processes and memory B-cell infiltration [ 221 ]. Similarly, the MuMo model, built on a transformer architecture, integrates 3D non-contrast CT and WSIs using modality-agnostic feature alignment and learnable embedding strategies, improving feature alignment and robustness under real-world incomplete data conditions. The model identified higher risk scores in male patients, those with poorly differentiated tumors, Lauren diffuse-type, and peritoneal metastasis. Consistent with clinical consensus, the authors also find that tumor-infiltrating lymphocyte abundance was negatively correlated with treatment response [ 222 ]. In colorectal cancer research, the rapid development of multimodal models is accelerating the understanding of molecular mechanisms and improving clinical prognosis prediction. Based on the CAN-Scan model, a multi-omics phenotype-driven precision oncology platform have revealed gene expression signatures associated with 5-FU resistance and Chr 7 copy number gains, and identified vemurafenib and regorafenib as potential alternative therapies [ 223 ]. To enhance multimodal integration, the pre-trained TMO-Net model developed by using pan-cancer multi-omics data, employed autoencoders and cross-modal inference to perform risk stratification and survival prediction even with missing modalities. It achieved significant performance in downstream tasks, including cancer type classification (acc:82.4%, AUC:98.9%) and survival risk prediction (C-index:65.5%). Furthermore, it identified mitochondrial metabolism and immune response-related genes as prognostic factors [ 224 ]. For imaging-genomics fusion, studies have utilized vision transformers and CNNs to extract morphological features from WSIs and integrate them with gene expression, mutation, and MSI status. For example, the Magic model uses contrastive learning and momentum distillation to accurately predict spatial gene expression and identify key prognostic genes [ 225 ]. Based on the TCGA-COAD cohort, another multimodal deep learning framework incorporated imaging, clinical, and molecular variables, which significantly improved risk stratification in CRC patients with the accuracy of 86.8%, AUC of 90.7%, MCC of 76.4%, and AUPRC of 88.9% [ 226 ]. Integration of spatial transcriptomics and pathology images further deciphered systematic characterization of the TME. Models such as IGI-DL and StereoMM, based on GNN and VAE, employed dual connectivity or cross-modal attention to capture spatial cell distribution and gene expression relationships, revealing MYC target regions, immune regulatory pathways, and microenvironmental differences among dMMR/pMMR subtypes [ 226 , 227 ]. Additionally, the application of social network analysis with cellular networks in a multi-center cohort comprising stage I-IV colorectal cancer patients (predominantly Caucasian and Asian), has shown that cellular connectivity is significantly correlated with the prediction of molecular pathways ( e.g ., CIN, HM) and key gene mutations ( e.g. , TP53, BRAF), allowing these models to resolve tumor heterogeneity [ 228 ]. Multiple studies have also demonstrated that multimodal models can identify and quantify distinct tissue types and niche structures in colorectal cancer liver metastases to predict patient responses to targeted therapies such as bevacizumab and post-metastasis prognosis. For instance, interpretability analyses revealed that elevated GLUT1 expression, an SPP1 + MTRNR2L12 + myeloid cell cluster, and the enrichment of cancer-associated fibroblasts are significantly associated with immunosuppression, tumor cell adaptation, and invasive behavior [ 229 – 231 ]. Driven by single-cell and spatial multi-omics technologies, cancer research is moving from tissue-level analysis to the cellular and microenvironmental interactions assessment. Key progress has been made in several areas, including the discovery of functional modules, elucidation of molecular mechanisms, and inference of cellular crosstalk. Functional association networks constructed from proteomic and transcriptomic data have identified critical modules such as extracellular matrix remodeling and angiogenesis, which are positively associated with poor prognosis [ 232 ]. To further decipher molecular regulatory mechanisms, the MSI-XGNN model was built by integrating multi-omics data, such as transcriptomics and DNA methylation from pan-cancer TCGA cohorts, which has not only achieved high-accuracy prediction of MSI status (acc:93.6%, AUC:97.4%, MCC:82%, F1-score:89.2%, AUPRC:97.4% in the external test set) but also confirmed key factors like MLH1 hypermethylation and genes such as MSH4 in immunotherapy response and prognosis [ 233 ]. At the cellular level, models combining single-cell RNA sequencing with multiplex immunofluorescence imaging have linked apical membrane-like cells to resistance against fluorouracil and oxaliplatin chemotherapy due to close interactions with tissue-resident macrophages [ 234 ]. Furthermore, approaches integrating spatial transcriptomics and histopathology images using hierarchical long short-term memory networks and graph attention mechanisms have enabled systematic inference of functional states, like pro-angiogenic and pro-inflammatory phenotypes, and spatial network architecture (cell-type prediction performance: acc:89.1%, dice:89.8%, AUC:93.2%, F1-score:86.7%, MCC:81.5%, and AUPRC:91.4%), providing refined tools for dissecting the TME [ 235 ]. The strategy of leveraging features extracted from pre-trained large models ( e.g. , UNI, CTransPath) followed by fine-tuning downstream modules has demonstrated considerable value in multimodal prediction models. For instance, Brim, a transformer-based multiple instance learning framework, incorporates self-normalizing networks and a bidirectional autoencoder bridging module to accurately predict the prognosis of colorectal cancer patients using WSIs alone, even in the absence of genomic or transcriptomic data [ 236 ]. Another model, GIST, integrates WSI with spatial transcriptomic features to overcome common limitations such as high missing values and limited resolution in conventional spatial transcriptomics, which precisely identify key biomarkers associated with colorectal cancer prognosis, including MKI67, KRAS, BRAF, and MUTYH, and captures the spatial distribution patterns of continuous tumor proliferation hotspots [ 103 ]. Furthermore, the CATfusion model with WSI, genomic, and epigenomic data, constructs a multimodal fusion system for prognostic prediction in colorectal cancer, extending the applicability of pre-trained feature-based multi-omics frameworks, which has been validated in 12 North American and European multi-centers, with an accuracy of 87.6% and an AUC of 91.2% for 5-year survival prediction and macro-F1 of 82.3%, AUPRC of 88.9% [ 237 ]. Multimodal models integrating radiomics and molecular biomarkers have enabled accurate prediction of patient prognosis and treatment response across multiple cancer types and metastatic sites, while uncovering relevant biomarkers and their underlying biological mechanisms using XAI methods (Fig. 4 ) (Fig. 5 ). Recently, multimodal frameworks based on fine-tuning strategies for pre-trained large models have emerged as a key direction in GI cancer research (Fig. 4 ). Nevertheless, their clinical translation was confined by several challenges. For instance, limitations in multimodal feature alignment algorithms, together with inherent data heterogeneity and frequent missing values, complicate the acquisition and standardization of multi-source, multi-type patient data requiring cross-modal matching, thereby reducing training efficiency and increasing computational demands. Furthermore, the integration of diverse data types often results in poor feature-level interpretability, rendering the decision-making process opaque. Although recent studies suggest that even limited or single-modality data may suffice for predicting survival and therapy response, the generalizability and robustness of such approaches require validation through large-scale, multi-center, prospective clinical trials. Fig. 5. Open in a new tab Model Interpretability for Evaluating the Biological Behavior of GI tumors. Contemporary AI models, through interpretability analyses, can reveal diverse tumor-specific molecular and biological features associated with downstream predictive tasks in GI malignancies, such as cancer cell plasticity, metabolic reprogramming, the immune microenvironment, and tumor angiogenesis. GI, Gastrointestinal; EMT, Epithelial-Mesenchymal Transition While AI demonstrates considerable potential in stratifying patients, predicting prognosis and guiding treatment decisions for GI cancers, several limitations are existed. Currently, there is still a lack of targeted models capable of effectively identifying cancer type-specific biomarkers, such as CDH1 or CTNNA1 mutations and CDKN2A promoter methylation [ 238 , 239 ]. Moreover, it remains challenging to systematically integrate conventional pathological classifications ( e.g. , the Lauren classification) with clinical outcomes and therapeutic responses, particularly in the context of neoadjuvant therapy. Although several studies have reported that gastric cancer patients with the intestinal-type Lauren classification, serving as a histological biomarker, exhibit pathological regression after neoadjuvant immunochemotherapy, this treatment appears to reverse the transition from the intestinal to the diffuse type and reprogram the immunosuppressive TME [ 240 , 241 ]. In addition, despite numerous models predict metastasis, few prognostic systems are built upon features of the pre-metastatic niche, such as immune infiltration, stromal composition, or metabolic state [ 242 ]. At the therapeutic level, most existing models are confined to single or limited treatment regimens and fail to provide systematic guidance tailored to molecular subtypes, hindering clinical translation. Model training often relies heavily on TCGA data, and generalizability needs further validation across multi-center external cohorts. High-contribution features identified by these models also require wet-lab validation to elucidate their biological mechanisms and enhance interpretability and credibility. Further interdisciplinary efforts are needed to optimize model architectures, expand data sources, and strengthen biological verification, ultimately improving the accuracy, interpretability, and generalizability of AI in GI cancer prognosis and therapeutic decision-making for patients. Challenges and prospects AI algorithms have shown considerable promise in GI oncology, including screening, patient stratification, prognosis evaluation, treatment decision-making and mechanistic exploration. Nevertheless, the clinical application of these advances remains challenging. Major barriers comprise the complexity of multimodal data integration, limited interpretability of models, difficulties in obtaining high-quality training datasets, clinical risks arising from false negative/positive results, restricted human–computer interaction capabilities, and patient privacy, rights, and ethical considerations. Challenges In multimodal foundation model research, stylistic heterogeneity within the same modality presents a major challenge for model recognition and generalization. For example, WSIs can differ substantially across staining techniques, including conventional H&E, immunohistochemistry, and multiplex fluorescence staining. Similarly, medical imaging data encompass multiple acquisition protocols and dimensional formats. In addition, intraoperative frozen sections, the gold standard for rapid pathological diagnosis, still lack a reliable model to support robust real-time analysis. Current multimodal approaches often merge heterogeneous data directly into encoders without excluding intra-modal stylistic variations, which make models overly sensitive to certain data types, thereby causing significant prediction bias. Model interpretability is another critical barrier to clinical application, especially to clinician acceptance. Current interpretability techniques, such as image feature attribution, are largely confined to evaluating the importance of data feature, which provide limited insights into how deep semantic representations relate to underlying pathological mechanisms. Moreover, generative large language models delivering concept-level explanations of model reasoning and addressing natural language queries are still lacking. This gap substantially impacts both the scientific credibility and the clinical reliability of AI systems. On the other hand, the foundation models in GI oncology are heavily dependent on large-scale, high-quality datasets, which remain scarce in this field. In addition, the complex and time-intensive preprocessing procedures of various data impose substantial burdens on both model training and validation. The robustness and reliability of these models under few-shot scenarios also require further investigation in large datasets. False negatives and false positives represent another critical bottleneck for hindering clinical deployment. While false negatives are typically regarded as higher-risk in binary classification tasks, risk assessment becomes more nuanced in multiclass contexts, such as drug selection. Failure to effectively control both types of errors can compromise model reliability and raise concerns regarding medical safety and ethical responsibility. Human–computer interaction remains an underdeveloped aspect of AI systems. Although several general-purpose large language models have demonstrated potential in certain contexts such as early screening for GI tumors [ 243 , 244 ], generation of cancer-related linguistic information [ 245 ], and support for dataset and model construction [ 246 , 247 ], their integration into clinical workflows remains limited. Most systems rely on code-driven interfaces, which restricts practical usability, particularly for clinicians lacking programming expertise. Moreover, limited computational resources and insufficient capability for real-time deployment of multiple modules often hinder the integration of large models into clinical practice. The training and inference of large-scale models typically require high-performance computing infrastructure, and the deployment of multi-functional clinical workflows further increases implementation costs-posing significant challenges for healthcare systems, particularly preliminary medical institutions. Even in settings equipped for cloud-based deployment, factors such as network stability and platforms compatibility with should be carefully considered. Finally, data privacy and access control pose substantial challenges to the use of real-world clinical oncology data. Although datasets from existing studies are typically de-identified in accordance with institutional and regulatory standards, residual risks persist, particularly when integrating multi-center data with heterogeneous governance policies, consent frameworks, and access privileges. Such heterogeneity complicates the consistent protection of patient confidentiality and increases the risk of unintended information leakage, including membership inference or reconstruction attacks [ 248 ], which also raised broader ethical concerns related to stakeholder access asymmetry, patient autonomy, accountability of AI-assisted decision-making, algorithmic bias, and the potential erosion of clinicians’ interpretability and control. Privacy-preserving approaches such as federated learning [ 249 ], secure aggregation, and differential privacy have been proposed to alleviate data-sharing constraints and support collaborative model training. However, their adoption in gastrointestinal oncology remains limited and introduces additional technical and regulatory challenges, including performance trade-offs, communication overhead, and difficulties in validating models trained on decentralized data. Collectively, these issues underscore that ethical governance and privacy protection remain critical and unresolved challenges for the clinical deployment of AI models in oncology. Prospects Facing these challenges, future research should prioritize the development of efficient fusion algorithms capable of handling heterogeneous data styles within the same modality. Approaches such as multimodal contrastive learning and generative modeling ( e.g. , VAE, GAN, and diffusion models) can be employed for feature transformation and integration to improve model perception and generalization across diverse datasets. Accumulating evidence suggests that deployment efficiency can be further improved by leveraging pre-trained large models to extract generalizable features, followed by fine-tuning on downstream tasks [ 39 , 98 , 179 , 236 , 237 ]. Furthermore, knowledge transfer strategies, particularly knowledge distillation, have been shown to reduce inference time and improve computational efficiency across a range of applications [ 180 , 225 ]. Therefore, it is critical for building transfer learning and knowledge distillation frameworks specifically optimized for few-shot learning and multi-center collaborative settings [ 250 ]. Regarding interpretability, future efforts should focus on developing concept-aware generative large language models to produce clinically and pathologically meaningful explanations in response to natural language queries. Furthermore, the integration of uncertainty quantification, continual learning, domain adaptation, and synthetic data augmentation can reduce false-negative and false-positive rates, enhancing both robustness and clinical applicability. Reinforcement learning offers a mechanism to improve clinical event discrimination through reward-penalty optimization, whereas federated learning enhances model generalization and protects patient privacy by enabling distributed training across multiple centers without direct data sharing. These approaches have been applied for guiding first-line Helicobacter pylori treatment strategies in gastric cancer and supporting subtype differentiation for colorectal and esophageal adenocarcinomas [ 208 , 251 ]. Looking forward, reinforcement learning and federated learning could also be leveraged to strengthen human–computer interaction. In parallel, optimizing interactive interfaces with advanced technologies such as AI agents, allowing clinicians to operate models through natural language commands, would markedly improve usability and clinical integration [ 252 ]. In general, key priorities for promoting the application of AI in GI oncology include optimizing multimodal fusion algorithms, establishing transfer learning strategies tailored for few-shot and multi-center datasets, and developing innovative approaches to reduce diagnostic errors. Other important improvement depends on advancing concept-level interpretability, enhancing clinical human–computer interaction (including model quantization and containerization, standardized interfaces for clinical AI platforms, and integrated deployment workflows), incorporating cutting-edge technologies such as AI agents, and introducing controllable noise into model parameters, employing GAN-generated synthetic data for training, enforcing dynamic access control mechanisms to safeguard patient privacy, data rights, and ethical integrity. Taken together, these efforts will pave the way for an integrated, real-time, and interpretable AI-assisted system that spans the entire clinical workflow, from specimen acquisition and diagnosis to prognosis evaluation and treatment decision-making. By integrating multi-dimensional information, such a system would deliver efficient, transparent, and reliable clinical support, ultimately raising the standard of care for patients with GI cancers. Conclusion The integration of heterogeneous data sources and the advancement of foundation models are accelerating the application of AI in gastrointestinal oncology. These technologies enhance early screening of cancer, improve individualized treatment, and offer new insights into tumor biology. However, challenges remain in model interpretability, generalizability, and clinical adaptability. Future directions should focus on developing cross-modal, lightweight, and interpretable AI frameworks to support robust and scalable deployment. Ultimately, AI is poised to become a central pillar of precision oncology, bridging algorithmic innovation with clinical treatment. Acknowledgements We thank all the lab members for critical reading and comments on the manuscript. Abbreviations LightGBM Light-Gradient Boosting Machine XGBoost Extreme Gradient Boosting LR Logistic Regression DT Decision Tree RF Random Forest SVM Support Vector Machine MLP Multilayer Perceptron CNN Convolutional Neural Network LLM Large Language Model EC Esophageal Cancer ESCC Esophageal Squamous Cell Carcinoma AEG Adenocarcinoma of Esophagogastric Junction BE Barrett’s Esophagus CRC Colorectal Cancer GC Gastric Cancer LAGC Locally Advanced Gastric Cancer EAC Esophageal Adenocarcinoma SSLs Sessile Serrated Lesions WLI White Light Imaging NBI Narrow Band Imaging ME Magnifying Endoscopy WCE Wireless Capsule Endoscopy FI Fluorescence Imaging PCI Pseudocolor Imaging OS Overall Survival DFS Disease-Free Survival ctDNA Circulating Tumor DNA cfDNA Cell-Free DNA EV Extracellular Vesicle CEA Carcinoembryonic Antigen AG Atrophic Gastritis FSP Fragmentation Size Profile CNV Copy Number Variation FBM Fragment Based Methylation FSC Fragmentation Size Coverage FSD Fragment Size Distribution NP Nucleosome Positioning FSR Fragment Size Ratio NF Nucleosome Footprint MC Mutational Context NK cell Natural Killer cell DC Dendritic Cell ECM Extracellular Matrix VOCs Volatile Organic Compounds SERS Surface-Enhanced Raman Spectroscopy VAE Variational Auto-Encoders GNN Graph Neural Network HE Hematoxylin and Eosin WSI Whole Slide Image dMMR Mismatch Repair deficiency MSI Microsatellite Instability TFF3 Trefoil Factor 3 IHC Immunohistochemical mIHC Multiple-Immunohistochemical PR Peritoneal Recurrence T2W T2-Weighted Imaging DWI Diffusion Weighted Imaging NAT Neoadjuvant Therapy TIL Tumor-Infiltrating Lymphocyte EMT Epithelial-Mesenchymal Transition TME Tumor Microenvironment Treg Regulatory T cell mIF Multiplex Immunofluorescence RFS Recurrence-Free Survival DSS Disease-Specific Survival PFS Progression-Free Survival TSR Tumor-Stroma Ratio HPC Histomorphological Phenotypic Cluster LNM Lymph Node Metastasis TMB Tumor Mutational Burden HGB Histopathological Growth Patterns CLM Colorectal Liver Metastases MRD Molecular Residual Disease GAN Generative Adversarial Network WGAN Wasserstein Generative Adversarial Network GCN Graph Convolutional Network GAT Graph Attention Network TTR Time To Recurrence TRG Tumor Regression Grade pCR Pathological Complete Response nICT Neoadjuvant Immunochemotherapy ICI Immune Checkpoint Inhibitor MDSC Myeloid-Derived Suppressor Cell ST Spatial Transcriptomics ecDNA Extrachromosomal DNA WES Whole-Exome Sequencing WGS Whole Genome Sequencing VAF Variant Allele Frequency ICB Immune Checkpoint Blockade NCRT Neoadjuvant Chemoradiotherapy LARC Locally Advanced Rectal Cancer SNV Single Nucleotide Variants TCN Tissue Cell Neighborhood CTL Cytotoxic T Lymphocyte GIST Gastrointestinal Stromal Tumor LVI Lymphovascular Invasion ADC Apparent Diffusion Coefficient CNA Copy Number Alterations ORR Overall Response Rate CIN Chromosomal Instability CAFs Cancer-associated Fibroblasts MNV Multi-nucleotide Variants Indels Short insertions and deletions SV Structural Variant MEI Mobile Element Insertions SBS Single-base Substitution DBS Doublet-base Substitution ID Insertion-and-deletion acc Accuracy AUC/AUROC Area Under the Receiver Operating Characteristic Curve AUPRC Area Under the Precision-Recall Curve MAE Mean Absolute Error MCC Matthews Correlation Coefficient Authors’ contributions Kaijie Liu and Wenjie Zhang collected, organized, and analyzed references and results, wrote the manuscript. Kaijie Liu, Wenjie Zhang, Zeyu Luo and Zhouyu Yang designed the figures and tables. Qiaoqiao Zhang Zeyu Luo provided technical support. Kaijie Liu, Zeyu Luo, Wenjie Zhang, Bin Wang, Bo Tang, Zongsheng He, and Jinjun Guo proposed the review topic and supervised the literature review. Zeyu Luo, Bin Wang, Zongsheng He and Jinjun Guo edited the manuscript and approved the submission. All the authors contributed to the final version of the manuscript. Funding This work was sponsored by the grants from the National Key Research and Development Program of China (Nos. 2023YFC3402100 and 2022YFA1105300 to Bin Wang), the National Natural Science Foundation of China (32500508 to Qiaoqiao Zhang, 82203318 to Zongsheng He), Science and Technology Innovation Key R&D Program of Chongqing (CSTB2024TIAD-KPX0028 to Qiaoqiao Zhang, CSTB2023TIAD-STX0002 to Bin Wang), Natural Science Foundation of Chongqing (Nos. CSTB2023NSCQ-LZX0156 to Bin Wang and CSTB2022NSCQ-MSX0880 and CSTB2022NSCQ-MSX0952 to Zongsheng He), Chongqing Medical Talent Leadership Program (YXLJ202401 to Jinjun Guo), and Jinfeng Lab Fundamental Research Funds (JFLKYXM202203AZ-215 to Bin Wang). Data availability No datasets were generated or analysed during the current study. Declarations Ethics approval and consent to participate Not applicable. Competing interests The authors declare no competing interests. Footnotes Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Kaijie Liu, Zeyu Luo and Wenjie Zhang are co-first authors. Contributor Information Bin Wang, Email: [email protected]. Bo Tang, Email: [email protected]. Zongsheng He, Email: [email protected]. Jinjun Guo, Email: [email protected]. References 1. Bray F, Laversanne M, Sung H, Ferlay J, Siegel RL, Soerjomataram I, et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2024;74:229–63. [ DOI ] [ PubMed ] [ Google Scholar ] 2. Diao X, Guo C, Jin Y, Li B, Gao X, Du X, et al. Cancer situation in China: an analysis based on the global epidemiological data released in 2024. Cancer Commun. 2025;45:178–97. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 3. Zhan T, Betge J, Schulte N, Dreikhausen L, Hirth M, Li M, et al. Digestive cancers: mechanisms, therapeutics and management. Signal Transduct Target Ther. 2025;10:24. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 4. Zhang P, Zhang C, Li X, Chang C, Gan C, Ye T, et al. Immunotherapy for gastric cancer: advances and challenges. MedComm – Oncology. 2024;3:e92. [ Google Scholar ] 5. Hanahan D. Hallmarks of cancer: new dimensions. Cancer Discov. 2022;12:31–46. [ DOI ] [ PubMed ] [ Google Scholar ] 6. Zhang C, Xu J, Tang R, Yang J, Wang W, Yu X, et al. Novel research and future prospects of artificial intelligence in cancer diagnosis and treatment. J Hematol Oncol. 2023;16:114. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 7. You Y, Lai X, Pan Y, Zheng H, Vera J, Liu S, et al. Artificial intelligence in cancer target identification and drug discovery. Signal Transduct Target Ther. 2022;7:156. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 8. Piras A, Corso R, Benfante V, Ali M, Laudicella R, Alongi P, et al. Artificial intelligence and statistical models for the prediction of radiotherapy toxicity in prostate cancer: a systematic review. Appl Sci. 2024;14(23):10947. [ Google Scholar ] 9. Nagaraju GP, Sandhya T, Srilatha M, Ganji SP, Saddala MS, El-Rayes BF. Artificial intelligence in gastrointestinal cancers: diagnostic, prognostic, and surgical strategies. Cancer Lett. 2025;612:217461. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 10. Yates J, Van Allen EM. New horizons at the interface of artificial intelligence and translational cancer research. Cancer Cell. 2025;43:708–27. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 11. Fahrner LJ, Chen E, Topol E, Rajpurkar P. The generative era of medical AI. Cell. 2025;188:3648–60. [ DOI ] [ PubMed ] [ Google Scholar ] 12. Hosmer DW, Lemeshow S. Applied logistic regression: Applied logistic regression. 3rd ed. Hoboken (NJ): Wiley; 2013. 13. Chang C-C, Lin C-J. LIBSVM: A library for support vector machines. ACM Trans Intell Syst Technol. 2011;2:27. [ Google Scholar ] 14. Breiman L. Random forests. Mach Learn. 2001;45:5–32. [ Google Scholar ] 15. Chen T, Guestrin C. XGBoost: A Scalable Tree Boosting System. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; San Francisco, California, USA: Association for Computing Machinery; 2016. p. 785–94. 16. Meng Q, editor LightGBM: A Highly Efficient Gradient Boosting Decision Tree. Neural Information Processing Systems. 2017;30:1–9. 17. Fard MM, Thonet T, Gaussier E. Deep k-means: jointly clustering with k-means and learning representations. Pattern Recognit Lett. 2020;138:185–92. [ Google Scholar ] 18. Lever J, Krzywinski M, Altman N. Principal component analysis. Nat Methods. 2017;14:641–2. [ Google Scholar ] 19. McInnes L, Healy J, Melville J. UMAP: uniform manifold approximation and projection for dimension reduction. arXiv. 2018;abs/1802.03426:1–63. 10.48550/arXiv.1802.03426. 20. Muthukrishnan R, Rohini R. LASSO: A feature selection technique in predictive modeling for machine learning. 2016 IEEE International Conference on Advances in Computer Applications (ICACA). 2016;18–20. 21. Chen Xw, Jeong JC. Enhanced recursive feature elimination. Sixth International Conference on Machine Learning and Applications (ICMLA 2007). 2007;429–35. 22. Popescu M-C, Balas V, Perescu-Popescu L, Mastorakis N. Multilayer perceptron and neural networks. WSEAS Transactions on Circuits and Systems. 2009;8(7):579–88. 23. Christensen MH, Drue SO, Rasmussen MH, Frydendahl A, Lyskjær I, Demuth C, et al. DREAMS: deep read-level error model for sequencing data applied to low-frequency variant calling and circulating tumor DNA detection. Genome Biol. 2023;24:99. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 24. Bouzid K, Sharma H, Killcoyne S, Castro DC, Schwaighofer A, Ilse M, et al. Enabling large-scale screening of Barrett’s esophagus using weakly supervised deep learning in histopathology. Nat Commun. 2024;15:2026. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 25. Bai F, Liao L, Tang Y, Wu Y, Wang Z, Zhao H, et al. RCMIX model based on pre-treatment MRI imaging predicts T-downstage in MRI-cT4 stage rectal cancer. Cancer Lett. 2025;628:217871. [ DOI ] [ PubMed ] [ Google Scholar ] 26. KimY. Convolutional neural networks for sentence classification. arXiv preprint arXiv:14085882. 2014. 10.48550/arXiv.1408.5882. 27. He K, Zhang X, Ren S, Sun J, editors. Deep residual learning for image recognition Proceedings of the IEEE conference on computer vision and pattern recognition; 2016. 10.48550/arXiv.1512.03385. 28. Mahmood T, Li J, Pei Y, Akhtar F, Rehman MU, Wasti SH. Breast lesions classifications of mammographic images using a deep convolutional neural network-based approach. PLoS ONE. 2022;17:e0263126. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 29. Mahmood T, Li J, Pei Y, Akhtar F. An Automated In-Depth Feature Learning Algorithm for Breast Abnormality Prognosis and Robust Characterization from Mammography Images Using Deep Transfer Learning. Biology [Internet]. 2021; 10(9):[859 p.]. [ DOI ] [ PMC free article ] [ PubMed ] 30. Jain S, Atale R, Gupta A, Mishra U, Seal A, Ojha A, et al. Coinnet: a convolution-involution network with a novel statistical attention for automatic polyp segmentation. IEEE Trans Med Imaging. 2023;42:3987–4000. [ DOI ] [ PubMed ] [ Google Scholar ] 31. Cao M, Hu C, Li F, He J, Li E, Zhang R, et al. Development and validation of a deep learning model for predicting gastric cancer recurrence based on CT imaging: a multicenter study. International Journal of Surgery. 2024;110(12):7598–606. [ DOI ] [ PMC free article ] [ PubMed ] 32. Scarselli F, Gori M, Tsoi AC, Hagenbuchner M, Monfardini G. The graph neural network model. IEEE Trans Neural Networks. 2009;20:61–80. [ DOI ] [ PubMed ] [ Google Scholar ] 33. Graham S, Minhas F, Bilal M, Ali M, Tsang YW, Eastwood M, et al. Screening of normal endoscopic large bowel biopsies with interpretable graph learning: a retrospective study. Gut. 2023;72:1709. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 34. Gao R, Yuan X, Ma Y, Wei T, Johnston L, Shao Y, et al. Harnessing TME depicted by histological images to improve cancer prognosis through a deep learning system. Cell Rep Med. 2024;5:101536. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 35. Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is all you need. Advances in neural information processing systems. 2017;30:1–11. 36. Xiao B, Hu J, Li W, Pun CM, Bi X. Ctnet: contrastive transformer network for polyp segmentation. IEEE Trans Cybern. 2024;54:5040–53. [ DOI ] [ PubMed ] [ Google Scholar ] 37. Yuan L, Yang L, Zhang S, Xu Z, Qin J, Shi Y, et al. Development of a tongue image-based machine learning tool for the diagnosis of gastric cancer: a prospective multicentre clinical cohort study. eClinicalMedicine. 2023;57:101834. [ DOI ] [ PMC free article ] [ PubMed ] 38. Huang YQ, Chen XB, Cui YF, Yang F, Huang SX, Li ZH, et al. Enhanced risk stratification for stage II colorectal cancer using deep learning-based CT classifier and pathological markers to optimize adjuvant therapy decision. Ann Oncol. 2025;36:1178–89. [ DOI ] [ PubMed ] [ Google Scholar ] 39. Loeffler CML, Bando H, Sainath S, Muti HS, Jiang X, van Treeck M, et al. HIBRID: histology-based risk-stratification with deep learning and ctDNA in colorectal cancer. Nat Commun. 2025;16:7561. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 40. Ronneberger O, Fischer P, Brox T. U-net: Convolutional networks for biomedical image segmentation. Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015. 2015;9351:234–41. 41. Hu C, Xia Y, Zheng Z, Cao M, Zheng G, Chen S, et al. AI-based large-scale screening of gastric cancer from noncontrast CT imaging. Nat Med. 2025;31(9):3011–19. [ DOI ] [ PMC free article ] [ PubMed ] 42. Liao Y, Chen X, Hu S, Chen B, Zhuo X, Xu H, et al. Artificial intelligence for predicting HER2 status of gastric cancer based on whole-slide histopathology images: a retrospective multicenter study. Adv Sci. 2025;12:2408451. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 43. Yao L, Li S, Tao Q, Mao Y, Dong J, Lu C, et al. Deep learning for colorectal cancer detection in contrast-enhanced CT without bowel preparation: a retrospective, multicentre study. eBioMedicine. 2024;104:105183. [ DOI ] [ PMC free article ] [ PubMed ] 44. Xia Z, Li H, Lan L. MedFormer: hierarchical medical vision transformer with content-aware dual sparse selection attention. Physics in Medicine & Biology. 2025;70:195005. [ DOI ] [ PubMed ] 45. Chen J, Lu Y, Yu Q, Luo X, Adeli E, Wang Y, et al. Transunet: Transformers make strong encoders for medical image segmentation. arXiv preprint arXiv:210204306. 2021. 10.48550/arXiv.2102.04306. 46. Kingma DP, Welling M. Auto-encoding variational bayes. arXiv preprint arXiv:13126114. 2013. 10.48550/arXiv.1312.6114. 47. Goodfellow I, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, et al. Generative adversarial networks. Commun ACM. 2020;63:139–44. [ Google Scholar ] 48. Rombach R, Blattmann A, Lorenz D, Esser P, Ommer B, editors. High-resolution image synthesis with latent diffusion models. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition; 2022. 10.48550/arXiv.2112.10752. 49. Liu S, Chen W, Zheng B, Pan W, Li X, Yuan Y. TumorGen: Boundary-Aware Tumor-Mask Synthesis with Rectified Flow Matching. arXiv preprint arXiv:250524687. 2025. 10.48550/arXiv.2505.24687. 50. Wang W, Cui ZX, Cheng G, Cao C, Xu X, Liu Z, et al. A Two-Stage Generative Model with CycleGAN and Joint Diffusion for MRI-based Brain Tumor Detection. IEEE J Biomed Health Inf. 2024;28:3534–44. [ DOI ] [ PubMed ] [ Google Scholar ] 51. Wang X, Yang S, Zhang J, Wang M, Zhang J, Yang W, et al. Transformer-based unsupervised contrastive learning for histopathological image classification. Med Image Anal. 2022;81:102559. [ DOI ] [ PubMed ] [ Google Scholar ] 52. Lu MY, Williamson DFK, Chen TY, Chen RJ, Barbieri M, Mahmood F. Data-efficient and weakly supervised computational pathology on whole-slide images. Nat Biomed Eng. 2021;5:555–70. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 53. Chen RJ, Ding T, Lu MY, Williamson DFK, Jaume G, Song AH, et al. Towards a general-purpose foundation model for computational pathology. Nat Med. 2024;30:850–62. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 54. Wang X, Zhao J, Marostica E, Yuan W, Jin J, Zhang J, et al. A pathology foundation model for cancer diagnosis and prognosis prediction. Nature. 2024;634:970–8. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 55. Xu H, Usuyama N, Bagga J, Zhang S, Rao R, Naumann T, et al. A whole-slide foundation model for digital pathology from real-world data. Nature. 2024;630:181–8. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 56. Jong MR, Boers TGW, Fockens KN, Jukema JB, Kusters CHJ, Jaspers TJM, et al. GastroNet-5M: A multicenter dataset for developing foundation models in gastrointestinal endoscopy. Gastroenterology. 2026;170(1):174–87. 10.1053/j.gastro.2025.07.030. [ DOI ] [ PubMed ] 57. Yang J, Cai D, Liu J, Zhuang Z, Zhao Y, Wang F-a, et al. CRCFound: A Colorectal Cancer CT Image Foundation Model Based on Self-Supervised Learning. Adv Sci. 2025;n/a:e07339. [ DOI ] [ PMC free article ] [ PubMed ] 58. Wang X, Jiang Y, Yang S, Wang F, Zhang X, Wang W, et al. Foundation Model for Predicting Prognosis and Adjuvant Therapy Benefit From Digital Pathology in GI Cancers. J Clin Oncol. 2025;0:JCO-24-01501. 10.1200/JCO.24.01501. [ DOI ] [ PubMed ] 59. Xiang J, Wang X, Zhang X, Xi Y, Eweje F, Chen Y, et al. A vision–language foundation model for precision oncology. Nature. 2025;638:769–78. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 60. Hua S, Yan F, Shen T, Ma L, Zhang X. Pathoduet: Foundation models for pathological slide analysis of H&E and IHC stains. Med Image Anal. 2024;97:103289. [ DOI ] [ PubMed ] [ Google Scholar ] 61. Mahmood T, Saba T, Rehman A. Breast cancer diagnosis with MFF-HistoNet: a multi-modal feature fusion network integrating CNNs and quantum tensor networks. J Big Data. 2025;12:60. [ Google Scholar ] 62. Chang Y, Jung C, Ke P, Song H, Hwang J. Automatic contrast-limited adaptive histogram equalization with dual gamma correction. IEEE Access. 2018;6:11782–92. [ Google Scholar ] 63. Valanarasu JMJ, Xu H, Usuyama N, Kim C, Wong C, Argaw P, et al. Multimodal AI generates virtual population for tumor microenvironment modeling. Cell. 2026;189(2):386–400.e19. 10.1016/j.cell.2025.11.016. [ DOI ] [ PubMed ] 64. Wang T, Luo Z. Large language models transform biological research: from architecture to utilization. Sci China Inf Sci. 2025;68:170101. [ Google Scholar ] 65. Dalla-Torre H, Gonzalez L, Mendoza-Revilla J, Lopez Carranza N, Grzywaczewski AH, Oteri F, et al. Nucleotide Transformer: building and evaluating robust foundation models for human genomics. Nat Methods. 2025;22:287–97. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 66. Ji Y, Zhou Z, Liu H, Davuluri RV. DNABERT: pre-trained bidirectional encoder representations from transformers model for DNA-language in genome. Bioinformatics. 2021;37:2112–20. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 67. Nguyen E, Poli M, Durrant MG, Kang B, Katrekar D, Li DB, et al. Sequence modeling and design from molecular to genome scale with Evo. Science. 2024;386:eado9336. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 68. Avsec Ž, Agarwal V, Visentin D, Ledsam JR, Grabska-Barwinska A, Taylor KR, et al. Effective gene expression prediction from sequence by integrating long-range interactions. Nat Methods. 2021;18:1196–203. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 69. UhlM, Tran VD, Heyl F, Backofen R. GraphProt2: A graph neural network-based method for predicting binding sites of RNA-binding proteins. BioRxiv. 2019:850024. 10.1101/850024. 70. Elnaggar A, Heinzinger M, Dallago C, Rehawi G, Wang Y, Jones L, et al. Prottrans: toward understanding the language of life through self-supervised learning. IEEE Trans Pattern Anal Mach Intell. 2022;44:7112–27. [ DOI ] [ PubMed ] [ Google Scholar ] 71. Abramson J, Adler J, Dunger J, Evans R, Green T, Pritzel A, et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature. 2024;630:493–500. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 72. Lin Z, Akin H, Rao R, Hie B, Zhu Z, Lu W, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. 2023;379:1123–30. [ DOI ] [ PubMed ] [ Google Scholar ] 73. Yang F, Wang W, Wang F, Fang Y, Tang D, Huang J, et al. ScBERT as a large-scale pretrained deep language model for cell type annotation of single-cell RNA-seq data. Nat Mach Intell. 2022;4:852–66. [ Google Scholar ] 74. Theodoris CV, Xiao L, Chopra A, Chaffin MD, Al Sayed ZR, Hill MC, et al. Transfer learning enables predictions in network biology. Nature. 2023;618:616–24. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 75. Cui H, Wang C, Maan H, Pang K, Luo F, Duan N, et al. ScGPT: toward building a foundation model for single-cell multi-omics using generative AI. Nat Methods. 2024;21:1470–80. [ DOI ] [ PubMed ] [ Google Scholar ] 76. QiC, Fang H, Hu T, Jiang S, Zhi W. Bidirectional Mamba for Single-Cell Data: Efficient Context Learning with Biological Fidelity. arXiv preprint arXiv:250416956. 2025. 10.48550/arXiv.2504.16956. 77. GuA, Dao T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:231200752. 2023. 10.48550/arXiv.2312.00752. 78. Avsec Ž, Latysheva N, Cheng J, Novati G, Taylor KR, Ward T, et al. AlphaGenome: advancing regulatory variant effect prediction with a unified DNA sequence model. bioRxiv. 2025:2025.06. 25.661532. 10.1101/2025.06.25.661532. 79. Murphy AE, Beardall W, Rei M, Phuycharoen M, Skene NG. Predicting cell type-specific epigenomic profiles accounting for distal genetic effects. Nat Commun. 2024;15:9951. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 80. Gao Z, Liu Q, Zeng W, Jiang R, Wong WH. EpiGePT: a pretrained transformer-based language model for context-specific human epigenomics. Genome Biol. 2024;25:310. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 81. Zhang G, Song C, Yin M, Liu L, Zhang Y, Li Y, et al. TRAPT: a multi-stage fused deep learning framework for predicting transcriptional regulators based on large-scale epigenomic data. Nat Commun. 2025;16:3611. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 82. Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, et al. Learning Transferable Visual Models From Natural Language Supervision. In: Proceedings of the 38th International Conference on Machine Learning (ICML); 2021 Jul 18–24; Virtual. PMLR 139:8748–63. 83. Lipkova J, Chen RJ, Chen B, Lu MY, Barbieri M, Shao D, et al. Artificial intelligence for multimodal data integration in oncology. Cancer Cell. 2022;40:1095–110. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 84. Brixi G, Durrant MG, Ku J, Poli M, Brockman G, Chang D, et al. Genome modeling and design across all domains of life with Evo 2. BioRxiv. 2025:2025.02. 18.638918. 10.1101/2025.02.18.638918. [ DOI ] [ PMC free article ] [ PubMed ] 85. Shoshan Y, Raboh M, Ozery-Flato M, Ratner V, Golts A, Weber JK, et al. MAMMAL--Molecular Aligned Multi-Modal Architecture and Language. arXiv preprint arXiv:241022367. 2024. 10.48550/arXiv.2410.22367. 86. Wang G, Zhao J, Lin Y, Liu T, Zhao Y, Zhao H. Scmodal: a general deep learning framework for comprehensive single-cell multi-omics data alignment with feature links. Nat Commun. 2025;16:4994. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 87. Madhu H, Rocha JF, Huang T, Viswanath S, Krishnaswamy S, Ying R. HEIST: A Graph Foundation Model for Spatial Transcriptomics and Proteomics Data. arXiv preprint arXiv:250611152. 2025. 10.48550/arXiv.2506.11152. 88. Chen W, Zhang P, Tran TN, Xiao Y, Li S, Shah VV, et al. A visual–omics foundation model to bridge histopathology with spatial transcriptomics. Nat Methods. 2025;22:1568–82. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 89. Moor M, Huang Q, Wu S, Yasunaga M, Dalmia Y, Leskovec J, et al. Med-Flamingo: a Multimodal Medical Few-Shot Learner. Proc. Machine Learning for Health (ML4H), PMLR. 2023;225:353–67. 90. Zhang S, Xu Y, Usuyama N, Xu H, Bagga J, Tinn R, et al. Biomed CLIP: A multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs. arXiv 2023. arXiv preprint arXiv:230300915. 2023. 10.48550/arXiv.2303.00915. 91. Luo Y, Zhang J, Fan S, Yang K, Hong M, Wu Y, et al. BioMedGPT: An Open Multimodal Large Language Model for BioMedicine. IEEE J Biomed Health Inf. 2024:1–12. 10.1109/JBHI.2024.3505955. [ DOI ] [ PubMed ] 92. Lu MY, Chen B, Williamson DFK, Chen RJ, Zhao M, Chow AK, et al. A multimodal generative AI copilot for human pathology. Nature. 2024;634:466–73. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 93. Rehman A, Mahmood T, Saba T. Robust kidney carcinoma prognosis and characterization using Swin-ViT and DeepLabV3+ with multi-model transfer learning. Appl Soft Comput. 2025;170:112518. [ Google Scholar ] 94. Mahmood T, Rehman A, Saba T, Wang Y, Alamri FS. Alzheimer’s disease unveiled: cutting-edge multi-modal neuroimaging and computational methods for enhanced diagnosis. Biomed Signal Process Control. 2024;97:106721. [ Google Scholar ] 95. Zheng Y, Qiu B, Liu S, Song R, Yang X, Wu L, et al. A transformer-based deep learning model for early prediction of lymph node metastasis in locally advanced gastric cancer after neoadjuvant chemotherapy using pretreatment CT images. EClinicalMedicine. 2024;75:102805. 10.1016/j.eclinm.2024.102805. [ DOI ] [ PMC free article ] [ PubMed ] 96. Jiang X, Zhao H, Saldanha OL, Nebelung S, Kuhl C, Amygdalos I, et al. An MRI deep learning model predicts outcome in rectal cancer. Radiology. 2023;307:e222223. [ DOI ] [ PubMed ] [ Google Scholar ] 97. Feng B, Shi J, Huang L, Yang Z, Feng S-T, Li J, et al. Robustly federated learning model for identifying high-risk patients with postoperative gastric cancer recurrence. Nat Commun. 2024;15:742. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 98. Hu Y, Rong J, Xu Y, Xie R, Peng J, Gao L, et al. Unsupervised and supervised discovery of tissue cellular neighborhoods from cell phenotypes. Nat Methods. 2024;21:267–78. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 99. Wang C, Yuan C, Wang Y, Chen R, Shi Y, Zhang T, et al. MPI-VGAE: protein–metabolite enzymatic reaction link learning by variational graph autoencoders. Brief Bioinform. 2023;24:bbad189. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 100. Jiang Y, Immadi MS, Wang D, Zeng S, On Chan Y, Zhou J, et al. IRnet: immunotherapy response prediction using pathway knowledge-informed graph neural network. J Adv Res. 2025;72:319–31. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 101. Zohora FT, Paliwal D, Flores-Figueroa E, Li J, Gao T, Notta F, et al. Cellnest reveals cell–cell relay networks using attention mechanisms on spatial transcriptomics. Nat Methods. 2025;22:1505–19. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 102. Gao P, Xiao Q, Tan H, Song J, Fu Y, Xu J, et al. Interpretable multi-modal artificial intelligence model for predicting gastric cancer response to neoadjuvant chemotherapy. Cell Rep Med. 2024;5:101848. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 103. Ge Y, Leng J, Tang Z, Wang K, U K, Zhang SM, et al. Deep learning-enabled integration of histology and transcriptomics for tissue spatial profile analysis. Research. 2025;8:0568. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 104. Lundberg SM, Lee S-I. A unified approach to interpreting model predictions. Adv Neural Inf Process Syst. 2017;30:4765–74. 105. Yiwen Z, Xuan Z, Xi Z, Liebin H, Wenjing J, Caixia Z, et al. Immunophenotype-guided interpretable radiomics model for predicting neoadjuvant anti-PD-1 response in stage III-IV d-MMR/MSI-H colorectal cancer. J ImmunoTher Cancer. 2025;13:e011569. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 106. Wang R, Liu Q, You W, Wang H, Chen Y. A transformer-based deep learning survival prediction model and an explainable XGBoost anti-PD-1/PD-L1 outcome prediction model based on the cGAS-STING-centered pathways in hepatocellular carcinoma. Brief Bioinform. 2025;26:bbae686. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 107. Vorontsov E, Bozkurt A, Casson A, Shaikovski G, Zelechowski M, Severson K, et al. A foundation model for clinical-grade computational pathology and rare cancers detection. Nat Med. 2024;30:2924–35. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 108. Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D. Grad-CAM: Visual explanations from deep networks via gradient-based localization. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV); 2017. pp. 618–26. 10.1109/ICCV.2017.74. 109. Spartivento G, Benfante V, Ali M, Yezzi A, Di Raimondo D, Tuttolomondo A, et al. Revolutionizing periodontal care: the role of artificial intelligence in diagnosis, treatment, and prognosis. Appl Sci. 2025;15(6):3295. [ Google Scholar ] 110. Yu Y, Ren W, Mao L, Ouyang W, Hu Q, Yao Q, et al. MRI-based multimodal AI model enables prediction of recurrence risk and adjuvant therapy in breast cancer. Pharmacol Res. 2025;216:107765. [ DOI ] [ PubMed ] [ Google Scholar ] 111. Li H, Han Z, Sun Y, Wang F, Hu P, Gao Y, et al. CGMega: explainable graph neural network framework with attention mechanisms for cancer gene module dissection. Nat Commun. 2024;15:5997. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 112. Chereda H, Bleckmann A, Menck K, Perera-Bel J, Stegmaier P, Auer F, et al. Explaining decisions of graph convolutional neural networks: patient-specific molecular subnetworks responsible for metastasis prediction in breast cancer. Genome Med. 2021;13:42. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 113. Chatzianastasis M, Vazirgiannis M, Zhang Z. Explainable multilayer graph neural network for cancer gene prediction. Bioinformatics. 2023;39:btad643. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 114. Liu Z, Sun Y, Li Y, Ma A, Willaims NF, Jahanbahkshi S, et al. An explainable graph neural framework to identify cancer-associated intratumoral microbial communities. Adv Sci. 2024;11:2403393. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 115. Sundararajan M, Taly A, Yan Q. Axiomatic Attribution for Deep Networks. In: Proceedings of the 34th International Conference on Machine Learning (ICML); 2017. Proceedings of Machine Learning Research 70:3319–28. 116. Mahmood T, Li J, Pei Y, Akhtar F, Imran A, Rehman KU. A brief survey on breast cancer diagnostic with deep learning schemes using multi-image modalities. IEEE Access. 2020;8:165779–809. [ Google Scholar ] 117. Mahmood T, Rehman A, Saba T, Nadeem L, Bahaj SAO. Recent advancements and future prospects in active deep learning for medical image segmentation and classification. IEEE Access. 2023;11:113623–52. [ Google Scholar ] 118. Zhou S, Xie Y, Feng X, Li Y, Shen L, Chen Y. Artificial intelligence in gastrointestinal cancer research: image learning advances and applications. Cancer Lett. 2025;614:217555. [ DOI ] [ PubMed ] [ Google Scholar ] 119. Thiruvengadam NR, Coté GA, Gupta S, Rodrigues M, Schneider Y, Arain MA, et al. An evaluation of critical factors for the cost-effectiveness of real-time computer-aided detection: sensitivity and threshold analyses using a microsimulation model. Gastroenterology. 2023;164:906–20. [ DOI ] [ PubMed ] [ Google Scholar ] 120. Thiruvengadam NR, Solaimani P, Shrestha M, Buller S, Carson R, Reyes-Garcia B, et al. The Efficacy of Real-time Computer-aided Detection of Colonic Neoplasia in Community Practice: A Pragmatic Randomized Controlled Trial. Clin Gastroenterol Hepatol. 2024;22:2221-30.e15. [ DOI ] [ PubMed ] [ Google Scholar ] 121. Seager A, Sharp L, Neilson LJ, Brand A, Hampton JS, Lee TJW, et al. Polyp detection with colonoscopy assisted by the GI genius artificial intelligence endoscopy module compared with standard colonoscopy in routine colonoscopy practice (COLO-DETECT): a multicentre, open-label, parallel-arm, pragmatic randomised controlled trial. Lancet Gastroenterol Hepatol. 2024;9:911–23. [ DOI ] [ PubMed ] [ Google Scholar ] 122. Houwen BB, Hazewinkel Y, Giotis I, Vleugels JL, Mostafavi NS, van Putten P, et al. Computer-aided diagnosis for optical diagnosis of diminutive colorectal polyps including sessile serrated lesions: a real-time comparison with screening endoscopists. Endoscopy. 2023;55:756–65. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 123. Xu H, Tang RSY, Lam TYT, Zhao G, Lau JYW, Liu Y, et al. Artificial Intelligence–Assisted Colonoscopy for Colorectal Cancer Screening: A Multicenter Randomized Controlled Trial. Clin Gastroenterol Hepatol. 2023;21:337-46.e3. [ DOI ] [ PubMed ] [ Google Scholar ] 124. Li B, Dong Y-L, Liu J-Y, Song R, Yang X, Wu L, et al. Application of artificial intelligence lesion labeling system-assisted endoscopic submucosal dissection for the treatment of esophageal lesions in a low-volume center: a prospective cohort study. Int J Surg. 2025;111:102797. 10.1016/j.ijsu.2025.102797. [ DOI ] [ PMC free article ] [ PubMed ] 125. Fockens KN, Jong MR, Jukema JB, Boers TGW, Kusters CHJ, van der Putten JA, et al. A deep learning system for detection of early Barrett’s neoplasia: a model development and validation study. The Lancet Digital Health. 2023;5:e905–16. [ DOI ] [ PubMed ] [ Google Scholar ] 126. Rondonotti E, Hassan C, Tamanini G, Antonelli G, Andrisani G, Leonetti G, et al. Artificial intelligence-assisted optical diagnosis for the resect-and-discard strategy in clinical practice: the Artificial intelligence BLI Characterization (ABC) study. Endoscopy. 2023;55:14–22. [ DOI ] [ PubMed ] [ Google Scholar ] 127. Ahmad A, Wilson A, Haycock A, Humphries A, Monahan K, Suzuki N, et al. Evaluation of a real-time computer-aided polyp detection system during screening colonoscopy: AI-DETECT study. Endoscopy. 2023;55:313–9. [ DOI ] [ PubMed ] [ Google Scholar ] 128. Patel HK, Mori Y, Hassan C, Rizkala T, Radadiya DK, Nathani P, et al. Lack of Effectiveness of Computer Aided Detection for Colorectal Neoplasia: A Systematic Review and Meta-Analysis of Nonrandomized Studies. Clin Gastroenterol Hepatol. 2024;22:971-80.e15. [ DOI ] [ PubMed ] [ Google Scholar ] 129. Yuan X-L, Liu W, Lin Y-X, Deng Q-Y, Gao Y-P, Wan L, et al. Effect of an artificial intelligence-assisted system on endoscopic diagnosis of superficial oesophageal squamous cell carcinoma and precancerous lesions: a multicentre, tandem, double-blind, randomised controlled trial. Lancet Gastroenterol Hepatol. 2024;9:34–44. [ DOI ] [ PubMed ] [ Google Scholar ] 130. Nakao E, Yoshio T, Kato Y, Namikawa K, Tokai Y, Yoshimizu S, et al. Randomized controlled trial of an artificial intelligence diagnostic system for the detection of esophageal squamous cell carcinoma in clinical practice. Endoscopy. 2025;57:210–7. [ DOI ] [ PubMed ] [ Google Scholar ] 131. Ortiz O, Daca-Alvarez M, Rivero-Sanchez L, Gimeno-Garcia AZ, Carrillo-Palau M, Alvarez V, et al. An artificial intelligence-assisted system versus white light endoscopy alone for adenoma detection in individuals with Lynch syndrome (TIMELY): an international, multicentre, randomised controlled trial. Lancet Gastroenterol Hepatol. 2024;9:802–10. [ DOI ] [ PubMed ] [ Google Scholar ] 132. Li H, Liu D, Zeng Y, Liu S, Gan T, Rao N, et al. Single-image-based deep learning for segmentation of early Esophageal cancer lesions. IEEE Trans Image Process. 2024;33:2676–88. [ DOI ] [ PubMed ] [ Google Scholar ] 133. Du X, Xu X, Chen J, Zhang X, Li L, Liu H, et al. UM-net: rethinking ICGNet for polyp segmentation with uncertainty modeling. Med Image Anal. 2025;99:103347. [ DOI ] [ PubMed ] [ Google Scholar ] 134. Wang J, Li Y, Chen B, Cheng D, Liao F, Tan T, et al. A real-time deep learning-based system for colorectal polyp size estimation by white-light endoscopy: development and multicenter prospective validation. Endoscopy. 2024;56:260–70. [ DOI ] [ PubMed ] [ Google Scholar ] 135. Li S-w, Zhang L-h, Cai Y, Zhou X-b, Fu X-y, Song Y-q, et al. Deep learning assists detection of esophageal cancer and precursor lesions in a prospective, randomized controlled study. Sci Transl Med. 2024;16:eadk5395. [ DOI ] [ PubMed ] [ Google Scholar ] 136. Gong EJ, Bang CS, Lee JJ, Baik GH, Lim H, Jeong JH, et al. Deep learning-based clinical decision support system for gastric neoplasms in real-time endoscopy: development and validation study. Endoscopy. 2023;55:701–8. [ DOI ] [ PubMed ] [ Google Scholar ] 137. Wang H, Wang K-N, Hua J, Tang Y, Chen Y, Zhou G-Q, et al. Dynamic spectrum-driven hierarchical learning network for polyp segmentation. Med Image Anal. 2025;101:103449. [ DOI ] [ PubMed ] [ Google Scholar ] 138. Gao Y, Xin L, Lin H, Yao B, Zhang T, Zhou A-J, et al. Machine learning-based automated sponge cytology for screening of oesophageal squamous cell carcinoma and adenocarcinoma of the oesophagogastric junction: a nationwide, multicohort, prospective study. Lancet Gastroenterol Hepatol. 2023;8:432–45. [ DOI ] [ PubMed ] [ Google Scholar ] 139. Yu J, Zhu Y, Fu P, Chen T, Huang J, Li Q, et al. Robust Polyp Detection and Diagnosis through Compositional Prompt-Guided Diffusion Models. IEEE Trans Med Imaging. 2025;44(12):5245–57. 10.1109/TMI.2025.3589456. [ DOI ] [ PubMed ] 140. Gouda MA, Janku F, Wahida A, Buschhorn L, Schneeweiss A, Abdel Karim N, et al. Liquid biopsy response evaluation criteria in solid tumors (LB-RECIST). Ann Oncol. 2024;35:267–75. [ DOI ] [ PubMed ] [ Google Scholar ] 141. Li C, Wang Z, Ding P, Zhou Z, Chen R, Hu Y, et al. A digital score based on circulating-tumor-cells-derived mRNA quantification and machine learning for early colorectal cancer detection. ACS Nano. 2025;19:18117–28. [ DOI ] [ PubMed ] [ Google Scholar ] 142. Cheng S, Luo Y, Dong X, Liu M-y, Wu Z, Xu L, et al. Advanced ensemble staking model employing cfDNA fragmentation for early detection of esophageal and gastric cancer. Cancer Lett. 2025;631:217945. [ DOI ] [ PubMed ] [ Google Scholar ] 143. Cao Y, Wang N, Wu X, Tang W, Bao H, Si C, et al. Multidimensional fragmentomics enables early and accurate detection of colorectal cancer. Cancer Res. 2024;84:3286–95. [ DOI ] [ PubMed ] [ Google Scholar ] 144. Michel M, Heidary M, Mechri A, Da Silva K, Gorse M, Dixon V, et al. Noninvasive Multicancer Detection Using DNA Hypomethylation of LINE-1 Retrotransposons. Clin Cancer Res. 2025;31:1275–91. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 145. Jiao Z, Zhang X, Xuan Y, Shi X, Zhang Z, Yu A, et al. Leveraging cfDNA fragmentomic features in a stacked ensemble model for early detection of esophageal squamous cell carcinoma. Cell Rep Med. 2024;5:101664. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 146. Momen-Roknabadi A, Karimzadeh M, Chen N-C, Cavazos TB, Wang J, Ku J, et al. Detection of Early-Stage Colorectal Cancer Using Cell-Free oncRNA Biomarkers and Artificial Intelligence. Clin Cancer Res. 2025;31:3229–38. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 147. Yu D, Li Y, Wang M, Gu J, Xu W, Cai H, et al. Exosomes as a new frontier of cancer liquid biopsy. Mol Cancer. 2022;21:56. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 148. Zhang Z, Liu X, Peng C, Du R, Hong X, Xu J, et al. Machine learning-aided identification of fecal extracellular vesicle microRNA signatures for noninvasive detection of colorectal cancer. ACS Nano. 2025;19:10013–25. [ DOI ] [ PubMed ] [ Google Scholar ] 149. Li P, Chen J, Chen Y, Song S, Huang X, Yang Y, et al. Construction of exosome SORL1 detection platform based on 3d porous microfluidic chip and its application in early diagnosis of colorectal cancer. Small. 2023;19:2207381. [ DOI ] [ PubMed ] [ Google Scholar ] 150. Yin H, Xie J, Xing S, Lu X, Yu Y, Ren Y, et al. Machine learning-based analysis identifies and validates serum exosomal proteomic signatures for the diagnosis of colorectal cancer. Cell Rep Med. 2024;5:101689. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 151. Guo Y, Luo S, Liu S, Yang C, Lv W, Liang Y, et al. Bimodal in situ analyzer for circular RNA in extracellular vesicles combined with machine learning for accurate gastric cancer detection. Adv Sci. 2025;12:2409202. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 152. Cai Z-R, Zheng Y-Q, Hu Y, Ma M-Y, Wu Y-J, Liu J, et al. Construction of exosome non-coding RNA feature for non-invasive, early detection of gastric cancer patients by machine learning: a multi-cohort study. Gut. 2025;74:884. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 153. Shin H, Choi BH, Shim O, Kim J, Park Y, Cho SK, et al. Single test-based diagnosis of multiple cancer types using exosome-SERS-AI for early stage cancers. Nat Commun. 2023;14:1644. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 154. Bu F, Shen X, Zhan H, Wang D, Min L, Song Y, et al. Efficient metabolomics profiling from plasma extracellular vesicles enables accurate diagnosis of early gastric cancer. J Am Chem Soc. 2025;147:8672–86. [ DOI ] [ PubMed ] [ Google Scholar ] 155. Wang F, Wang C, Chen S, Wei C, Ji J, Liu Y, et al. Identification of blood-derived exosomal tumor RNA signatures as noninvasive diagnostic biomarkers for multi-cancer: a multi-phase, multi-center study. Mol Cancer. 2025;24:60. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 156. Koo B, Kim YI, Lee M, Lim S-B, Shin Y. Enhanced early detection of colorectal cancer via blood biomarker combinations identified through extracellular vesicle isolation and artificial intelligence analysis. J Extracell Vesicles. 2025;14:e70088. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 157. Iwamoto Y, Salmon B, Yoshioka Y, Kojima R, Krull A, Ota S. High throughput analysis of rare nanoparticles with deep-enhanced sensitivity via unsupervised denoising. Nat Commun. 2025;16:1728. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 158. Jo K, Linh VTN, Yang J-Y, Heo B, Kim JY, Mun NE, et al. Machine learning-assisted label-free colorectal cancer diagnosis using plasmonic needle-endoscopy system. Biosens Bioelectron. 2024;264:116633. [ DOI ] [ PubMed ] [ Google Scholar ] 159. Liu Y, Ji Y, Chen J, Zhang Y, Li X, Li X. Pioneering noninvasive colorectal cancer detection with an AI-enhanced breath volatilomics platform. Theranostics. 2024;14:4240–55. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 160. Liu S, Hu N, Hu J, Li W, Wu S, Ma X, et al. DNA-mediated bioinspired MXene gas sensor array with machine learning for noninvasive cancer recognition. ACS Nano. 2025;19:25363–84. [ DOI ] [ PubMed ] [ Google Scholar ] 161. Chen Y, Wang B, Zhao Y, Shao X, Wang M, Ma F, et al. Metabolomic machine learning predictor for diagnosis and prognosis of gastric cancer. Nat Commun. 2024;15:1657. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 162. Du L, Li S, Xiao X, Li J, Sun Y, Ji S, et al. Metabolomic profiling of plasma reveals potential biomarkers for screening and early diagnosis of gastric cancer and precancerous stages. MedComm – Oncology. 2023;2:e32. [ Google Scholar ] 163. Novielli P, Baldi S, Romano D, Magarelli M, Diacono D, Di Bitonto P, et al. Personalized colorectal cancer risk assessment through explainable AI and gut microbiome profiling. Gut Microbes. 2025;17:2543124. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 164. Ding Pa, Yang J, Guo H, Wu J, Wu H, Li T, et al. Multimodal artificial intelligence-based virtual biopsy for diagnosing abdominal lavage cytology-positive gastric cancer. Adv Sci. 2025;12:2411490. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 165. Bilal M, Tsang YW, Ali M, Graham S, Hero E, Wahab N, et al. Development and validation of artificial intelligence-based prescreening of large-bowel biopsies taken in the UK and Portugal: a retrospective cohort study. Lancet Digit Health. 2023;5:e786–97. [ DOI ] [ PubMed ] [ Google Scholar ] 166. Saillard C, Dubois R, Tchita O, Loiseau N, Garcia T, Adriansen A, et al. Validation of MSIntuit as an AI-based pre-screening tool for MSI detection from colorectal cancer histology slides. Nat Commun. 2023;14:6695. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 167. Tolkach Y, Wolgast LM, Damanakis A, Pryalukhin A, Schallenberg S, Hulla W, et al. Artificial intelligence for tumour tissue detection and histological regression grading in oesophageal adenocarcinomas: a retrospective algorithm development and validation study. Lancet Digit Health. 2023;5:e265–75. [ DOI ] [ PubMed ] [ Google Scholar ] 168. Jiang G, Wang Z, Cheng Z, Wang W, Lu S, Zhang Z, et al. The integrated molecular and histological analysis defines subtypes of esophageal squamous cell carcinoma. Nat Commun. 2024;15:8988. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 169. Wang C-W, Liu T-C, Lai P-J, Muzakky H, Wang Y-C, Yu M-H, et al. Ensemble transformer-based multiple instance learning to predict pathological subtypes and tumor mutational burden from histopathological whole slide images of endometrial and colorectal cancer. Med Image Anal. 2025;99:103372. [ DOI ] [ PubMed ] [ Google Scholar ] 170. Chang X, Wang J, Zhang G, Yang M, Xi Y, Xi C, et al. Predicting colorectal cancer microsatellite instability with a self-attention-enabled convolutional neural network. Cell Rep Med. 2023;4:100914. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 171. Gustav M, van Treeck M, Reitsam NG, Carrero ZI, Loeffler CML, Rabasco Meneghetti A, et al. Assessing genotype-phenotype correlations in colorectal cancer with deep learning: a multicentre cohort study. The Lancet Digital Health. 2025;7(8):100891. [ DOI ] [ PMC free article ] [ PubMed ] 172. Niehues JM, Quirke P, West NP, Grabsch HI, van Treeck M, Schirris Y, et al. Generalizable biomarker prediction from cancer pathology slides with self-supervised deep learning: a retrospective multi-centric study. Cell Rep Med. 2023;4:100980. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 173. Jin D, Liang S, Shmatko A, Arnold A, Horst D, Grünewald TGP, et al. Teacher-student collaborated multiple instance learning for pan-cancer PDL1 expression prediction from histopathology slides. Nat Commun. 2024;15:3063. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 174. Jiang X, Hoffmeister M, Brenner H, Muti HS, Yuan T, Foersch S, et al. End-to-end prognostication in colorectal cancer by deep learning: a retrospective, multicentre study. The Lancet Digital Health. 2024;6:e33–43. [ DOI ] [ PubMed ] [ Google Scholar ] 175. Claudio Quiros A, Coudray N, Yeaton A, Yang X, Liu B, Le H, et al. Mapping the landscape of histomorphological cancer phenotypes using self-supervised learning on unannotated pathology slides. Nat Commun. 2024;15:4596. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 176. El Nahhas OSM, Loeffler CML, Carrero ZI, van Treeck M, Kolbinger FR, Hewitt KJ, et al. Regression-based deep-learning predicts molecular biomarkers from pathology slides. Nat Commun. 2024;15:1253. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 177. Hoang D-T, Dinstag G, Shulman ED, Hermida LC, Ben-Zvi DS, Elis E, et al. A deep-learning framework to predict cancer treatment response from histopathology images through imputed transcriptomics. Nat Cancer. 2024;5:1305–17. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 178. Liu B, Polack M, Coudray N, Claudio Quiros A, Sakellaropoulos T, Le H, et al. Self-supervised learning reveals clinically relevant histomorphological patterns for therapeutic strategies in colon cancer. Nat Commun. 2025;16:2328. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 179. Wagner SJ, Reisenbüchler D, West NP, Niehues JM, Zhu J, Foersch S, et al. Transformer-based biomarker prediction from colorectal cancer histology: A large-scale multicentric study. Cancer Cell. 2023;41:1650-61.e4. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 180. Elforaici MEA, Montagnon E, Romero FP, Le WT, Azzi F, Trudel D, et al. Semi-supervised ViT knowledge distillation network with style transfer normalization for colorectal liver metastases survival prediction. Med Image Anal. 2025;99:103346. [ DOI ] [ PubMed ] [ Google Scholar ] 181. Huang J, Wang J, Shi J, Ni H, Xu S, Wu P, et al. Transformer optimization with meta learning on pathology images for breast cancer lymph node micrometastasis. NPJ Digit Med. 2025;8:421. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 182. Lin R, Chen Y, Li Y, Tan Y, Wang C, Wang Z, et al. Artificial Intelligence-powered copilots for precision diagnosis and surgical assessment of histological growth patterns in resectable colorectal liver metastases: a prospective study. International Journal of Surgery. 2025;111(11):7939–55. [ DOI ] [ PMC free article ] [ PubMed ] 183. Lecuelle J, Truntzer C, Basile D, Laghi L, Greco L, Ilie A, et al. Machine learning evaluation of immune infiltrate through digital tumour score allows prediction of survival outcome in a pooled analysis of three international stage III colon cancer cohorts. eBioMedicine. 2024;105:105207. [ DOI ] [ PMC free article ] [ PubMed ] 184. Foersch S, Glasner C, Woerl A-C, Eckstein M, Wagner D-C, Schulz S, et al. Multistain deep learning for prediction of prognosis and therapy response in colorectal cancer. Nat Med. 2023;29:430–9. [ DOI ] [ PubMed ] [ Google Scholar ] 185. Nowak M, Jabbar F, Rodewald A-K, Gneo L, Tomasevic T, Harkin A, et al. Single-cell AI-based detection and prognostic and predictive value of DNA mismatch repair deficiency in colorectal cancer. Cell Rep Med. 2024;5:101727. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 186. Bao X, Li Q, Chen D, Dai X, Liu C, Tian W, et al. A multiomics analysis-assisted deep learning model identifies a macrophage-oriented module as a potential therapeutic target in colorectal cancer. Cell Rep Med. 2024;5:101399. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 187. Wu E, Bieniosek M, Wu Z, Thakkar N, Charville GW, Makky A, et al. ROSIE: AI generation of multiplex immunofluorescence staining from histopathology images. Nat Commun. 2025;16:7633. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 188. Zhou Z, Jiang Y, Sun Z, Zhang T, Feng W, Li G, et al. Virtual multiplexed immunofluorescence staining from non-antibody-stained fluorescence imaging for gastric cancer prognosis. eBioMedicine. 2024;107:105287. [ DOI ] [ PMC free article ] [ PubMed ] 189. Sang S, Sun Z, Zheng W, Wang W, Islam MT, Chen Y, et al. TME-guided deep learning predicts chemotherapy and immunotherapy response in gastric cancer with attention-enhanced residual Swin transformer. Cell Rep Med. 2025;6:102242. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 190. Hu C, Chen W, Li F, Zhang Y, Yu P, Yang L, et al. Deep learning radio-clinical signatures for predicting neoadjuvant chemotherapy response and prognosis from pretreatment CT images of locally advanced gastric cancer patients. International Journal of Surgery. 2023;109(7):1980–92. [ DOI ] [ PMC free article ] [ PubMed ] 191. Cui Y, Zhao K, Meng X, Mao Y, Han C, Shi Z, et al. A computed tomography-based multitask deep learning model for predicting tumour stroma ratio and treatment outcomes in patients with colorectal cancer: a multicentre cohort study. International Journal of Surgery. 2024;110(5):2845–54. [ DOI ] [ PMC free article ] [ PubMed ] 192. Qiu B, Zheng Y, Liu S, Song R, Wu L, Lu C, et al. Multitask deep learning based on longitudinal CT images facilitates prediction of lymph node metastasis and survival in chemotherapy-treated gastric cancer. Cancer Res. 2025;85:2527–36. [ DOI ] [ PubMed ] [ Google Scholar ] 193. Jiang Y, Zhang Z, Wang W, Huang W, Chen C, Xi S, et al. Biology-guided deep learning predicts prognosis and cancer immunotherapy response. Nat Commun. 2023;14:5135. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 194. Jiang Y, Zhou K, Sun Z, Wang H, Xie J, Zhang T, et al. Non-invasive tumor microenvironment evaluation and treatment response prediction in gastric cancer using deep learning radiomics. Cell Rep Med. 2023;4:101146. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 195. Sun Z, Wang W, Huang W, Zhang T, Chen C, Yuan Q, et al. Noninvasive imaging evaluation of peritoneal recurrence and chemotherapy benefit in gastric cancer after gastrectomy: a multicenter study. International Journal of Surgery. 2023;109(7):2010–24. [ DOI ] [ PMC free article ] [ PubMed ] 196. Guo X, Chen M, Zhou L, Zhu L, Liu S, Zheng L, et al. Predicting early recurrence in locally advanced gastric cancer after gastrectomy using CT-based deep learning model: a multicenter study. International Journal of Surgery. 2025;111(2):2089–2100. [ DOI ] [ PubMed ] 197. Xie C, Ning Z, Guo T, Yao L, Chen X, Huang W, et al. Multimodal data integration for biologically-relevant artificial intelligence to guide adjuvant chemotherapy in stage II colorectal cancer. eBioMedicine. 2025;117:105789. [ DOI ] [ PMC free article ] [ PubMed ] 198. Zhen Z, Tianchen L, Meng Y, Haixia S, Kaiyi T, Jian Z, et al. Voxel-level radiomics and deep learning for predicting pathologic complete response in esophageal squamous cell carcinoma after neoadjuvant immunotherapy and chemotherapy. J ImmunoTher Cancer. 2025;13:e011149. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 199. Zhu X, Sun H, Wang Y, Hu G, Shao L, Zhang S, et al. Prediction of lymph node metastasis in colorectal cancer using intraoperative fluorescence multi-modal imaging. IEEE Trans Med Imaging. 2025;44:1568–80. [ DOI ] [ PubMed ] [ Google Scholar ] 200. Li D, Liu X, Gao W, Zhao W, Ji S, Wang S, et al. An immune subtype classification system enables the development of strategies to predict and enhance immunotherapy responses in colorectal cancer. Cancer Res. 2025;85:1441–58. [ DOI ] [ PubMed ] [ Google Scholar ] 201. Bahrambeigi V, Lee JJ, Branchi V, Rajapakshe KI, Xu Z, Kui N, et al. Transcriptomic Profiling of Plasma Extracellular Vesicles Enables Reliable Annotation of the Cancer-Specific Transcriptome and Molecular Subtype. Can Res. 2024;84:1719–32. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 202. Domingo E, Rathee S, Blake A, Samuel L, Murray G, Sebag-Montefiore D, et al. Identification and validation of a machine learning model of complete response to radiation in rectal cancer reveals immune infiltrate and TGFβ as key predictors. eBioMedicine. 2024;106:105228. [ DOI ] [ PMC free article ] [ PubMed ] 203. Thrift WJ, Lounsbury NW, Broadwell Q, Heidersbach A, Freund E, Abdolazimi Y, et al. Towards designing improved cancer immunotherapy targets with a peptide-MHC-I presentation model, hlapollo. Nat Commun. 2024;15:10752. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 204. Zhu G, Rahman CR, Getty V, Odinokov D, Baruah P, Carrié H, et al. A deep-learning model for quantifying circulating tumour DNA from the density distribution of DNA-fragment lengths. Nat Biomed Eng. 2025;9:307–19. [ DOI ] [ PubMed ] [ Google Scholar ] 205. Sanjaya P, Maljanen K, Katainen R, Waszak SM, Ambrose JC, Arumugam P, et al. Mutation-Attention (MuAt): deep representation learning of somatic mutations for tumour typing and subtyping. Genome Med. 2023;15:47. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 206. Wang S, Wu C-Y, He M-M, Yong J-X, Chen Y-X, Qian L-M, et al. Machine learning-based extrachromosomal DNA identification in large-scale cohorts reveals its clinical implications in cancer. Nat Commun. 2024;15(1):1515. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 207. Widman AJ, Shah M, Frydendahl A, Halmos D, Khamnei CC, Øgaard N, et al. Ultrasensitive plasma-based monitoring of tumor burden using machine-learning-guided signal enrichment. Nat Med. 2024;30:1655–66. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 208. Cai Z, Boys EL, Noor Z, Aref AT, Xavier D, Lucas N, et al. Federated Deep Learning Enables Cancer Subtyping by Proteomics. Cancer Discovery. 2025;15(9):1803–18. [ DOI ] [ PMC free article ] [ PubMed ] 209. Liu Y, Shi J, Liu W, Tang Y, Shu X, Wang R, et al. A deep neural network predictor to predict the sensitivity of neoadjuvant chemoradiotherapy in locally advanced rectal cancer. Cancer Lett. 2024;589:216641. [ DOI ] [ PubMed ] [ Google Scholar ] 210. Wang Y, Yue X, Lou S, Feng P, Cui B, Liu Y. Gene swin transformer: new deep learning method for colorectal cancer prognosis using transcriptomic data. Brief Bioinform. 2025;26:bbaf275. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 211. Zeng D, Yu Y, Qiu W, Ou Q, Mao Q, Jiang L, et al. Immunotyping the tumor microenvironment reveals molecular heterogeneity for personalized immunotherapy in cancer. Adv Sci. 2025;12:2417593. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 212. Ding X, Zhang L, Fan M, Li L. TME-NET: an interpretable deep neural network for predicting pan-cancer immune checkpoint inhibitor responses. Brief Bioinform. 2024;25:bbae410. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 213. Zhong Z, Hou J, Yao Z, Dong L, Liu F, Yue J, et al. Domain generalization enables general cancer cell annotation in single-cell and spatial transcriptomics. Nat Commun. 2024;15:1929. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 214. Wang ZJ, Farooq AS, Chen Y-J, Bhargava A, Xu AM, Thomson MW. Identifying perturbations that boost T-cell infiltration into tumours via counterfactual learning of their spatial proteomic profiles. Nat Biomed Eng. 2025;9:390–404. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 215. Yu C, Bian Y, Gao Y, Jiao Y, Xu Y, Wang W, et al. Machine learning-based lactate-related genes signature predicts clinical outcomes and unveils novel therapeutic targets in esophageal squamous cell carcinoma. Cancer Lett. 2025;613:217458. [ DOI ] [ PubMed ] [ Google Scholar ] 216. Xu Z, Huang Y, Hu C, Du L, Du Y-A, Zhang Y, et al. Efficient plasma metabolic fingerprinting as a novel tool for diagnosis and prognosis of gastric cancer: a large-scale, multicentre study. Gut. 2023;72:2051. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 217. Huang X, Wang Q, Xu W, Liu F, Pan L, Jiao H, et al. Machine learning to predict lymph node metastasis in T1 esophageal squamous cell carcinoma: a multicenter study. International Journal of Surgery. 2024;110(12):7852–59. [ DOI ] [ PMC free article ] [ PubMed ] 218. Chen Q, Chen J, Deng Y, Bi X, Zhao J, Zhou J, et al. Personalized prediction of postoperative complication and survival among colorectal liver metastases patients receiving simultaneous resection using machine learning approaches: a multi-center study. Cancer Lett. 2024;593:216967. [ DOI ] [ PubMed ] [ Google Scholar ] 219. Guo J, Zhang K, Ji G, Wang W, Li G, Liu Z, et al. Artificial neural network model enhancing the accuracy of clinical evaluation for high-risk population of lymph node metastasis in non-intestinal type early gastric cancer: a multicenter real-world study in China. International Journal of Surgery. 2025;111(6):4068–73. [ DOI ] [ PMC free article ] [ PubMed ] 220. Bertsimas D, Margonis GA, Sujichantararat S, Koulouras A, Ma Y, Antonescu CR, et al. Interpretable artificial intelligence to optimise use of imatinib after resection in patients with localised gastrointestinal stromal tumours: an observational cohort study. Lancet Oncol. 2024;25:1025–37. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 221. Huang W, Wang X, Zhong R, Li Z, Zhou K, Lyu Q, et al. Multimodal radiopathomics signature for prediction of response to immunotherapy-based combination therapy in gastric cancer using interpretable machine learning. Cancer Lett. 2025;631:217930. [ DOI ] [ PubMed ] [ Google Scholar ] 222. Chen Z, Chen Y, Sun Y, Tang L, Zhang L, Hu Y, et al. Predicting gastric cancer response to anti-HER2 therapy or anti-HER2 combined immunotherapy based on multi-modal data. Signal Transduct Target Ther. 2024;9:222. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 223. Chia S, Wen Seow JJ, Peres da Silva R, Suphavilai C, Shirgaonkar N, Murata-Hori M, et al. CAN-Scan: A multi-omic phenotype-driven precision oncology platform identifies prognostic biomarkers of therapy response for colorectal cancer. Cell Rep Med. 2025;6(4):102053. [ DOI ] [ PMC free article ] [ PubMed ] 224. Wang F-a, Zhuang Z, Gao F, He R, Zhang S, Wang L, et al. TMO-net: an explainable pretrained multi-omics model for multi-task learning in oncology. Genome Biol. 2024;25:149. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 225. Yan Q, Li X, Cui J, Rong J, Zhang J, Gao P, et al. Spatial histology and gene-expression representation and generative learning via online self-distillation contrastive learning. Brief Bioinform. 2025;26:bbaf317. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 226. Zhou J, Foroughi pour A, Deirawan H, Daaboul F, Aung TN, Beydoun R, et al. Integrative deep learning analysis improves colon adenocarcinoma patient stratification at risk for mortality. eBioMedicine. 2023;94:104726. [ DOI ] [ PMC free article ] [ PubMed ] 227. Luo B, Teng F, Tang G, Cen W, Liu X, Chen J, et al. Stereomm: a graph fusion model for integrating spatial transcriptomic data and pathological images. Brief Bioinform. 2025;26:bbaf210. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 228. Zamanitajeddin N, Jahanifar M, Bilal M, Eastwood M, Rajpoot N. Social network analysis of cell networks improves deep learning for prediction of molecular pathways and key mutations in colorectal cancer. Med Image Anal. 2024;93:103071. [ DOI ] [ PubMed ] [ Google Scholar ] 229. Katiyar P, Schwenck J, Frauenfeld L, Divine MR, Agrawal V, Kohlhofer U, et al. Quantification of intratumoural heterogeneity in mice and patients via machine-learning models trained on PET–MRI data. Nat Biomed Eng. 2023;7:1014–27. [ DOI ] [ PubMed ] [ Google Scholar ] 230. Zhou S, Sun D, Mao W, Liu Y, Cen W, Ye L, et al. Deep radiomics-based fusion model for prediction of bevacizumab treatment response and outcome in patients with colorectal cancer liver metastases: a multicentre cohort study. eClinicalMedicine. 2023;65:102271. [ DOI ] [ PMC free article ] [ PubMed ] 231. Zuo C, Xia J, Xu Y, Xu Y, Gao P, Zhang J, et al. Stclinic dissects clinically relevant niches by integrating spatial multi-slice multi-omics data in dynamic graphs. Nat Commun. 2025;16:5317. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 232. Shi Z, Lei JT, Elizarraras JM, Zhang B. Mapping the functional network of human cancer through machine learning and pan-cancer proteogenomics. Nat Cancer. 2025;6:205–22. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 233. Cao Y, Wang D, Wu J, Yao Z, Shen S, Niu C, et al. MSI-XGNN: an explainable GNN computational framework integrating transcription- and methylation-level biomarkers for microsatellite instability detection. Brief Bioinform. 2023;24:bbad362. [ DOI ] [ PubMed ] [ Google Scholar ] 234. Che G, Yin J, Wang W, Luo Y, Chen Y, Yu X, et al. Circumventing drug resistance in gastric cancer: a spatial multi-omics exploration of chemo and immuno-therapeutic response dynamics. Drug Resist Updat. 2024;74:101080. [ DOI ] [ PubMed ] [ Google Scholar ] 235. Zhang P, Gao C, Zhang Z, Yuan Z, Zhang Q, Zhang P, et al. Systematic inference of super-resolution cell spatial profiles from histology images. Nat Commun. 2025;16:1838. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 236. Gao F, Ding J, Gai B, Cai D, Hu C, Wang F-A, et al. Interpretable Multimodal Fusion Model for Bridged Histology and Genomics Survival Prediction in Pan-Cancer. Adv Sci. 2025;12:2407060. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 237. Hu Y, Li X, Yi Y, Huang Y, Wang G, Wang D. Deep learning-driven survival prediction in pan-cancer studies by integrating multimodal histology-genomic data. Brief Bioinform. 2025;26:bbaf121. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 238. Decourtye-Espiard L, Guilford P. Hereditary diffuse gastric cancer. Gastroenterology. 2023;164:719–35. [ DOI ] [ PubMed ] [ Google Scholar ] 239. Jiang W, Zhang B, Xu J, Xue L, Wang L. Current status and perspectives of esophageal cancer: a comprehensive review. Cancer Commun. 2025;45:281–331. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 240. Wang L, Wan L, Chen X, Gao P, Hou Y, Wu L, et al. Reduced intestinal-to-diffuse conversion and immunosuppressive responses underlie superiority of neoadjuvant immunochemotherapy in gastric adenocarcinoma. MedComm. 2024;5:e762. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 241. Wang L, Sun M, Li J, Wan L, Tan Y, Tian S, et al. Intestinal Subtype as a Biomarker of Response to Neoadjuvant Immunochemotherapy in Locally Advanced Gastric Adenocarcinoma: Insights from a Prospective Phase II Trial. Clin Cancer Res. 2025;31:74–86. [ DOI ] [ PubMed ] [ Google Scholar ] 242. Wang Y, Jia J, Wang F, Fang Y, Yang Y, Zhou Q, et al. Pre-metastatic niche: formation, characteristics and therapeutic implication. Signal Transduct Target Ther. 2024;9:236. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 243. Chang PW, Amini MM, Davis RO, Nguyen DD, Dodge JL, Lee H, et al. ChatGPT4 Outperforms Endoscopists for Determination of Postcolonoscopy Rescreening and Surveillance Recommendations. Clin Gastroenterol Hepatol. 2024;22:1917-25.e17. [ DOI ] [ PubMed ] [ Google Scholar ] 244. Maida M, Ramai D, Mori Y, Dinis-Ribeiro M, Facciorusso A, Hassan C. The role of generative language systems in increasing patient awareness of colon cancer screening. Endoscopy. 2025;57:262–8. [ DOI ] [ PubMed ] [ Google Scholar ] 245. Pan A, Musheyev D, Bockelman D, Loeb S, Kabarriti AE. Assessment of artificial intelligence chatbot responses to top searched queries about cancer. JAMA Oncol. 2023;9:1437–40. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 246. Huang Z, Yang E, Shen J, Gratzinger D, Eyerer F, Liang B, et al. A pathologist–AI collaboration framework for enhancing diagnostic accuracies and efficiencies. Nat Biomed Eng. 2025;9:455–70. [ DOI ] [ PubMed ] [ Google Scholar ] 247. Jee J, Fong C, Pichotta K, Tran TN, Luthra A, Waters M, et al. Automated real-world data integration improves cancer outcome prediction. Nature. 2024;636:728–36. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 248. Feng S, Li S, Chen L, Chen S. Unveiling potential threats: backdoor attacks in single-cell pre-trained models. Cell Discov. 2024;10:122. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 249. Dayan I, Roth HR, Zhong A, Harouni A, Gentili A, Abidin AZ, et al. Federated learning for predicting clinical outcomes in patients with COVID-19. Nat Med. 2021;27:1735–43. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 250. Mahmood T, Saba T, Rehman A, Alamri FS. Harnessing the power of radiomics and deep learning for improved breast cancer diagnosis with multiparametric breast mammography. Expert Syst Appl. 2024;249:123747. [ Google Scholar ] 251. Higgins K, Nyssen OP, Southern J, Laponogov I, Marco AM, Cabeza-Segura M, et al. The Helicobacter pylori AI-clinician harnesses artificial intelligence to personalise H. pylori treatment recommendations. Nat Commun. 2025;16:6472. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 252. Ferber D, El Nahhas OSM, Wölflein G, Wiest IC, Clusmann J, Leßmann M-E, et al. Development and validation of an autonomous artificial intelligence agent for clinical decision-making in oncology. Nat Cancer. 2025;6:1337–49. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Data Availability Statement No datasets were generated or analysed during the current study. Articles from Molecular Cancer are provided here courtesy of BMC ACTIONS View on publisher site PDF (5.0 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top

Record · ID 1408 · SHA-256 75e2b51609d35950
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.