ConceptioArchiveNCBI PubMed Central
NCBI PubMed Centralopen access

SpaHE-Infil: A spatial heterogeneity framework for decoding TME infiltration from H&E-stained slides.

Wang F et al. · ncbi_pmc
NCBI PubMed Central · Papers · License: Open Access
Open Source ↗Direct PDF ↓
machine learning systems

Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice iScience . 2026 Feb 26;29(4):115155. doi: 10.1016/j.isci.2026.115155 Search in PMC Search in PubMed View in NLM Catalog Add to search SpaHE-Infil: A spatial heterogeneity framework for decoding TME infiltration from H&E-stained slides Fang Wang Fang Wang 1 Department of Cancer Diagnosis and Treatment Center, Affiliated Hospital of Jiangnan University, Wuxi, China 2 School of Public Health, Zhejiang University School of Medicine, Hangzhou, China Find articles by Fang Wang 1, 2, 6, ∗ , Hejia Xu Hejia Xu 1 Department of Cancer Diagnosis and Treatment Center, Affiliated Hospital of Jiangnan University, Wuxi, China 3 Wuxi Medical College of Jiangnan University, Wuxi, China Find articles by Hejia Xu 1, 3 , Xue Wang Xue Wang 3 Wuxi Medical College of Jiangnan University, Wuxi, China Find articles by Xue Wang 3 , Mengyin Wu Mengyin Wu 2 School of Public Health, Zhejiang University School of Medicine, Hangzhou, China 4 Division of Noncommunicable Diseases and Injury, Shanghai Municipal Center for Disease Control and Prevention, Shanghai, China Find articles by Mengyin Wu 2, 4 , Sisi Ding Sisi Ding 3 Wuxi Medical College of Jiangnan University, Wuxi, China Find articles by Sisi Ding 3 , Lian Duan Lian Duan 1 Department of Cancer Diagnosis and Treatment Center, Affiliated Hospital of Jiangnan University, Wuxi, China 3 Wuxi Medical College of Jiangnan University, Wuxi, China Find articles by Lian Duan 1, 3 , Chen Yang Chen Yang 3 Wuxi Medical College of Jiangnan University, Wuxi, China Find articles by Chen Yang 3 , Jiaxuan Li Jiaxuan Li 1 Department of Cancer Diagnosis and Treatment Center, Affiliated Hospital of Jiangnan University, Wuxi, China 3 Wuxi Medical College of Jiangnan University, Wuxi, China Find articles by Jiaxuan Li 1, 3 , Xiaosong Ge Xiaosong Ge 1 Department of Cancer Diagnosis and Treatment Center, Affiliated Hospital of Jiangnan University, Wuxi, China 3 Wuxi Medical College of Jiangnan University, Wuxi, China Find articles by Xiaosong Ge 1, 3 , Yan Qin Yan Qin 5 Department of Pathology, Affiliated Hospital of Jiangnan University, Wuxi, China Find articles by Yan Qin 5 , Xiaowei Qi Xiaowei Qi 3 Wuxi Medical College of Jiangnan University, Wuxi, China 5 Department of Pathology, Affiliated Hospital of Jiangnan University, Wuxi, China Find articles by Xiaowei Qi 3, 5, ∗∗ , Yong Mao Yong Mao 1 Department of Cancer Diagnosis and Treatment Center, Affiliated Hospital of Jiangnan University, Wuxi, China 3 Wuxi Medical College of Jiangnan University, Wuxi, China Find articles by Yong Mao 1, 3, ∗∗∗ Author information Article notes Copyright and License information 1 Department of Cancer Diagnosis and Treatment Center, Affiliated Hospital of Jiangnan University, Wuxi, China 2 School of Public Health, Zhejiang University School of Medicine, Hangzhou, China 3 Wuxi Medical College of Jiangnan University, Wuxi, China 4 Division of Noncommunicable Diseases and Injury, Shanghai Municipal Center for Disease Control and Prevention, Shanghai, China 5 Department of Pathology, Affiliated Hospital of Jiangnan University, Wuxi, China ∗ Corresponding author [email protected] ∗∗ Corresponding author [email protected] ∗∗∗ Corresponding author [email protected] 6 Lead contact Received 2025 Aug 27; Revised 2026 Jan 29; Accepted 2026 Feb 23; Collection date 2026 Apr 17. © 2026 Published by Elsevier Inc. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). PMC Copyright notice PMCID: PMC13068657  PMID: 41971992 Summary The spatial distribution of immune cells in the tumor microenvironment (TME) is a key determinant of immunotherapy response, while current methods are limited by sequencing dependence and restricted spatial resolution. We developed SpaHE-Infil, a multimodal computational framework integrating spatial transcriptomics and whole-slide images to train a Random forest model, extracting morphological, textural, and density features to identify 12 core TME cell types in situ , with dynamic calibration correcting immune cell proportion biases in clinical samples. Cross-cancer validation confirmed its accurate spatial cell distribution prediction, consistent with mIHC and canonical deconvolution algorithms. In clinical cohorts, TME immune infiltration stratification via H&E images predicted enhanced immunotherapy response and prolonged recurrence-free survival across multiple cancers. This framework provides a clinically applicable tool for spatial TME characterization, supporting tumor immunology research, and precision immunotherapy practice. Subject areas: Medicine, Microenvironment, Bioinformatics, Cancer Graphical abstract Open in a new tab Highlights • SpaHE-Infil decodes spatial patterns of 12 TME cells from H&E slides via multimodal features • SpaHE-Infil’s TME calibration mechanism dynamically corrects immune cell proportion imbalance • SpaHE-Infil is validated against spatial transcriptomics and mIHC across pan-cancer cohorts • SpaHE-Infil predicts ICI response and survival outcomes in high immune-infiltration patients Medicine; Microenvironment; Bioinformatics; Cancer Introduction The spatial distribution characteristics of immune cells within the tumor microenvironment (TME) serve as crucial indicators for predicting immunotherapy response and patient prognosis. In recent years, the advent of spatial transcriptomics (ST) technology has enabled researchers to analyze the association between cellular heterogeneity and spatial architecture within intact tissue contexts. 1 , 2 However, the high cost and complex experimental procedures associated with this technology have limited its widespread clinical application. Currently, most immune deconvolution algorithms (such as TIMER, CIBERSORT, etc.) still rely on bulk RNA-seq data to infer cellular composition. While these methods can provide overall abundance information, they fail to capture the in situ spatial distribution features of cells within the tissue. 3 , 4 Concurrently, analysis models based on pathological images, although applied to histological classification and prognosis prediction, exhibit significant limitations in the precise identification of immune cell types. Traditional models primarily focus on the association between tissue morphological features and clinical outcomes, lacking the capability for fine-grained analysis of specific immune cells (such as CD8 + T cells, macrophages, etc.) within the TME. 5 , 6 More importantly, existing methods struggle to effectively integrate molecular-level cell type annotations with tissue-level spatial morphological features, leading to incomplete and mechanistically ambiguous characterization of the immune microenvironment. Although technologies like spatial transcriptomics, spatial proteomics, single-cell RNA-seq, and bulk RNA-seq combined with deconvolution algorithms provide multi-dimensional tools for TME analysis, their common limitation lies in relying on sequencing or high-throughput detection data as core inputs. 7 , 8 , 9 , 10 , 11 , 12 They fail to fully leverage the vast reservoir of routinely available pathological resources in clinical settings, such as conventional pathology slides and diagnostic tests like IHC and FISH. This reliance on high-throughput sequencing not only significantly increases experimental costs but also suffers from lengthy sample processing times, leading to insufficient timeliness. This makes it difficult to meet the immediate demands of clinical diagnosis and treatment, ultimately creating a substantial translational barrier between basic research and clinical application. Against this backdrop, we developed SpaHE-Infil : a novel tool based directly on pathological slide images for dissecting immune cell distribution within the TME. Compared to existing methods, the core breakthrough of SpaHE-Infil manifests in two key aspects: First, unlike immune deconvolution algorithms dependent on bulk RNA-seq, this method uses (Hematoxylin and Eosin, H&E)-stained sections as input. By capturing in situ features such as nuclear morphology and cellular arrangement density, it achieves simultaneous resolution of immune cell types and their spatial distribution. Second, it innovatively employs spatial transcriptomics data as the “anchoring standard”, utilizing machine learning to establish a mapping relationship between molecular annotations and pathological image features. 13 , 14 This enables the model to inherit the cell-type annotation precision of ST data while leveraging the spatial resolution advantages of pathological images. SpaHE-Infil can identify 12 key TME cell types and optimizes prediction accuracy for heterogeneous samples through its unique TME_adjust dynamic calibration mechanism. In validation across multiple cancer types, its predictions demonstrated high concordance with the in situ cellular distribution revealed by spatial transcriptomics (with significantly high Pearson correlation coefficients) and exhibited clinical utility value in predicting immune checkpoint inhibitor (ICI) treatment response and patient prognosis. Crucially, this method eliminates the need for sequencing data, enabling comprehensive panoramic analysis of the TME immune landscape solely through routine pathology slides, thereby providing an efficient technical pathway for translating basic research into clinical practice. Results SpaHE-Infil: A spatial heterogeneity reader for reconstructing TME infiltration from H&E-stained slides The workflow of SpaHE-Infil begins with the integration and analysis of spatial transcriptomic data. First, sequencing information was extracted from spatial transcriptomic datasets, and cell-type annotations were performed on corresponding tissue sections ( Figure 1 A). Subsequently, high-resolution images of these spatially resolved pathological sections were acquired to analyze markers within the tumor microenvironment (TME) ( Figure 1 B). The core predictive model was constructed using a machine learning approach with Random forest (RF)-based feature selection, which effectively pruned non-informative features and reinforced the discriminative power of morphological and spatial parameters ( Figure 1 C). Trained via cross-validation, this model accurately identifies morphological features of cells in H&E images and their corresponding TME cell types ( Figure 1 D). This trained model is exportable as RDS files for the R programming environment, and allows for continuous performance improvement by incorporating additional annotated datasets. The trained model was then applied to H&E slides of test samples for spatial distribution prediction. To optimize prediction accuracy, the system innovatively incorporates a dynamic TME proportion calibration module (optional), refining predicted cellular abundances ( Figure 1 E). Finally, SpaHE-Infil outputs a visual report with spatially resolved maps and performance evaluation metrics. Figure 1. Open in a new tab Overview of SpaHE-Infil (A) Cell-type annotation of tissue sections using spatial transcriptomic datasets. (B) High-resolution imaging and TME marker analysis of spatially resolved pathological sections. (C) Core predictive model construction via random forest-based feature selection. (D) Cross-validation training for identifying morphological features of cells in H&E images and corresponding TME cell types; dynamic TME proportion calibration refines predicted cellular abundances. (E) Output of spatially resolved maps and prediction performance metrics. Dynamic calibration-enabled TME spatial mapping from routine histology The SpaHE-Infil model integrates spatial transcriptomic data per sample, including sequencing-based cell-type annotations (with molecular markers) and paired pathological sections capturing nuclear morphology, tissue texture, and spatial density features. These parameters collectively serve as inputs for the training module ( Figure 2 A). Using TME marker expression levels, cell-type annotation can be performed via either AUCell automated labeling or manual curation ( Figure 2 B); in this study, AUCell was used as the representative method for the corresponding results. SpaHE-Infil currently supports spatial distribution analysis of 12 key TME cell types: tumor cells, normal epithelial cells, CD8 T cells, CD4 T cells, macrophages, neutrophils, NK cells, Tregs, fibroblasts, endothelial cells, B cells, and dendritic cells ( Figure 2 C). Spatial information was extracted by segmenting training set tissue sections to analyze nuclear morphology, cellular distribution patterns, and density features ( Figures 2 D–2F). For each tumor type (e.g., pancreatic and breast cancer), multiple spatially annotated sections were used for model training ( Figures 2 G, S2 A, and S2B). Validation of the prediction matrix confirmed reliable identification of tumor-dominant regions ( Figures 2 H, 2I, and S2 C–S2E; Table S7 ). To address tumor immune heterogeneity, the algorithm incorporates a TME_adjust module, allowing users to calibrate microenvironmental cell proportions via parameter tuning ( Figure 2 J). These steps constitute the complete SpaHE-Infil training architecture ( Figure 2 K). Figure 2. Open in a new tab Dynamic calibration-enabled TME spatial mapping from routine histology (A) Integration of ST data and pathological sections as training inputs. (B) Cell-type annotation via AUCell auto-labeling with manual correction. (C) Twelve key TME cell types supported by SpaHE-Infil. (D–F) Nuclear morphology, cellular distribution patterns, and density features extracted from segmented training sections. (G) Selection of ST-annotated sections (e.g., pancreatic cancer) for model training. (H–I) Validation of tumor-dominant region predictions using the prediction matrix. (J) TME_adjust module for microenvironmental cell proportion calibration. (K) Architecture of the SpaHE-Infil training module. Multi-cancer validation of TME mapping against spatial transcriptomics The prediction module of SpaHE-Infil applies pre-trained tumor-type models to H&E slides from new samples (same species and cancer type), measuring cellular features and using deep learning to classify TME cell types ( Figure 3 A). Testing involved H&E slides and spatial transcriptomic (ST) annotations from tumors not used in training. Upon input, the algorithm generated cell-type prediction maps and abundance percentages ( Figures 3 B–3D). Comparison between SpaHE-Infil-predicted and ST-derived cell type distributions was evaluated at the grid-spot level (each dot corresponds to a single spatial grid spot): the majority of grid spots showed high concordance and positive PCC, with negative PCC values exclusive to a small fraction of spots with extreme TME heterogeneity ( Figures 3 E and 3F). Principal component distance analysis further confirmed the overall spatial distribution consistency between predictions and ST annotations. Stratified analysis of qualified spots (with normal TME heterogeneity) validated reliable performance for major cell types including tumor cells, CD8 + T cells, and macrophages ( Figures 3 G–3I), with a median positive PCC of 0.68. Full testing across human breast, colon, gastric, ovarian, and pancreatic cancers, as well as murine lung and pancreatic tissues, demonstrated consistent PCC distributions ( Figures S2 A–S2G), confirming SpaHE-Infil’s accuracy in TME cell identification and abundance detection. Figure 3. Open in a new tab Multi-cancer validation of TME mapping against spatial transcriptomics (A) Prediction module classifying cell types in H&E slides from test samples (same species/cancer type). (B–D) Cell-type prediction maps and abundance percentages generated from input H&E slides. (E and F) Grid-level Pearson correlation (PCC) and principal component distance analysis comparing predicted vs. ST-derived distributions. (G–I) Validation results for tumor cells, CD8 + T cells, and macrophages. (J) Workflow of the SpaHE-Infil prediction module. Microscopy-optimized spatial reconstruction with mIHC benchmarking To validate and optimize SpaHE-Infil, experiments used colon cancer H&E sections ( Figure 4 A). Adjusting the immune cell calibration parameter (TME_ADJ) significantly improved detection sensitivity ( Figures 4 B, 4C, and S3 A–S3C), reproducible in pancreatic cancer tissues ( Figure 4 D). The algorithm successfully identified spatial signals of tumor and microenvironmental cells ( Figures 4 E and 4F). Multiplex immunohistochemistry (mIHC) on consecutive tissue sections validated strong consistency between the spatially predicted distributions of key cell types (tumor cells, macrophages, fibroblasts, CD8 + T cells, NK cells) and their corresponding protein expression patterns ( Figures 4 G–4I, S3 D, and S3E). Protein detection of EpCAM, CD68, vimentin, CD8, and CD56 in colon cancer further validated these results ( Figure S3 C). We further applied the SpaHE-Infil algorithm to analyze consecutive sections from ovarian cancer ( Figures 4 J and 4K) and gastric cancer ( Figures 4 L–4M) samples. The algorithm generated spatial distribution maps of tumor-infiltrating immune and stromal cells, which exhibited clear spatial correspondence with the histological features observed in H&E staining. Quantitative analysis via pie charts revealed that both tumor types were dominated by tumor cells (63.5% in ovarian cancer, 66.3% in gastric cancer), with fibroblasts being the most abundant stromal component (26.1% and 30.7%, respectively). Minor populations of immune cells, including macrophages, neutrophils, and T cells, were also consistently identified across both cancer types, demonstrating the robustness of SpaHE-Infil in profiling the TME across distinct malignancies. Notably, microscope field-of-view (FOV) impacted outcomes: 100× magnification (FOV≈0.785 mm 2 ) covered larger areas (more cells) at lower resolution, while 200× (FOV ≈0.196 mm 2 ) offered higher resolution with smaller coverage ( Figures S3 F–S3H). Optimal FOV should be adjusted based on tissue characteristics, with input resolution meeting algorithmic requirements. Figure 4. Open in a new tab Microscopy-optimized spatial reconstruction with mIHC benchmarking (A) Algorithm validation and optimization using colon cancer H&E sections. Scale bars, 100 μm. (B and C) Enhanced immune cell detection sensitivity after TME_ADJ calibration. (D) Reproducibility of sensitivity improvement in pancreatic cancer. Scale bars, 100 μm. (E and F) Spatial signals of tumor and microenvironmental cells identified by the algorithm. (G) Immunomarker expression detected by mIHC on consecutive sections. Scale bars, 100 μm. (H) mIHC signal distribution. Each dot represents a cell, colored by its mIHC marker (DAPI: gray; CD8: red; CD68: green; Vimentin: cyan; Vimentin: yellow; CD56: purple) and signal intensity (gradient: 2×10 4 to 10×10 4 ), showing spatial heterogeneity of marker expression across the tissue section. (I) Cell type proportion comparison. Bar chart showing relative proportions of NK cells, macrophages, CD8 + T cells, fibroblasts, and tumor cells, quantified by SpaHE-Infil and validated by mIHC, demonstrating consistency between computational inference and experimental staining. (J and K) Spatial signals in ovarian cancer H&E sections. Scale bars, 100 μm. (L and M) Spatial signals in gastric cancer H&E sections. Scale bars, 100 μm. Concordant quantification of immune cells from H&E and transcriptomic deconvolution To evaluate SpaHE-Infil’s accuracy in immune cell quantification, we collected fresh surgical specimens (colon cancer, healthy controls, gastric/ovarian/pancreatic cancers), prepared OCT-embedded frozen sections for H&E staining, and extracted RNA from adjacent tissues for bulk RNA-seq ( Figures 5 A, S4 A, and S4B). Immune cell predictions from H&E were compared with six deconvolution algorithms (TIMER, CIBERSORT, CIBERSORT-ABS, xCELL, EPIC, and MCP-counter). Across five tissues, SpaHE-Infil showed high concordance with TIMER, CIBERSORT, CIBERSORT-ABS, xCELL, and EPIC, while MCP-counter diverged ( Figure 5 B). Abundance comparisons revealed SpaHE-Infil aligned closely with these five algorithms and was consistently lower than xCELL ( Figure 5 C). Trends for major immune lineages (T cells, B cells, myeloid cells, NK cells, stromal cells) were identical: SpaHE-Infil agreed with the five algorithms and was significantly lower than xCELL ( Figure 5 D). Unidentified cell types by SpaHE-Infil stemmed from training set exclusions but can be flexibly extended via script adjustments ( Figure S4 C). In contrast to comparator tools, SpaHE-Infil supports custom parameters and scripting. Figure 5. Open in a new tab Concordant quantification of immune cells from H&E and transcriptomic deconvolution (A) H&E staining of OCT-embedded frozen sections and bulk RNA-seq of adjacent tissues from multiple cancers. (B) Consistency between SpaHE-Infil and six deconvolution algorithms across five tissues. (C) Comparison of immune cell abundance estimates across algorithms. (D) Stacked bar plot (proportion) depicting recognition trends of major immune cell types. (E) Scatterplot of major immune cell types distributions with discrepancy validation, ( N.S. , no significant). Clinical utility of SpaHE-Infil in predicting ICI response and patient prognosis To assess clinical value, we analyzed H&E images of pathological specimens from CRC, GC, pancreatic, and ovarian cancer cohorts ( Figures 6 A–6D). Patients were stratified into high/low immune infiltration groups based on TME abundance (median split). In ICI-treated CRC patients, high-infiltration showed significantly superior response rates ( p < 0.05; Figures 6 A–6E), a trend replicated in GC ( Figure 6 F). Prognostically, high-infiltration groups exhibited longer recurrence-free survival (RFS) across CRC, GC, ovarian, and pancreatic cohorts ( Figures 6 G–6J). Time-dependent ROC curves further validated these findings: CRC high-infiltration had significantly higher 12/24-month AUCs ( Figure S4 D; no events at 6 months precluded AUC calculation). Ovarian and pancreatic cohorts confirmed high infiltration as an independent predictor of favorable outcomes ( Figures S4 E and S4F). Figure 6. Open in a new tab Clinical utility of SpaHE-Infil in predicting ICI response and patient prognosis (A–D) Pathological specimens from colorectal cancer, gastric cancer, pancreatic, and ovarian cancer cohorts; patient stratification into high/low immune infiltration groups based on median TME abundance. Scale bars, 100 μm. (E–F) Significantly higher objective response rates (ORR) in high-infiltration groups of ICI-treated CRC/GC patients ( p < 0.05). (G–J) Prolonged recurrence-free survival (RFS) in high-infiltration groups across postoperative cohorts ( p < 0.05). Discussion The SpaHE-Infil tool developed in this study pioneers in situ resolution of immune cell spatial distributions within the TME using routine H&E-stained histopathological slides. 15 By establishing a mapping relationship between pathological image features and molecular-level cell-type annotations, this method innovatively eliminates the dependency of spatial omics technologies on sequencing data, transforming H&E slides into a direct medium for decoding TME spatial information. 16 Central to this approach is a multimodal feature-driven “spatial reader”: this model integrates 21-dimensional parameters—including nuclear morphology, local texture, spatial density, and color intensity—through a random forest (RF) classifier. The RF was prioritized and selected after benchmarking against comparable machine learning methods (e.g., support vector machines, logistic regression, convolutional neural network) for its inherent advantages tailored to our TME cell typing task. Specifically, its superiority lies in three key aspects that align with our research needs: first, it exhibits robust performance on high-dimensional, noisy multimodal data (a common characteristic of integrated pathological image and phenotypic features), avoiding overfitting more effectively than linear models even with limited clinical samples. Second, its ability to capture non-linear feature interactions enables reliable integration of heterogeneous parameters, outperforming competing methods in resolving subtle phenotypic differences among rare or low-abundance cell populations. Third, it offers higher computational efficiency and interpretability (via feature importance quantification) than complex models like deep learning, facilitating result validation and clinical translation. Leveraging these strengths, the RF achieves precise identification of 12 key TME cell types (e.g., CD8 + T cells, macrophages, NK cells). 17 Particularly noteworthy is its TME_adjust dynamic calibration mechanism. This module employs adaptive sampling strategies (oversampling/undersampling) to effectively address challenges posed by immune cell proportion imbalances and “cold tumor” phenomena, synergizing with the RF’s intrinsic robustness to significantly enhance model generalizability in heterogeneous tissues-especially for low-infiltrating samples (immune cells <5%). This integrated approach not only preserves the reliability of cell type identification but also leverages the spatial resolution advantages of pathological images, offering a cost-effective, high-throughput TME analysis solution for clinical practice. In cross-cancer and cross-species validations, SpaHE-Infil demonstrated high concordance with ST and mIHC. Predicted cellular spatial distributions showed significant grid-level PCC with ST data across seven tissue types (breast, gastric, etc.), and key immune cells (e.g., CD8 + T cells) exhibited precise alignment with in situ molecular annotations. Notably, this robust performance was achieved under the premise of species-matched and cancer type-matched model training and validation. Inherent differences in cell composition, pathological features and immune infiltration patterns across distinct cancer types will significantly compromise the accuracy of cross-cancer type prediction. We thus recommend optimizing the model with large-scale public spatial transcriptomics and digital pathology datasets of the same species and cancer type to enhance its generalization capability. SpaHE-Infil is designed as a flexible analytical method that can be continuously optimized by incorporating additional training data, rather than a static, unchangeable model. To validate cell type identification, consecutive-section mIHC experiments further confirmed its spatial resolution: predicted distributions of tumor cells (EpCAM + ), macrophages (CD68 + ), and others closely matched fluorescence protein signal curves. In immune cell abundance quantification, SpaHE-Infil agreed closely with mainstream deconvolution algorithms (TIMER, CIBERSORT, etc.) in estimating TME composition and distinctly differed from xCELL, which overestimated infiltration levels. Collectively, these results validate that morphological-spatial information embedded in H&E slides suffices for refined TME characterization, while the method’s iterable nature enables long-term performance improvement with expanded datasets. The clinical translational value of SpaHE-Infil was preliminarily demonstrated in independent cohorts. Of note, these analyses primarily validated that SpaHE-Infil can be directly applied to routine clinicopathological H&E-stained sections, rather than being limited to laboratory-based research settings. In ICI-treated colorectal and gastric cancer patients, the high immune infiltration group identified by this tool showed significantly higher objective response rates (ORR) than the low infiltration group. Across postoperative cohorts of colorectal, gastric, pancreatic, and ovarian cancers, high infiltration consistently correlated with prolonged RFS, further validated by time-dependent ROC curves. These findings highlight SpaHE-Infil’s potential as a clinical decision-support tool—generating TME spatial maps from routine pathology slides, with pretrained models (RDS format) for different cancers directly applicable as standardized templates. Such features provide immediate, accessible molecular insights for immunotherapy patient stratification and postoperative recurrence risk assessment. SpaHE-Infil establishes a novel image-driven paradigm for TME spatial profiling, enabling effective identification of 12 key cell types from routine H&E-stained slides. Integrating 21-dimensional multimodal features and a dynamic calibration module, it addresses immune cell proportion imbalance in clinical samples. Cross-cancer validation shows high concordance with spatial transcriptomics and mIHC, and independent clinical cohorts confirm its utility in predicting immunotherapy response and prognosis, providing a cost-effective tool for precision immunotherapy stratification. The model is currently optimized for epithelial-derived tumors with limited sensitivity for rare cell subtypes, which can be improved by expanding the training dataset and integrating high-resolution spatial transcriptomic data and more clinical samples to enhance tumor marker recognition for clinical pathological auxiliary application. Limitations of the study Currently, the method employed is AUCell-based automatic annotation combined with manual correction. Although it covers 12 types of TME cell populations, it has limited sensitivity for rare subtypes (e.g., γδ T cells); Morphological differences between FFPE and OCT-embedded samples may impact feature generalizability, though current sample sizes are insufficient to quantify this effect; Slide artifacts such as folds, fixation defects, and uneven staining are not automatically detected, requiring manual inspection and annotation prior to analysis; High resolution image processing relies on GPU acceleration, requiring substantially longer computation times on consumer-grade hardware-efficiency optimization is needed for deployment in primary hospitals; Clinical validation of SpaHE-Infil’s predictive power for ICI response and prognosis warrants larger cohorts due to current sample constraints. Future work will focus on: expanding the cell-type annotation system; optimizing algorithms by integrating high-resolution spatial transcriptomic platforms (e.g., Visium HD, Xenium, and MERFISH); establishing multicenter standardized validation platforms; and exploring 3D TME modeling via radiology-data fusion to deepen mechanistic insights into tumor immunotherapies (e.g., CAR-T, ICIs, and adoptive cell therapy). Resource availability Lead contact Further information and request for resources should be directed to the lead contact, Fang Wang ( [email protected] ). Materials availability SpaHE-Infil can be made available upon request through a collaborative arrangement. Data and code availability Detailed information on the omics datasets utilized in all analyses is summarized in Table S1 . Any additional information and materials required to replicate or reanalyze the study findings are also available from the lead contact upon request. The code base and comprehensive documentation for SpaHE-Infil, including the full analytical workflow and operating manual, are publicly and freely accessible at https://github.com/FangWangLab/SpaHE-Infil with no access restrictions. No new protein structural analysis was performed in this study, and thus no PDB validation reports are applicable. For other items, contact the lead contact upon reasonable request. Acknowledgments This work was supported by the Wu Jieping Medical Foundation (no. 320.6750.2023-05-57), the National Natural Science Foundation of China (no. 82504176), and Jiangsu Provincial Dual Creative Ph.D. Training Program, the Wuxi Science and Technology Innovation and Entrepreneurship Project (no. K20241013), the Medical Research Project Plan of Research Hospital Affiliated Hospital of Jiangnan University (no. YJZ202302). We also thank Dr. Hunan Wang from the Affiliated Hospital of Jiangnan University for her assistance during this research. And this work was supported by research guidance and funding from Professor Dajing Xia at the Zhejiang University School of Public Health (supported by the National Natural Science Foundation of China, grant no. 32370584). Author contributions F.W., Y.M., and X.Q. conceived the idea for this study. C.Y., L.D., and J.L. performed tissue section staining and immunological experiments. X.W., S.D., and X.G. were responsible for the evaluation and supervision of the TME cell detection data. Y.Q. and X.Q. were responsible for the pathological diagnosis. H.X. collected the publicly available ST sequencing datasets. F.W. and M.W. collected and analyzed the data. All authors contributed to drafting and review of the paper. All authors agreed to the final version of this manuscript. Declaration of interests The authors declare no competing interests. Declaration of generative AI and AI-assisted technologies in the writing process During the preparation of this work, the authors used ChatGPT (OpenAI) and AJE (Research Square) to assist with language polishing and manuscript structuring. After using those tools, the authors reviewed and edited the content as needed and take full responsibility for the publication’s content. STAR★Methods Key resources table REAGENT or RESOURCE SOURCE IDENTIFIER Antibodies CD8 primary antibody Proteintech #66868-1-Ig; RRID: AB_2882205 CD56 primary antibody Cell Signaling Technology #99746; RRID: AB_2868490 EpCAM primary antibody Cell Signaling Technology #2929; RRID: AB_2098657 CD68 primary antibody Proteintech #66231-2-Ig; RRID: AB_2881622 Vimentin primary antibody Affinity #AF7013; RRID: AB_2835318 DAPI Beyotime #C1006; RRID: AB_3712077 Opal 6-Plex Manual Detection Kit Akoya Biosciences RRID: AB_3665660 Biological samples Human FFPE tissue sections (CRC, GC, pancreatic, ovarian cancer) Affiliated Hospital of Jiangnan University Ethics Approval: LS2020058 & LS2025018 Human OCT-embedded tissues (CRC, GC, pancreatic, ovarian cancer) Affiliated Hospital of Jiangnan University Ethics Approval: LS2020058 & LS2025018 Public spatial transcriptomics samples (human/mouse tumor tissues) Public databases 18 , 19 , 20 , 21 , 22 , 23 , 24 See original publications Chemicals, peptides, and recombinant proteins RNeasy Kit Qiagen # 74104 Citrate buffer (pH 6.0) Abcam # SP-0001; RRID: AB_11029728 Opal fluorescent dyes (520, 540, 570, 480, 780) Akoya Biosciences RRID: AB_3674065 Deposited data Raw RNA-seq data This study Tables S2 and S3 Spatial transcriptomics processed data This study Table S1 Model feature importance/ranking data This study Table S6 Bias mitigation calibration data This study Table S8 Patient demographics/clinicopathological data This study Tables S3 and S4 Experimental models: Organisms/strains Mouse lung cancer model Public datasets 18 , 19 , 20 , 21 , 22 , 23 , 24 See original publications Mouse pancreatic tissue model Public datasets 18 , 19 , 20 , 21 , 22 , 23 , 24 See original publications Software and algorithms STAR v2.7.10 Seurat v5.1.0; RRID: SCR_016341 R R Foundation for Statistical Computing v4.3.3 AUCell v1.20.0; RRID: SCR_021327 inForm Akoya Biosciences v2.4; RRID: SCR_019155 ggplot2 R package V3.5.2 pheatmap R package V1.0.12 EBImage R package V4.42.0 Harmony R package V1.2.3 HoVer-Net NA Image segmentation architecture Other Vectra Polaris system Akoya Biosciences RRID: SCR_025508 NanoPhotometer® IMPLEN #N60 Bioanalyzer 2100 system Agilent Technologies – Illumina sequencing platform Illumina – Deep Learning Server (2× AMD EPYC 9654, 8× NVIDIA A100) Custom-built Ubuntu 22.04 LTS Image Processing Workstation (Intel i9-13950HX, NVIDIA RTX 4000) Custom-built Windows 11 Open in a new tab Experimental model and study participant details Human subjects This study utilized retrospective cohorts of human subjects. The treatment cohort comprised colorectal cancer (CRC, n = 32) and gastric cancer (GC, n = 16) patients treated with PD-1 inhibitors (ICI). The prognostic cohort included surgical patients with CRC, GC, pancreatic cancer ( n = 28), and ovarian cancer ( n = 26). Patients were dichotomized into high/low infiltration groups based on the median total immune cell abundance in the tumor microenvironment (TME) for survival analysis. Formalin-Fixed Paraffin-Embedded (FFPE) sections from these patients were obtained after acquiring signed informed consent. This study was approved by the Ethics Committee of Affiliated Hospital of Jiangnan University (Approval Nos. LS2020058 & LS2025018) and strictly adhered to all relevant ethical regulations. Detailed patient demographics and clinicopathological parameters are provided in Tables S3 and S4 . The influence of sex or gender was not specifically analyzed in this study; this is acknowledged as a limitation. Animal models This study incorporated publicly available spatial transcriptomics datasets from mouse models, including samples from mouse lung cancer ( n = 4) and mouse pancreatic tissue ( n = 4). 18 , 19 , 20 , 21 , 22 , 23 , 24 Specific details regarding the strain, age, sex, and maintenance conditions of these mice are available in the original publications associated with the referenced datasets. All animal experiments in the source studies were conducted with institutional permission and oversight. Method details Acquisition of spatial transcriptomics datasets This study integrated spatial transcriptomics (ST) datasets sourced from public databases, encompassing samples from breast cancer ( n = 14), pancreatic cancer ( n = 13), colon cancer ( n = 4), gastric cancer ( n = 10), ovarian cancer ( n = 8), mouse lung cancer ( n = 4), and mouse pancreatic tissue ( n = 4). 18 , 19 , 20 , 21 , 22 , 23 , 24 Raw data were aligned to reference genomes (GRCh38/mm10) using STAR (v2.7.10) and then analyzed using the Seurat toolkit (v5.1.0; RRID: SCR_016341 ): low-quality spots were filtered (total genes <100, mitochondrial gene percentage >10%), followed by normalization (SCTransform) and batch correction (Harmony). Included H&E-stained sections required a high-resolution image and a low-resolution image, with tissue integrity confirmed by pathologists. Details are provided in Table S1 . TME cell annotation Spatial transcriptomics data were processed using R (v4.3.3) and the Seurat package (v5.1.0). Spatial data (H5/MTX format) were loaded, and quality control was performed (mitochondrial gene percentage calculation and visualization). After data normalization (SCTransform), dimensionality reduction was conducted via PCA and t-SNE, followed by cell clustering based on shared nearest neighbors (SNN; resolution = 1.2). Cell type annotation employed the AUCell algorithm (v1.20.0, RRID: SCR_021327 ), which calculated gene set enrichment scores (Area Under the Curve, AUC) using predefined marker gene sets for 12 TME cell types (including tumor cells, immune cells, and stromal cells). Notably, the AUC score for CD8 + T cells was weighted by a factor of 1.3 to enhance sensitivity. Each cell was assigned to the cell type with the highest AUC score. 25 Annotation results were visualized via t-SNE plots and spatial distribution maps, and cell type proportions were quantified. The final annotated data were saved in RDS format, and high-resolution spatial annotation maps were exported as 300 DPI TIFF files. Two pathologists independently reviewed the annotations, with discrepancies resolved by a third pathologist. Twelve TME cell types were ultimately annotated: Tumor cells, Normal epithelial cells, CD8 + T cells, CD4 + T cells, Macrophages, Neutrophils, NK cells, Tregs, Fibroblasts, Endothelial cells, B cells, and Dendritic cells. Unassigned cells were labeled as “Unidentified”. SpaHE-Infil multimodal feature extraction and processing The following features were extracted from ST data and paired H&E sections. Molecular features Cell type-specific molecular signature scores based on differentially expressed genes (|logFC| > 1, FDR <0.05). To mitigate the impact of potential cell-type admixture within individual ST spots, we prefiltered spots to retain only those with high cell-type purity. Morphological features Nuclear morphological parameters (e.g., area, perimeter, ellipticity) were extracted via a nucleus-cytoplasm separation strategy. This method leverages H&E staining properties: nuclei are identified via adaptive thresholding of the blue channel (EBImagethresh), while cytoplasm is isolated by subtracting nucleus masks from low-threshold grayscale cell regions, outputting binary masks for subsequent feature calculation. Textural features Texture parameters were calculated using a modified Haralick method on grayscale-converted local regions. Following H&E feature extraction, three key morphological feature categories were extracted per cell based on spatial annotations, employing a cell-centered “visual territory” strategy (radius = 5 pixels) instead of fixed grids. This included: nuclear morphological features (area, perimeter, circularity, aspect ratio, compactness); textural features (contrast, correlation, energy, homogeneity) computed using the Haralick method; and spatial density features (local cell density, nearest neighbor distance, and its standard deviation) calculated using Euclidean distance. 26 Multiple features, including color/intensity, Haralick texture, morphology, and spatial distribution, were extracted. A Random Forest (RF) model was then used to screen and rank these features by importance ( Table S6 ). To mitigate bias from potential TME cell proportion imbalance, an adaptive TME_adjust parameter (0–1 scale) was implemented. This process employed stratified resampling (oversampling/undersampling) to adjust class proportions while preserving feature distributions, with sample protection rules (≥1 cell per type, ≥5 total cells) and a post-adjustment proportion error target of <0.1%. The effectiveness of bias mitigation was evaluated by comparing ROC and precision–recall curves before and after calibration ( Table S8 ). Robustness measures were integrated throughout feature processing, including adaptive thresholding, correction for fully black images (+0.01 offset), backup mask activation upon segmentation failure, jittering of zero-value features (jitter factor = 0.1), and filtering of constant features prior to analysis. Calibrated data were used for final RF modeling. For validation, a unified color scheme (12 cell-type colors) was applied to generate visualizations: boxplots of nuclear morphology, spatial heatmaps of textural features (viridis colormap), scatterplots of density–distance correlations, and PCA clustering with 90% confidence ellipses. Final reports were exported as 300 DPI PDF/TIFF and CSV files. The entire pipeline was executed via the HE_feature_Get command in the Sup.train_model-TMEadj_HE_feature_get.R script. SpaHE-Infil machine learning architecture Tumor-specific training sets were constructed individually for each tissue/tumor type, utilizing spatial transcriptomic annotations and multimodal features extracted from matched H&E images. Specifically, we used a total of 63 spatial transcriptomic samples (stratified by tissue type: 14 breast, 13 pancreatic, 4 colon, 10 gastric, 8 ovarian, 6 murine colon, 4 murine lung, 4 murine pancreatic samples), with each tumor-specific training set including spots from its corresponding tissue type only. After strict quality control, a total of 142,382 spots were incorporated across all training sets, with an average of 2,260 spots per sample. Cell coordinates and annotation information were extracted from ST data (RDS format) to build its respective training set. Paired H&E sections were used to extract four categories totaling 21 dimensions of features: Nuclear & Cytoplasmic Morphology (10 dim): Nuclear features (area, perimeter, circularity, aspect ratio, compactness) and cytoplasmic features (area, perimeter, circularity, aspect ratio, compactness), extracted using an improved H&E segmentation algorithm. Texture (4 dim): Haralick features (contrast, correlation, energy, homogeneity) calculated via texture analysis (32 bins, dual-scale computation) on grayscale images. Spatial Density (3 dim): Local cell density, nearest neighbor distance, and its standard deviation. Color Intensity (4 dim): Mean and standard deviation of RGB channels, mean and standard deviation of grayscale intensity. All features were integrated into the training set. A Random Forest model was employed for training (adaptive parameters: Leave-One-Out Cross-Validation (LOOCV) for sample size <10, 3-fold cross-validation for sample size >10; number of trees = min (200, sample size × 5)). The final output included the trained prediction model and a feature importance ranking. The prediction of a random forest integrates the voting results of all decision trees: y ˆ = argmax k ∑ t = 1 T I ( h t ( x ) = k ) ; ht is the prediction of the t-th tree, I is the indicator function, and κ is the cell type category. 27 The process was executed by running the 02.train_model-TMEadj.R script using the train_celltype_model command. Dynamic calibration module (TME_adjust) To address cell proportion imbalances within the TME, an adjustable parameter TME_adjust (scale 0–1; 0 corresponds to 5% TME cells, 1 corresponds to 50%) was introduced. When calibration is enabled, the original TME cell proportion (R_orig) in the training set is calculated. A target proportion (R_target = 0.05 + 0.45 × TME_adjust) is used to generate a balanced dataset. For TME cells, stratified oversampling is applied if R_orig < R_target, while undersampling is used if R_orig > R_target; non-TME cells undergo the reverse sampling strategy synchronously. Stratified sampling ensures that the inherent feature distribution of each cell type is maintained. Sampled subsets are merged and randomly reshuffled. This process preserves cellular heterogeneity, ensures post-adjustment proportion error <0.1%, and incorporates sample protection mechanisms (minimum 1 cell per type, total ≥5 cells). Original/adjusted proportions, sampling strategies, and bias assessment metrics (ROC/precision-recall curves) are exported to model metadata (CSV format). Model evaluation and visualization module This module systematically evaluates cell classification model performance using five core diagnostic plots: (1) Cross-validation boxplots displaying accuracy and Kappa distributions; (2) Confusion matrix heatmap; (3) Class performance bar plots presenting sensitivity/specificity/precision per cell type; (4) Feature importance ranking highlighting the top 20 key features; and (5) Sample distribution plots and TME calibration comparison plots. All plots are output as 300 DPI PDF files. SpaHE-Infil H&E image-based cell type prediction architecture Inputs to the prediction pipeline include: a pre-trained model (celltype_model.rds) for the sample type, an H&E image (≥300 dpi, ∼84.67 μm/pixel), and key parameters (TME_adjust, step_size, point_size, point_alpha). Multimodal features are extracted as described above. The Random Forest model then predicts cell types. Based on the TME_adjust parameter, predictions are optimized using either the original training set proportions or calibrated proportions. Standardization of cell type proportions ensures sum-to-1 consistency: P normalized ( c ) = P raw ( c ) ∑ k P raw ( k ) . Outputs include calibrated predictions (cell ID, coordinates, predicted type, probability) and visualization files. This is executed by running the 03.predict_celltypes-TMEadj.R script with the generate_predictions command. ST dataset validation and reporting To validate the accuracy of the SpaHE model, this study employed independent ST sequencing data. Validation comprised: (1) Data Loading & Alignment of ST data and SpaHE predictions. (2) Grid-based Aggregation onto identical spatial grids to calculate cell type proportions. (3) Correlation Analysis using Pearson Correlation Coefficients (PCC) for cell type proportions at the grid level. (4) Visualization of PCC spatial maps, density distributions, and scatterplots. This module is executed by running the 04.spa-HE_evaluate_Y-correction.R script with the evaluate_spahe_accuracy command. Multiplex immunohistochemistry (mIHC) FFPE tissue sections underwent deparaffinization, rehydration, antigen retrieval (citrate buffer, pH 6.0, microwave), and blocking (3% BSA). Sequential multicolor antibody labeling (Opal Polaris, Akoya Biosciences) was performed: primary antibody incubation, followed by Opal Polymer HRP-conjugated secondary antibody and fluorescent dye incubation. Antibody complexes were removed via microwave stripping (95°C, pH 6.0 buffer) before the next cycle. Markers and dilutions used: CD8 (Proteintech, #66868-1-Ig; 1:600, RRID: AB_2882205 ) labeled with Opal 570; CD56 (Cell Signaling Technology, #99746; 1:100, RRID: AB_2868490 ) labeled with Opal 780; EpCAM (Cell Signaling Technology, #2929; 1:1000, AB_2098657) labeled with Opal 540; CD68 (Proteintech, #66231-2-Ig; 1:1000,RRID: AB_2881622 ) labeled with Opal 520; Vimentin (Affinity, #AF7013; 1:300, RRID: AB_2835318 ) labeled with Opal 480. Nuclei were stained with DAPI (1 μg/mL, Beyotime, #C1006, RRID: AB_3712077 ). Slides were scanned using the Vectra Polaris system (Akoya Biosciences, RRID: SCR_025508 ). Acquired images were processed using inForm software (v2.4, Akoya, RRID: SCR_019155 ) for spectral unmixing and cell phenotyping. 28 Transcriptome sequencing Total RNA was extracted from OCT-embedded tissues (one paired set of colon cancer and adjacent normal, one pancreatic cancer, one gastric cancer, one ovarian cancer) using the RNeasy Kit (Qiagen). RNA quality control was assessed by gel electrophoresis, NanoPhotometer (IMPLEN), and Bioanalyzer 2100 system (Agilent Technologies). cDNA libraries were prepared (average insert size 300 ± 50 bp) and paired-end sequenced on an Illumina platform. Raw RNA-seq data are available in Tables S2 and S3 . Technical implementation details Image segmentation was primarily based on a thresholding strategy, with the HoVer-Net architecture employed as an auxiliary module to enhance segmentation performance. Morphological features underwent Z score normalization prior to PCA. Spatial features were subjected to log(x+1) transformation. For model deployment, R interfaces integrated with GPU libraries were utilized for GPU-accelerated inference. Computational environment A heterogeneous computing architecture was employed. Deep Learning Server: 2× AMD EPYC 9654 processors, 8× NVIDIA A100 80 GB GPUs, 1.5 TB RAM, Ubuntu 22.04 LTS. Image Processing Workstation: Intel Core i9-13950HX processor, NVIDIA RTX 4000 Ada 12 GB GPU, 128 GB RAM, Windows 11. Quantification and statistical analysis All statistical analyses were performed in R (v4.3.3). Inter-group comparisons for categorical variables employed the χ 2 test or Fisher’s exact test. Survival analysis used the Kaplan-Meier method with log rank tests for group comparisons. Time-dependent Receiver Operating Characteristic (ROC) curves assessed predictive performance. For model evaluation, accuracy and Kappa statistics were derived from cross-validation. Pearson Correlation Coefficients (PCC) were used to quantify spatial concordance between SpaHE predictions and ST validation data. For bulk RNA-seq deconvolution algorithm comparison, Spearman correlation analysis was used to quantify inter-algorithm concordance. The exact statistical test applied, the number of biologically independent units (N), and all relevant statistical parameters (e.g., χ 2 values, log rank χ 2 values, ROC area under the curve (AUC) with 95% confidence intervals, degrees of freedom, exact unadjusted p -values) are reported in the legend of each relevant figure. Statistical significance for non-omics analyses was set at p < 0.05 , with significance asterisks uniformly defined across all figure legends as: ∗p < 0.05 , ∗∗p < 0.01 , ∗∗∗p < 0.001 , ∗∗∗∗p < 0.0001 . All analyses underwent rigorous data cleaning and quality control. Visualizations were generated using R packages including ggplot2 and pheatmap. Additional resources Our study does not include clinical trials, so this section is not applicable. Published: February 26, 2026 Footnotes Supplemental information can be found online at https://doi.org/10.1016/j.isci.2026.115155 . Contributor Information Fang Wang, Email: [email protected]. Xiaowei Qi, Email: [email protected]. Yong Mao, Email: [email protected]. Supplemental information Document S1. Figures S1–S4 mmc1.pdf (45.4MB, pdf) Table S1. Summary of spatial transcriptomics datasets integrated in this study, including tissue types, public access numbers, sequencing platforms, and file formats mmc2.xlsx (9.9KB, xlsx) Table S2. Bulk RNA-seq expression profiles of tumor and adjacent normal tissues across four cancer types mmc3.xlsx (2MB, xlsx) Table S3. Clinical characteristics of retrospective cohorts (treatment and prognostic cohorts) for immunotherapy response and survival analysis mmc4.xlsx (18.7KB, xlsx) Table S4. Supplementary clinical data of study participants, including sample allocation and follow-up information mmc5.xlsx (10.7KB, xlsx) Table S5. Marker gene sets for 12 TME cell types used in AUCell annotation mmc6.xlsx (13KB, xlsx) Table S6. Feature importance ranking of multimodal features (molecular, morphological, textural, spatial) screened by Random Forest model mmc7.xlsx (10.2KB, xlsx) Table S7. Performance metrics of SpaHE-Infil model across different tumor types mmc8.xlsx (10.3KB, xlsx) Table S8. Comparison of ROC and precision-recall curves before and after TME cell proportion calibration mmc9.xlsx (10.6KB, xlsx) References 1. Elhanani O., Ben-Uri R., Keren L. Spatial profiling technologies illuminate the tumor microenvironment. Cancer Cell. 2023;41:404–420. doi: 10.1016/j.ccell.2023.01.010. [ DOI ] [ PubMed ] [ Google Scholar ] 2. Larsson L., Frisén J., Lundeberg J. Spatially resolved transcriptomics adds a new dimension to genomics. Nat. Methods. 2021;18:15–18. doi: 10.1038/s41592-020-01038-7. [ DOI ] [ PubMed ] [ Google Scholar ] 3. Zaitsev A., Chelushkin M., Dyikanov D., Cheremushkin I., Shpak B., Nomie K., Zyrin V., Nuzhdina E., Lozinsky Y., Zotova A., et al. Precise reconstruction of the TME using bulk RNA-seq and a machine learning algorithm trained on artificial transcriptomes. Cancer Cell. 2022;40:879–894.e16. doi: 10.1016/j.ccell.2022.07.006. [ DOI ] [ PubMed ] [ Google Scholar ] 4. Wang Y., Liu B., Zhao G., Lee Y., Buzdin A., Mu X., Zhao J., Chen H., Li X. Spatial transcriptomics: Technologies, applications and experimental considerations. Genomics. 2023;115 doi: 10.1016/j.ygeno.2023.110671. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 5. Campanella G., Kumar N., Nanda S., Singi S., Fluder E., Kwan R., Muehlstedt S., Pfarr N., Schüffler P.J., Häggström I., et al. Real-world deployment of a fine-tuned pathology foundation model for lung cancer biomarker detection. Nat Med. 2025;31:3002–3010. doi: 10.1038/s41591-025-03780-x. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 6. Shamai G., Livne A., Polónia A., Sabo E., Cretu A., Bar-Sela G., Kimmel R. Deep learning-based image analysis predicts PD-L1 status from H&E-stained histopathology images in breast cancer. Nat. Commun. 2022;13 doi: 10.1038/s41467-022-34275-9. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 7. Li T., Fu J., Zeng Z., Cohen D., Li J., Chen Q., Li B., Liu X.S. TIMER2.0 for analysis of tumor-infiltrating immune cells. Nucleic Acids Res. 2020;48:W509–W514. doi: 10.1093/nar/gkaa407. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 8. Newman A.M., Liu C.L., Green M.R., Gentles A.J., Feng W., Xu Y., Hoang C.D., Diehn M., Alizadeh A.A. Robust enumeration of cell subsets from tissue expression profiles. Nat. Methods. 2015;12:453–457. doi: 10.1038/nmeth.3337. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 9. Aran D., Hu Z., Butte A.J. xCell: digitally portraying the tissue cellular heterogeneity landscape. Genome Biol. 2017;18:220. doi: 10.1186/s13059-017-1349-1. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 10. Racle J., Gfeller D. EPIC: A Tool to Estimate the Proportions of Different Cell Types from Bulk Gene Expression Data. Methods Mol. Biol. 2020;2120:233–248. doi: 10.1007/978-1-0716-0327-7_17. [ DOI ] [ PubMed ] [ Google Scholar ] 11. Plattner C., Finotello F., Rieder D. Deconvoluting tumor-infiltrating immune cells from RNA-seq data using quanTIseq. Methods Enzymol. 2020;636:261–285. doi: 10.1016/bs.mie.2019.05.056. [ DOI ] [ PubMed ] [ Google Scholar ] 12. Becht E., Giraldo N.A., Lacroix L., Buttard B., Elarouci N., Petitprez F., Selves J., Laurent-Puig P., Sautès-Fridman C., Fridman W.H., et al. Estimating the population abundance of tissue-infiltrating immune and stromal cell populations using gene expression. Genome Biol. 2016;17:218. doi: 10.1186/s13059-016-1070-5. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 13. You Y., Fu Y., Li L., Zhang Z., Jia S., Lu S., Ren W., Liu Y., Xu Y., Liu X., et al. Systematic comparison of sequencing-based spatial transcriptomic methods. Nat. Methods. 2024;21:1743–1754. doi: 10.1038/s41592-024-02325-3. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 14. Liu X., Tang G., Chen Y., Li Y., Li H., Wang X. SpatialDeX Is a Reference-Free Method for Cell-Type Deconvolution of Spatial Transcriptomics Data in Solid Tumors. Cancer Res. 2025;85:171–182. doi: 10.1158/0008-5472.CAN-24-1472. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 15. Bergstrom E.N., Abbasi A., Díaz-Gay M., Galland L., Ladoire S., Lippman S.M., Alexandrov L.B. Deep Learning Artificial Intelligence Predicts Homologous Recombination Deficiency and Platinum Response From Histologic Slides. J. Clin. Oncol. 2024;42:3550–3560. doi: 10.1200/JCO.23.02641. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 16. Pitino E., Pascual-Reguant A., Segato-Dezem F., Wise K., Salvador-Martinez I., Crowell H.L., Marção M., Ruiz M., Courtois E., Flynn W.F., et al. STAMP: Single-cell transcriptomics analysis and multimodal profiling through imaging. Cell. 2025;188:5100–5117.e26. doi: 10.1016/j.cell.2025.05.027. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 17. Duan Q., Zhang H., Zheng J., Zhang L. Turning Cold into Hot: Firing up the Tumor Microenvironment. Trends Cancer. 2020;6:605–618. doi: 10.1016/j.trecan.2020.02.022. [ DOI ] [ PubMed ] [ Google Scholar ] 18. Coutant A., Cockenpot V., Muller L., Degletagne C., Pommier R., Tonon L., Ardin M., Michallet M.C., Caux C., Laurent M., et al. Spatial Transcriptomics Reveal Pitfalls and Opportunities for the Detection of Rare High-Plasticity Breast Cancer Subtypes. Lab. Invest. 2023;103 doi: 10.1016/j.labinv.2023.100258. [ DOI ] [ PubMed ] [ Google Scholar ] 19. Park S.S., Lee Y.K., Choi Y.W., Lim S.B., Park S.H., Kim H.K., Shin J.S., Kim Y.H., Lee D.H., Kim J.H., Park T.J. Cellular senescence is associated with the spatial evolution toward a higher metastatic phenotype in colorectal cancer. Cell Rep. 2024;43 doi: 10.1016/j.celrep.2024.113912. [ DOI ] [ PubMed ] [ Google Scholar ] 20. Lee S.H., Lee D., Choi J., Oh H.J., Ham I.H., Ryu D., Lee S.Y., Han D.J., Kim S., Moon Y., et al. Spatial dissection of tumour microenvironments in gastric cancers reveals the immunosuppressive crosstalk between CCL2+ fibroblasts and STAT3-activated macrophages. Gut. 2025;74:714–727. doi: 10.1136/gutjnl-2024-332901. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 21. Ju H.Y., Youn S.Y., Kang J., Whang M.Y., Choi Y.J., Han M.R. Integrated analysis of spatial transcriptomics and CT phenotypes for unveiling the novel molecular characteristics of recurrent and non-recurrent high-grade serous ovarian cancer. Biomark. Res. 2024;12:80. doi: 10.1186/s40364-024-00632-7. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 22. Pei G., Min J., Rajapakshe K.I., Branchi V., Liu Y., Selvanesan B.C., Thege F., Sadeghian D., Zhang D., Cho K.S., et al. Spatial mapping of transcriptomic plasticity in metastatic pancreatic cancer. Nature. 2025;642:212–221. doi: 10.1038/s41586-025-08927-x. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 23. Dhainaut M., Rose S.A., Akturk G., Wroblewska A., Nielsen S.R., Park E.S., Buckup M., Roudko V., Pia L., Sweeney R., et al. Spatial CRISPR genomics identifies regulators of the tumor microenvironment. Cell. 2022;185:1223–1239.e20. doi: 10.1016/j.cell.2022.02.015. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 24. Tindall R.R., Yang Y., Hernandez I., Qin A., Li J., Zhang Y., Gomez T.H., Younes M., Shen Q., Bailey-Lundberg J.M., et al. Aging- and alcohol-associated spatial transcriptomic signature in mouse acute pancreatitis reveals heterogeneity of inflammation and potential pathogenic factors. J. Mol. Med. 2024;102:1051–1061. doi: 10.1007/s00109-024-02460-6. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 25. Aibar S., González-Blas C.B., Moerman T., Huynh-Thu V.A., Imrichova H., Hulselmans G., Rambow F., Marine J.C., Geurts P., Aerts J., et al. SCENIC: single-cell regulatory network inference and clustering. Nat. Methods. 2017;14:1083–1086. doi: 10.1038/nmeth.4463. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 26. Shah M., Polónia A., Curado M., Vale J., Janowczyk A., Eloy C. Impact of Tissue Thickness on Computational Quantification of Features in Whole Slide Images for Diagnostic Pathology. Endocr. Pathol. 2025;36:10. doi: 10.1007/s12022-025-09855-2. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 27. Jiang Y., Chen Y., Cheng Q., Lu W., Li Y., Zuo X., Wu Q., Wang X., Zhang F., Wang D., et al. A random survival forest-based pathomics signature classifies immunotherapy prognosis and profiles TIME and genomics in ES-SCLC patients. Cancer Immunol. Immunother. 2024;73:241. doi: 10.1007/s00262-024-03829-9. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 28. Li S., Jiang B., Zhou H., Yang S., Yang L., Hong Y. Development of a prognostic immune cell-based model for ovarian cancer using multiplex immunofluorescence. J. Transl. Med. 2025;23:688. doi: 10.1186/s12967-025-06745-3. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials Document S1. Figures S1–S4 mmc1.pdf (45.4MB, pdf) Table S1. Summary of spatial transcriptomics datasets integrated in this study, including tissue types, public access numbers, sequencing platforms, and file formats mmc2.xlsx (9.9KB, xlsx) Table S2. Bulk RNA-seq expression profiles of tumor and adjacent normal tissues across four cancer types mmc3.xlsx (2MB, xlsx) Table S3. Clinical characteristics of retrospective cohorts (treatment and prognostic cohorts) for immunotherapy response and survival analysis mmc4.xlsx (18.7KB, xlsx) Table S4. Supplementary clinical data of study participants, including sample allocation and follow-up information mmc5.xlsx (10.7KB, xlsx) Table S5. Marker gene sets for 12 TME cell types used in AUCell annotation mmc6.xlsx (13KB, xlsx) Table S6. Feature importance ranking of multimodal features (molecular, morphological, textural, spatial) screened by Random Forest model mmc7.xlsx (10.2KB, xlsx) Table S7. Performance metrics of SpaHE-Infil model across different tumor types mmc8.xlsx (10.3KB, xlsx) Table S8. Comparison of ROC and precision-recall curves before and after TME cell proportion calibration mmc9.xlsx (10.6KB, xlsx) Data Availability Statement Detailed information on the omics datasets utilized in all analyses is summarized in Table S1 . Any additional information and materials required to replicate or reanalyze the study findings are also available from the lead contact upon request. The code base and comprehensive documentation for SpaHE-Infil, including the full analytical workflow and operating manual, are publicly and freely accessible at https://github.com/FangWangLab/SpaHE-Infil with no access restrictions. No new protein structural analysis was performed in this study, and thus no PDB validation reports are applicable. For other items, contact the lead contact upon reasonable request. Articles from iScience are provided here courtesy of Elsevier ACTIONS View on publisher site PDF (21.9 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top

Record · ID 4381 · SHA-256 ca66fff9c07c4cf3
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.