Spatial multi-omics integration by cross-modal graph contrastive learning - PMC Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Brief Bioinform . 2026 Apr 19;27(2):bbag192. doi: 10.1093/bib/bbag192 Search in PMC Search in PubMed View in NLM Catalog Add to search Spatial multi-omics integration by cross-modal graph contrastive learning Yang Gui Yang Gui 1 School of Mathematics and Physics, University of Science and Technology Beijing, 30 Xueyuan Road, Haidian District, Beijing 100083, China Find articles by Yang Gui 1 , Yan Xu Yan Xu 2 School of Mathematics and Physics, University of Science and Technology Beijing, 30 Xueyuan Road, Haidian District, Beijing 100083, China Find articles by Yan Xu 2, ✉ , Chao Li Chao Li 3 School of Statistics and Applied Mathematics, Anhui University of Finance and Economics, 962 Caoshan Road, Longzihu District, Bengbu 233041, Anhui, China Find articles by Chao Li 3, ✉ Author information Article notes Copyright and License information 1 School of Mathematics and Physics, University of Science and Technology Beijing, 30 Xueyuan Road, Haidian District, Beijing 100083, China 2 School of Mathematics and Physics, University of Science and Technology Beijing, 30 Xueyuan Road, Haidian District, Beijing 100083, China 3 School of Statistics and Applied Mathematics, Anhui University of Finance and Economics, 962 Caoshan Road, Longzihu District, Bengbu 233041, Anhui, China ✉ Corresponding authors. Yan Xu, E-mail: [email protected] ; Chao Li, E-mail: [email protected] . Received 2025 Dec 3; Revised 2026 Feb 22; Accepted 2026 Mar 24; Collection date 2026 Mar. © The Author(s) 2026. Published by Oxford University Press. This is an Open Access article distributed under the terms of the Creative Commons Attribution-NonCommercial License ( https://creativecommons.org/licenses/by-nc/4.0/ ), which permits non-commercial re-use, distribution, and reproduction in any medium, provided the original work is properly cited. For commercial re-use, please contact [email protected] for reprints and translation rights for reprints. All other permissions can be obtained through our RightsLink service via the Permissions link on the article page on our site—for further information please contact [email protected]. PMC Copyright notice PMCID: PMC13092272 PMID: 42001472 Abstract Recent advances in spatial multi-omics technologies have enabled high-resolution profiling of cellular heterogeneity while preserving spatial context, offering unprecedented opportunities to decipher tissue architecture and intercellular communication. Although existing spatial transcriptomics tools have been effective for single modal analysis, integrated interpretation of multi omics layers including spatial transcriptome, spatial proteome, and spatial epigenome remains limited due to modality specific technical biases and biological complexity. To address this, we present CoMo, a graph-based framework that synergizes multi-modal feature learning through cross attention mechanisms, coupled with dual optimization via neighbor-aware contrastive loss for cross-omics feature fusion and cluster-aware contrastive loss for spatially coherent domain identification. Evaluations on five spatial omics datasets demonstrate superior performance in spatial domain identification compared with state-of-the-art methods. CoMo provides a robust computational tool for multi-omics studies and supports comprehensive characterization of tissue through synergistic feature learning. Keywords: spatial multi-omics, multi-omics integration, graph contrastive learning Introduction The development of spatial multi-omics technologies, including spatial transcriptome, spatial proteome, and spatial epigenome, has provided opportunities to decode cellular ecosystems through complementary molecular layers. Nevertheless, the inherent heterogeneity among different modalities in both technical noise and biological variation presents challenges for integration and requires advanced computational approaches for data harmonization. Recent years have witnessed significant advances in multi-omics integration across both single-cell and spatial domains. Among single-cell multi-omics integration methods, MultiVi [ 1 ] probabilistically integrates paired RNA and chromatin accessibility data using variational inference. However, such approaches neglect spatial information, resulting in inconsistent domain identification compared with spatially informed multi-omics frameworks, such as SpatialGlue [ 2 ], COSMOS [ 3 ], and CANDIES [ 4 ]. Concurrently, a wide range of spatial omics tools [ 5–7 ], such as SpaGCN [ 8 ], STAGATE [ 9 ], GraphST [ 10 ], and STAIG [ 11 ] have been developed for modeling spatial transcriptomics. Although limited to single modalities, these methods have been instrumental in characterizing spatial domain organization and now serve as key baselines for guiding the design and benchmarking of spatial multi-omics integration strategies. Building upon these single-modality foundations, spatial multi-omics approaches extend the analysis to jointly capture cross-modal spatial dependencies. SpatialGlue employs dual-attention graph neural networks to adaptively weigh cross-modal contributions. COSMOS leverages graph-based architectures to fuse epigenome–transcriptome spatial correlations. CANDIES integrates heterogeneous information in spatial multi-omics data through a collaborative framework combining conditional diffusion models with contrastive learning. However, existing methods potentially underutilize cross-omics dependencies and lack the ability to identify both spatially coherent domains and dispersed yet transcriptionally distinctive domains. To overcome the aforementioned limitations in multi-omics data integration, we developed CoMo, a novel computational framework that leverages graph contrastive learning to unify multi-omics data. CoMo employs multi-objective contrastive loss to eliminate inter-data heterogeneity while capturing cross-modal synergy across omics layers, enabling robust integration of diverse molecular modalities, thereby achieving more accurate identification of spatial domains. In particular, during the pretraining stage, CoMo employs a graph autoencoder architecture based on cross-attention to integrate latent correspondences from different omics, while utilizing a reconstruction loss to eliminate single-omics noise. In the fused embedding extraction stage, CoMo appends a clustering head to the embedding of each omics and designs cluster-aware contrastive loss to further enhance inter-omics consistency, while a projection head and neighbor-aware contrastive losses are introduced to enhance the discriminative power of the fused embedding. Ultimately, downstream tasks such as spatial domain identification, pseudo-spatiotemporal trajectory mapping, and differential expression analysis can be performed based on the fused embedding. Systematic benchmarking across multiple experimental scenarios, such as spatial transcriptome–epigenome and spatial transcriptome–proteome data, demonstrates that CoMo consistently outperforms six state-of-the-art (SOTA) methods. These results validate the effectiveness of CoMo in mitigating inter-omics heterogeneity, integrating information across modalities, and enhancing the discriminative nature of the fused embedding, thereby offering a robust solution for spatial multi-omics data integration and downstream analysis. Materials and methods Data description We curated and analyzed various spatial omics datasets spanning diverse biological contexts, species, and molecular modalities. All datasets were retrieved from publicly available sources and uniformly processed using the CoMo framework to ensure consistency and comparability across analyses. To evaluate the performance of multi-omics integration algorithms, both simulated and real-world datasets were employed. The simulated datasets (1296 spots) were generated to approximate transcriptomic and epigenomic data. For the real datasets, the following resources were used: (i) the Human Lymph dataset (3484 spots) containing transcriptomic and proteomic data, and (ii) the Mouse Brain dataset (9215 spots) comprising transcriptomic and epigenomic data. To reconstruct embryonic cell fate trajectories, the Mouse Embryo dataset (2187 spots) was utilized. For exploring human immune niche organization, the Human Tonsil dataset (4194 spots) was processed to obtain jointly profiled transcriptomes and proteins. More details of spatial omics datasets can be found in the Supplementary Section 2.1 . Overview of CoMo The CoMo framework implements a four-stage workflow comprising data preprocessing, pretraining, fused embedding extraction, and downstream analysis. In the data preprocessing module, omics-specific preprocessing steps are performed to extract features. In the pretraining stage, CoMo employs a graph autoencoder framework to integrate shared and complementary information across omics layers. For multi-omics fusion, CoMo leverages graph convolutional networks (GCNs) to capture molecular features and physical interactions from each omics layer, followed by refined cross-attention and weighted nearest neighbor (WNN) integration to consolidate inter-omics complementarity. Simultaneously, CoMo reconstructs the original features of both omics layers from the fused embedding. During training, a reconstruction loss is optimized to denoise individual omics data while promoting cross-omics information exchange ( Fig. 1a ). In the fused embedding extraction stage, for each omics-specific embedding, CoMo introduces a cluster head to infer category matrix. The fused embedding is then processed by a projection head, which maps the embedding to a new latent space. Training at this stage involves the joint optimization of reconstruction loss, cluster-aware contrastive loss, and neighbor-aware contrastive loss, which together enhance alignment across omics and promote discriminative representation learning ( Fig. 1b ). Finally, the refined fused embedding obtained from the fused embedding extraction stage are applied to downstream tasks such as spatial domain identification, pseudo-spatiotemporal map, and differential expression analysis ( Fig. 1c ). Figure 1. Open in a new tab Overview of CoMo. (a) In the data preprocessing and pretraining stage, CoMo first performs omics-specific preprocessing before constructing graphs using feature matrices and spatial coordinates. These omics-specific graphs serve as inputs to GCN-based methods to extract omics-specific embeddings. Cross-attention and WNN are then employed to combine the low-dimensional representations from both omics layers, generating a fused embedding. During training, CoMo optimizes the framework using a reconstruction loss. (b) During the fused embedding extraction stage, cluster heads and projection heads are applied to both omics-specific embeddings and the integrated fused embedding in CoMo. The framework is optimized through a composite loss function comprising reconstruction loss, cluster-aware contrastive loss, and neighbor-aware contrastive loss during training. (c) The resulting fused embedding obtained from (b) are critical for three downstream analytical tasks: spatial domain identification, pseudo-spatiotemporal mapping, and differential expression analysis. Data preprocessing and graph construction The CoMo framework data preprocessing pipeline begins with the extraction of omics-specific features through tailored normalization and dimensionality reduction techniques. For transcriptomic data, highly variable genes are first selected using the Seurat v3 method with 3000 top genes, followed by total count normalization to 10 000 reads per cell and log1p transformation. The processed transcriptomic features are then reduced to principal components. For epigenomic data, the top 3000 highly variable features are selected using the Seurat v3 method, and Latent Semantic Indexing (LSI) is applied to project the sparse peak-count matrix into a low-dimensional representation. The dimensionality of the LSI space is adaptively scaled to the complexity of each dataset, capturing the essential chromatin accessibility landscape to initialize node features for the construction of epigenomic graphs. Proteomic data receive centered log-ratio normalization at the single-cell level prior to principal component analysis [ 2 ]. The spatial coordinates and omics-specific features are utilized to construct graphs , , , and (or alternatively and ). For graphs that incorporate both spatial and omics features, the spatial coordinates are used to build the adjacency matrix , while the corresponding omics features serve as node attributes . In contrast, for omics-only graphs , the same omics features are employed to construct both the adjacency matrix and node features . For any adjacency matrix , indicates an edge between node and node , whereas denotes its absence Fig. 1a . Multi-omics feature extraction The CoMo method advances multi-omics integration through a graph contrastive learning paradigm. Graph representation learning [ 12 , 13 ] have demonstrated remarkable success in modeling complex biological systems by explicitly encoding molecular interactions as graph structures, while contrastive learning techniques [ 14–16 ] have proven particularly effective for learning discriminative representations by maximizing the similarity between positive views and minimizing the similarity between positive and negative views [ 17 ]. Building upon these foundations, CoMo implements a multi-stage architecture based on GCNs that progressively refines omics integration through two specialized training phases. The pretraining phase establishes initial omics-specific and cross-omics embedding using graph autoencoder and cross attention. The subsequent fused embedding extraction phase introduces cluster and projection heads, employing cluster- and neighbor-aware contrastive learning to enhance both discriminative power and cross-omics consistency Fig. 1a . Pretraining stage In the pretraining stage of CoMo, each omics is projected into a unified low-dimensional embedding space through architecturally consistent encoders. We employ GCNs as the universal encoder backbone for all modalities, leveraging their ability to preserve microenvironmental context through neighborhood aggregation. While the GCN architecture is shared across modalities to ensure structural symmetry, the weight matrices remain non-shared to adapt to the specific feature distributions of each omics type. For individual modalities, dual GCN encoders process two graphs: (i) a spatial adjacency graph encoding physical proximity between spots, and (ii) a feature similarity graph capturing functional relationships among transcriptionally similar spots regardless of spatial distance. The GCN propagation rule is formalized as: (1) (2) where denotes the normalized adjacency matrix, represents the degree matrix, and is the adjacency matrix with self-loops. is the activation function, such as the Parametric Rectified Linear Unit. The weight matrix and bias vector of the th layer in the encoder are denoted as and , respectively. The low-dimensional embedding at the th layer is represented by , with as the initial input. All definitions and formulations for apply analogously to . For each individual omics, GCNs generate two distinct graph-specific representations denoted as and . To achieve optimal unification of these dual embeddings, we implement an adaptive attention mechanism that dynamically weights their contributions based on local context. The attention mechanism operates through a learned parametric function that evaluates the relative importance of each graph-specific representation. For a given spot , we first apply a linear transformation to each graph-specific representation (where ), followed by a nonlinear activation: (3) where denotes the trainable weight vector. and are trainable parameters shared across all spots within the modality. These weights are normalized via softmax to produce the final attention scores: (4) The omics-specific representation is computed as the weighted combination: (5) This formulation ensures the integrated representation preserves both structural relationships from the spatial graph and functional patterns from the feature similarity graph, with the attention mechanism automatically adjusting their relative contributions based on local context. Omics-specific embeddings and , generated through dedicated graph-based encoders, are integrated via a reversed softmax-based cross-attention mechanism. This design specifically addresses the limitation of conventional attention in multi-omics fusion, where strong cross-modality correlations tend to dominate the attention weights, thereby suppressing biologically meaningful but weakly correlated complementary signals. For each modality, we first compute query, key, and value projections using learnable matrices. For , we compute query vectors , key vectors , and value vectors . Similarly for , we obtain , , and . The cross-attended embeddings are then computed as: (6) (7) where represents the dimension of key vectors. Here, denotes the initial state of the cross-attention update, which can be initialized as . The negative exponent ( ) inverts standard attention by amplifying low-affinity associations rather than high-affinity correlations. This design explicitly prioritizes uncorrelated, omics-specific features that provide complementary biological insights, effectively suppressing redundant information while enhancing synergistic inter-modal relationships. The representation for the other modality is obtained through an analogous procedure by exchanging the roles of and . Following the generation of cross-attended embeddings and , we employ a modified WNN algorithm to compute omics-specific weights and integrate embeddings. The final fused embedding is then computed as the convex combination: (8) where denotes element-wise multiplication. and are the omics-specific weights obtained through the WNN. CoMo achieves unified representation learning of multi-omics data through synergistic consideration of cross-omics information exchange, feature reconstruction fidelity, and spatial topological constraints. The dual reconstruction decoders form the core of the first stage. The decoders employ GCNs that simultaneously integrate the embedded representation and spatial adjacency matrix . The initial hidden representation is defined as , which serves as the input to the decoder layers. The reconstruction of omics layer at layer is computed as: (9) where is updated through graph convolution across layers. The reconstruction loss plays a critical role in eliminating intra-modality noise and enhancing signal fidelity within each omics layer. The loss function in the pretraining stage is defined as: (10) where and represent the original feature matrices of the two omics layers, and and denote their reconstructed matrices. and are the trade-off hyperparameters. This formulation ensures that the integration process retains shared biological signals across modalities. Fused embedding extraction stage Omics-specific cluster heads independently process the embedded representations of each omics (e.g. transcriptome embedding and proteome embedding ), mapping them into categorical probability spaces through non-shared weight matrices, such that and . The cluster contrastive loss aligns cross-modality cluster centers, enforcing consistent semantics across different omics for the same domain: (11) Here, denotes the th column of the probability matrix , representing the global distribution of the th cluster in modality . The parameter indicates the total number of latent spatial domains, which is predefined based on biological annotations or estimated via community detection. This loss is supported by semantic alignment theory in cross-modal retrieval. Spatial contrastive learning guided by a projection head further constrains the geometric structure of the fused embedding. The projection matrix reduces to a lower dimensional space and applies a neighborhood contrastive loss: (12) where represents the spatial neighborhood of spot . This loss encodes the microenvironment-dependency principle, where cells within spatially adjacent spots exhibit functional synergy in tissue microenvironments, and contrastive learning reinforces local consistency. The overall loss function in the fused embedding extraction stage is defined as: (13) where , , and are hyperparameters that balance the contributions of different modules. The biological significance of the three-loss synergy manifests as follows: the reconstruction loss effectively denoises raw data while preserving essential biological features; the cluster-aware contrastive loss ensures cross-omics consistency in domain type identification; and the neighbor-aware contrastive loss maintains microenvironment-dependent spatial relationships. Downstream tasks Spatial domain identification. For multi-omics data data, CoMo generates fused embeddings that subsequently undergo clustering analysis. CoMo supports three clustering algorithms: (i) the mclust algorithm from the mclust R package [ 18 ], and (ii) the Louvain and (iii) Leiden algorithms from the Python package scanpy [ 19 ]. These algorithms are suitable for datasets where the number of domains is either known or unknown. Pseudo-spatiotemporal mapping. To delineate dynamic biological processes across spatial domains, CoMo employs diffusion pseudotime (DPT) analysis by using “ scanpy.tl.dpt ” function implemented by scanpy. The root cell for DPT initialization is manually specified to correspond to biologically relevant progenitor states, ensuring alignment with known biological hierarchies. Pseudotime values are visualized as spatial heatmaps, highlighting transition zones and differentiation gradients across domains. Differential expression analysis. To detect domain-specific marker genes through differential expression analysis across distinct spatial domains, we employed the “ correlation_matrix ” and “ rank_genes_groups_dotplot ’ functions implemented in the scanpy. For the “ rank_genes_groups_dotplot ” method, a minimum log2 fold change threshold of 2 was applied to ensure the selection of biologically relevant and statistically significant marker genes. Evaluation metrics We evaluated multi-omics spatial domain identification using both the Adjusted Rand Index [ 20 ] (ARI) and the Normalized Mutual Information [ 21 ] (NMI), along with four additional metrics: Mutual Information [ 22 ] (MI), Adjusted Mutual Information [ 23 ] (AMI), Homogeneity, and V-measure [ 24 ]. In the case of datasets lacking manual annotation, Moran’s I [ 25 ] was utilized as the key metric for assessment. More information about the evaluation metrics can be found in Supplementary Section 2.2 . Results CoMo-driven benchmarking demonstrates superior cross-modal integration in spatial multi-omics To demonstrate the effectiveness of CoMo, comprehensive benchmarking was conducted across three spatial multi-omics datasets: a simulated dataset (transcriptome–epigenome), the Human Lymph dataset (transcriptome–proteome), and the Mouse Brain P22 dataset (transcriptome–epigenome). We compared CoMo with single-omic spatial domain identification algorithms (SpaGCN, STAGATE, GraphST, and STAIG), single-cell multi-omics integration algorithm MultiVi, and Spatial multi-omics integration algorithms (SpatialGlue, COSMOS, and CANDIES) across multi-omics benchmark datasets. Detailed information about these algorithms is provided in Supplementary Section 2.3 . On the simulated benchmark, CoMo achieved near-optimal performance with an ARI of 0.9853 and NMI of 0.9807. This represented a improvement in ARI over SpatialGlue (ARI=0.9280) and a gain versus CANDIES (ARI=0.9429). While GraphST demonstrated competitive performance among single-omic methods (ARI=0.9669), CoMo maintained a advantage ( Fig. 2a ). Other single-omic spatial methods, including SpaGCN, STAGATE, and STAIG, exhibited significantly lower performance compared with CoMo ( Supplementary Table S6 ). Specifically, the CoMo result was closer to the annotation in data distribution, with sharper spatial domain boundaries. Whereas, SpatialGlue and CANDIES contained some discrete points, presumably due to the absence of specific processing for spatial neighborhoods. COSMOS generated blurred boundaries because of excessive focus on local neighborhood information. SpaGCN, STAGATE, GraphST, and STAIG, meanwhile, required more modal information to enhance the discriminative power of their embeddings. Lastly, MultiVi produced highly disorganized identification results due to a lack of spatial information ( Fig. 2b ). Figure 2. Open in a new tab Benchmarking performance of CoMo and eight competing methods on a simulated spatial multi-omics dataset. (a) Quantitative evaluation of the nine methods using multiple metrics, presented as boxplots. (b) Spatial domain identification results, from left to right: ground truth annotation and the predictions generated by CoMo, SpatialGlue, CANDIES, COSMOS, SpaGCN, STAGATE, GraphST, STAIG, and MultiVi. Next, CoMo was benchmarked against SOTA methods using a human lymph dataset profiled by 10 Genomics Visium spatial transcriptomics, with expert manual annotation defining major anatomical domains—including the cortex, medulla, and follicular areas—as ground truth. Quantitative evaluation across six clustering metrics consistently ranked CoMo first, achieving the highest ARI (0.2989), Homogeneity (0.3930), and V-measure (0.3997), outperforming the second-best method, SpatialGlue, by in ARI, in Homogeneity, and in V-measure. Compared with the single-cell multi-omics integration method MultiVi, CoMo demonstrated improvements of in ARI, in Homogeneity, and in V-measure. When compared with single-omic spatial domain identification methods, CoMo consistently maintained a significant lead. Among these, GraphST and SpaGCN showed relatively competitive performance with ARIs of 0.2493 and 0.2491, respectively, followed by STAGATE and STAIG ( Fig. 3a and Supplementary Table S7 ). Biologically, CoMo, SpatialGlue, and CANDIES showed closer alignment with ground truth by successfully identifying major lymphoid structures including follicles, cortex, and medulla (sinuses and cords). Notably, CoMo distinguished the capsule from pericapsular adipose tissue. In contrast, COSMOS and GraphST suffered from over-smoothing that obscured discontinuous spatial structures, while MultiVi generated excessively fragmented domain patterns ( Fig. 3b ). Notably, low-dimensional embeddings from CoMo were visualized using Uniform Manifold Approximation and Projection (UMAP), revealing spatial trajectories that aligned with established developmental patterns. In contrast to alternative spatial multi-omics integration methods, the UMAP of the embeddings successfully reconstructed the developmental trajectory from pericapsular adipose tissue to capsule ( Fig. 3c ). Figure 3. Open in a new tab Performance evaluation of spatial domain clustering across multiple methods on a real-world spatial multi-omics datasets. (a) Six metrics boxplots of the nine methods on the human lymph dataset. (b) Manual annotation and spatial domain identification results for the nine methods on the human lymph dataset. (c) UMAP visualizations in CoMo, SpatialGlue, CANDIES, and COSMOS for human lymph dataset. (d–f) Corresponding results for the mouse brain (postnatal day 22) dataset, presented in the same order: (d) metric boxplots, (e) manual annotation and spatial domain identification results, and (f) UMAP visualizations from four selected methods. Benchmarking on the Mouse Brain P22 dataset revealed that while COSMOS achieved the highest quantitative scores due to its strong spatial constraints, CoMo delivered the most competitive and balanced results among the remaining methods. The spatial prioritization inherent in COSMOS was found to cause extensive oversmoothing, which merged biologically distinct domains into broad spatial regions. CoMo achieved a highly precise identification of major domains within the mouse brain tissue, notably demonstrating a superior ability to accurately resolve the boundary between Layer 5 (L5) and Layers 6a/6b (L6a/b). This fine-grained anatomical resolution underscores the effectiveness of the integrative framework in capturing intricate cortical structures that are often consolidated by other benchmarking methods. CoMo demonstrated the most balanced performance with an ARI of 0.3896 and NMI of 0.5804, outperforming SpatialGlue by in ARI and in NMI. Significant performance gaps were observed when comparing CoMo to single-modality methods, including SpaGCN, STAGATE, GraphST, and STAIG. Among this group, SpaGCN achieved the highest ARI (0.3403), followed by STAGATE (0.3238), GraphST (0.3099), and STAIG (0.2343). This disparity underscores the critical advantage of multi-omics integration in capturing the complex architecture of the brain. Notably, single-cell multi-omics method MultiVi showed limited effectiveness in this context, achieving only 0.1050 in ARI ( Fig. 3d - and e and Supplementary Table S8 ). The UMAP visualization demonstrates that CoMo produces a developmental trajectory, i.e. both clearer and more biologically accurate than those generated by other integration methods ( Fig. 3f ). CoMo-enabled reconstruction uncovers spatial cell fate dynamics To evaluate integration performance on spatially resolved transcriptome–epigenome data, comprehensive benchmarking was performed using mouse embryo (E13) dataset. Quantitative evaluation revealed superior performance by CoMo, which achieved an ARI of 0.3658 and significantly outperformed the second-best method by . Optimal scores were obtained across all clustering metrics, including homogeneity (0.5131), V-measure (0.5044), MI (1.2152), NMI (0.5044), and AMI (0.4958) demonstrating consistent advantages over both multi-omics and single-omics methods. Specifically, CoMo, SpatialGlue, and MultiVi accurately identified key developmental domains including the ventricular zone, stratum subpallium, and stratum pallium (corresponding to domain j8, domain j7, and domain j10, respectively). However, spatial domain boundaries were delineated with improved clarity in CoMo compared with both alternative methods ( Fig. 4a–c and Supplementary Table S9 ). To validate the biological relevance of domains identified by CoMo, highly variable genes were examined across distinct spatial domains. Specifically, genes including Snhg11 , Ttyh1 , Zbtb20, and Six3 showed statistically significant expression patterns that corresponded to domains j1, j6, j8, and j9. Both RNA expression and ATAC accessibility signals for these marker genes demonstrated strong spatial enrichment within their respective domains, confirming that CoMo effectively captures biologically meaningful spatial patterns across transcriptomic and epigenomic modalities ( Fig. 4d and Supplementary Table S10 ). Pseudo-spatiotemporal analysis further demonstrated the spatiotemporal patterns of CoMo, revealing clearly stratified structures that outperformed COSMOS. The pseudotime progression from ventricular zone to subpallium to pallium precisely recapitulated the established developmental hierarchy. This biologically consistent trajectory reconstruction underscores the advantage of CoMo in preserving developmental dynamics within fused embeddings ( Fig. 4e ). Figure 4. Open in a new tab Analysis of a spatial transcriptome-epigenome data from the mouse embryo at E13. (a) Manual annotation of mouse embryo E13 and spatial domain identification result of CoMo. (b) Six metrics boxplots of the nine methods on the mouse embryo dataset. (c) Eight spatial domain identification results. (d) Spatial gene expression plots of marker genes. (e) Pseudo spatiotemporal maps (pSM) generated by CoMo and COSMOS. CoMo-based mapping uncovers human tonsil immune niches To further assess performance in complex human tissue, CoMo was evaluated on a human tonsil dataset integrating transcriptomic and proteomic data. While CoMo, SpatialGlue, and CANDIES all successfully detected PCNA-positive germinal center (GC) regions, CoMo uniquely delineated a more integral and continuous representation of the GC core. In contrast, COSMOS failed to identify these distinct functional domains entirely. Spatial autocorrelation was quantified using Moran’s I , where COSMOS achieved the highest score (0.7803) due to its inherent over-smoothing characteristics, though this came at the cost of poor domain discrimination. In contrast, CoMo maintained strong spatial coherence (Moran’s ) while preserving accurate domain boundaries, demonstrating an optimal balance between spatial continuity and biological specificity that surpassed both SpatialGlue (0.4454) and CANDIES (0.5217) ( Fig. 5a and b ). Figure 5. Open in a new tab Analysis of a spatial transcriptome–proteome data from the human tonsil. (a) H&E-stained histology image of tissue slice. (b) Spatial domain results generated by the four multi-omics integration algorithms. (c) Dot plots of each spatial domain in the domain-specific genes. The size of the dots represents the proportion of cells expressing the gene. The color represents the average expression of the gene in the domain, with red indicating high expression and blue indicating low expression. (d) The expression pattern of marker gene IL7R and CD22 . (e) Pseudo spatiotemporal maps (pSM) generated by CoMo. Further investigation of the identification results reveals that CoMo uniquely uncovers the intricate spatial organization of functional compartments. Specifically, domain 6 was characterized by the high expression of IL7R , indicating these regions function as the T-cell zone or paracortical areas responsible for lymphocyte recruitment [ 26 ]. Conversely, Domain 7 exhibited strong enrichment of CD22 , which are canonical markers for germinal center (GC) B cells and early-stage activation [ 27 ]. Notably, the precise overlap of Domain 7 with PCNA-positive regions indicates that CoMo accurately delineates the highly proliferative GC core ( Fig. 5c and d and Supplementary Table S11 ). To further investigate the dynamic biological processes preserved by CoMo, we performed pseudotime analysis ( Fig. 5 e). By defining Domain 6 (the T-cell zone rich in IL7R ) as the developmental root, the inferred trajectory clearly radiates toward the PCNA-positive Domain 7 (the germinal center). This transition aligns with the biological paradigm of B-cell activation, where naive B cells migrate from the T-cell boundary into the follicles to undergo rapid proliferation and somatic hypermutation upon antigen encounter. The pseudotime trajectory reveals a seamless transition from the paracortical regions to the follicular cores, demonstrating the capability of CoMo to capture the underlying biological manifold that remains unresolved by alternative methods. Ablation studies and analysis Systematic ablation studies and sensitivity analyses were conducted to evaluate the contribution of key components and hyperparameters within the CoMo framework. To provide a comprehensive validation, ablation experiments were performed across all datasets with ground-truth annotations, including the Simulation, Human Lymph, Mouse Brain P22, and Mouse Embryo E13 datasets. Ablation experiments involved sequentially removing critical modules, such as the refined cross-attention mechanism, the cluster-aware contrastive loss ( ), and the neighbor-aware contrastive loss ( ). Across all evaluated datasets, the removal of any single component resulted in measurable performance degradation. While the main analysis emphasizes the ARI, the robustness of each module was further confirmed through a holistic evaluation involving five additional metrics: NMI, MI, AMI, Homogeneity, and V-measure ( Fig. 6a and Supplementary Fig. S4 ). These results confirm that each architectural element contributes non-redundantly to the overall integration performance. Hyperparameter sensitivity was specifically evaluated by varying the embedding dimension across values of 32, 64, 128, and 256. The Simulation dataset demonstrated stable value at lower dimensions, but exhibited a significant performance drop at dimension 256. The Human Lymph dataset maintained consistent performance across all tested dimensions, indicating robust performance across biologically distinct tissue architectures ( Fig. 6b ). In the Supplementary Section 2.4 , we provide a detailed explanation of the key hyperparameters setting. Figure 6. Open in a new tab Ablation study and hyperparameter sensitivity analysis of CoMo. (a) Performance of ablated models on the Simulation and Human Lymph datasets. (b) Effect of embedding dimension on clustering performance on two datasets. Discussion The development of CoMo addresses a critical computational challenge in spatial biology: the integrated analysis of multiple molecular modalities. Our framework demonstrates that graph contrastive learning, when coupled with cross-attention mechanisms and multi-objective optimization, can effectively reconcile the biological heterogeneity inherent in spatial multi-omics data. Instead of relying on simple feature concatenation, CoMo captures subtle inter-modality variations, exemplified by its precise delineation of the T-cell zone and the highly proliferative germinal center core in human lymph nodes. This robust performance suggests that the synergistic learning of shared and modality-specific features is essential for accurate spatial domain identification in complex tissues. The architecture of CoMo incorporates several key innovations that contribute to its effectiveness. The dual-phase training strategy, which separates feature reconstruction from discriminative embedding learning, proves crucial for eliminating technical noise while preserving biologically relevant signals. Specifically, the reversed cross-attention mechanism prioritizes low-affinity associations. The integration of cluster-aware and neighbor-aware contrastive losses specifically addresses the challenge of maintaining both transcriptional similarity and spatial coherence, enabling the identification of both compact domains and dispersed cell populations sharing molecular profiles. Beyond these contributions, the design of cross-attention architectures warrants further consideration in the broader context of multi-omics integration. Different omics assays, including transcriptomics, epigenomics, and proteomics, exhibit substantial heterogeneity in data distributions, sparsity patterns, and noise characteristics. This raises an important question: should a unified cross-attention architecture be universally applied across omics combinations, or should models adopt modality-specific attention designs that explicitly account for the unique statistical and biological properties of each omics type? Although CoMo currently employs a unified cross-attention module, incorporating more flexible or modality-adaptive attention components may further improve its adaptability, particularly for datasets characterized by extreme sparsity or pronounced modality-specific signals [ 28 , 29 ]. Investigating how to best balance architectural generality with modality-tailored specialization represents a meaningful direction for future research in multi-omics data fusion. As spatial multi-omics technologies continue to evolve, generating increasingly complex and high-dimensional data, computational frameworks like CoMo will be essential for unlocking the full potential of these technologies. The model demonstrates excellent performance in multi-omics integration and clustering, though its effectiveness may be compromised when confronted with significant inter-modality heterogeneity arising from technical noise. Future work will focus on incorporating advanced denoising techniques, modality-adaptive attention architectures, and refined information complementation strategies to enhance the robustness of CoMo in high-noise and highly heterogeneous data. Key Points Unified spatial multi-omics integration. CoMo addresses the heterogeneity across spatial transcriptome, proteome, and epigenome by introducing a graph contrastive learning framework that harmonizes multi-modal signals. Cross-attention-enhanced feature fusion. A cross-attention graph autoencoder captures complementary information across omics layers during pretraining, while reconstruction loss effectively suppresses modality-specific noise. Multi-objective contrastive optimization. CoMo integrates neighbor-aware and cluster-aware contrastive losses to strengthen inter-omics consistency and improve the discriminative power of fused embeddings. Supplementary Material CoMo_Sup_bbag192 como_sup_bbag192.pdf (1.1MB, pdf) Contributor Information Yang Gui, School of Mathematics and Physics, University of Science and Technology Beijing, 30 Xueyuan Road, Haidian District, Beijing 100083, China. Yan Xu, School of Mathematics and Physics, University of Science and Technology Beijing, 30 Xueyuan Road, Haidian District, Beijing 100083, China. Chao Li, School of Statistics and Applied Mathematics, Anhui University of Finance and Economics, 962 Caoshan Road, Longzihu District, Bengbu 233041, Anhui, China. Author contributions Y.G. and Y.X. designed this research. Y.G. and C.L. participated in the discussion of the model. Y.G. proposed CoMo and conducted the experiments. Y.X. and C.L. oversaw the entire project. Y.G. drafted the manuscript. Y.X. and C.L. reviewed the manuscript. Conflicts of interest The authors declare no competing interests. Funding This work was supported by the National Natural Science Foundation of China (grant no. 12071024), the Fundamental Research Funds for the Central Universities (grant no. FRF-BRB-25-007), and Anhui Provincial Quality Engineering Projects for Institutions of Higher Education (grant no. 2024jxgl029). Data availability The spatial multi-omics datasets supporting the findings of this study are all publicly available. (i) The simulated transcriptome-epigenome dataset is available at https://zenodo.org/records/10362607 . (ii) The Human Lymph dataset containing transcriptomic and proteomic data is accessed at https://zenodo.org/records/10362607 . (iii) The Mouse Brain dataset comprising transcriptomic and epigenomic data is available at https://cells.ucsc.edu/?ds=brain-spatial-omics+p22-atac . (iv) The Mouse Embryo dataset comprising transcriptomic and epigenomic data is accessed at https://cells.ucsc.edu/?ds=brain-spatial-omics+e13 . (v) The Human Tonsil dataset co-profiling transcriptomes and proteins is available at https://www.10xgenomics.com/resources/datasets/gene-protein-expression-library-of-human-tonsil-cytassist-ffpe-2-standard . All source codes used in our experiments have been deposited at https://github.com/Lab-Xu/CoMo/ . References 1. Ashuach T, Gabitto MI, Koodli RV et al. MultiVI: deep generative model for the integration of multimodal data. Nat Methods 2023; 20:1222–31. 10.1038/s41592-023-01909-9 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 2. Long Y, Ang KS, Sethi R et al. Deciphering spatial domains from spatial multi-omics with spatialglue. Nat Methods 2024; 21:1658–67. 10.1038/s41592-024-02316-4 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 3. Zhou Y, Xiao X, Dong L et al. Cooperative integration of spatially resolved multi-omics data with cosmos. Nat Commun 2025; 16:27. 10.1038/s41467-024-55204-y [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 4. Liu Y, Zou W, Li Y et al. Cross-modal denoising and integration of spatial multi-omics data with candies. bioRxiv 2025. 10.1101/2025.04.17.649333v1 [ DOI ] [ Google Scholar ] 5. Li J, Chen S, Pan X et al. Cell clustering for spatial transcriptomics data with graph neural networks. Nat Comput Sci 2022; 2:399–408. 10.1038/s43588-022-00266-5 [ DOI ] [ PubMed ] [ Google Scholar ] 6. Gui Y, Li C, Yan X. Spatial domains identification in spatial transcriptomics using modality-aware and subspace-enhanced graph contrastive learning. Comput Struct Biotechnol J 2024; 23:3703–13. 10.1016/j.csbj.2024.10.029 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 7. Gui Y, Tan Z, Yan X et al. Heterogeneous graph contrastive learning for integration and alignment of spatial transcriptomics data. Brief Bioinform 2025; 26:bbaf497. 10.1093/bib/bbaf497 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 8. Jian H, Li X, Coleman K et al. SpaGCN: integrating gene expression, spatial location and histology to identify spatial domains and spatially variable genes by graph convolutional network. Nat Methods 2021; 18:1342–51. [ DOI ] [ PubMed ] [ Google Scholar ] 9. Dong K, Zhang S. Deciphering spatial domains from spatially resolved transcriptomics with an adaptive graph attention auto-encoder. Nat Commun 2022; 13:1739. 10.1038/s41467-022-29439-6 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 10. Long Y, Ang KS, Li M et al. Spatially informed clustering, integration, and deconvolution of spatial transcriptomics with GraphST. Nat Commun 2023; 14:1155. 10.1038/s41467-023-36796-3 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 11. Yang Y, Cui Y, Zeng X et al. STAIG: spatial transcriptomics analysis via image-aided graph contrastive learning for domain exploration and alignment-free integration. Nat Commun 2025; 16:1067. 10.1038/s41467-025-56276-0 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 12. Wei J, Fang Z, Yiyang G et al. A comprehensive survey on deep graph representation learning. Neural Netw 2024; 173:106207. 10.1016/j.neunet.2024.106207 [ DOI ] [ PubMed ] [ Google Scholar ] 13. Shou Y, Cao X, Liu H et al. Masked contrastive graph representation learning for age estimation. Pattern Recognit 2025; 158:110974. 10.1016/j.patcog.2024.110974 [ DOI ] [ Google Scholar ] 14. Velickovic P, Fedus W, Hamilton WL et al. Deep graph infomax ICLR (Poster) 2019; 2:4. [ Google Scholar ] 15. Mavromatis C, Karypis G. Graph InfoClust: leveraging cluster-level node information for unsupervised graph representation learning. arXiv preprint arXiv:2009.06946. 2020. https://arxiv.org/abs/2009.06946 16. Chen J, Kou G. Attribute and structure preserving graph contrastive learning. In: Proceedings of the AAAI Conference on Artificial Intelligence 2023, Washington, DC, USA: AAAI Press, 2023; 37:7024–32. 10.1609/aaai.v37i6.25858 [ DOI ] [ Google Scholar ] 17. Trivedi P, Lubana ES, Yan Y et al. Augmentations in graph contrastive learning: current methodological flaws & towards better practices. In: Proceedings of the ACM Web Conference, New York, NY, USA: ACM Press, 2022; pp. 1538–49. https://dl.acm.org/doi/abs/10.1145/3485447.3512200 [ Google Scholar ] 18. Fraley C, Raftery AE, Murphy TB et al. Mclust version 4 for r: normal mixture modeling for model-based clustering, classification, and density estimation. Technical Report no. 597, Department of Statistics, University of Washington, June 2012. [ Google Scholar ] 19. wolf FA, Angerer P, Theis FJ. SCANPY: large-scale single-cell gene expression data analysis. Genome Biol 2018; 19:1–5. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 20. Rand WM. Objective criteria for the evaluation of clustering methods. J Am Stat Assoc 1971; 66:846–50. 10.1080/01621459.1971.10482356 [ DOI ] [ Google Scholar ] 21. Strehl A, Ghosh J. Cluster ensembles—a knowledge reuse framework for combining multiple partitions. J Mach Learn Res 2002; 3:583–617. [ Google Scholar ] 22. Shannon CE. A mathematical theory of communication. Bell Syst Tech J 1948; 27:379–423. 10.1002/j.1538-7305.1948.tb01338.x [ DOI ] [ Google Scholar ] 23. Vinh NX, Epps J, Bailey J. Information theoretic measures for clusterings comparison: variants, properties, normalization and correction for chance. J Mach Learn Res 2010, 2010; 11:2837–54. [ Google Scholar ] 24. Rosenberg A, Hirschberg J. V-measure: a conditional entropy-based external cluster evaluation measure. In: Proceedings of the 2007 joint conference on empirical methods in natural language processing and computational natural language learning (EMNLP-CoNLL), Prague, Czech Republic: Association for Computational Linguistics, 2007; pp. 410–20. https://aclanthology.org/D07-1043/ 25. Moran PAP. Notes on continuous stochastic phenomena. Biometrika 1950; 37:17–23. 10.1093/biomet/37.1-2.17 [ DOI ] [ PubMed ] [ Google Scholar ] 26. Wakisaka N, Moriyama-Kita M, Kondo S et al. Immune-related gene expression profile at peritumoral tonsillar tissue is modified by oropharyngeal cancer nodal status. Am J Pathol 2023; 193:1006–12. 10.1016/j.ajpath.2023.04.010 [ DOI ] [ PubMed ] [ Google Scholar ] 27. Macauley MS, Kawasaki N, Peng W et al. Unmasking of CD22 co-receptor on germinal center B-cells occurs by alternative mechanisms in mouse and man. J Biol Chem 2015; 290:30066–77. 10.1074/jbc.M115.691337 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 28. Ge S, Sun S, Xu H et al. Deep learning in single-cell and spatial transcriptomics data analysis: advances and challenges from a data science perspective. Brief Bioinform 2025; 26:bbaf136. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 29. Li G, Shaliu F, Wang S et al. A deep generative model for multi-view profiling of single-cell RNA-seq and ATAC-seq data. Genome Biol 2022; 23:20. 10.1186/s13059-021-02595-6 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials CoMo_Sup_bbag192 como_sup_bbag192.pdf (1.1MB, pdf) Data Availability Statement The spatial multi-omics datasets supporting the findings of this study are all publicly available. (i) The simulated transcriptome-epigenome dataset is available at https://zenodo.org/records/10362607 . (ii) The Human Lymph dataset containing transcriptomic and proteomic data is accessed at https://zenodo.org/records/10362607 . (iii) The Mouse Brain dataset comprising transcriptomic and epigenomic data is available at https://cells.ucsc.edu/?ds=brain-spatial-omics+p22-atac . (iv) The Mouse Embryo dataset comprising transcriptomic and epigenomic data is accessed at https://cells.ucsc.edu/?ds=brain-spatial-omics+e13 . (v) The Human Tonsil dataset co-profiling transcriptomes and proteins is available at https://www.10xgenomics.com/resources/datasets/gene-protein-expression-library-of-human-tonsil-cytassist-ffpe-2-standard . All source codes used in our experiments have been deposited at https://github.com/Lab-Xu/CoMo/ . Articles from Briefings in Bioinformatics are provided here courtesy of Oxford University Press ACTIONS View on publisher site PDF (2.0 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top