ConceptioArchivearXiv CS
arXiv CSopen access

PathMoG: A Pathway-Centric Modular Graph Neural Network for Multi-Omics Survival Prediction

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

PathMoG: A Pathway-Centric Modular Graph Neural Network for Multi-Omics Survival Prediction Di Wang,1 Chupei Tang,2 Junxiao Kong,3 Jixiu Zhai,4 Moyu Tang5 and Tianchi Lu6,∗ 1

arXiv:2604.24371v1 [cs.LG] 27 Apr 2026

Cuiying Honors College, Lanzhou University, China, 2 School of Mathematics and Statistics, Lanzhou University, China, 3 School of Mathematics and Statistics, Lanzhou University, China, 4 School of Mathematics and Statistics, Lanzhou University, China and Shanghai Innovation Institute, 5 School of Mathematics and Statistics, Lanzhou University, China and 6 Department of Computer Science, City University of Hong Kong, China ∗

Corresponding author. Email: [email protected]

Abstract Motivation: Cancer survival prediction from multi-omics data remains challenging because prognostic signals are high dimensional, heterogeneous, and distributed across interacting genes and pathways. Existing survival models either ignore biological structure, rely on genome-scale monolithic graphs, or fuse expression, mutation, and copy number variation (CNV) too coarsely for the low-sample, high-dimensional setting. Results: In this article, we propose PathMoG (Pathway-centric Modular Omics Graph), a biologically structured multi-omics graph neural network for cancer survival prediction. PathMoG first reorganizes genome-scale inputs into 354 KEGG-informed pathway modules, introducing functional priors that constrain representation learning in the high-dimensional, low-sample regime. We then use a Hierarchical Omics Modulation (HOM) module to condition gene-expression representations on mutation, CNV, pathway, and clinical context, enabling semantically aligned cross-omics integration rather than coarse feature concatenation. A dual-level attention mechanism further captures both intra-pathway driver signals and inter-pathway clinical relevance to generate patient-level Cox risk estimates with multi-level interpretability. We comprehensively evaluated PathMoG on 5,650 patients across 10 TCGA cancer types, and the results demonstrate clear superiority over representative survival baselines, while the framework provided interpretable gene-level, pathway-level, and patient-level readouts for biologically grounded and clinically actionable risk stratification. Availability and implementation: TCGA data were accessed through UCSC Xena and external validation used METABRIC. Source code is available at https://github.com/wangzoyou/pathmog. Key words: Multi-omics integration; Graph Neural Networks; Survival Analysis; Cancer Genomics; Interpretability; Modular Design

1. Introduction

designed to directly capture the high-order cross-gene interactions and cross-omics dependencies that dominate modern cancer data. Deep survival and multimodal fusion models emerged to address this representational bottleneck. CAMR introduced crossaligned multimodal representation learning to improve crossomics consistency in prognosis prediction (11); HFBSurv used hierarchical fusion with factorized bilinear interactions to model cross-modal feature coupling (15); and PCLSurv incorporated prototypical contrastive learning to strengthen class-aware multiomics representations for survival modeling (17). However, many of these methods still rely on coarse fusion, where omics views are treated as parallel feature blocks and merged by concatenation or weakly constrained attention, which can blur biological directionality between upstream genomic alterations and downstream transcriptional readouts. Graph-based prognosis models were introduced as a secondstage correction to this limitation by injecting prior biological structure. Representative studies follow a clear progression:

Cancer prognosis is intrinsically a multi-factor problem, because clinically similar patients can follow markedly different trajectories driven by distinct molecular programs (22; 23). In modern cohorts, prognostic signals are rarely confined to one modality and instead span expression, mutation, and copy-number variation (CNV), making multi-omics integration essential for risk modeling (8). At the same time, cancer datasets typically operate under a high-dimensional low-sample regime (p ≫ n), where robust generalization is difficult. Methodologically, survival modeling still builds on three classical families: non-parametric, parametric, and semi-parametric formulations (2). Cox proportional hazards and its extensions remain central in biomedical applications (3), and random survival forest provides a strong nonlinear baseline (9). These approaches are statistically grounded and interpretable, but they are not

1

2 GraphSurv combined GCN encoding with a Cox survival objective for multi-omics prognosis (26); PathGNN introduced pathwayoriented interpretable graph modeling for risk stratification and pathway analysis (18); FGCNSurv jointly learned omics feature representations and feature-relation structure for pancancer survival prediction (27); and prior knowledge-guided multilevel GNNs further strengthened biologically constrained prognosis modeling (28). In adjacent but highly relevant graphcentric settings, DeepKEGG used biologically informed graph modeling to capture cell heterogeneity from single-cell and spatial transcriptomics data, while PathHDNN used pathway-hierarchical deep architecture for immunotherapy response prediction and mechanism interpretation (14; 16). In adjacent tasks such as multi-omics cancer subtyping, graph contrastive learning has also reinforced the value of structure-aware representations (29). Yet many graph pipelines remain monolithic at genome scale, and this “one-big-graph” strategy often introduces three practical issues in p ≫ n cohorts: noise propagation through task-irrelevant longrange edges, explanatory dilution of patient-specific drivers, and excessive hypothesis complexity that increases overfitting risk. Taken together, the field has progressed from black-box fusion toward biologically informed graphs, but a key tension remains: prior biology is necessary, while unconstrained global graphs are often too blunt for stable survival learning. This tension defines two gaps addressed in this work: a graph structure gap (how to encode biology as a useful constraint rather than a massive graph prior) and an omics semantics gap (how to preserve directional relations across omics layers rather than treating them as symmetric channels). To resolve these gaps, we propose PathMoG, a pathway-centric modular graph neural network for multi-omics survival prediction. Instead of a monolithic graph, PathMoG decomposes the interactome into 354 KEGG-informed predefined pathway modules; applies a Hierarchical Omics Modulation (HOM) mechanism to condition expression on mutation, CNV, pathway, and clinical context via FiLM-style modulation; and aggregates evidence through dual-level attention for patient-level Cox risk estimation with interpretable outputs. We evaluate PathMoG across 10 TCGA cancer types and further examine transferability on METABRIC, clinical relevance in BRCA, and mechanistic validity through ablation and interpretability analyses. The end-to-end workflow is shown in Figure 1.

2. Materials and methods 2.1. Datasets We evaluated PathMoG on a harmonized pan-cancer cohort of 5,650 patients from 10 TCGA cancer types, where each patient was represented by gene expression, somatic mutation, CNV, and survival outcome information. These cohorts span diverse tissue origins, event rates, and sample sizes, making them suitable for assessing generalizability across heterogeneous disease settings. We additionally reserved the METABRIC breast cancer cohort as an external transfer cohort for independent testing without model finetuning. Cohort-wise sample sizes and event rates are summarized in Table 1. Expression profiles were log-transformed and standardized, somatic mutation calls were binarized at the gene level, and GISTIC CNV calls were encoded on a five-level scale from −2 to +2 (6; 20). Let I denote the patient set used for modeling. In

PathMoG: A Pathway-Centric Modular GNN Table 1. Dataset characteristics for 10 TCGA cancer types. The cohort covers diverse tissue origins and survival event rates, providing a rigorous testbed for generalizability. All cancer types use 7,595 genes from 354 KEGG pathways, and the total event rate is reported as a sample-size-weighted average across cohorts. Cancer Type

Samples

Event Rate (%)

BRCA (Breast Invasive Carcinoma) COAD (Colon Adenocarcinoma) GBM (Glioblastoma Multiforme) KIRC (Kidney Renal Clear Cell) LGG (Lower Grade Glioma) LIHC (Liver Hepatocellular) LUAD (Lung Adenocarcinoma) UCEC (Uterine Corpus Endometrial) HNSC (Head & Neck Squamous) OV (Ovarian Serous Cystadenocarcinoma)

1,082 430 580 534 511 371 506 541 525 570

14.0 22.2 80.6 33.1 24.5 35.6 36.1 16.5 35.1 59.7

Total

5,650

34.4

the implementation, I is defined by the intersection of patients with both clinical and survival annotations, after which all omics matrices are reindexed to a common patient order. Likewise, let U = {g1 , . . . , gM } denote the aligned gene universe, where M = 7,595 genes are induced by the curated gene–pathway mapping used throughout the study. Expression, CNV, and mutation matrices are reindexed to this common gene order before pathwaylevel graph instantiation. Missing modality-specific entries are preserved during alignment and converted to zero only at the final tensor-materialization stage, which avoids mixing patient alignment with imputation. Clinical covariates were compressed into a numerically stable feature vector containing age, sex, and aggregated stage variables, with cancer-specific markers retained when clinically standard information was available, such as receptor status and PAM50 subtype in BRCA. All preprocessing parameters were estimated within training folds to avoid information leakage. Detailed preprocessing, cohort descriptors, and leakage-prevention rules are provided in Supplementary Sections S1, S3, and S4.

2.2. Overview of PathMoG PathMoG follows a pathway-first pipeline. Multi-omics inputs are first mapped onto KEGG pathway graphs, then gene expression is modulated by genomic and clinical context through HOM, pathway-specific node embeddings are updated with heterogeneous graph message passing, pathway vectors are aggregated through gene-to-pathway and pathway-to-patient attention, and the final patient representation is fused with clinical covariates for Cox-based survival prediction. A key implementation feature is that the model does not operate on a single fixed graph tensor. Instead, it uses a nested representation in which pathway topologies are precomputed once and each patient is instantiated as a variable-length collection of pathway graphs plus one clinical vector. This hierarchical data structure is central to both the computational tractability and the biological interpretability of the framework.

2.3. Pathway-centric modular graph construction The central design choice in PathMoG is to replace a single global gene graph with pathway modules. This modular design introduces an explicit biological prior, constrains the search space, and reduces spurious interactions in the p ≫ n regime, rather than asking the model to infer all genome-wide dependencies from limited patient samples (5; 13; 12). In this sense, pathway modularity acts as an

3

PathMoG: A Pathway-Centric Modular GNN

Fig. 1: Workflow overview of PathMoG. The framework combines pathway-centric graph construction, HOM-based omics modulation, heterogeneous message passing within each pathway module, dual-level attention, and Cox-based survival prediction with clinical fusion. inductive bias for survival modeling rather than as a visualization convenience. Formally, let P = {1, . . . , K} denote the pathway catalogue with K = 354 modules. The implementation first constructs a patient-invariant library of local pathway topologies. For pathway k ∈ P, the static graph is  (r) Gktopo = Vk , {Ek }r∈Rk , where Vk ⊆ U is the set of genes belonging to pathway k after intersection with the global aligned gene universe, and (r)

Ek

2.4. Hierarchical Omics Modulation (HOM) Rather than treating multi-omics inputs as interchangeable channels, we model gene expression as the target signal modulated by genomic and clinical context. This reflects the biological intuition that mutation and CNV states act as regulatory conditions that shape the functional meaning of transcript abundance, while pathway identity and patient context provide higher-order information about the state in which a gene is being interpreted. In the implemented model, HOM conditions each gene on three nested contexts. For gene u ∈ Vk in patient i, define the local genomic context

= {(u, v) : u, v ∈ Vk , rel(u, v) = r}

collects all directed edges of relation type r whose two endpoints both remain inside the pathway. Relation labels are normalized from the curated pathway interaction file and stored explicitly as typed edges in HeteroData. Pathways retaining fewer than two mapped genes after alignment are not instantiated. This static precomputation step fixes a pathway-local node ordering and avoids rebuilding the same biological topology for every patient. For patient i ∈ I and pathway k, omics measurements are then injected into the precomputed topology through a pathway-local feature matrix Xk,i ∈ R|Vk |×3 ,

cnv mut Xk,i (u, :) = [xexpr u,i , xu,i , xu,i ],

where the column order follows the implementation exactly: expression, CNV, then mutation. If all three modalities are absent for every gene in pathway k for patient i, that pathway is omitted from the patient’s package; otherwise missing entries are mapped to 0 only when the tensor is materialized. The resulting patient-level input is therefore not a monolithic graph but a hierarchical object  Di = {Gk,i }k∈Ki , ci ,

mut 2 mu,k,i = [xcnv u,i , xu,i ] ∈ R ,

the encoded patient context zi = fclin (ci ), and the pathway context m̄k,i =

1 X mu,k,i , |Vk | u∈V

where ek is a learnable embedding representing pathway identity. HOM then forms the context vector ctxu,k,i = [mu,k,i ; pk,i ; zi ] and feeds it to a modulation network. Following a FiLM-style parameterization (21), the network predicts raw coefficients (γ̂u,k,i , β̂u,k,i ), which are range-controlled in the implementation as γu,k,i = 1 + 0.5 tanh(γ̂u,k,i ),

where Ki ⊆ P is the set of pathway graphs instantiated for patient i, Gk,i carries the typed pathway edges together with node features Xk,i , and ci is the processed clinical feature vector. Supplementary Section S4 provides the full graph-construction rules, patient-package definition, and code-level batching details.

pk,i = fpath ([ek ; m̄k,i ]),

k

βu,k,i = 2 tanh(β̂u,k,i ),

so that the modulated expression remains numerically stable while still being patient- and pathway-specific. The actual modulation is expr x̃expr u,k,i = γu,k,i xu,i + βu,k,i .

4

PathMoG: A Pathway-Centric Modular GNN

The initial node representation passed to the graph encoder is then the fused vector (0)

expr cnv mut hu,k,i = [x̃expr u,k,i , xu,i , xu,i , xu,i , γu,k,i , βu,k,i , ek , m̄k,i ],

which has dimension 16 in the current implementation. Hence, HOM does not merely rescale one channel; it creates an explicit, inspectable representation in which modulation coefficients, pathway identity, and pathway state all remain accessible to downstream analysis. Full mathematical details are provided in Supplementary Section S4.2.

to pathways and then to genes without padding every patient to a fixed number of pathways, which is essential because different patients instantiate different subsets of pathway modules after modality-aware filtering.

2.6. Survival objective and evaluation protocol Model parameters were optimized with Cox partial likelihood and L2 regularization,  L(θ) = −

 h(ℓ+1) = σ v

θi − log

i:Ei =1

2.5. Heterogeneous graph learning and dual-level attention Within each pathway module, PathMoG applies Heterogeneous Graph Transformer (HGT) layers to account for typed regulatory edges such as activation, inhibition, and phosphorylation (7). Each instantiated pathway graph contains one node type (gene) and multiple relation types inherited from the pathway interaction (ℓ) file. If hv denotes the embedding of gene v at layer ℓ, then the relation-aware update can be written abstractly as

X

 X

θj 

e

+ λ∥W∥22 ,

j∈R(ti )

where θi is the predicted risk score for patient i, Ei is the event indicator, and R(ti ) is the risk set at time ti (2). We used stratified 5-fold cross-validation within each TCGA cohort. The concordance index (C-index) served as the primary metric, with time-dependent AUC used as a secondary metric. Detailed implementation, evaluation metrics, and reproducibility settings are provided in Supplementary Sections S2–S4.

 X

X

, α(ℓ,r) Wr(ℓ) h(ℓ) uv u

3. Results

r∈Rk u∈N (r) (v)

where N (r) (v) denotes the neighbors of v connected by relation (ℓ,r) type r, and the attention coefficients αuv depend jointly on source state, target state, and relation type. In the current implementation, two HGT layers are used inside each pathway module. After message passing, node embeddings are summarized through two attention stages. For pathway k in patient i, gene-to-pathway attention computes (L)

au,k,i = wg⊤ tanh(Wg hu,k,i ), αu,k,i = P

exp(au,k,i ) v∈Vk exp(av,k,i )

,

and forms the pathway vector gk,i =

X

(L)

αu,k,i hu,k,i .

u∈Vk

Pathway-to-patient attention then aggregates the variable-length pathway set {gk,i }k∈Ki by ⊤ bk,i = wp tanh(Wp gk,i ),

πk,i = P Hi =

exp(bk,i ) ℓ∈Ki exp(bℓ,i )

X

,

πk,i gk,i .

k∈Ki

This representation is fused with encoded clinical features and passed to a multilayer perceptron to produce the final survival risk score. The corresponding batching strategy is also hierarchical. A custom collate function concatenates all pathway graphs from all patients in a minibatch into a single PyG batch, while retaining three mappings: the gene-to-pathway batch index generated automatically by PyG, an explicit pathway-to-patient map, and the global pathway index vector used to retrieve pathway embeddings. These mappings allow patient-level clinical context to be broadcast

3.1. Benchmark performance We benchmarked PathMoG against eight representative methods under the same stratified 5-fold protocol: Cox-PH, RSF, GraphSurv (Wang et al., BIBM 2021), CAMR, PathGNN (Liang et al., BMC Bioinformatics 2022), HFBSurv, FGCNSurv, and PCLSurv. This comparison spans classical survival analysis, deep survival learning, multi-omics integration, and graph-based survival modeling, providing a broad view of how PathMoG behaves relative to established baselines. All methods were trained under a consistent evaluation protocol (Supplementary Sections S2–S4). Under this controlled benchmark, PathMoG delivered the best overall discrimination performance across TCGA cohorts (Table 2), ranking first in C-index in all 10 cancer types and showing especially strong gains in GBM, UCEC, BRCA, and COAD. These results provide direct evidence that pathway-constrained modular learning is more effective than unconstrained global interaction modeling when prognostic signals are noisy, heterogeneous, and distributed across biological programs. The superiority of PathMoG also transferred to external testing: a model trained on TCGA BRCA maintained significant risk separation in METABRIC without fine-tuning, indicating that the learned representations generalize beyond a single cohort distribution. Across TCGA cohorts, the PathMoG score likewise achieved significant high-versus-low risk stratification in Kaplan– Meier analysis (all log-rank p < 0.05), supporting clinical utility beyond point-estimate ranking metrics. Full external-validation results and cohort-wise Kaplan–Meier curves are provided in Supplementary Sections S5 and S6, respectively. Having established the overall performance advantage of PathMoG, we next examined which architectural choices produced that gain.

3.2. Architecture validation To explain why PathMoG performs well, we next tested whether pathway modularity acts as a genuine regularizer rather than a cosmetic architectural choice.

5

PathMoG: A Pathway-Centric Modular GNN

Table 2. Concordance Index (C-index) and Time-dependent AUC across 10 TCGA cancer types. Per-cohort entries are mean test-set metrics from 5-fold cross-validation, and the Overall row reports sample-size-weighted averages across the 10 cohorts. PathMoG ranks first in C-index for all 10 cohorts and in AUC for 8 of 10 cohorts. Cancer

C-index/AUC

N Cox

RSF

CAMR

PathG

GraphS

FGCN

HFB

PCL

PathMoG

BRCA

1082 0.712/0.663 0.695/0.642 0.667/0.717 0.685/0.776 0.640/0.687 0.756/0.690 0.671/0.627 0.760/0.740 0.784/0.811

KIRC

534 0.726/0.626 0.718/0.619 0.702/0.756 0.658/0.675 0.687/0.652 0.742/0.738 0.574/0.581 0.748/0.765 0.751/0.780

GBM

580 0.520/0.550 0.542/0.573 0.659/0.650 0.569/0.605 0.553/0.592 0.610/0.592 0.563/0.443 0.654/0.665 0.724/0.810

LUAD

506 0.607/0.547 0.615/0.560 0.651/0.676 0.605/0.660 0.555/0.618 0.605/0.702 0.584/0.563 0.683/0.630 0.709/0.757

COAD

430 0.517/0.546 0.538/0.567 0.650/0.660 0.597/0.667 0.504/0.623 0.650/0.700 0.673/0.477 0.650/0.680 0.694/0.699

LGG

511 0.713/0.738 0.702/0.727 0.704/0.732 0.692/0.718 0.711/0.732 0.720/0.725 0.698/0.735 0.715/0.748 0.753/0.814

HNSC

525 0.590/0.673 0.582/0.655 0.595/0.672 0.600/0.658 0.585/0.672 0.605/0.660 0.588/0.638 0.602/0.650 0.610/0.668

OV

570 0.564/0.617 0.572/0.625 0.645/0.671 0.576/0.637 0.617/0.671 0.609/0.598 0.594/0.548 0.600/0.680 0.629/0.691

UCEC

541 0.503/0.531 0.519/0.548 0.592/0.624 0.568/0.585 0.503/0.602 0.598/0.602 0.582/0.532 0.600/0.623 0.680/0.714

LIHC

371 0.626/0.598 0.642/0.615 0.625/0.661 0.647/0.653 0.597/0.638 0.652/0.580 0.524/0.563 0.640/0.650 0.654/0.666

Overall (weighted) 5650 0.618/0.615 0.620/0.616 0.651/0.686 0.625/0.674 0.601/0.653 0.664/0.662 0.612/0.576 0.675/0.689 0.708/0.751

Cox: Cox-PH; RSF: random survival forest; CAMR: cross-modal attention; PathG: PathGNN (Liang et al., BMC Bioinformatics 2022); GraphS: GraphSurv (Wang et al., BIBM 2021); FGCN: FGCNSurv; HFB: HFBSurv; PCL: PCLSurv. AUC: Area Under ROC Curve at median survival time. PathMoG achieves the highest C-index in all 10 cancer types. $EODWLRQ6WXG\&RPSRQHQW&RQWULEXWLRQ$QDO\VLV$FURVV7&*$&DQFHU7\SHV

3.2.2. Component ablation Beyond the modular graph design, we evaluated which internal components drive performance once pathway modularity is fixed. We therefore tested four ablations across all 10 TCGA cohorts (Figure 2): removing HOM, removing intra-pathway attention,



&OLQLFDO7KUHVKROG 

)XOO0RGHO

ZR+20

ZR,QWUD$WWQ

ZR,QWHU$WWQ

6LPSOH3RRO





&,QGH[

3.2.1. Modular versus monolithic graph design We first compared PathMoG with a monolithic graph model built on the same 7,595 pathway genes but merged into a single topology. Under this matched comparison, the difference between the two designs lies not in the gene set itself but in the structural constraints imposed on learning. In the monolithic model, all genes are embedded within a single global interaction network, allowing message passing to propagate across distant and potentially unrelated interactions. Because each patient then has only one graph instance, the pathwaylevel aggregation stage is removed in this baseline and the single graph representation is fed directly to the patient-level survival head. In contrast, PathMoG organizes the same genes into biologically curated pathway modules, restricting information propagation to functionally coherent neighborhoods defined by pathway topology. This modular organization introduces a biologically informed inductive bias that constrains the hypothesis space of the model. Rather than exploring genome-scale interactions indiscriminately, PathMoG limits representation learning to pathway-structured contexts, which helps suppress spurious correlations that frequently arise in high-dimensional survival modeling (p ≫ n). Consistent with this hypothesis, the monolithic baseline underperformed overall on a sample-size-weighted basis (C-index 0.643 vs. 0.708; +10.1% relative gain for PathMoG; p < 0.001), with the largest performance gaps observed in GBM and BRCA (Figure 3C). Detailed cohort-wise comparisons and the exact monolithic-baseline construction (including removal of pathway-level aggregation) are provided in Supplementary Section S8. Taken together, these results suggest that the advantage of PathMoG does not arise from reducing model size, but from replacing an unconstrained genome-scale interaction space with a biologically structured learning framework.









%5&$

&2$'

*%0

.,5&

/**

/,+&

/8$'

+16&

29

8&(&

Fig. 2: Ablation study of PathMoG components. Most ablations reduce performance relative to the full model; cohortspecific exceptions are reported from traceable result files.

removing inter-pathway attention, and replacing dual attention with simple pooling. Full model remains strongest overall. The complete model outperformed ablated variants in most cohorts, indicating complementary contributions from modular design, HOM, and hierarchical attention. Intra-pathway attention contributes most among modular components. Within the ablation benchmark, replacing gene-level attention with mean pooling produced the largest samplesize-weighted drop among non-graph ablations (weighted C-index 0.655 vs. 0.708 for the full model; absolute drop 0.053, relative -7.5%), with especially large declines in BRCA (-24.0%), GBM (-8.1%), and LUAD (-7.4%). This pattern indicates that weighting genes within each pathway is critical for capturing patient-specific molecular heterogeneity. Component importance is cancer specific. KIRC and LIHC were most sensitive to inter-pathway attention, LGG depended most on HOM, and UCEC/LUAD showed more balanced dependence across components. These differences support the view that PathMoG learns disease-specific multi-omics structure rather than a single universal pattern. Together with the modularversus-monolithic comparison, these results explain why PathMoG achieves stronger benchmark performance.

6

PathMoG: A Pathway-Centric Modular GNN Monolithic global graph vs. PathMoG modular pathway grid

A

B

Traditional Global Graph

C

PathMoG's Modular Pathway Grid JAK-STAT...

Citrate cycle...

Hedgehog...

MAPK signaling

Cell cycle

TGF-beta...

p53 signaling

Wnt signaling

Notch signaling

Performance comparison

0.9 PathMoG Monolithic

C-index

0.8 0.7 0.6 0.5 0.4

BRCA

COAD

GBM

KIRC

LGG

LIHC

LUAD

HNSC

OV

UCEC

Fig. 3: Comparison of monolithic and modular graph designs. (A) Traditional global graph architecture with all 7,595 pathway genes merged into a single topology. (B) PathMoG’s modular pathway grid organizing the same genes into biologically curated pathway modules. (C) Performance comparison showing PathMoG consistently outperforms the monolithic baseline across all 10 TCGA cancer types. Table 3. Univariate and multivariate Cox regression for BRCA. The PathMoG risk score remains independently prognostic after adjustment for routine clinical variables. Univariate

Multivariate

Variable

HR (95% CI)

P-value

HR (95% CI)

P-value

Risk score (per SD) Age (per SD) Grade (G3/G4 vs G1/G2) T stage (T3/T4 vs T1/T2) N stage (N1–N3 vs N0) M stage (M1 vs M0)

1.62 (1.47–1.78) 1.53 (1.30–1.80) 1.66 (1.07–2.58) 1.51 (1.03–2.21) 1.81 (1.30–2.52) 4.81 (2.88–8.03)

<0.001*** <0.001*** 0.023* 0.036* <0.001*** <0.001***

1.57 (1.20–2.07) 1.20 (0.95–1.53) 1.10 (0.53–2.29) 0.93 (0.51–1.73) 1.21 (0.77–1.90) 0.77 (0.26–2.26)

0.001** 0.130 0.794 0.828 0.414 0.628

Fig. 4: Treatment response across PathMoG risk strata in BRCA. Benefit from treatment is strongest in the mediumand high-risk groups, whereas the low-risk group shows weaker separation.

3.4. Treatment stratification 3.3. Clinical prognostic analysis Having established both benchmark gains and their architectural basis, we next asked whether the PathMoG risk score captured information beyond conventional clinical staging. BRCA was used as the main case study because it provided the largest cohort and the clearest downstream treatment-stratification analysis. We fit univariate and multivariate Cox models including age, grade, and TNM variables together with the model-derived risk score. The PathMoG score was significant in univariate analysis and remained independently prognostic after clinical adjustment (Table 3). Notably, T stage and N stage were no longer significant after the molecular risk score was added, indicating that PathMoG captures prognostic variation not explained by anatomical staging alone. This is the core clinical-statistical claim of the paper: the learned risk score is not a restatement of routine clinicopathological variables.

We then evaluated whether the PathMoG risk score could support treatment-oriented stratification in BRCA. Patients were divided into low-, medium-, and high-risk groups according to the modelderived score, and treatment benefit was compared between treated and untreated patients within each subgroup. Kaplan–Meier subgroup curves are shown in Figure 4. Treatment benefit was strongest in the medium-risk group (HR = 0.43, p = 0.008) and remained significant in the highrisk group (HR = 0.53, p = 0.004), whereas the low-risk group showed only borderline benefit (HR = 0.49, p = 0.061). This pattern suggests that the PathMoG score may help distinguish patients more likely to benefit from systemic therapy from those whose outcomes may not justify equally aggressive intervention. At the same time, this analysis should be interpreted cautiously. It is retrospective, potentially confounded by treatment-selection bias, and should be treated as hypothesis-generating rather than

PathMoG: A Pathway-Centric Modular GNN

7

Fig. 5: Multi-level interpretability analysis of PathMoG. (A) Top key driver genes identified by attention weights in cell cycle and p53 pathways. (B) Cross-cancer comparison of gene importance between BRCA and KIRC. (C) Heatmap of patient heterogeneity showing distinct gene expression patterns. (D) Statistical validation correlating attention weights with differential expression (log2 Fold Change). (E) Cross-cancer pathway attention heatmap revealing universal and specific pathway dysregulations. (F-G) Top ranked pathways in BRCA and their statistical significance (volcano plot). (H) PathMoG Molecular Diagnostic Report illustrating personalized risk assessment, dominant genes, active pathways, and survival prognosis for representative patients. causal evidence. The value of this section is therefore not to claim treatment guidance as a solved problem, but to show that the learned risk score aligns with clinically meaningful subgroup structure.

3.5. Biological interpretability Finally, PathMoG supports comprehensive multi-level interpretability through its dual-level attention mechanism, enabling insights at the gene, pathway, and patient levels (Figure 5). To keep the main text concise, we show a compact integrative panel here, while expanded interpretability evidence is provided in Supplementary Section S11. 3.5.1. Gene- and pathway-level insights By analyzing the learned gene-level attention weights (intrapathway attention, αg,pw,i ), we extracted nonlinear featureimportance rankings. Although attention weights should not be treated as a complete explanation on their own (10), they provide a useful starting point when combined with differential expression analysis. In the p53 signaling pathway, pro-apoptotic genes such as BAX, APAF1, and TP53 received the highest attention weights, whereas E2F3, CCNE1, and CDK4 emerged as dominant drivers in the Cell Cycle pathway. These findings are consistent with established cancer biology and suggest that

PathMoG captures survival-relevant dysregulation rather than arbitrary feature salience (24; 1). Cross-cancer comparison further revealed disease-specific molecular signatures. For example, EGFR and AKT1 received higher attention weights in BRCA than in KIRC, consistent with the stronger contribution of receptor tyrosine kinase signaling to breast cancer progression. Statistical validation within the interpretability panel also showed that high attention-weighted genes, including MDM2, RB1, and ERBB2, were strongly aligned with differential expression patterns, reinforcing the biological relevance of the extracted features (25). At the pathway level, PathMoG provides a macroscopic view of cancer progression through pathway attention weights (βpw,i ). Across cohorts, Cell Cycle and p53 signaling appeared as recurrent pan-cancer pathways, whereas BRCA showed stronger emphasis on HIF-1, MAPK, and mTOR signaling. This pattern suggests that the model captures both shared oncogenic programs and tissue-specific pathway dependencies, which is exactly the type of structured readout expected from a pathway-centric design (19; 11). 3.5.2. Patient-level insights At the patient level, PathMoG can consolidate model outputs into an individualized molecular diagnostic profile. For high-risk patients, the model highlights dominant genes, active pathways, and the survival context associated with the predicted score;

8 for lower-risk patients, the same report exposes comparatively quiescent pathway activity and less aggressive molecular signatures. In the original interpretability analysis, this patient-level view was illustrated with a representative report in which elevated EGFR, CDK2, and AKT1 expression co-occurred with strong HIF-1 and PI3K-Akt pathway activity, whereas low-risk patients showed comparatively weaker activation. This individualized readout is important because it links prediction to actionability. Rather than returning only a scalar risk estimate, PathMoG offers a route toward patient-specific explanation: which genes dominate the risk score, which pathways are most active, and how the patient sits within the survival landscape implied by the cohort. That combination of gene-level, pathway-level, and patient-level explanation is the core reason we keep the interpretability module in the main manuscript instead of reducing it to a supplementary-only note. Detailed interpretability analysis, including cross-cancer pathway heatmaps, cancer-specific pathway rankings, pathway interaction networks, pathway-to-gene decompositions, patientlevel diagnostic reports, driver-gene recurrence summaries, and additional gene-level statistical validation, is provided in Supplementary Section S11.

PathMoG: A Pathway-Centric Modular GNN conceptual guidance, and contributed to manuscript revision. C.P.T., J.X.K., J.X.Z., and M.Y.T. participated in data collection and experimental validation.

Funding Statement This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.

Key Points •

• •

4. Conclusion PathMoG addresses multi-omics survival prediction with a pathway-centric graph design that is explicitly tuned to the p ≫ n regime. By replacing a monolithic genome-scale graph with pathway modules, modulating expression with genomic and clinical context, and aggregating molecular evidence through hierarchical attention, the model achieves strong benchmark performance while preserving clinically relevant and biologically interpretable outputs. Several limitations remain. The present implementation relies mainly on KEGG pathways, uses three omics layers, and currently has external validation from a single major breast cancer cohort. Future work should extend the framework to broader pathway resources, additional molecular modalities, richer external cohorts, and prospective validation settings.

5. Availability and implementation Data availability: TCGA multi-omics and clinical data were obtained from the UCSC Xena Pan-Cancer resource (4). External validation used the METABRIC cohort. Code availability: Source code for preprocessing, model training, and figure generation is available at https://github. com/wangzoyou/pathmog.

6. Supplementary data Supplementary data are organized as Sections S1–S12 and are provided separately. They contain the detailed implementation notes, external validation analyses, all-cohort Kaplan–Meier curves, modularity analyses, additional ablation statistics, and extended interpretability figures referenced throughout the main text.

Author Contributions D.W. conceptualized the study, developed the methodology, implemented the models, performed the data analysis, and drafted the manuscript. T.C.L. provided overall research supervision,

PathMoG reframes multi-omics survival prediction as a pathway-centric modular graph problem rather than a genomescale monolithic graph problem. The HOM module uses mutation, CNV, pathway, and clinical context to modulate gene expression instead of simply concatenating omics features. Across 10 TCGA cancer types, PathMoG achieves a samplesize-weighted mean C-index of 0.708 and the highest C-index in all 10 cohorts. The model-derived risk score remains independently prognostic in BRCA after adjustment for routine clinical variables. Extended evidence, including external validation, all-cohort KM curves, design-choice analyses, and richer interpretability panels, is reorganized into a targeted Supplementary file.

References 1. Justin M. Balko, Jacqueline M. Giltnane, Kuan-Hung Wang, Lindsay J. Schwarz, Christopher D. Young, Rebecca S. Cook, Philip Owens, Melinda E. Sanders, Marta G. Kuba, Maciej Szklarczyk, Mindy Red-Brewer, Ana M. Gonzalez-Angulo, Gordon B. Mills, Joseph A. Pinto, Hernan Gomez, and Carlos L. Arteaga. Molecular profiling of the residual disease of triple-negative breast cancers after neoadjuvant chemotherapy identifies actionable therapeutic targets. Cancer Discovery, 4(2):232–245, 2014. 2. Michael J. Bradburn, Timothy G. Clark, Simon B. Love, and Douglas G. Altman. Survival analysis part ii: multivariate data analysis – an introduction to concepts and methods. British Journal of Cancer, 89(3):431–436, 2003. 3. D. R. Cox. Regression models and life-tables. Journal of the Royal Statistical Society: Series B (Methodological), 34(2):187– 220, 1972. 4. Mary J. Goldman, Brian Craft, Mark Hastie, Kristupas Repecka, Fan McDade, Akhil Kamath, Anup Banerjee, Yuming Luo, David Rogers, Angela N. Brooks, Jingchun Zhu, and David Haussler. Visualizing and interpreting cancer genomics data via the xena platform. Nature Biotechnology, 38(6):675–678, 2020. 5. Leland H. Hartwell, John J. Hopfield, Stanislas Leibler, and Andrew W. Murray. From molecular to modular cell biology. Nature, 402(6761):C47–C52, 1999. 6. Katherine A. Hoadley, Christina Yau, Toshinori Hinoue, Denise M. Wolf, Alexander J. Lazar, Elizabeth Drill, Ronglai Shen, Allyson M. Taylor, Andrew D. Cherniack, Vésteinn Thorsson, Rehan Akbani, Reanne Bowlby, Christina K. Wong, Maciej Wiznerowicz, et al. Cell-of-origin patterns dominate the molecular classification of 10,000 tumors from 33 types of cancer. Cell, 173(2):291–304.e6, 2018.

PathMoG: A Pathway-Centric Modular GNN 7. Ziniu Hu, Yuxiao Dong, Kuansan Wang, and Yizhou Sun. Heterogeneous graph transformer. In Proceedings of The Web Conference 2020, pages 2704–2710, 2020. 8. Shengli Huang, Kumardeep Chaudhary, and Lana X. Garmire. More is better: recent progress in multi-omics data integration methods. Frontiers in Genetics, 8:84, 2017. 9. Hemant Ishwaran, Udaya B. Kogalur, Eugene H. Blackstone, and Michael S. Lauer. Random survival forests. The Annals of Applied Statistics, 2(3):841–860, 2008. Attention is not 10. Sarthak Jain and Byron C. Wallace. explanation. In Proceedings of NAACL-HLT 2019, pages 3543–3556, 2019. 11. Boren Jiang, Xiaopan Lin, Xie Ning, Ji Du, Xi Luo, Liyue Wang, and Jiyue Wang. Camr: cross-aligned multimodal representation learning for cancer survival prediction. Bioinformatics, 39:btad025, 2023. 12. Minoru Kanehisa, Miho Furumichi, Yoko Sato, Masayuki Kawashima, and Mari Ishiguro-Watanabe. Kegg for taxonomybased analysis of pathways and genomes. Nucleic Acids Research, 51(D1):D587–D592, 2023. 13. Minoru Kanehisa and Susumu Goto. Kegg: Kyoto encyclopedia of genes and genomes. Nucleic Acids Research, 28:27–30, 2000. 14. Wei Lan, Siyuan Liu, Lei Zhang, Han Liang, Yuhang Deng, Yang Ji, Xing Zhou, and Jinping Xu. Deepkegg: a graph neural network framework to capture cell heterogeneity from single-cell and spatial transcriptomics data. Briefings in Bioinformatics, 25(3):bbae185, 2024. 15. Ruilong Li, Xiang Wang, Xiaoyi Hu, Han Jin, Di Wu, and Lei Chen. Hfbsurv: hierarchical multimodal fusion with factorized bilinear models for cancer survival prediction. Bioinformatics, 38(9):2587–2594, 2022. 16. Xiangmei Li, Bingyue Pan, Yalan He, Zhixuan Wang, Yu Tang, Ya Zhang, Lei Wang, and Junwei Han. Pathhdnn: a pathway hierarchical-informed deep neural network framework for predicting immunotherapy response and mechanism interpretation. Genome Medicine, 17:152, 2025. 17. Zheng Li, Weifeng Shi, Zhijie Wang, Ling Li, Yanhui Deng, Xin Liu, Xuelian Luo, Fan Jiang, and Jian Wang. Pclsurv: a prototypical contrastive learning-based multi-omics data integration model for cancer survival prediction. Briefings in Bioinformatics, 26(2):bbaf124, 2025. 18. B. Liang, H. Gong, L. Lu, and J. Xu. Risk stratification and pathway analysis based on graph neural network and interpretable algorithm. BMC Bioinformatics, 23(1):394, 2022. 19. Andriy Marusyk, Vanessa Almendro, and Kornelia Polyak. Intra-tumour heterogeneity: a looking glass for cancer. Nature Reviews Cancer, 12(5):323–334, 2012. 20. Craig H. Mermel, Steven E. Schumacher, Barbara Hill, Matthew L. Meyerson, Rameen Beroukhim, and Gad Getz. Gistic2.0 facilitates sensitive and confident localization of the targets of focal somatic copy-number alteration in human cancers. Genome Biology, 12(4):R41, 2011. 21. Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron Courville. Film: visual reasoning with a general conditioning layer. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32, 2018. 22. Rebecca L. Siegel, Kimberly D. Miller, Nikita S. Wagle, and Ahmedin Jemal. Cancer statistics, 2023. CA: A Cancer Journal for Clinicians, 73:17–48, 2023. 23. Christopher D. Steele, Sandeep Kumar, Ross J. Edmondson, et al. Signatures of copy number alterations in human cancer.

9 Nature, 606(7916):984–991, 2022. 24. Ryan C. Taylor, Stephen P. Cullen, and Seamus J. Martin. Apoptosis: controlled demolition at the cellular level. Nature Reviews Molecular Cell Biology, 9(3):231–241, 2008. 25. Meredith Wade, Yi-Chin Li, and Geoffrey M. Wahl. Mdm2, mdmx and p53 in oncogenesis and cancer therapy. Nature Reviews Cancer, 13(2):83–96, 2013. 26. Yi Wang, Zhongyue Zhang, Hua Chai, and Yuedong Yang. Multi-omics cancer prognosis analysis based on graph convolution network. In 2021 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), pages 1564–1568, 2021. 27. Guangyu Wen, Ling Li, Jian Wang, and Fan Jiang. Fgcnsurv: fine-grained cancer survival prediction by jointly learning feature representation and feature relations from multi-omics data. Bioinformatics, 39(8):btad472, 2023. 28. Hongyan Yan, Xiaobo Yang, Yuting Wang, Jian Wang, Yun Pan, and Xiaohua Hu. Prior knowledge-guided multilevel graph neural network model for cancer survival prediction from multi-omics data. Briefings in Bioinformatics, 25(3):bbae184, 2024. 29. Bo Yang, Chenxi Cui, Meng Wang, Hong Ji, and Feiyue Gao. Multi-view multi-level contrastive graph convolutional network for cancer subtyping on multi-omics data. Briefings in Bioinformatics, 26:bbaf043, 2025.

Record · ID 138949 · SHA-256 6eea8a71b320b2e1
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.