ConceptioArchivearXiv CS
arXiv CSopen access

Benchmarking Pathology Foundation Models for Breast Cancer Survival Prediction

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Benchmarking Pathology Foundation Models for Breast Cancer Survival Prediction Fredrik K. Gustafsson1,2 Constance Boissin1 Johan Vallon-Christersson3 2,4 David A. Clifton Mattias Rantalainen∗,1

arXiv:2604.24679v1 [cs.CV] 27 Apr 2026

1

Department of Medical Epidemiology and Biostatistics, Karolinska Institutet, Stockholm, Sweden 2 Department of Engineering Science, University of Oxford, Oxford, UK 3 Division of Oncology, Department of Clinical Sciences Lund, Lund University, Lund, Sweden 4 Oxford Suzhou Centre for Advanced Research, University of Oxford, Suzhou, China ∗

Corresponding Author, [email protected]

Abstract

pathology foundation models (PFMs) [4, 10, 11, 22, 25, 37, 44, 48]. Trained on hundreds of thousands to millions of whole-slide images (WSIs), these models aim to learn generalizable visual representations that can be transferred across a wide range of downstream tasks, including tumor grading [8, 47], biomarker prediction [1, 33], and patient prognosis [16, 18, 43]. This paradigm shift mirrors the success of foundation models in natural language processing and computer vision [2, 6, 31], positioning PFMs as a key component of modern computational pathology workflows.

Pathology foundation models (PFMs) have recently emerged as powerful pretrained encoders for computational pathology, enabling transfer learning across a wide range of downstream tasks. However, systematic comparisons of these models for clinically meaningful prediction problems remain limited, especially in the context of survival prediction under external validation. In this study, we benchmark widely used and recently proposed PFMs for breast cancer survival prediction from whole-slide histopathology images. Using a standardized pipeline based on patch-level feature extraction and a unified survival modeling framework, we evaluate model representations across three independent clinical cohorts comprising more than 5,400 patients with long-term follow-up. Models are trained on one cohort and evaluated on two independent external cohorts, enabling a rigorous assessment of cross-dataset generalization. Overall, H-optimus-1 achieves the strongest survival prediction performance. More broadly, we observe consistent generational improvements across model families, with second-generation PFMs outperforming their first-generation counterparts. However, absolute performance differences between many recent PFMs remain modest, suggesting diminishing returns from further scaling of pretraining data or model size alone. Notably, the compact distilled model H0-mini slightly outperforms its larger teacher model H-optimus-0, despite using fewer than 8% of the parameters and enabling significantly faster feature extraction. Together, these results provide the first largescale, externally validated benchmark of PFMs for breast cancer survival prediction, and offer practical guidance for efficient deployment of PFMs in clinical workflows.

While numerous PFMs have recently been proposed, their relative performance on clinically meaningful prediction tasks remains incompletely understood. Recent benchmarking efforts have begun to address this gap, but with certain limitations. Marza et al. [29] provide a comprehensive comparison of more than 20 models but exclusively focus on patch-level analysis, which limits relevance for downstream clinical prediction tasks. Kasireddy et al. [20] evaluate PFMs on both patch- and slide-level kidney pathology tasks, but the slide-level cohorts remain relatively small (at most ≈200 WSIs per task). Breen et al. [7] provide a rigorous single-task evaluation of PFMs for ovarian cancer subtyping, but focus exclusively on a classification setting rather than broader clinical endpoints. Similarly, Bareja et al. [3] conduct a large-scale comparison of models across various datasets and tasks, but restrict their evaluation to classification-based endpoints, without considering time-toevent modeling for prognosis. Campanella et al. [9] present a large-scale benchmark of publicly available PFMs across multiple datasets, but their evaluation is limited to binary detection and biomarker prediction tasks, and relies on internal cross-validation rather than systematic external validation across independent cohorts, limiting insight into real-world generalization performance. Neidlinger et al. [32] perform a comprehen-

Computational pathology has recently undergone a rapid transformation driven by the emergence of large-scale 1

C-index (↑)

0.70 0.65 0.60 0.55 0.50

RFS – All Patients

RFS – Patient Subgroup (ER+ & HER2-)

PFS – All Patients

PFS – Patient Subgroup (ER+ & HER2-)

C-index (↑)

0.70 0.65 0.60 0.55 0.50 Resnet-IN

CTransPath

RetCCL

UNI

UNI2-h

H-optimus-0

H-optimus-1

H0-mini

Prov-GigaPath

Virchow

Virchow2

CONCH

CONCHv1.5

Figure 1. Main model comparison across evaluation settings, showing performance in terms of C-index (↑) for all thirteen evaluated models. Models are evaluated for recurrence-free survival (RFS) and progression-free survival (PFS), each assessed both for the full cohort (‘All Patients’) and the ‘ER+ & HER2-’ patient subgroup. Bars show the bootstrap mean C-index with 95% confidence intervals. Table 1. Main model comparison with numerical results and model rankings across evaluation settings, reporting performance in terms of C-index (bootstrap mean with 95% confidence intervals) for all thirteen evaluated models. Models are separately ranked based on the bootstrap mean within each of the four evaluation settings. Rank

Model Name

1 2 3 4 5 6 6 8 9 10 11 12 13

H-optimus-1 H0-mini Virchow2 UNI2-h H-optimus-0 CONCH Prov-GigaPath CONCHv1.5 UNI Virchow RetCCL CTransPath Resnet-IN

C-index (↑) 0.678 0.676 0.675 0.667 0.664 0.663 0.663 0.660 0.657 0.656 0.645 0.632 0.612

(0.656 – 0.700) (0.653 – 0.698) (0.653 – 0.698) (0.644 – 0.689) (0.642 – 0.686) (0.641 – 0.685) (0.640 – 0.686) (0.637 – 0.683) (0.635 – 0.680) (0.633 – 0.679) (0.622 – 0.668) (0.609 – 0.655) (0.588 – 0.635)

Rank

Model Name

1 2 3 4 5 6 6 8 9 10 11 12 13

H-optimus-1 H0-mini Virchow2 UNI2-h H-optimus-0 CONCH Prov-GigaPath CONCHv1.5 UNI Virchow RetCCL CTransPath Resnet-IN

(a) RFS – All Patients.

Rank

Model Name

1 2 2 4 5 6 7 8 9 10 11 12 13

H-optimus-1 Virchow CONCHv1.5 H0-mini H-optimus-0 UNI2-h Virchow2 CONCH Prov-GigaPath UNI CTransPath Resnet-IN RetCCL

0.670 0.668 0.665 0.663 0.660 0.657 0.657 0.651 0.648 0.639 0.636 0.627 0.600

(0.645 – 0.695) (0.642 – 0.693) (0.639 – 0.690) (0.637 – 0.688) (0.633 – 0.685) (0.631 – 0.681) (0.631 – 0.682) (0.625 – 0.676) (0.622 – 0.674) (0.613 – 0.665) (0.609 – 0.662) (0.601 – 0.652) (0.572 – 0.627)

(b) RFS – Patient Subgroup (ER+ & HER2-).

C-index (↑) 0.702 0.695 0.695 0.692 0.686 0.683 0.681 0.675 0.673 0.671 0.647 0.628 0.627

C-index (↑)

(0.666 – 0.737) (0.660 – 0.729) (0.660 – 0.729) (0.655 – 0.728) (0.651 – 0.722) (0.647 – 0.719) (0.645 – 0.718) (0.639 – 0.710) (0.636 – 0.710) (0.634 – 0.708) (0.609 – 0.686) (0.590 – 0.666) (0.589 – 0.666)

(c) PFS – All Patients.

Rank

Model Name

1 2 3 3 5 6 7 8 9 10 11 12 13

CONCHv1.5 Virchow H-optimus-1 H-optimus-0 H0-mini UNI2-h CONCH Virchow2 Prov-GigaPath UNI CTransPath Resnet-IN RetCCL

C-index (↑) 0.680 0.678 0.670 0.670 0.666 0.657 0.656 0.655 0.641 0.638 0.622 0.606 0.604

(0.636 – 0.721) (0.634 – 0.720) (0.626 – 0.715) (0.624 – 0.714) (0.619 – 0.711) (0.611 – 0.702) (0.612 – 0.700) (0.608 – 0.700) (0.594 – 0.687) (0.592 – 0.683) (0.573 – 0.670) (0.557 – 0.655) (0.555 – 0.653)

(d) PFS – Patient Subgroup (ER+ & HER2-).

2

Table 3. Overview of all thirteen evaluated models, including architecture, model size (number of parameters), feature dimension, and pretraining data scale. The models span a natural-image baseline, two early pathology-specific models, seven state-of-the-art PFMs, a compact distilled PFM, and two vision-language PFMs.

Table 2. Main model ranking, aggregated as the mean model rank (↓) across the four evaluation settings based on Table 1. Rank

Model Name

Model Ranks

Mean Model Rank (↓)

1 2 3 4 5 6 7 8 9 10 11 12 13

H-optimus-1 H0-mini H-optimus-0 CONCHv1.5 UNI2-h Virchow2 Virchow CONCH Prov-GigaPath UNI CTransPath RetCCL Resnet-IN

1,1,1,3 2,2,4,5 5,5,5,3 8,8,2,1 4,4,6,6 3,3,7,8 10,10,2,2 6,6,8,7 6,6,9,9 9,9,10,10 12,12,11,11 11,11,13,13 13,13,12,12

1.5 3.25 4.5 4.75 5 5.25 6 6.75 7.5 9.5 11.5 12 12.5

sive multi-cohort evaluation across a wide range of weakly supervised tasks, including prognostic endpoints. However, these are formulated as binary classification problems rather than time-to-event modeling, thereby not fully capturing the complexity of survival prediction. Moreover, many of their evaluations are conducted in relatively small or low-sample settings. PathBench [26] is a more comprehensive framework that extends benchmarking across multiple cancer types and a wide spectrum of tasks, including diagnosis, molecular prediction, and survival prediction. However, survival prediction is evaluated exclusively using cross-validation, without independent external validation, and with moderate cohort sizes for all considered cancer types (at most 451 patients for breast cancer, 260 for gastric cancer, and 608 for colorectal cancer, across different endpoints). While external datasets are used for several diagnostic and classification tasks in PathBench, they are not applied to survival prediction. Consequently, despite these important advances, there remains a lack of large-scale, multi-cohort evaluations specifically targeting survival prediction from WSIs under rigorous external validation. Existing benchmarks provide only a partial view of model utility, since clinically relevant applications require robust generalization across different patient populations, institutions, and scanning conditions [4, 17, 23, 41]. Survival prediction from histopathology images represents a particularly challenging and clinically important task: prognostic signals are often subtle, spatially heterogeneous, and confounded by clinical and biological variability. At the same time, accurate survival prediction has clear potential for improving risk stratification, treatment planning, and clinical decision support. Systematic evaluation of PFMs in this setting is therefore essential for understanding their practical value. To address this gap, we benchmark a diverse set of widely used and recently proposed PFMs (Table 3) for breast cancer survival prediction from WSIs. Our eval-

Model Name

Architecture

Size

Feature Dimension

Pretraining Data

Resnet-IN [15]

ResNet-50

25M

1024

1.3M natural images

CTransPath [45] RetCCL [46]

CNN + Swin-T ResNet-50

22M 25M

768 2048

30K WSIs 32K WSIs

UNI [10] UNI2-h [28] H-optimus-0 [37] H-optimus-1 [5] Prov-GigaPath [48] Virchow [44] Virchow2 [49]

ViT-L ViT-H ViT-G ViT-G ViT-G ViT-H ViT-H

307M 682M 1.1B 1.1B 1.1B 632M 632M

1024 1536 1536 1536 1536 2560 2560

100K WSIs 350K WSIs 500K WSIs 1M WSIs 170K WSIs 1.5M WSIs 3.1M WSIs

H0-mini [13]

ViT-B

86M

768

500K + 6K WSIs

CONCH [25] CONCHv1.5 [11]

ViT-B ViT-L

86M 307M

512 768

21K WSIs + 1.1M image-text pairs

N/A

uation is conducted on three independent clinical cohorts comprising 5,434 patients with long-term follow-up. Models are trained on one cohort and evaluated on two independent external cohorts, enabling a rigorous assessment of cross-dataset generalization. To the best of our knowledge, this constitutes the largest multi-cohort benchmark for histopathology-based survival prediction with independent external validation. The benchmark spans multiple generations of pathology representation learning, including early pathology-specific models trained using selfsupervised learning, and state-of-the-art PFMs trained on more than one million WSIs. Our study makes four main contributions: (1) We present a large-scale, multi-cohort benchmark for breast cancer survival prediction, leveraging over 5,400 patients with independent external validation to enable robust assessment of model generalization. (2) We provide a systematic head-to-head evaluation of a natural-image baseline, early pathology-specific models, state-of-the-art PFMs, and vision-language PFMs. (3) We characterize performance trends across model families, showing consistent generational improvements but only modest absolute gains, suggesting diminishing returns from further scaling of pretraining data or model size alone. (4) We demonstrate that compact distilled models can match or even exceed the prognostic risk-stratification performance of significantly larger teacher models, highlighting knowledge distillation as a promising approach for efficient deployment of PFMs.

Results We evaluate thirteen representative models spanning a natural-image baseline (Resnet-IN), two early pathologyspecific models (CTransPath [45], RetCCL [46]), seven state-of-the-art PFMs (Prov-GigaPath [48], UNI [10], UNI2-h [28], Virchow [44], Virchow2 [49], H-optimus0 [37], H-optimus-1 [5]), a compact distilled PFM (H03

mini [13], distilled from H-optimus-0), and two visionlanguage PFMs (CONCH [25], CONCHv1.5 [11]). Models are evaluated using a unified pipeline for WSIbased survival prediction (see Methods for details). Each model is used as a frozen feature extractor to compute patch-level feature vectors for the given WSI, which are aggregated into a predicted patient-level risk score using the PANTHER [40] survival modelling framework. Survival models are trained on a dataset of 2,315 patients (SöS-BC4) and evaluated on two independent external datasets (KSSolna and SCAN-B-Lund) comprising 3,119 patients in total. None of the evaluated PFMs were pretrained on any of these datasets, ensuring a strict separation between pretraining data and the benchmark. Performance is measured using the concordance index (C-index) and Kaplan-Meier survival analysis. Models are evaluated under four complementary settings: recurrencefree survival (RFS) and progression-free survival (PFS), each assessed both for the full cohort (‘All Patients’) and for the clinically relevant ‘ER+ & HER2-’ patient subgroup. Model rankings are computed for each of the four settings and aggregated to obtain a final overall ranking. The combined evaluation set of 3,119 patients contains 615 RFS events and 233 PFS events, with 2,524 patients (80.9% of the full set), 475 RFS events (77.2%) and 157 PFS events (67.4%) in the ‘ER+ & HER2-’ patient subgroup.

1, CONCHv1.5, UNI2-h, Virchow2) slightly outperforms its corresponding first-generation counterpart (H-optimus0, CONCH, UNI, Virchow) in the aggregated ranking. Despite these consistent trends, absolute performance differences between many models are relatively small. Confidence intervals for the C-index estimates show substantial overlap across most models. For example, in the RFS – All Patients setting, the top-ranked model exhibits nonoverlapping 95% confidence intervals only when compared to the two lowest-ranked models (CTransPath and ResnetIN). A similar pattern is observed for RFS – Patient Subgroup, where non-overlapping intervals are only observed for the lowest-ranked model. For PFS, confidence intervals overlap across all models. Overall, these results indicate that differences in predictive performance between many models are modest despite consistent ranking patterns. While model rankings are broadly consistent across the two survival endpoints (RFS and PFS), some variability is observed. In particular, CONCHv1.5 and Virchow rank quite low for RFS (8th and 10th in both settings, respectively) but substantially higher for PFS (top-two positions in both settings). Virchow2 also shows a shift in ranking, achieving 3rd place for RFS but 7th and 8th place for PFS. These deviations contrast with the otherwise stable ordering of models across endpoints, indicating that relative performance can vary depending on the specific survival task.

Main Model Comparison

Survival Analysis Using the Kaplan-Meier Estimator

Figure 1 shows the C-index performance of all thirteen models across the four evaluation settings, with corresponding numerical results and model rankings in Table 1. The aggregated overall ranking is summarized in Table 2. H-optimus-1 achieves the highest C-index in three out of four settings (RFS – All Patients, RFS – Patient Subgroup, PFS – All Patients) and attains the best overall model ranking. In absolute terms, this corresponds to a C-index of 0.678 for RFS – All Patients and 0.702 for PFS – All Patients. Notably, the compact distilled model H0-mini achieves the second-best overall ranking, and slightly outperforms its teacher model H-optimus-0 in three out of four settings. CONCHv1.5, UNI2-h and Virchow2 also demonstrate strong overall performance. In contrast, the natural-image baseline Resnet-IN achieves the lowest overall performance and ranks in the bottom two across all four settings. The two early pathology-specific models CTransPath and RetCCL also perform poorly, consistently occupying the bottom three ranks together with Resnet-IN. Among recent PFMs, UNI achieves the lowest overall performance, only beating Resnet-IN, CTransPath and RetCCL. Across model families, a consistent pattern is observed in which each second-generation PFM (H-optimus-

Kaplan-Meier (KM) survival curves illustrate a survival function while accounting for right-censoring, and is used to demonstrate the ability of models to stratify patients into distinct risk groups. Figure 2 shows stratification into two risk groups across all four evaluation settings, for ResnetIN, UNI, and H-optimus-1. For each model and setting, the KM plots report log-rank test p-values, along with the number of patients at risk and the number of events over time (0-14 years). These three models are selected to represent the overall best-performing model (H-optimus-1), the overall worstperforming model (Resnet-IN), and the worst-performing model among PFMs (UNI), enabling direct visual comparison of risk stratification quality. KM plots for the remaining PFMs show broadly similar patterns across settings, with differences between models which are less pronounced than those observed between UNI and H-optimus-1. Across all three models, consistent trends are observed in the degree of separation between risk groups across evaluation settings. The strongest separation is observed for RFS – All Patients, followed by RFS – Patient Subgroup, PFS – All Patients, and finally PFS – Patient Subgroup, as reflected by progressively larger log-rank p-values. Within each evaluation setting, differences between models are also evident. 4

0.8 0.7 Low risk score High risk score

0 2 4 6 8 Time (years) Low risk score At risk 1560 1534 1484 1244 973 Events 0 25 74 118 160 High risk score At risk 1559 1500 1406 1227 996 Events 0 57 148 229 313

1.0

12

14

520 189

183 207

9 211

667 360

238 394

19 404

Survival probability

0.8 0.7 Low risk score High risk score

0 2 4 6 8 Time (years) Low risk score At risk 1262 1242 1206 998 769 Events 0 20 55 94 128 High risk score At risk 1262 1230 1160 1011 820 Events 0 30 98 160 230

10

12

14

409 151

146 165

6 169

548 272

203 298

21 306

0.8 0.7 0.6

10

12

14

566 163

191 182

13 185

621 386

230 419

15 430

1.0

0.7 Low risk score High risk score 10

12

14

452 122

155 137

12 139

505 301

194 326

15 336

0 2 4 6 8 10 Time (years) Low risk score At risk 1560 1535 1478 1248 994 608 Events 0 11 36 51 71 78 High risk score At risk 1559 1499 1412 1223 975 579 Events 0 37 81 117 137 147

0.80 12

14

183 80

10 80

238 152

18 153

Survival probability

0.95

Survival probability

0.95

0.85 Low risk score High risk score 12

14

168 66

11 66

253 166

17 167

0.95

0.95

0.80

0 2 4 6 8 10 Time (years) Low risk score At risk 1262 1242 1205 1013 804 498 Events 0 7 22 32 48 54 High risk score At risk 1262 1230 1161 996 785 459 Events 0 15 45 71 88 96

0.90 0.85 0.80

12

14

144 56

8 56

205 100

19 101

Survival probability

0.95

Survival probability

1.00

Low risk score High risk score

Low risk score High risk score

Low risk score High risk score 10

12

14

482 120

159 135

14 136

475 303

190 328

13 339

Low risk score High risk score 12

14

169 59

11 59

252 173

17 174

0.90 0.85 0.80

0 2 4 6 8 10 Time (years) Low risk score At risk 1262 1242 1200 997 800 471 Events 0 6 21 35 43 50 High risk score At risk 1262 1230 1166 1012 789 486 Events 0 16 46 68 93 100

12 442

H-optimus-1 C-index: 0.670 | Logrank test, p = 3.9e-09

1.00

0.85

226 431

0 2 4 6 8 10 Time (years) Low risk score At risk 1560 1535 1489 1245 1001 575 Events 0 5 19 37 52 56 High risk score At risk 1559 1499 1401 1226 968 612 Events 0 43 98 131 156 169

1.00

0.90

591 399

0.85

(c) PFS – All Patients. UNI C-index: 0.638 | Logrank test, p = 2.0e-05

Resnet-IN C-index: 0.606 | Logrank test, p = 1.8e-04

16 173

0.90

0.80

0 2 4 6 8 10 Time (years) Low risk score At risk 1560 1535 1482 1237 987 578 Events 0 7 27 43 57 64 High risk score At risk 1559 1499 1408 1234 982 609 Events 0 41 90 125 151 161

195 170

H-optimus-1 C-index: 0.702 | Logrank test, p = 1.0e-14

0.95

Low risk score High risk score

596 150

H-optimus-1 C-index: 0.670 | Logrank test, p = 7.6e-22

0 2 4 6 8 Time (years) Low risk score At risk 1262 1247 1218 1019 818 Events 0 14 40 69 102 High risk score At risk 1262 1225 1148 990 771 Events 0 36 113 185 256

1.00

0.85

14

0.7

1.00

0.90

12

0.8

(b) RFS – Patient Subgroup (ER+ & HER2-). UNI C-index: 0.671 | Logrank test, p = 1.4e-11

0.90

10

0.9

0.6

0 2 4 6 8 Time (years) Low risk score At risk 1262 1244 1218 1006 779 Events 0 17 41 76 105 High risk score At risk 1262 1228 1148 1003 810 Events 0 33 112 178 253

Low risk score High risk score

0 2 4 6 8 Time (years) Low risk score At risk 1560 1544 1503 1264 1022 Events 0 15 52 88 127 High risk score At risk 1559 1490 1387 1207 947 Events 0 67 170 259 346

(a) RFS – All Patients. UNI C-index: 0.648 | Logrank test, p = 4.6e-19

0.8

H-optimus-1 C-index: 0.678 | Logrank test, p = 4.2e-30

0.9

1.00

0.80

Survival probability

Low risk score High risk score

0.9

0.6

Resnet-IN C-index: 0.628 | Logrank test, p = 6.0e-07 Survival probability

0.7

1.0

0.9

0.6

0.8

0 2 4 6 8 Time (years) Low risk score At risk 1560 1539 1498 1248 982 Events 0 20 59 101 139 High risk score At risk 1559 1495 1392 1223 987 Events 0 62 163 246 334

Resnet-IN C-index: 0.600 | Logrank test, p = 1.2e-08

1.0

0.9

0.6 10

UNI C-index: 0.657 | Logrank test, p = 1.6e-23

Survival probability

0.9

0.6

Survival probability

1.0

Survival probability

Resnet-IN C-index: 0.612 | Logrank test, p = 5.8e-14 Survival probability

Survival probability

1.0

12

14

133 52

9 52

216 104

18 105

Low risk score High risk score

0 2 4 6 8 10 Time (years) Low risk score At risk 1262 1240 1203 1002 805 470 Events 0 5 16 30 38 40 High risk score At risk 1262 1232 1163 1007 784 487 Events 0 17 51 73 98 110

12

14

136 42

11 42

213 114

16 115

(d) PFS – Patient Subgroup (ER+ & HER2-).

Figure 2. Two-group Kaplan-Meier risk stratification for three representative models. KM survival curves showing stratification into low- and high-risk groups for RFS and PFS, each assessed for the full cohort (‘All Patients’) and the ‘ER+ & HER2-’ patient subgroup. Results for Resnet-IN (left column), UNI (middle), and H-optimus-1 (right). Each plot includes the C-index, log-rank test p-value, and the number of patients at risk and events over time (0-14 years). Note the difference in range of the y-axis between RFS and PFS.

5

Resnet-IN | C-index: 0.612

1.0

0.7 0.6

0

2

4

6 8 Time (years)

0.8 0.7 0.6

Low risk score Medium low risk score Medium high risk score High risk score

0.5

0.9 Survival probability

0.8

12

14

0

2

4

0.8 0.7 0.6

Low risk score Medium low risk score Medium high risk score High risk score

0.5 10

H-optimus-1 | C-index: 0.678

1.0

0.9 Survival probability

0.9 Survival probability

UNI | C-index: 0.657

1.0

Low risk score Medium low risk score Medium high risk score High risk score

0.5

6 8 Time (years)

10

12

14

0

2

4

6 8 Time (years)

10

12

14

12

14

12

14

12

14

(a) RFS – All Patients. Resnet-IN | C-index: 0.600

1.0

0.7 0.6

0

2

4

6 8 Time (years)

0.8 0.7 0.6

Low risk score Medium low risk score Medium high risk score High risk score

0.5

0.9 Survival probability

0.8

12

14

0

2

4

0.8 0.7 0.6

Low risk score Medium low risk score Medium high risk score High risk score

0.5 10

H-optimus-1 | C-index: 0.670

1.0

0.9 Survival probability

0.9 Survival probability

UNI | C-index: 0.648

1.0

Low risk score Medium low risk score Medium high risk score High risk score

0.5

6 8 Time (years)

10

12

14

0

2

4

6 8 Time (years)

10

(b) RFS – Patient Subgroup (ER+ & HER2-). UNI | C-index: 0.671

H-optimus-1 | C-index: 0.702

1.00

1.00

0.95

0.95

0.95

0.90 0.85

Low risk score Medium low risk score Medium high risk score High risk score

0.80 0

2

4

6 8 Time (years)

Survival probability

1.00

Survival probability

Survival probability

Resnet-IN | C-index: 0.628

0.90 0.85

Low risk score Medium low risk score Medium high risk score High risk score

0.80 10

12

14

0

2

4

0.90 0.85

Low risk score Medium low risk score Medium high risk score High risk score

0.80

6 8 Time (years)

10

12

14

0

2

4

6 8 Time (years)

10

(c) PFS – All Patients. UNI | C-index: 0.638

H-optimus-1 | C-index: 0.670

1.00

1.00

0.95

0.95

0.95

0.90 0.85

Low risk score Medium low risk score Medium high risk score High risk score

0.80 0

2

4

6 8 Time (years)

Survival probability

1.00

Survival probability

Survival probability

Resnet-IN | C-index: 0.606

0.90 0.85

Low risk score Medium low risk score Medium high risk score High risk score

0.80 10

12

14

0

2

4

0.90 0.85

Low risk score Medium low risk score Medium high risk score High risk score

0.80

6 8 Time (years)

10

12

14

0

2

4

6 8 Time (years)

10

(d) PFS – Patient Subgroup (ER+ & HER2-).

Figure 3. Four-group Kaplan-Meier risk stratification for three representative models. KM survival curves showing stratification into four risk groups for RFS and PFS, each assessed for the full cohort (‘All Patients’) and the ‘ER+ & HER2-’ patient subgroup. Results for Resnet-IN (left column), UNI (middle), and H-optimus-1 (right). Note the difference in range of the y-axis between RFS and PFS.

6

H-optimus-1 consistently shows the clearest separation between low- and high-risk groups, with the smallest log-rank p-values across all settings. UNI exhibits intermediate separation, while Resnet-IN shows the weakest separation, with larger p-values and more overlap between survival curves. Figure 3 shows stratification into four risk groups across the same settings for the three representative models, providing a more detailed view of model behavior. Similar trends are observed across evaluation settings, with the clearest separation for RFS – All Patients and progressively weaker separation for RFS – Patient Subgroup, PFS – All Patients, and PFS – Patient Subgroup. In the latter setting, survival curves show increased overlap and less consistent ordering between adjacent risk groups, particularly at earlier time points. Differences between models remain apparent in this setting. H-optimus-1 shows more distinct and consistently ordered risk groups over time, with clearer separation between curves. UNI exhibits moderate separation with some overlap between adjacent groups, while Resnet-IN shows the least consistent separation, with greater overlap and less stable ordering of risk groups. Notably, H-optimus-1 maintains clear separation across four risk groups in multiple evaluation settings, with well-ordered survival curves and sustained differences in survival probability over time. These patterns are consistent with the two-group analysis (Figure 2) and further highlight differences in the strength and stability of risk stratification across models. Four-group KM plots including the number of patients at risk and the number of events over time are provided in Figure S1 - S3 in the supplementary material, with corresponding plots for three risk groups shown in Figure S4 - S6. Overall, KM analysis confirms that all models capture prognostic signal to some extent, including the lowestperforming model Resnet-IN. Clear differences in risk stratification are observed when comparing the three representative models, with the most pronounced contrast between Hoptimus-1 and Resnet-IN, while differences between PFMs are more subtle. These results complement the C-index analysis by providing a qualitative view of model performance that is not fully captured by ranking metrics alone.

of training data varies, indicating that performance differences are robust across data regimes. Performance improves for all models as the amount of training data increases. While the performance gap between models remains relatively stable across subset sizes, a slight widening of the gap between H-optimus-1 and Resnet-IN is observed. For example, in the RFS – All Patients setting, H-optimus-1 improves from a C-index of 0.578 at 10% of the training data to 0.674 at 100%, corresponding to a relative improvement of 16.6%. In comparison, Resnet-IN improves from 0.533 to 0.612 (14.7% relative improvement). A similar trend is observed for PFS – All Patients, where H-optimus-1 improves from 0.597 to 0.696 (16.6%), while Resnet-IN improves from 0.562 to 0.627 (11.6%). Comparable patterns are observed across the remaining settings. Across all evaluation settings, H-optimus-1 trained on 25% of the data achieves performance matching or exceeding that of Resnet-IN trained on the full dataset. Similarly, H-optimus-1 trained on 50% of the data matches or exceeds the performance of UNI trained on 100% of the data. Figure 5 shows the same analysis of C-index performance across subset sizes for the three highest-ranked models overall: H-optimus-1, H0-mini, and H-optimus-0. These models exhibit very similar performance across all subset sizes and evaluation settings. Notably, the compact distilled model H0-mini consistently performs on par with, and in most cases even slightly better than, its teacher model Hoptimus-0. Overall, the performance curves of these three models are closely aligned across all settings, with only minor differences in C-index values. Finally, to further investigate the structure of learned feature representations, Figure 6 shows UMAP [30] projections of the mean patch-level feature vectors for ResnetIN, UNI, and H-optimus-1 on the combined KS-Solna and SCAN-B-Lund evaluation set, with one point per patient. In the upper row, points are colored by dataset (KS-Solna vs SCAN-B-Lund), while in the lower row, points are colored according to survival outcome. Specifically, points are colored red for patients with a PFS event within three years, blue for patients with no PFS/RFS event and a follow-up time of at least 12 years, and gray otherwise. Clear clustering by dataset is observed for both UNI and H-optimus-1 in the first two UMAP components, indicating that these models capture systematic differences between WSIs from different cohorts, while Resnet-IN shows no clear separation of the two datasets. In contrast, clustering by survival outcome is not clearly observed for any of the models, suggesting that survival-related variability is not dominating the variability in the data. Despite the substantially better survival prediction performance of H-optimus-1 compared to Resnet-IN, early-event and event-free patients are not clearly separated in the feature space, as visualized by the first two UMAP components, of either model.

Extended Model Analysis Figure 4 shows C-index performance across all four evaluation settings for Resnet-IN, UNI, and H-optimus-1 as a function of the amount of training data used from the SöSBC-4 dataset (10%, 25%, 50%, 75%, 100%). Across all subset sizes, consistent trends are observed in both absolute performance and relative ranking. H-optimus-1 achieves the highest performance across all settings and training set sizes, followed by UNI and Resnet-IN. Importantly, the relative ordering of models remains unchanged as the amount 7

0.70

0.65

0.65 C-index (↑)

C-index (↑)

0.70

0.60

H-optimus-1 UNI Resnet-IN

0.55

0.50

10

25 50 75 % of SöS-BC-4 used for training

0.55

0.50

100

(a) RFS – All Patients. 0.70

0.70

0.65

0.65

0.60

0.55

0.50

10

25 50 75 % of SöS-BC-4 used for training

100

(b) RFS – Patient Subgroup (ER+ & HER2-).

C-index (↑)

C-index (↑)

0.60

0.60

0.55

10

25 50 75 % of SöS-BC-4 used for training

0.50

100

(c) PFS – All Patients.

10

25 50 75 % of SöS-BC-4 used for training

100

(d) PFS – Patient Subgroup (ER+ & HER2-).

Figure 4. Effect of training data size on survival prediction performance for three representative models. C-index performance for Resnet-IN, UNI, and H-optimus-1 as a function of the fraction of training data used (10%, 25%, 50%, 75%, 100%) across all four evaluation settings (RFS and PFS, ‘All Patients’ and ‘ER+ & HER2-’). Results are reported as mean ± standard deviation (std) over five random seeds, where a new random subset of training patients is sampled for each seed at each data fraction. The PANTHER prototype estimation step is performed once on the full SöS-BC-4 dataset for each seed, after which the survival model is retrained on the corresponding subset.

Discussion

A particularly notable finding is the strong performance of the compact distilled model H0-mini, which slightly outperforms its teacher model H-optimus-0 while using fewer than 8% of the parameters and enabling significantly faster feature extraction. This result is consistent across both fulldata and subset analyses (Figure 1 & 5), where H0-mini matches or exceeds H-optimus-0 across evaluation settings. Given the substantial computational cost associated with WSI processing, this highlights knowledge distillation as a highly promising strategy for developing efficient PFMs without sacrificing predictive performance. Notably, H0mini is distilled using a small public dataset of just 6,000 WSIs, further emphasizing the potential of distillation to transfer representational knowledge efficiently. Taken together, these findings suggest that compact models such as H0-mini offer a favorable trade-off between performance and efficiency, and may represent a strong candidate for resource-constrained settings. Across model families, we observe consistent generational improvements, with second-generation PFMs

In this study, we present a large-scale, multi-cohort benchmark of PFMs for breast cancer survival prediction, providing a systematic comparison across widely used models. Several key insights emerge from our results. First, while H-optimus-1 achieves the strongest overall performance, the absolute differences between topperforming models are relatively small. Across all four evaluation settings, multiple recent PFMs including Hoptimus-1, H0-mini, H-optimus-0, CONCHv1.5, UNI2h, and Virchow2 achieve C-index values within a narrow range, with substantially overlapping confidence intervals. This suggests that although architectural and training improvements lead to consistent ranking gains, these improvements are incremental rather than transformative. From a practical perspective, this implies that model selection should not be based solely on marginal performance differences, but also on factors such as computational efficiency, robustness, and access to models. 8

0.70

0.65

0.65 C-index (↑)

C-index (↑)

0.70

0.60

H-optimus-0 H-optimus-1 H0-mini

0.55

0.50

10

25 50 75 % of SöS-BC-4 used for training

0.55

0.50

100

0.70

0.65

0.65 C-index (↑)

0.70

0.60

0.55

0.50

10

25 50 75 % of SöS-BC-4 used for training

100

(b) RFS – Patient Subgroup (ER+ & HER2-).

(a) RFS – All Patients.

C-index (↑)

0.60

0.60

0.55

10

25 50 75 % of SöS-BC-4 used for training

0.50

100

(c) PFS – All Patients.

10

25 50 75 % of SöS-BC-4 used for training

100

(d) PFS – Patient Subgroup (ER+ & HER2-).

Figure 5. Effect of training data size on survival prediction performance for the three top-performing models. C-index performance for H-optimus-0, H-optimus-1, and H0-mini as a function of the fraction of training data used (10%, 25%, 50%, 75%, 100%) across all four evaluation settings. The experimental setup is identical to Figure 4, with results reported as mean ± std over five random seeds.

(H-optimus-1, CONCHv1.5, UNI2-h, Virchow2) outperforming their respective first-generation counterparts (Hoptimus-0, CONCH, UNI, Virchow) according to the aggregated ranking. However, these gains remain modest despite substantial increases in pretraining scale. For example, Hoptimus-1 is trained on twice as many WSIs as H-optimus-0 (1 million vs 0.5 million WSIs, Table 3), and Virchow2 similarly scales pretraining data compared to Virchow, yet both yield only incremental improvements in downstream performance. This suggests that scaling pretraining data alone may yield diminishing returns. Similarly, model size alone is not a reliable predictor of performance. Despite being among the very largest models evaluated (1.1 billion parameters), Prov-GigaPath places only 9th in the overall ranking, while substantially smaller models such as H0-mini and CONCH (86 million parameters) achieve competitive or superior performance. This further reinforces the notion that aspects such as pretraining data quality and training strategy are more critical than raw model capacity. Our results also provide tentative evidence regarding vision-language pretraining. CONCH and CONCHv1.5

perform competitively for prognostic stratification despite relatively smaller model sizes, suggesting that incorporating image-caption pairs and multimodal alignment objectives in the pretraining may improve representation learning efficiency. However, it remains unclear whether these gains arise from the multimodal training paradigm itself or from the scale and diversity of the underlying pretraining data. Furthermore, given the limited number of visionlanguage models evaluated and the lack of controlled comparisons, no definitive conclusions can be drawn. Future work should explore this direction more systematically, for example through ablation studies comparing vision-only and multimodal pretraining under controlled settings. Consistent with expectations, the natural-image baseline Resnet-IN and early pathology-specific models (CTransPath, RetCCL) perform consistently worse than recent PFMs, confirming the importance of large-scale domain-specific pretraining. At the same time, it is notable that even Resnet-IN achieves non-trivial performance and can stratify patients into two risk groups (Figure 2), indicating that some prognostic morphological signals are sufficiently generic to be captured by natural-image features. 9

Resnet-IN

UNI KS-Solna patients SCAN-B-Lund patients

11

KS-Solna patients SCAN-B-Lund patients

8

8

6

10 9

H-optimus-1

10

6

4

8 2

4

7 0

6

2

5

KS-Solna patients SCAN-B-Lund patients

2 2

0

2

4

6

8

4

2

Resnet-IN

2

4

UNI Other patients No PFS/RFS event after 12 years PFS event within 3 years

11

0

Other patients No PFS/RFS event after 12 years PFS event within 3 years

8

9

6

8

10

H-optimus-1

10

8

6

10

4

6

4

8 2

4

7 0

6

2

5

Other patients No PFS/RFS event after 12 years PFS event within 3 years

2 2

0

2

4

6

8

4

2

0

2

4

4

6

8

10

Figure 6. UMAP visualizations of learned feature representations for three representative models. UMAP projections of the mean patch-level feature vectors for Resnet-IN (left column), UNI (middle), and H-optimus-1 (right) on the combined KS-Solna and SCAN-BLund evaluation set, with one point per patient. Upper row: points are colored by dataset, green for KS-Solna and orange for SCANB-Lund. Lower row: points are colored by survival outcome, red for patients with a PFS event within three years (82 patients), blue for patients with no PFS/RFS event and a follow-up time of at least 12 years (407 patients), and gray for all other patients (2,630 patients).

The data efficiency analysis (Figure 4 & 5) provides additional insight into model behavior. Performance improves steadily with increasing training data across models, with similar relative gains observed for both strong and weak feature extractors. This suggests that downstream survival modeling benefits consistently from increased supervision, largely independent of the choice of encoder. At the same time, stronger PFMs maintain a consistent performance advantage across all data regimes, and notably achieve competitive performance even with limited training data. For example, H-optimus-1 trained on 25% of the data matches or exceeds the performance of Resnet-IN trained on the full dataset. These findings indicate that both better representations and larger labeled datasets contribute independently to improved survival prediction, and that further gains may be achievable by scaling either dimension.

However, while Kaplan-Meier analysis shows that all models capture prognostic signal to some extent, they differ in the strength and stability of risk stratification. In particular, the advantage of PFMs becomes more pronounced in more demanding settings, such as multi-group risk stratification (Figure 3), where models like H-optimus-1 produce more clearly separated and consistently ordered survival curves. Notably, the ability to stratify patients into four distinct risk groups is non-trivial and provides a more stringent test of model quality than binary stratification. The fact that topperforming PFMs maintain clear separation in this setting suggests that they capture clinically meaningful gradations of risk. Despite broadly consistent rankings across survival endpoints, we observe some task-dependent variability. CONCHv1.5 and Virchow rank relatively low for RFS but among the top performers for PFS. One likely explanation is the smaller number of PFS events relative to RFS, leading to higher variance and wider confidence intervals. These findings may also reflect differences in the types of prognostic signals captured by different models, although the current results are not sufficient to draw firm conclusions. Regardless, these results highlight the importance of evaluating models across multiple clinically relevant endpoints.

From a data perspective, an important implication is that increasing the number of observed survival events, rather than simply the number of patients, is likely critical for improving model performance, particularly for endpoints such as PFS where event counts remain relatively low. This reflects the fact that time-to-event models rely primarily on observed events for learning when optimizing the Cox proportional hazards loss. Consequently, sufficiently long 10

follow-up time, together with larger studies or aggregation of multiple cohorts, is expected to contribute to improved performance. Finally, analysis of the learned feature space using UMAP (Figure 6) reveals that while PFMs capture datasetspecific structure, clear separation by survival outcome is not observed. Notably, even though models such as Hoptimus-1 substantially outperform Resnet-IN in survival prediction, this improvement is not reflected in obvious clustering patterns. This suggests that prognostic signals are likely subtle, distributed, and not easily separable in low-dimensional projections, particularly when compared to more pronounced site- or cohort-specific differences. Overall, our findings highlight both the progress and current limitations of pathology foundation models for survival prediction. While recent PFMs provide consistent improvements over earlier approaches, performance gains are mostly incremental, and multiple models achieve comparable results. In this context, efficiency emerges as a key consideration, particularly in resource-constrained settings, with compact distilled models such as H0-mini potentially offering a compelling balance between performance and scalability. This study has several limitations. Although multiple independent datasets were used for evaluation, all cohorts originate from similar healthcare systems within a single country (Sweden) and may not fully capture global variability in clinical practice, staining protocols, or scanner hardware. In addition, all experiments are conducted on breast cancer cohorts, and results may not generalize to other cancer types or disease settings, where different morphological patterns and prognostic signals may influence model performance and relative rankings. We evaluate models using frozen feature extractors, within a unified pipeline based on PANTHER aggregation and an MLP survival head, to ensure a fair and controlled comparison across models. While this isolates the quality of pretrained representations, results and model rankings may differ under task-specific fine-tuning, end-to-end training, or alternative survival modeling approaches. The relatively small performance differences observed between many state-of-theart PFMs, together with overlapping confidence intervals, suggest that further scaling of pretraining data or model size alone may yield diminishing returns. Alternative directions such as vision-language pretraining and knowledge distillation appear promising, with our results providing especially strong evidence for the effectiveness of distillation, but both require more systematic investigation in controlled settings. Distillation may be particularly valuable in resource-constrained settings due to its favorable efficiency–performance trade-off. At the same time, further work is needed to assess how well these gains generalize across tasks and applications, as distillation may in some

cases lead to a loss of more fine-grained or task-specific information in the learned representations, even if such effects are not observed in the present evaluation. In particular, the strong performance of H0-mini, distilled using a small dataset of just 6,000 WSIs, raises important questions about the role of distillation data scale and composition, which should be explored in future work. The main takeaways from our study are: (1) H-optimus1 achieves the strongest overall performance, but absolute differences between top-performing PFMs are small and confidence intervals substantially overlap. While architectural and training improvements lead to consistent performance gains, these improvements are incremental rather than transformative. (2) Across model families, consistent generational improvements are observed, with secondgeneration PFMs (H-optimus-1, CONCHv1.5, UNI2-h, Virchow2) outperforming their respective first-generation counterparts. However, these gains remain modest despite substantial increases in pretraining data scale, suggesting diminishing returns from scaling alone. (3) Model size is not a reliable predictor of performance, as smaller and more efficient models can match or exceed much larger architectures, emphasizing the importance of training strategy and pretraining data quality over model scaling. (4) The compact distilled model H0-mini achieves the second-best overall ranking and slightly outperforms its much larger teacher model H-optimus-0, while using less than 8% of the parameters and enabling significantly faster feature extraction. This demonstrates that knowledge distillation can yield highly efficient PFMs without sacrificing predictive performance, making it a particularly promising approach for resource-constrained settings. (5) Strong pretrained feature extractors and large labeled datasets contribute independently to improved survival prediction performance, with high-quality PFMs maintaining advantages also in low-data regimes. Increasing the number of observed survival events, for example through longer follow-up time, is likely to be key for further improvements.

Methods We benchmark pretrained PFMs for WSI-based breast cancer survival prediction using a unified experimental setup. The workflow consists of WSI preprocessing and patch extraction, patch-level feature extraction using frozen pretrained models, and slide-level survival prediction using the prototype-based aggregation approach PANTHER [40]. Survival models are trained on SöS-BC-4 and evaluated on the independent external datasets KS-Solna and SCANB-Lund to assess generalization. The following sections describe the survival endpoints, datasets, preprocessing pipeline, survival model, evaluation framework, and the evaluated PFMs and baselines. 11

Survival Endpoints

Train & Evaluation Sets Models are trained using SöSBC-4 and evaluated on the combined KS-Solna + SCAN-BLund datasets. This combined evaluation set contains 3,119 patients with 615 RFS events and 233 PFS events. Models are evaluated both on the full set and on the ‘ER+ & HER2-’ patient subgroup, consisting of 2,524 patients (80.9% of the full set) with 475 RFS events (77.2%) and 157 PFS events (67.4%). All models are trained on the full set of 2,315 patients in SöS-BC-4.

We evaluate two clinically relevant survival endpoints: recurrence-free survival (RFS) and progression-free survival (PFS). RFS is defined as the time from initial diagnosis to disease recurrence or death from any cause. Recurrence includes local recurrence, distant metastasis, or detection of contralateral tumors. Patients without an event are censored at the date of last follow-up. In contrast, PFS is defined as the time from initial diagnosis to disease recurrence. Death without documented recurrence is not counted as an event and is treated as censoring. Accordingly, RFS captures both recurrence and mortality events, whereas PFS focuses specifically on recurrence events. Evaluating both endpoints provides complementary perspectives: PFS focuses specifically on disease progression, while RFS includes mortality events and therefore yields a larger number of events for statistical analysis.

WSI Processing & Model Overview Each WSI is preprocessed using a standardized workflow. Tissue regions are first identified using Otsu’s thresholding [34], after which non-overlapping image patches of size 256 × 256 pixels are extracted at a resolution of 0.4536 µm/pixel (corresponding to 20× equivalent magnification for a reference slide scanner). Blurry patches are removed using the variance of Laplacian (VL) metric [35], discarding patches with VL < 300. All patches are also color normalized using the Macenko method [27], adapted for WSI-level color correction following Wang et al. [47]. After preprocessing, each WSI x is represented as a set of P image patches {x̃i }P i=1 , where P varies between slides. A pretrained PFM is then used as a frozen feature extractor to compute a feature vector f (x̃i ) for each patch x̃i . These patch-level feature vectors {f (x̃i )}P i=1 are used as input to a slide-level survival model, which aggregates the patch-level features into a single WSI-level feature vector f (x) and predicts a patient risk score r(x).

Datasets We use three independent breast cancer datasets collected at Swedish clinical centers. In total, the three datasets comprise 5,434 patients, each with a corresponding H&Estained WSI and clinical follow-up data. WSIs were generated at 40× magnification from archived clinical routine resected tumor slides, using Hamamatsu NanoZoomer (S360 or XR) or Aperio GT 450 DX whole-slide scanners. SöS-BC-4 This is a retrospective observational cohort including patients diagnosed at Södersjukhuset (South General Hospital) in Stockholm, Sweden between 2012 and 2018 [39, 47], with clinical outcome data retrieved from the Swedish National Registry for Breast Cancer (NKBC) in 2025. We use a subset of 2,315 patients with available follow-up data. This subset contains 351 RFS events and 144 PFS events, with a mean follow-up time of 7.7 years.

Survival Model We perform slide-level survival prediction using PANTHER [40], a prototype-based aggregation approach for constructing compact, fixed-length slide representations from patch-level features. Given the set of P patch-level feature vectors {f (x̃i )}P i=1 extracted from a WSI x, PANTHER fits a Gaussian mixture model to estimate a small set of C = 16 morphological prototypes, where each prototype represents a recurring morphological pattern within the tissue. The parameters of the mixture model (C mixture probabilities, means, and diagonal covariance matrices) summarize the distribution of patch-level features and are concatenated to form a WSI-level representation f (x) of dimension DW SI = C(1 + 2Dp ), where Dp is the patchlevel feature dimension. For example, for a PFM with Dp = 1536, the resulting representation f (x) has dimension DW SI = 49,168. The dimensionality DW SI is constant across WSIs and independent of the number of patches P extracted from each slide. This WSI-level feature vector f (x) captures both the appearance of morphological patterns and their relative abun-

KS-Solna CHIME breast KS-Solna is a populationrepresentative retrospective cohort of primary breast cancer patients treated at the Karolinska University Hospital in Stockholm and diagnosed between 2009 and 2018 [38], with clinical outcome data retrieved from NKBC in 2025. We use a subset of 1,857 patients with available follow-up data, including 389 RFS events and 145 PFS events. The mean follow-up time is 9.1 years. SCAN-B-Lund This is a subset of 1,262 patients enrolled in the prospective SCAN-B study [42], diagnosed between 2010 and 2019 in Lund, Sweden [38, 39]. Clinical followup data was updated in 2025. The subset includes 226 RFS events and 88 PFS events, with a mean follow-up time of 8.1 years. 12

dance within the WSI x, and is used as input to a downstream survival predictor. The survival prediction model is implemented as a structured multilayer perceptron (MLP) head that processes each prototype representation independently before combining them into a final risk prediction, outputting a continuous risk score r(x) for each patient. We train separate survival models for RFS and PFS. Training is performed using the SöS-BC-4 dataset with random 5-fold cross-validation to determine the number of training epochs for each evaluated feature extractor. This is the only hyperparameter that is tuned individually for each feature extractor, while all other survival model hyperparameters are kept fixed to ensure a fair comparison. After selecting the number of training epochs, the model is retrained on the full SöS-BC-4 dataset. The AdamW optimizer [24] with a cosine learning rate schedule is used, optimizing the Cox proportional hazards loss [21]. To improve stability, both the prototype estimation and survival model training of PANTHER are repeated with five random seeds, and the resulting survival prediction models are ensembled by averaging standardized risk scores across seeds. The feature extractors are kept frozen, and only the MLP head of the survival model is updated during training.

evaluation settings based on their bootstrap mean C-index. The final overall model ranking (Table 2) is computed as the mean rank across the four settings, providing a robust comparison of model performance across both survival endpoints and patient populations. In addition, we evaluate the effect of training data size by training survival models on random subsets of the SöSBC-4 dataset comprising 75%, 50%, 25%, and 10% of the training data (Figure 4 & 5). For each subset size, results are reported as the mean ± standard deviation over five random seeds, where a new random subset of training patients is sampled for each seed. The PANTHER prototype estimation step is performed once on the full SöS-BC-4 dataset for each seed, after which the survival model is retrained using the corresponding sampled data subset. This analysis provides insight into the data efficiency and robustness of different PFMs. Evaluated Models We evaluate thirteen pretrained patch-level feature extractors spanning multiple generations of representation learning in computational pathology. Table 3 summarizes the architecture, parameter count, feature dimensionality, and pretraining datasets of all evaluated models.

Evaluation Framework Model performance is primarily evaluated using the concordance index (C-index) [14], which measures the agreement between predicted risk scores and observed survival times. Higher C-index values indicate better alignment between predicted risk ordering and actual patient outcomes. To estimate uncertainty, C-index values are computed using 10,000 bootstrap resamples with 95% confidence intervals. We also perform Kaplan-Meier (KM) [19] survival analysis to assess the ability of models to stratify patients into two or more risk groups (Figure 2 & 3). Patients are divided into risk groups based on predicted risk scores, and differences between groups are assessed using the log-rank test. All models are evaluated under four complementary settings: (1) RFS for the full patient cohort (‘All Patients’). (2) RFS for the subgroup of patients which are oestrogen receptor (ER)-positive and human epidermal growth factor receptor 2 (HER2)-negative (‘ER+ & HER2-’). (3) PFS for ‘All Patients’. (4) PFS for ‘ER+ & HER2-’. The patient subgroup ‘ER+ & HER2-’ represents the most common breast cancer subtype and has distinct biological and clinical characteristics, making it an important subgroup for prognostic modeling. We train separate survival models for RFS and PFS, but all models are trained on the full SöSBC-4 dataset and then evaluated for both ‘All Patients’ and ‘ER+ & HER2-’ on the combined KS-Solna + SCAN-BLund evaluation dataset. Models are ranked separately within each of the four

Natural-Image Baseline We include Resnet-IN in the evaluation, a Resnet-50 [15] model pretrained on the ImageNet dataset [36] of natural images, which has historically been used in early computational pathology pipelines. Resnet-IN serves as a simple reference baseline to quantify the benefit of large-scale pathology-specific pretraining, and is expected to be significantly outperformed by PFMs. Early Pathology Models We also evaluate two early pathology-specific models trained directly on histopathology images using self-supervised learning: CTransPath [45] and RetCCL [46]. Both models were trained on approximately 30,000 WSIs, which is substantially smaller than the datasets used to train state-of-the-art PFMs. These models represent an intermediate stage between generic naturalimage encoders and more recent large-scale PFMs. State-of-the-art PFMs The majority of evaluated models are PFMs trained on large collections of WSIs (ranging from 100,000 to 3.1 million WSIs) using self-supervised learning. These include Prov-GigaPath [48], UNI [10], Virchow [44] and H-optimus-0 [37], as well as their more recent second-generation counterparts UNI2-h [28], Virchow2 [49], and H-optimus-1 [5]. The second-generation models are trained on substantially larger datasets and incorporate improved training strategies, reflecting recent ad13

vances in large-scale representation learning for computational pathology.

2024.0159) from the Knut & Alice Wallenberg Foundation to SciLifeLab for research in Data-Driven Life Science (DDLS) and the Wallenberg AI, Autonomous Systems and Software Program (WASP). DAC was funded by an NIHR Research Professorship; a Royal Academy of Engineering Research Chair; and the InnoHK Hong Kong Centre for Cerebro-cardiovascular Engineering (COCHE); and was supported by the National Institute for Health and Care Research (NIHR) Oxford Biomedical Research Centre (BRC) and the Pandemic Sciences Institute at the University of Oxford. The authors acknowledge patients, clinicians, and hospital staff participating in the SCAN-B study; the staff at the central SCAN-B laboratory at the Division of Oncology, Lund University; the Swedish National Breast Cancer Quality Registry (NKBC); the Regional Cancer Center South; and the South Swedish Breast Cancer Group (SSBCG). SCAN-B was funded by the Swedish Cancer Society, the Mrs. Berta Kamprad Foundation, the Lund-Lausanne L2Bridge/Biltema Foundation, the Mats Paulsson Foundation, and Swedish governmental funding (ALF).

Distilled Model We further evaluate H0-mini [13], a compact distilled variant of H-optimus-0 designed to significantly reduce model size and computational cost while preserving representation quality. H0-mini is a ViT-Base vision transformer [12] with 86 million parameters, corresponding to less than 8% of the original ViT-Giant H-optimus-0 (1.1 billion parameters). The distillation is performed on a dataset of 43 million image patches extracted from 6,093 WSIs from the public TCGA dataset, covering 16 cancer types. This evaluation allows us to assess whether compact distilled models can retain the performance of substantially larger PFMs. Vision-Language PFMs Finally, we include two multimodal vision-language PFMs, CONCH [25] and its secondgeneration version CONCHv1.5 [11]. These models are trained using paired image-text supervision in addition to histopathology WSIs, representing an alternative training paradigm that leverages multimodal data with the aim of learning more generalizable representations.

Author Contributions FKG was responsible for project conceptualization, software implementation, preparation of figures and tables, and manuscript drafting. CB processed the survival datasets. JVC supported the use of the SCAN-B-Lund dataset. DAC acquired funding. MR supervised the project, contributed to the design of experiments and interpretation of results, and acquired funding. All authors contributed to manuscript revision and finalization.

Ethics Statement The study has approval from the Swedish Ethical Review Authority (2017/2106-31, with amendments 2018/1462-32 and 2019–02336). The study was performed in accordance with the Declaration of Helsinki. No additional informed consent was required in accordance with ethical approval in this non-interventional collection and analysis of data from patient records.

Competing Interests

Data Availability

MR is a co-founder and shareholder of Stratipath AB. All other authors declare no competing interests.

The SöS-BC-4, KS-Solna, and SCAN-B-Lund datasets cannot be made publicly available due to restrictions relating to sensitive patient-related information.

Declaration of Generative AI Use During the preparation of this manuscript, the authors used ChatGPT 5.3 to assist with language editing and drafting. All content produced using this tool was critically reviewed, edited and validated by the authors, who take full responsibility for the final content of the manuscript.

Code Availability The code for this study is based on PANTHER which is available at https://github.com/mahmoodlab/ Panther. Further implementation details are available from FKG upon reasonable request.

References

Acknowledgments

[1] Salim Arslan, Julian Schmidt, Cher Bass, Debapriya Mehrotra, Andre Geraldes, Shikha Singhal, Julius Hense, Xiusi Li, Pandu Raharja-Liu, Oscar Maiques, et al. A systematic pan-cancer study on deep learning-based prediction of multiomic biomarkers from routine pathology images. Communications Medicine, 4(1):48, 2024. 1 [2] Bobby Azad, Reza Azad, Sania Eskandari, Afshin Bozorgpour, Amirhossein Kazerouni, Islem Rekik, and Dorit

The project was supported by funding from the Swedish Cancer Society (23 2905 Pj 01 H), VINNOVA (SwAIPP2), Swedish e-science Research Centre (SeRC) (eMPHasis project), Bröstcancerförbundet, Swedish Research Council (2024–06634, 2025-03411), and the AID4BC consortium supported by a WASP/DDLS NEST grant (KAW 14

tions (ICLR), 2021. URL https://openreview.net/ forum?id=YicbFdNTTy. 14 [13] Alexandre Filiot, Nicolas Dop, Oussama Tchita, Auriane Riou, Rémy Dubois, Thomas Peeters, Daria Valter, Marin Scalbert, Charlie Saillard, Geneviève Robin, et al. Distilling foundation models for robust and efficient models in digital pathology. In International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), pages 162–172. Springer, 2025. 3, 4, 14 [14] Frank E Harrell, Robert M Califf, David B Pryor, Kerry L Lee, and Robert A Rosati. Evaluating the yield of medical tests. Jama, 247(18):2543–2546, 1982. 13 [15] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016. 3, 13 [16] Julia Höhn, Eva Krieghoff-Henning, Christoph Wies, Lennard Kiehl, Martin J Hetz, Tabea-Clara Bucher, Jitendra Jonnagaddala, Kurt Zatloukal, Heimo Müller, Markus Plass, et al. Colorectal cancer risk stratification on histological slides based on survival curves predicted by deep learning. npj Precision Oncology, 7(1):98, 2023. 1 [17] Mostafa Jahanifar, Manahil Raza, Kesi Xu, Trinh Thi Le Vuong, Robert Jewsbury, Adam Shephard, Neda Zamanitajeddin, Jin Tae Kwak, Shan E. Ahmed Raza, Fayyaz Minhas, and Nasir Rajpoot. Domain generalization in computational pathology: Survey and guidelines. ACM Computing Surveys, Just Accepted, April 2025. doi: 10.1145/3724391. URL https://doi.org/10.1145/3724391. 3 [18] Xiaofeng Jiang, Michael Hoffmeister, Hermann Brenner, Hannah Sophie Muti, Tanwei Yuan, Sebastian Foersch, Nicholas P West, Alexander Brobeil, Jitendra Jonnagaddala, Nicholas Hawkins, et al. End-to-end prognostication in colorectal cancer by deep learning: a retrospective, multicentre study. The Lancet Digital Health, 6(1):e33–e43, 2024. 1 [19] Edward L Kaplan and Paul Meier. Nonparametric estimation from incomplete observations. Journal of the American statistical association, 53(282):457–481, 1958. 13 [20] Harishwar Reddy Kasireddy, Patricio S La Rosa, Akshita Gupta, Anindya S Paul, Jamie L Fermin, William L Clapp, Meryl A Waldman, Tarek M El-Ashkar, Sanjay Jain, Luis Rodrigues, et al. A comprehensive benchmark of histopathology foundation models for kidney digital pathology images. arXiv preprint arXiv:2603.15967, 2026. 1 [21] Jared L Katzman, Uri Shaham, Alexander Cloninger, Jonathan Bates, Tingting Jiang, and Yuval Kluger. DeepSurv: personalized treatment recommender system using a cox proportional hazards deep neural network. BMC Medical Research Methodology, 18(1):24, 2018. 13 [22] Dong Li, Guihong Wan, Xintao Wu, Xinyu Wu, Ajit J Nirmal, Christine G Lian, Peter K Sorger, Yevgeniy R Semenov, and Chen Zhao. A survey on computational pathology foundation models: Datasets, adaptation strategies, and evaluation tasks. arXiv preprint arXiv:2501.15724, 2025. 1 [23] Weiping Lin, Shen Liu, Runchen Zhu, and Liansheng Wang. Unveiling institution-specific bias in pathology foundation models: Detriments, causes, and potential solutions. arXiv preprint arXiv:2502.16889, 2025. 3

Merhof. Foundational models in medical imaging: A comprehensive survey and future vision. arXiv preprint arXiv:2310.18689, 2023. 1 [3] Rohan Bareja, Francisco Carrillo-Perez, Yuanning Zheng, Marija Pizurica, Tarak Nath Nandi, Jeanne Shen, Ravi Madduri, and Olivier Gevaert. Evaluating vision and pathology foundation models for computational pathology: a comprehensive benchmark study. medRxiv, pages 2025–05, 2025. 1 [4] Mohsin Bilal, Manahil Raza, Youssef Altherwy, Anas Alsuhaibani, Abdulrahman Abduljabbar, Fahdah Almarshad, Paul Golding, Nasir Rajpoot, et al. Foundation models in computational pathology: A review of challenges, opportunities, and impact. arXiv preprint arXiv:2502.08333, 2025. 1, 3 [5] Bioptimus. H-optimus-1, 2025. URL https : / / huggingface.co/bioptimus/H- optimus- 1. 3, 13 [6] Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021. 1 [7] Jack Breen, Katie Allen, Kieran Zucker, Lucy Godson, Nicolas M Orsi, and Nishant Ravikumar. A comprehensive evaluation of histopathology foundation models for ovarian cancer subtype classification. npj Precision Oncology, 9(1):33, 2025. 1 [8] Wouter Bulten, Kimmo Kartasalo, Po-Hsuan Cameron Chen, Peter Ström, Hans Pinckaers, Kunal Nagpal, Yuannan Cai, David F Steiner, Hester Van Boven, Robert Vink, et al. Artificial intelligence for diagnosis and gleason grading of prostate cancer: the PANDA challenge. Nature Medicine, 28(1):154– 163, 2022. 1 [9] Gabriele Campanella, Shengjia Chen, Manbir Singh, Ruchika Verma, Silke Muehlstedt, Jennifer Zeng, Aryeh Stock, Matt Croken, Brandon Veremis, Abdulkadir Elmas, et al. A clinical benchmark of public self-supervised pathology foundation models. Nature Communications, 16(1): 3640, 2025. 1 [10] Richard J Chen, Tong Ding, Ming Y Lu, Drew FK Williamson, Guillaume Jaume, Andrew H Song, Bowen Chen, Andrew Zhang, Daniel Shao, Muhammad Shaban, et al. Towards a general-purpose foundation model for computational pathology. Nature Medicine, 30(3):850–862, 2024. 1, 3, 13 [11] Tong Ding, Sophia J Wagner, Andrew H Song, Richard J Chen, Ming Y Lu, Andrew Zhang, Anurag J Vaidya, Guillaume Jaume, Muhammad Shaban, Ahrong Kim, et al. A multimodal whole-slide foundation model for pathology. Nature Medicine, pages 1–13, 2025. 1, 3, 4, 14 [12] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representa-

15

[24] Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations (ICLR), 2019. URL https : / / openreview.net/forum?id=Bkg6RiCqY7. 13 [25] Ming Y Lu, Bowen Chen, Drew FK Williamson, Richard J Chen, Ivy Liang, Tong Ding, Guillaume Jaume, Igor Odintsov, Long Phi Le, Georg Gerber, et al. A visuallanguage foundation model for computational pathology. Nature Medicine, 30:863–874, 2024. 1, 3, 4, 14 [26] Jiabo Ma, Yingxue Xu, Fengtao Zhou, Yihui Wang, Cheng Jin, Zhengrui Guo, Jianfeng Wu, On Ki Tang, Huajun Zhou, Xi Wang, et al. PathBench: A comprehensive comparison benchmark for pathology foundation models towards precision oncology. arXiv preprint arXiv:2505.20202, 2025. 3 [27] Marc Macenko, Marc Niethammer, James S Marron, David Borland, John T Woosley, Xiaojun Guan, Charles Schmitt, and Nancy E Thomas. A method for normalizing histology slides for quantitative analysis. In 2009 IEEE International Symposium on Biomedical Imaging: from Nano to Macro, pages 1107–1110. IEEE, 2009. 12 [28] MahmoodLab. Uni2-h, 2024. URL https : / / huggingface.co/MahmoodLab/UNI2-h. 3, 13 [29] Pierre Marza, Leo Fillioux, Sofiène Boutaj, Kunal Mahatha, Christian Desrosiers, Pablo Piantanida, Jose Dolz, Stergios Christodoulidis, and Maria Vakalopoulou. THUNDER: Tilelevel Histopathology image UNDERstanding benchmark. In The Thirty-ninth Annual Conference on Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, 2025. 1 [30] Leland McInnes, John Healy, and James Melville. UMAP: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426, 2018. 7 [31] Michael Moor, Oishi Banerjee, Zahra Shakeri Hossein Abad, Harlan M Krumholz, Jure Leskovec, Eric J Topol, and Pranav Rajpurkar. Foundation models for generalist medical artificial intelligence. Nature, 616(7956):259–265, 2023. 1 [32] Peter Neidlinger, Omar SM El Nahhas, Hannah Sophie Muti, Tim Lenz, Michael Hoffmeister, Hermann Brenner, Marko van Treeck, Rupert Langer, Bastian Dislich, Hans Michael Behrens, et al. Benchmarking foundation models as feature extractors for weakly supervised computational pathology. Nature Biomedical Engineering, pages 1–11, 2025. 1 [33] Jan Moritz Niehues, Philip Quirke, Nicholas P West, Heike I Grabsch, Marko van Treeck, Yoni Schirris, Gregory P Veldhuizen, Gordon GA Hutchins, Susan D Richman, Sebastian Foersch, et al. Generalizable biomarker prediction from cancer pathology slides with self-supervised deep learning: A retrospective multi-centric study. Cell Reports Medicine, 4 (4), 2023. 1 [34] Nobuyuki Otsu. A threshold selection method from graylevel histograms. IEEE Transactions on Systems, Man, and Cybernetics, 9(1):62–66, 1979. doi: 10.1109/TSMC.1979. 4310076. 12 [35] Jose Luis Pech-Pacheco, Gabriel Cristobal, Jesus ChamorroMartinez, and Joaquin Fernandez-Valdivia. Diatom autofocusing in brightfield microscopy: a comparative study. In

Proceedings of the 15th International Conference on Pattern Recognition (ICPR), volume 3, pages 314–317. IEEE, 2000. 12 [36] Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International Journal of Computer Vision (IJCV), 115:211–252, 2015. 13 [37] Charlie Saillard, Rodolphe Jenatton, Felipe Llinares-López, Zelda Mariet, David Cahané, Eric Durand, and Jean-Philippe Vert. H-optimus-0, 2024. URL https://github. com/bioptimus/releases/tree/main/models/ h-optimus/v0. 1, 3, 13 [38] Abhinav Sharma, Sandy Kang Lövgren, Kajsa Ledesma Eriksson, Yinxi Wang, Stephanie Robertson, Johan Hartman, and Mattias Rantalainen. Validation of an AI-based solution for breast cancer risk stratification using routine digital histopathology images. Breast Cancer Research, 26(1):123, 2024. 12 [39] Abhinav Sharma, Philippe Weitz, Yinxi Wang, Bojing Liu, Johan Vallon-Christersson, Johan Hartman, and Mattias Rantalainen. Development and prognostic validation of a three-level nhg-like deep learning-based model for histological grading of breast cancer. Breast Cancer Research, 26 (1):17, 2024. 12 [40] Andrew H Song, Richard J Chen, Tong Ding, Drew FK Williamson, Guillaume Jaume, and Faisal Mahmood. Morphological prototyping for unsupervised slide representation learning in computational pathology. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 4, 11, 12 [41] Erik Thiringer, Fredrik K Gustafsson, Kajsa Ledesma Eriksson, and Mattias Rantalainen. Scanner-induced domain shifts undermine the robustness of pathology foundation models. arXiv preprint arXiv:2601.04163, 2026. 3 [42] Johan Vallon-Christersson, Jari Häkkinen, Cecilia Hegardt, Lao H Saal, Christer Larsson, Anna Ehinger, Henrik Lindman, Helena Olofsson, Tobias Sjöblom, Fredrik Wärnberg, et al. Cross comparison and prognostic assessment of breast cancer multigene signatures in a large population-based contemporary clinical series. Scientific Reports, 9(1):12184, 2019. 12 [43] Sarah Volinsky-Fremond, Nanda Horeweg, Sonali Andani, Jurriaan Barkey Wolf, Maxime W Lafarge, Cor D de Kroon, Gitte Ørtoft, Estrid Høgdall, Jouke Dijkstra, Jan J Jobsen, et al. Prediction of recurrence risk in endometrial cancer with multimodal deep learning. Nature Medicine, pages 1– 12, 2024. 1 [44] Eugene Vorontsov, Alican Bozkurt, Adam Casson, George Shaikovski, Michal Zelechowski, Kristen Severson, Eric Zimmermann, James Hall, Neil Tenenholtz, Nicolo Fusi, et al. A foundation model for clinical-grade computational pathology and rare cancers detection. Nature Medicine, pages 1–12, 2024. 1, 3, 13 [45] Hongming Wang, Jie Wang, Liya Yu, and Dinggang Shen. Transformer-based unsupervised contrastive learning for histopathological image classification. Medical Image Analysis, 84:102710, 2023. 3, 13

16

[46] Xiyue Wang, Yuexi Du, Sen Yang, Jun Zhang, Minghui Wang, Jing Zhang, Wei Yang, Junzhou Huang, and Xiao Han. RetCCL: Clustering-guided contrastive learning for whole-slide image retrieval. Medical Image Analysis, 83: 102645, 2023. 3, 13 [47] Y Wang, B Acs, S Robertson, B Liu, Leslie Solorzano, Carolina Wählby, J Hartman, and M Rantalainen. Improved breast cancer histological grading using deep learning. Annals of Oncology, 33(1):89–98, 2022. 1, 12 [48] Hanwen Xu, Naoto Usuyama, Jaspreet Bagga, Sheng Zhang, Rajesh Rao, Tristan Naumann, Cliff Wong, Zelalem Gero, Javier González, Yu Gu, et al. A whole-slide foundation model for digital pathology from real-world data. Nature, pages 1–8, 2024. 1, 3, 13 [49] Eric Zimmermann, Eugene Vorontsov, Julian Viret, Adam Casson, Michal Zelechowski, George Shaikovski, Neil Tenenholtz, James Hall, Thomas Fuchs, Nicolo Fusi, et al. Virchow2: Scaling self-supervised mixed magnification models in pathology. arXiv preprint arXiv:2408.00738, 2024. 3, 13

17

Benchmarking Pathology Foundation Models for Breast Cancer Survival Prediction

Supplementary Material A. Supplementary Figures This section contains Figure S1 - S6.

Resnet-IN | C-index: 0.612

0.9 0.8 0.7 0.6 0.5

0 Low risk score At risk 780 Events 0 Medium low risk score At risk 779 Events 0 Medium high risk score At risk 780 Events 0 High risk score At risk 780 Events 0

Low risk score Medium low risk score Medium high risk score High risk score 2

4

771 8

Resnet-IN | C-index: 0.600

1.0 Survival probability

Survival probability

1.0

0.9 0.8 0.7 0.6 0.5

12

14

753 26

6 8 10 Time (years) 610 465 237 45 68 79

75 86

2 88

762 17

730 48

633 73

507 92

283 110

108 121

7 123

761 18

718 59

633 89

498 127

326 146

114 159

7 162

740 39

689 89

595 140

499 186

341 214

124 235

12 242

0 Low risk score At risk 631 Events 0 Medium low risk score At risk 631 Events 0 Medium high risk score At risk 631 Events 0 High risk score At risk 631 Events 0

(a) RFS – All Patients.

Low risk score Medium low risk score Medium high risk score High risk score 2

4

12

14

612 19

6 8 10 Time (years) 491 363 187 35 55 61

624 7

61 68

2 70

618 13

594 36

507 59

406 73

222 90

85 97

4 99

622 9

591 38

523 62

410 97

268 113

95 124

9 126

608 21

569 60

488 98

410 133

280 159

108 174

12 180

(b) RFS – Patient Subgroup (ER+ & HER2-).

Resnet-IN | C-index: 0.606 1.00

0.95

0.95

Survival probability

Survival probability

Resnet-IN | C-index: 0.628 1.00

0.90 0.85 0.80

0 Low risk score At risk 780 Events 0 Medium low risk score At risk 779 Events 0 Medium high risk score At risk 780 Events 0 High risk score At risk 780 Events 0

Low risk score Medium low risk score Medium high risk score High risk score 2

4

770 4

0.90 0.85 0.80

12

14

744 17

6 8 10 Time (years) 623 513 325 24 31 33

97 35

2 35

764 7

733 19

624 27

480 40

282 45

85 45

7 45

759 9

723 28

614 42

479 51

285 53

106 55

11 55

741 28

690 53

610 75

497 86

295 94

133 97

8 98

0 Low risk score At risk 631 Events 0 Medium low risk score At risk 631 Events 0 Medium high risk score At risk 631 Events 0 High risk score At risk 631 Events 0

(c) PFS – All Patients.

Low risk score Medium low risk score Medium high risk score High risk score 2

4

14

604 16

6 8 10 Time (years) 501 419 269 22 27 29

12

623 4

79 31

2 31

619 3

601 6

512 10

385 21

229 25

65 25

6 25

619 4

591 17

499 28

382 37

225 39

97 40

11 40

611 11

570 28

497 43

403 51

234 57

108 60

8 61

(d) PFS – Patient Subgroup (ER+ & HER2-).

Figure S1. Four-group Kaplan-Meier risk stratification with at-risk and event counts (Resnet-IN). KM survival curves corresponding to Figure 3 for Resnet-IN, including the number of patients at risk and the number of events over time (0-14 years) for each of the four risk groups. Note the difference in range of the y-axis between RFS and PFS.

18

UNI | C-index: 0.657

0.9 0.8 0.7 0.6 0.5

0 Low risk score At risk 780 Events 0 Medium low risk score At risk 779 Events 0 Medium high risk score At risk 780 Events 0 High risk score At risk 780 Events 0

Low risk score Medium low risk score Medium high risk score High risk score 2

4

773 7

UNI | C-index: 0.648

1.0 Survival probability

Survival probability

1.0

0.9 0.8 0.7 0.6 0.5

12

14

764 15

6 8 10 Time (years) 604 463 249 33 49 60

90 68

7 69

765 13

733 44

643 68

518 90

316 103

100 114

6 116

765 14

727 51

636 82

499 122

315 147

103 159

9 163

731 48

666 112

588 164

489 212

307 239

128 260

6 267

0 Low risk score At risk 631 Events 0 Medium low risk score At risk 631 Events 0 Medium high risk score At risk 631 Events 0 High risk score At risk 631 Events 0

(a) RFS – All Patients.

Low risk score Medium low risk score Medium high risk score High risk score 2

4

12

14

618 12

6 8 10 Time (years) 482 365 205 26 39 44

624 7

74 50

6 51

620 10

600 29

524 50

414 66

247 78

81 87

6 88

625 6

594 36

520 61

410 98

263 119

85 129

8 132

603 27

554 76

483 117

400 155

242 182

109 197

7 204

(b) RFS – Patient Subgroup (ER+ & HER2-).

UNI | C-index: 0.638 1.00

0.95

0.95

Survival probability

Survival probability

UNI | C-index: 0.671 1.00

0.90 0.85 0.80

0 Low risk score At risk 780 Events 0 Medium low risk score At risk 779 Events 0 Medium high risk score At risk 780 Events 0 High risk score At risk 780 Events 0

Low risk score Medium low risk score Medium high risk score High risk score 2

4

772 3

0.90 0.85 0.80

12

14

747 13

6 8 10 Time (years) 611 499 304 21 27 28

95 29

6 29

762 4

734 14

626 22

488 30

274 36

73 37

5 37

758 12

726 28

624 40

496 53

300 56

120 59

10 59

742 29

683 62

610 85

486 98

309 105

133 107

7 108

0 Low risk score At risk 631 Events 0 Medium low risk score At risk 631 Events 0 Medium high risk score At risk 631 Events 0 High risk score At risk 631 Events 0

(c) PFS – All Patients.

Low risk score Medium low risk score Medium high risk score High risk score 2

4

12

14

604 10

6 8 10 Time (years) 491 402 245 17 21 22

625 3

74 23

4 23

617 3

596 11

506 18

398 22

226 28

59 29

5 29

617 5

596 16

512 23

392 39

238 40

100 42

9 42

613 11

570 30

500 45

397 54

248 60

116 62

9 63

(d) PFS – Patient Subgroup (ER+ & HER2-).

Figure S2. Four-group Kaplan-Meier risk stratification with at-risk and event counts (UNI). KM survival curves corresponding to Figure 3 for UNI, including the number of patients at risk and the number of events over time (0-14 years) for each of the four risk groups. Note the difference in range of the y-axis between RFS and PFS.

19

H-optimus-1 | C-index: 0.678

0.9 0.8 0.7 0.6 0.5

0 Low risk score At risk 780 Events 0 Medium low risk score At risk 779 Events 0 Medium high risk score At risk 780 Events 0 High risk score At risk 780 Events 0

Low risk score Medium low risk score Medium high risk score High risk score 2

4

773 7

H-optimus-1 | C-index: 0.670

1.0 Survival probability

Survival probability

1.0

0.9 0.8 0.7 0.6 0.5

12

14

763 17

6 8 10 Time (years) 638 524 298 28 40 48

92 58

5 59

770 8

739 35

626 60

498 87

298 102

103 112

11 114

761 18

726 53

638 80

485 118

299 143

120 155

7 160

730 49

662 117

569 179

462 228

292 256

106 276

5 282

0 Low risk score At risk 631 Events 0 Medium low risk score At risk 631 Events 0 Medium high risk score At risk 631 Events 0 High risk score At risk 631 Events 0

(a) RFS – All Patients.

Low risk score Medium low risk score Medium high risk score High risk score 2

4

12

14

617 14

6 8 10 Time (years) 512 418 242 23 34 39

624 7

75 46

4 47

623 7

601 26

507 46

400 68

240 81

84 89

10 89

620 11

597 34

516 57

388 88

234 109

86 121

7 127

605 25

551 79

474 128

383 168

241 194

104 207

6 212

(b) RFS – Patient Subgroup (ER+ & HER2-).

H-optimus-1 | C-index: 0.670 1.00

0.95

0.95

Survival probability

Survival probability

H-optimus-1 | C-index: 0.702 1.00

0.90 0.85 0.80

0 Low risk score At risk 780 Events 0 Medium low risk score At risk 779 Events 0 Medium high risk score At risk 780 Events 0 High risk score At risk 780 Events 0

Low risk score Medium low risk score Medium high risk score High risk score 2

4

771 4

0.90 0.85 0.80

12

14

747 10

6 8 10 Time (years) 622 513 294 16 21 22

88 23

5 23

764 1

742 9

623 21

488 31

281 34

81 36

6 36

765 6

734 21

632 31

490 40

299 46

103 47

7 47

734 37

667 77

594 100

478 116

313 123

149 126

10 127

0 Low risk score At risk 631 Events 0 Medium low risk score At risk 631 Events 0 Medium high risk score At risk 631 Events 0 High risk score At risk 631 Events 0

(c) PFS – All Patients.

Low risk score Medium low risk score Medium high risk score High risk score 2

4

12

14

601 10

6 8 10 Time (years) 495 411 241 14 17 18

623 4

70 19

5 19

617 1

602 6

507 16

394 21

229 22

66 23

6 23

620 4

600 14

509 21

384 34

226 41

76 43

5 43

612 13

563 37

498 52

400 64

261 69

137 71

11 72

(d) PFS – Patient Subgroup (ER+ & HER2-).

Figure S3. Four-group Kaplan-Meier risk stratification with at-risk and event counts (H-optimus-1). KM survival curves corresponding to Figure 3 for H-optimus-1, including the number of patients at risk and the number of events over time (0-14 years) for each of the four risk groups. Note the difference in range of the y-axis between RFS and PFS.

20

Resnet-IN | C-index: 0.612

1.0

0.9 Survival probability

Survival probability

0.9 0.8 0.7 0.6 0.5

Resnet-IN | C-index: 0.600

1.0

Low risk score Medium risk score High risk score

0 2 Low risk score At risk 1029 1014 Events 0 14 Medium risk score At risk 1029 1010 Events 0 19 High risk score At risk 1061 1010 Events 0 49

4

0.8 0.7 0.6 0.5

12

14

986 42

6 8 10 Time (years) 799 616 320 72 101 119

106 129

4 131

957 69

844 104

667 139

399 164

136 180

9 183

947 111

828 171

686 233

468 266

179 292

15 301

0 Low risk score At risk 833 Events 0 Medium risk score At risk 833 Events 0 High risk score At risk 858 Events 0

(a) RFS – All Patients.

Low risk score Medium risk score High risk score 2

4

12

14

800 33

6 8 10 Time (years) 645 492 252 55 77 89

820 13

87 97

3 99

822 11

788 43

685 78

535 111

320 134

110 146

9 149

830 26

778 77

679 121

562 170

385 200

152 220

15 227

(b) RFS – Patient Subgroup (ER+ & HER2-).

Resnet-IN | C-index: 0.606 1.00

0.95

0.95

Survival probability

Survival probability

Resnet-IN | C-index: 0.628 1.00

0.90 0.85 0.80

Low risk score Medium risk score High risk score

0 2 Low risk score At risk 1029 1017 Events 0 4 Medium risk score At risk 1029 1007 Events 0 11 High risk score At risk 1061 1010 Events 0 33

4

0.90 0.85 0.80

12

14

983 20

6 8 10 Time (years) 827 682 420 29 39 44

118 46

4 46

962 31

818 45

620 63

374 67

138 69

13 69

945 66

826 94

667 106

393 114

165 117

11 118

0 Low risk score At risk 833 Events 0 Medium risk score At risk 833 Events 0 High risk score At risk 858 Events 0

(c) PFS – All Patients.

Low risk score Medium risk score High risk score 2

4

12

14

799 16

6 8 10 Time (years) 662 544 345 23 29 31

821 4

94 33

3 33

819 6

789 14

684 22

512 38

300 44

116 45

11 45

832 12

778 37

663 58

533 69

312 75

139 78

13 79

(d) PFS – Patient Subgroup (ER+ & HER2-).

Figure S4. Three-group Kaplan-Meier risk stratification with at-risk and event counts (Resnet-IN). KM survival curves corresponding to Figure S1, but with stratification into three risk groups instead of four, for Resnet-IN. The plots include the number of patients at risk and the number of events over time (0-14 years) for each risk group. Note the difference in range of the y-axis between RFS and PFS.

21

UNI | C-index: 0.657

1.0

0.9 Survival probability

Survival probability

0.9 0.8 0.7 0.6 0.5

UNI | C-index: 0.648

1.0

Low risk score Medium risk score High risk score

0 2 Low risk score At risk 1029 1017 Events 0 12 Medium risk score At risk 1029 1012 Events 0 16 High risk score At risk 1061 1005 Events 0 54

4

0.8 0.7 0.6 0.5

12

14

997 30

6 8 10 Time (years) 806 630 350 57 78 93

123 106

10 107

968 59

853 92

678 136

420 161

138 173

9 176

925 133

812 198

661 259

417 295

160 322

9 332

0 Low risk score At risk 833 Events 0 Medium risk score At risk 833 Events 0 High risk score At risk 858 Events 0

(a) RFS – All Patients.

Low risk score Medium risk score High risk score 2

4

12

14

809 22

6 8 10 Time (years) 644 501 281 46 64 73

823 10

98 81

10 82

820 12

788 43

694 65

544 99

333 121

112 133

7 134

829 28

769 88

671 143

544 195

343 229

139 249

10 259

(b) RFS – Patient Subgroup (ER+ & HER2-).

UNI | C-index: 0.638 1.00

0.95

0.95

Survival probability

Survival probability

UNI | C-index: 0.671 1.00

0.90 0.85 0.80

Low risk score Medium risk score High risk score

0 2 Low risk score At risk 1029 1013 Events 0 5 Medium risk score At risk 1029 1004 Events 0 10 High risk score At risk 1061 1017 Events 0 33

4

0.90 0.85 0.80

12

14

980 18

6 8 10 Time (years) 818 662 396 28 36 40

120 41

7 41

968 27

821 41

649 54

383 58

127 62

12 62

942 72

832 99

658 118

408 127

174 129

9 130

0 Low risk score At risk 833 Events 0 Medium risk score At risk 833 Events 0 High risk score At risk 858 Events 0

(c) PFS – All Patients.

Low risk score Medium risk score High risk score 2

4

12

14

797 14

6 8 10 Time (years) 664 539 325 22 28 32

825 3

95 33

7 33

812 5

788 16

662 26

517 34

301 37

102 39

8 39

835 14

781 37

683 55

533 74

331 81

152 84

12 85

(d) PFS – Patient Subgroup (ER+ & HER2-).

Figure S5. Three-group Kaplan-Meier risk stratification with at-risk and event counts (UNI). KM survival curves corresponding to Figure S2, but with stratification into three risk groups instead of four, for UNI. The plots include the number of patients at risk and the number of events over time (0-14 years) for each risk group. Note the difference in range of the y-axis between RFS and PFS.

22

H-optimus-1 | C-index: 0.678

1.0

0.9 Survival probability

Survival probability

0.9 0.8 0.7 0.6 0.5

H-optimus-1 | C-index: 0.670

1.0

Low risk score Medium risk score High risk score

0 2 4 Low risk score At risk 1029 1021 1003 Events 0 8 25 Medium risk score At risk 1029 1005 961 Events 0 22 63 High risk score At risk 1061 1008 926 Events 0 52 134

0.8 0.7 0.6 0.5

6 8 10 Time (years) 836 681 391 48 65 78

12

14

128 91

10 92

831 95

651 139

391 162

141 178

10 183

804 204

637 269

405 309

152 332

8 340

0 Low risk score At risk 833 Events 0 Medium risk score At risk 833 Events 0 High risk score At risk 858 Events 0

(a) RFS – All Patients.

Low risk score Medium risk score High risk score 2

4

12

14

812 20

6 8 10 Time (years) 673 550 318 37 51 61

826 7

104 70

7 71

816 16

783 47

671 74

517 108

319 125

110 141

11 144

830 27

771 86

665 143

522 199

320 237

135 252

9 260

(b) RFS – Patient Subgroup (ER+ & HER2-).

H-optimus-1 | C-index: 0.670 1.00

0.95

0.95

Survival probability

Survival probability

H-optimus-1 | C-index: 0.702 1.00

0.90 0.85 0.80

Low risk score Medium risk score High risk score

0 2 Low risk score At risk 1029 1016 Events 0 5 Medium risk score At risk 1029 1008 Events 0 3 High risk score At risk 1061 1010 Events 0 40

4

0.90 0.85 0.80

12

14

987 14

6 8 10 Time (years) 820 674 385 24 30 31

114 33

6 33

971 19

824 33

632 47

364 55

110 56

10 56

932 84

827 111

663 131

438 139

197 143

12 144

0 Low risk score At risk 833 Events 0 Medium risk score At risk 833 Events 0 High risk score At risk 858 Events 0

(c) PFS – All Patients.

Low risk score Medium risk score High risk score 2

4

12

14

798 11

6 8 10 Time (years) 665 546 317 19 24 25

824 4

91 27

6 27

812 3

789 14

660 24

498 35

293 39

91 40

7 40

836 15

779 42

684 60

545 77

347 86

167 89

14 90

(d) PFS – Patient Subgroup (ER+ & HER2-).

Figure S6. Three-group Kaplan-Meier risk stratification with at-risk and event counts (H-optimus-1). KM survival curves corresponding to Figure S3, but with stratification into three risk groups instead of four, for H-optimus-1. The plots include the number of patients at risk and the number of events over time for each risk group. Note the difference in range of the y-axis between RFS and PFS.

23

Record · ID 138916 · SHA-256 cfc6bbf9c1787f90
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.