Memorisation bias in medical AI Moritz A. Knolle1*, Martin J. Menten1,2 , Laurin Lux1,2 , Mélanie Roschewitz, Emma A.M. Stanley3 , Georgios Kaissis4† , Daniel Rueckert1,2,3† , Ben Glocker3†
arXiv:2609.17223v1 [cs.LG] 15 Sep 2026
1
Chair for AI in Healthcare and Medicine, Technical University of Munich (TUM) and TUM University Hospital, Germany. 2 Munich Center for Machine Learning, Germany. 3 Department of Computing, Imperial College London, United Kingdom. 4 Hasso Plattner Institute, Germany.
*Corresponding author(s). E-mail(s): [email protected]; † These authors contributed equally Abstract Medical AI models hold immense potential to improve patient outcomes. However, these models are also known to unintentionally memorise individual records from their training datasets [1–4]. While such memorisation has been linked to targeted privacy attacks [5–10], its consequences for clinical deployment, where patients may be assessed by a model that saw their historical data during training, remain poorly understood. Here we show that predictions on a patient’s unseen future data can change significantly if a model observed that same patient’s anonymised historical data during training, a phenomenon we term “memorisation bias”. We demonstrate that this bias exists across diverse data modalities and model architectures, and over prolonged time spans: in some cases, we find that memorisation bias can persist on future records acquired decades after the patient’s historical records used for model training. Moreover, in simulated prospective deployment, we find that memorisation bias has asymmetric effects on the diagnostic accuracy of returning data contributors. When a patient returned with a de novo condition that was absent from their historical records in the model’s training dataset, diagnostic sensitivity decreased significantly compared to an otherwise identical model not trained on their historical data. Conversely, when their health state was unchanged, both sensitivity and specificity were significantly inflated. Together, our findings reveal a previously uncharacterised risk in medical AI that arises when an AI model is deployed on patients that contributed to the model’s training data. This exposes a shortcoming of current model development practice: the de-identification measures
1
designed to protect patients’ privacy make it difficult to identify returning data contributors and exclude them from the AI-assisted interpretation of their own future data. Effectively mitigating memorisation risks may thus require changes to current model training and deployment protocols.
Main Artificial intelligence (AI) models are quickly moving from retrospective evaluation into routine use in population screening programs and clinical practice. For good reason. Given enough high-quality training data, AI models can often match, and in some cases even exceed, the performance of medical experts for diagnostic tasks [11– 15] and prospective trials now demonstrate clinical benefit at scale [16–18]. Most of these advances are enabled by large-scale training datasets of real patient data. Training data, however, is not only absorbed in aggregate. Besides learning general patterns, AI models also retain information about individual records in their training data: a model’s output for a given input can depend on whether that same input was present in its training dataset. This phenomenon, termed memorisation, is well documented across model families, from deep learning models [1–8] to classical machine learning models [19–22]. Indeed, research by Feldman [23] and Feldman and Zhang [2] suggests that such memorisation is not merely an artefact of over-parameterisation, but is in fact required to achieve optimal generalisation performance on the long-tailed data distributions common to many real-world settings. Previous research has studied memorisation almost exclusively in the context of privacy attacks, where an adversary aims to extract sensitive information from a trained model [4–10]. As a result, the consequences for patients whose records were memorised and who may encounter the same model again during their future care journey remain largely unexplored. With this study, we show that when models are deployed prospectively, memorisation creates a specific and under-appreciated risk. Because a patient’s records are highly self-similar over time, a future record may act as a partial cue for a memorised historical record, shifting the model’s predictions towards a patient’s previous health state. This is not a hypothetical scenario. Medical AI models are routinely deployed on the same population from which their training data were sourced. For example, Germany’s national breast cancer screening program uses an AI model trained on 1.2 million mammograms sourced from the same screening population it now serves [17, 24]. Comparable deployment is underway elsewhere, including national breast cancer screening programmes in the United Kingdom [25] and Sweden [16]. Regulatory guidance actively encourages this, with the EU AI Act [26] requiring that training data reflect the geographical and contextual setting of intended use (Art. 10(4)), and similar international guidelines for medical AI development [27] calling for the intended patient population to be sufficiently represented in a model’s training dataset. If memorisation were present, the risk for harm in AI-assisted population screening programs is substantial: hundreds of thousands of patients will return at regular intervals over years or decades, meaning that data contributors will repeatedly encounter a model
2
that may have memorised their earlier, typically healthy, records. Whether training data memorisation could systematically alter predictions on data contributors’ unseen future data encountered only during prospective deployment has not been systematically studied. Our study addresses this gap, marking the first time that longitudinal memorisation effects have been studied using real-world patient data. Here we show that medical AI models, trained to perform standard diagnostic (supervised classification) tasks, exhibit a phenomenon we term “memorisation bias”: a systematic shift in a patient’s future predictions towards the health state that the model observed in the same patient’s historical records during training. More specifically, using four large datasets spanning electrocardiograms, chest radiographs and electronic health records, we show that including a patient’s historical data in a training dataset can significantly alter the model’s future predictions for that same patient. Crucially, while memorisation bias becomes both rarer and weaker over time, we show that it can persist for multiple decades, affecting predictions on records acquired long after the historical data used for model training. We further show that memorisation bias can lead to systematic diagnostic errors. In simulated deployment scenarios, diagnostic sensitivity was significantly reduced for patients who returned with a condition not present in their historical data used to train the model. On the other hand, sensitivity and specificity were significantly artificially inflated for patients whose health status was unchanged compared to their historical training records. Our findings pose an immediate and practical dilemma. Patients are rarely informed that their personal data was used to train a specific model. This is because model training constitutes a secondary use of health data in most jurisdictions, and typically proceeds under broad consent or a consent waiver rather than study-specific consent. Moreover, because of standard de-identification procedures, it is not straightforward to determine whether any given patient is represented in a model’s training dataset once the model is deployed. Protecting data contributors from memorisation bias by excluding them from the AI-assisted interpretation of their future data is thus difficult in practice and also raises ethical questions about disparities in the provision of care.
Detecting longitudinal memorisation The established gold-standard definition of AI memorisation [2, 3] is the marginal change in a model’s predictions on an individual record caused by the inclusion (or exclusion) of that same record in its training dataset. In practice, measuring memorisation involves training a large number of models on random data subsets and comparing the predictions of models whose training subsets included the record of interest with those that did not. Inspired by prior research on AI memorisation [2, 3] and closely related research on membership inference attacks [5, 6], we propose a novel framework to detect memorisation bias. Specifically, our framework allows us to quantify whether, and to what extent, the inclusion of a patient’s historical records in a model training dataset causes a significant change in their future predictions. Our framework operates as follows (see Fig. 1 & Fig. 2a for an overview). First, we temporally split the data of patients in each follow-up cohort into historical and future records (see Fig. 2b-e for the resulting split sizes and the Methods for details
3
Historical data
(seen/not seen during model training) Future data
sinus rhythm sinus rhythm
Alice
1998
2003
(never seen)
sinus rhythm
anterior infarct
2013
today
p(infarct |
) = 15%
AI model trained on Alice's historical data
p(infarct |
) = 73%
AI model not trained on Alice's historical data
Fig. 1: Schematic of our longitudinal memorisation detection approach. Each patient’s records are split temporally: earlier (historical) records may be used for model training; later (future) records are used for model evaluation only and are never used for training or model selection. To detect longitudinal memorisation, we compare the predictions models make on a future record when trained on that patient’s historical records with those made by models trained without them. In the illustrated example, Alice contributed three historical electrocardiograms to the training dataset, all showing a sinus rhythm (a normal heart rhythm), and returns today with an anterior infarct (a heart attack involving the anterior wall of the heart). The model that saw her healthy historical records during training assigns a substantially lower probability to her anterior infarct, potentially leading to a missed diagnosis. For clarity, one model is shown per group; our analysis uses M = 100 models per group for each patient (see Methods for details).
of the split strategy which enriches the future datasets for de novo cases, i.e., positive cases of conditions not present in the respective patient’s historical data). Second, we train M = 200 models, each on historical data from a randomly selected 50% subset of patients. Note that these models were trained using state-of-the-art techniques (learning rate schedules, exponential moving parameter average and, where applicable, data augmentation) and using hyperparameters chosen to yield optimal generalisation performance on the validation set; future data was never used for model training or model selection. Third, for each patient of interest, we partition the models into two groups: those whose training subsets included that patient’s historical data and those whose training subsets did not. Notably, because all models are trained on random data subsets comprising roughly half of the available historical data, performance differences on held-out validation and test sets between models are negligible (Extended Data Fig. 1a), isolating patient inclusion as the variable of interest. Using these model groups, partitioned by whether each patient was included in their training data subset, we generate predictions for patients’ future records (predicted probabilities for all available annotations/classes). Feldman and Zhang [2] show that the influence of other patients’ data, included or excluded by chance here, is marginalised out quickly and is thus negligible across the large number of models we study. Fourth, we apply
4
a
1) Train many models on historical data from random patient subsets
2) Partition models for each patient
3) Test whether model predictions on unseen future record(s) differ significantly
Alice AI models trained on Alice's historical data
AI models not trained on Alice's historical data
p(infarct |
)
AI model
b
c
Follow-up cohort: 146,332 patients 725,340 ECGs
MIMIC-ECG: 161,332 patients 799,981 ECGs
MIMIC-CXR: 61,868 patients 357,522 CXRs
patient-wise split
patient-wise split
Validation: 5,000 patients 24,986 ECGs
Follow-up cohort: 46,868 patients 270,758 CXRs
Test: 10,000 patients 49,655 ECGs
patient-level, temporal split Historical: 146,332 patients 475,512 ECGs
Validation: 5,000 patients 29,235 CXRs
Test: 10,000 patients 57,529 CXRs
patient-level, temporal split
Future: 72,664 patients 249,828 ECGs
Historical: 46,868 patients 196,516 CXRs
d
Future: 18,298 patients 74,242 CXRs
e
Follow-up cohort: 186,213 patients 386,932 EHRs
MIMIC-IV-ED: 201,213 patients 418,007 EHRs
HEEDB: 1,802,104 patients 10,457,413 ECGs
patient-wise split
patient-wise split
Validation: 5,000 patients 10,203 EHRs
Follow-up cohort: 1,767,104 patients 10,256,486 ECGs
Test: 10,000 patients 20,872 EHRs
Test: 25,000 patients 143,062 ECGs
patient-level, temporal split
patient-level, temporal split Historical: 186,213 patients 321,530 EHRs
Validation: 10,000 patients 57,865 ECGs
Historical: 1,767,104 patients 6,147,511 ECGs
Future: 18,771 patients 65,402 EHRs
Future: 888,338 patients 4,108,975 ECGs
Fig. 2: Study profile. a, the three key steps of our longitudinal memorisation detection approach: (1) training M = 200 models on the historical data of randomly selected 50% patient subsets, (2) partitioning these models for each patient depending on whether a model’s training subset included the respective patient’s historical data or not (M = 100 models per group for each patient), and (3) testing whether the models’ predictions on the respective patient’s future records differ significantly between model groups. b-e, our dataset split strategy. For each investigated dataset, we perform a four-way dataset split in two stages: (1) a standard, randomised patientwise split into follow-up cohort, validation dataset and test dataset; (2) an additional patient-level, temporal split performed on the records of each patient in the follow-up cohort, yielding the historical and future record datasets. Patients with insufficient follow-up data are excluded from the temporal split, and their records are added only to the respective historical dataset. By construction, every patient represented in a future dataset also contributes at least one earlier record to the respective historical dataset. Sub-panels show resulting split sizes for MIMIC-ECG [28] (b), MIMIC-CXR [29] (c), MIMIC-IV-ED [30] (d) and HEEDB [31](e). Historical data is used for model training (grey boxes with solid lines in b-e), future, validation and test data are used for model evaluation only (white boxes with solid lines in b-e). ECGs, CXRs and EHRs refer to electrocardiograms, chest radiographs and electronic health records, respectively.
5
energy-based hypothesis testing to assess whether the two groups of predictions differ significantly (see Methods for details). A statistically significant difference in the prediction for a patient’s future record indicates that the patient’s historical data was memorised during model training and that this memorisation influenced the prediction for one of their future records.
Memorisation alters future predictions Across all investigated datasets, we find evidence for memorisation bias. Specifically, we find that a substantial share of future records show significant changes in their predictions when the respective patient’s historical data is included in the training dataset (Fig. 3a, left). We also measured the effect size of these prediction changes, i.e., the absolute difference in average predicted probability between models trained vs. not trained on a given patient’s historical data. Notably, while effect sizes are small on average, some records show large changes in their predicted probability for one of the available classes, in some cases over 70 percentage points. This is indicated by empirical complementary cumulative distribution function (ecCDF) plots of the effect sizes for all future records (Fig. 3a right). For a given percentage, the ecCDF plots show the share of records which show a memorisation-induced change in their average predicted probability of at least that percentage or higher. Corresponding energy test statistics are reported in Extended Data Fig. 1b. As a baseline comparison, we provide the same effect size measures and energy test statistics for models that were partitioned randomly (dashed lines in Fig. 3a right and Extended Data Fig. 1b). Crucially, across all investigated datasets, these randomly partitioned models yielded zero future records with significant differences in their predictions. To understand the root cause of these memorisation effects, we performed the same analysis on patients’ historical records (the data observed during training and thus directly responsible for any memorisation). As expected, memorisation effects are substantially stronger here. Compared to future records, a larger share of historical records show significant changes in their predictions when the respective patient was included in the training dataset (Fig. 3b left), and effect sizes are also correspondingly larger (Fig. 3b right). This is consistent with the intuition that memorisation is strongest on the data a model directly observed during training, and attenuates – but, as shown previously, can persist – when the model encounters a data contributor’s unseen future data. As before, randomly partitioned models yielded zero historical records with significant differences in their predictions across all datasets (corresponding effect sizes and test statistics are reported in Fig. 3b and Extended Data Fig. 1b, respectively). This confirms that the observed effects are driven by patientspecific memorisation rather than by random variation in predictions between model groups. Notably, the memorisation-induced prediction changes on patients’ (historical) training records that we quantify here are the signal that a membership inference attack exploits to determine whether a given record was used for model training [5–7]. Our results demonstrate that this membership signal can extend beyond the historical training records themselves and systematically alters predictions on contributors’ unseen future data.
6
MIMIC-ECG
a
MIMIC-CXR
HEEDB
MIMIC-IV-ED
b 100
10
−6
40
20
10−7
0
20
40
60
10−2 10−3 10−4 10−5 10−6 10−7
0 0
Dataset
10−1 1 - cumulative probability
10
−5
60
Random partitioning
3.0 × 10 5 of 3.2 × 10 5
10
−4
98,001 of 4.8 × 105
10−3
1.3 × 10 5 of 2.0 × 10 5
10−2
80
1.7 × 10 6 of 6.1 × 10 6
Historical records with significant change in prediction (%)
1 - cumulative probability
1
25,210 of 4.1 × 106
2
Random partitioning
3,518 of 74,242
3
100
100
10−1
1,469 of 65,402
4 2,848 of 2.5 × 105
Future records with significant change in prediction (%)
5
80
0
Dataset
20
40
60
80
100
Change in predicted probability (%)
Change in predicted probability (%)
c 20.0 Significance Threshold 17.5
−log 10 (p)
15.0 12.5 10.0 7.5 5.0 2.5 0.0 80
100
0
100
20,477 of 2.1 × 106
24 of 6,653
114 of 5,844
183 of 21,181
200
300
400
500
Acquisition interval (months)
11 of 64,095
60
102 of 2.7 × 105
40
473 of 4.8 × 10 5
20
Acquisition interval (months)
1,811 of 7.1 × 105
0
2,336 of 4.5 × 105
60
250 of 9,086
2,302 of 48,967
63 of 19,818
176 of 14,658
1
409 of 47,077
2
300 of 19,591
3
550 of 26,983
5 1,350 of 1.2 × 105
Future records with significant change in prediction (%)
6
4
40
186 of 7,026
20
Acquisition interval (months)
612
0
12
125
712 of 15,612
100
202 of 3,818
75
204 of 7,780
50
283 of 5,265
25
Acquisition interval (months)
527 of 8,412
0
d
Acquisition interval (months)
Acquisition interval (months)
Acquisition interval (months)
24
050
0
0
0
024 12
-6 0
-1 2 60
24
-2 4
012
12
-9 7 60
-2 4
-6 0 24
18
-1 8
06
-6 0 24
-1 8
-2 4 18
12
06
612
0 -1 3 60
-6 0
-2 4 18
24
-1 8 12
06
612
0
Acquisition interval (months)
Fig. 3: Historical training data alters future predictions. a-b, the relative share of records for which we detected a significant change in predicted probability due to the patient’s inclusion in the training data set (left), alongside empirical complementary cumulative distribution function (ecCDF) plots of effect sizes (right). Effect size was measured as the absolute difference in average predicted probability between models trained vs. not-trained on a given patient’s historical data (maximum across available classes). Panels show results for unseen future records (a) and historical records used for model training (b). c, Manhattan plots of multiple-comparison corrected p values versus the time in months between the acquisition of the patient’s most recent historical training record and the acquisition of the respective future record (hereafter, the acquisition interval). d, relative share of future records with a significant change in predicted probability, binned by the acquisition interval. Statistical significance was determined using two-sided, nonparametric, energy-based hypothesis tests comparing the predictions (predicted probabilities across all available classes) of models trained versus not trained on the respective patient’s historical data (M = 100 models per group for each patient). The threshold for statistical significance was p ≤ 0.05. Multiple-comparison correction was performed separately across each dataset’s historical/future records using the Benjamini-Hochberg method to control the false discovery rate at α = 0.05; p values were clipped to a minimum value of 10−20 to improve visibility.
7
For MIMIC-IV-ED [30, 32], a tabular dataset of electronic health records, the results reported above were obtained using a random forest model. To assess how tabular data memorisation varies for different types of classical machine learning and modern AI models, we also applied our memorisation detection framework to logistic regression and tabular ResNet models [33](Extended Data Fig. 1d-f). Note that, as before, model and optimisation hyperparameters were chosen to yield optimal generalisation performance on the validation set. The share of records for which we detected memorisation varied by orders of magnitude between model types: on unseen future records, we detected significant changes in predicted probability for 2.25% (1,469/65,402) of records for the random forest, compared to 0.04% (25/65,402) for the logistic regression model and 0.12% (83/65,402) for the tabular ResNet, with corresponding rates of 93.53%, 0.77% and 3.61% on historical training records. Effect sizes reached up to 16.7 percentage points on future records for the random forest, compared to 1.2 and 10.2 percentage points for the logistic regression model and the tabular ResNet, respectively. Together, these results suggest that memorisation bias is not limited to modern AI models and that random forests, a model class widely used for tabular clinical prediction problems [34], are particularly prone to memorisation bias.
Memorisation can persist for decades Next, we asked how long memorisation bias can persist for. To investigate this, we closely examined future records for which we detected memorisation and, for each, measured the time interval between the acquisition of the patient’s most recent historical training record and the acquisition of the respective future record. Plots of multiple-comparison-corrected p values from our per-record test for memorisation bias against this acquisition interval (Fig. 3c) reveal that memorisation bias can persist over remarkably long time spans. For example, for HEEDB [31], a large dataset comprising electrocardiograms from 1.8 million patients collected from the 1980s until 2025, we detect statistically significant changes in model predictions for future records acquired more than 25 years after the patient’s most recent historical training record. To further characterise how memorisation decays over time, we computed the relative share of memorised future records as a function of this acquisition interval (Fig. 3d). The rate at which future records exhibit memorisation bias, as well as the decay of this rate for increasing acquisition intervals, differs considerably between datasets. We hypothesise that this reflects differences in how strongly a patient’s future records resemble their historical records and natural differences in the intervals at which patients return. A survival-style analysis of effect sizes for increasing acquisition intervals is reported in Extended Data Fig. 2, ecCDF plots of the acquisition intervals over all future records, as well as the subset of future records for which we detected memorisation bias, are reported in Extended Data Fig. 1c. Across datasets, we find that memorisation bias generally becomes less pronounced and rarer with longer acquisition intervals. Nevertheless, we find that predictions on some future records with large acquisition intervals still exhibit significant memorisation-induced prediction changes. Together, these results suggest that memorisation bias is not a transient artefact that
8
resolves completely as a patient’s data changes over time, but a persistent phenomenon capable of affecting predictions decades into the future.
Memorisation affects diagnostic accuracy After establishing that including a patient’s historical data in a model’s training dataset can systematically alter their future predictions, we next investigated whether these changes could have diagnostic implications. To this end, we simulated the prospective deployment of AI models on patient populations that contributed to their training data and quantified the impact of memorisation bias on their diagnostic accuracy. First, we determined optimal decision thresholds for the annotated conditions in each investigated dataset by maximising Youden’s J statistic [35, 36] on the respective held-out test set. To remove patient-level inclusion effects and general random variation across models, this was performed with predictions averaged across all random-subset models for each dataset. Second, using these thresholds, we computed the corresponding clinical decisions for patients’ future record predictions from models trained vs not trained on each patient’s historical data. Third, we compared the sensitivity and specificity of these decisions. Notably, we stratified patients’ future records into two groups: future records where patients returned with a de novo condition – a positive case of a condition that was not present in any of their historical training records – and future records where patients returned with unchanged health states (comprising both positive and negative cases). Since the de novo group comprises only positive cases, we sampled an equal number of age- and sex-matched controls for each condition to enable standard sensitivity and specificity analyses. The results reveal a striking asymmetry: on future records of patients returning with de novo conditions, models trained on patients’ historical data showed significantly lower sensitivity than models not trained on that data (Fig. 4). Conversely, on future records of patients returning with a health state unchanged relative to their historical training records, models trained on patients’ historical data showed significantly increased sensitivity and specificity compared to models not trained on that data (Extended Data Fig. 3), artificially inflating apparent diagnostic performance. In other words, memorisation bias is associated with a higher rate of false-negative model predictions when contributors return with a condition that was absent from their historical training records. When contributors return with unchanged health states, memorisation bias artificially increases both: sensitivity and specificity.
Risk mitigation via differential privacy Next, we investigated which level of differential privacy [37] (DP) protection would be needed to prevent memorisation bias. When applied during model training [38], DP provably limits the amount of information a model can extract from an individual training record and thereby naturally reduces memorisation. Notably, DP can also be implemented at the patient-level, where it instead provides an upper bound on the information extracted from all records that a single patient contributes to a training dataset. We trained a range of MIMIC-ECG [28] and HEEDB models with increasing
9
p
N+
N-
Fascicular block
**
1.6e-03
p
1.0
48
48
Lateral ST-T changes
**
1.6e-03
1.0
68
68
Borderline ECG
**
1.6e-03
1.0
45
45
Abnormal ECG
**
1.6e-03
1.0
31
31
1.0
1.0
43
43
1.6e-03
1.0
73
73
Pacemaker
1.0
1.0
26
26
LBBB
1.0
1.0
24
24
Mean Difference (Sensitivity)
Condition Name
PVC
MIMIC-ECG
Infarct
**
RBBB
**
4.9e-03
1.0
43
43
AV block
**
3.3e-03
1.0
68
68
1.0e-01
1.0
40
40
1.6e-03
1.0
52
52
Atrial flutter Atrial fibrillation
**
Bradycardia
5.0e-01
1.0
48
48
Tachycardia
**
8.2e-03
1.0
60
60
Sinus
**
1.6e-03
1.0
29
29
Support devices
**
1.6e-03
1.0
82
82
Fracture
**
3.3e-03
1.0
43
43
1.0
1.9e-01
62
62
**
1.6e-03
1.0
118
118
Pneumothorax
*
2.3e-02
1.0
54
54
Atelectasis
**
1.6e-03
1.0
135
135
**
1.6e-03
1.0
134
134
1.0
1.0
103
103
**
1.6e-03
1.0
79
79
1.0
1.0
57
57
2.0e-02
141
141 153
Pleural other
MIMIC-CXR
Pleural effusion
Pneumonia Consolidation Oedema Lung lesion Lung opacity
**
1.6e-03
Cardiomegaly
**
1.6e-03
1.0
153
3.4e-01
1.0
73
73
MIMIC-IV-ED
Enlarged cardiomediastinum
*
No finding
**
1.6e-03
1.0
432
432
Critical Outcome
**
1.6e-03
1.0
338
338
Hospitalisation
**
1.6e-03
1.0
343
343
Normal sinus rhythm
**
1.6e-03
1.0
620
489
Sinus tachycardia
**
1.6e-03
1.8e-02
589
470
6.2e-02
1.0
319
268
1.6e-03
1.0
294
245
1.0
1.0
292
238
Premature atrial complexes Inferior infarct HEEDB
Mean Difference (Specificity)
**
Lateral infarct
*
1st degree AV block
**
1.6e-03
1.0
289
218
Anterior infarct
**
4.9e-03
1.0
286
239
Non-specific T-wave abnormality
**
1.6e-03
1.0
269
221
Septal infarct
**
1.6e-03
1.0
268
230
Normal ECG
**
1.6e-03
1.0
253
228
−7.5
−5.0
−2.5
0.0
2.5
−4
Sensitivity (%)
−2
0
2
Specificity (%)
Fig. 4: Memorisation bias reduces diagnostic sensitivity on future records of patients returning with de novo conditions. Memorisation-induced changes in diagnostic sensitivity and specificity on future records of patients returning with a de novo condition, i.e. a condition for which no positive cases were present in that patient’s historical training record(s). Each row represents a single condition: the lefthand panel shows the mean difference in sensitivity between models trained vs. not trained on patients’ historical data (M = 100 models per group for each patient), the right-hand panel the corresponding difference in specificity. N + and N − denote the number of positive and negative cases. Evaluated cases comprised patients returning with a de novo positive diagnosis, each matched to a randomly selected age- and sex-matched control where possible. Condition-specific decision thresholds maximised Youden’s J statistic on the test set. Statistical significance was determined using exact, two-sided permutation tests that reassign model indices to the corresponding patient subset masks (B = 100,000 permutations per comparison); error bars denote simultaneous 95% confidence intervals obtained by inverting the same test. p values and confidence intervals were corrected across all C = 82 illustrated comparisons using the Bonferroni method. * and ** denote p ≤ 0.05 and p ≤ 0.01, respectively; the resolution of the permutation test floors raw p values at 2/(B + 1) ≈ 2 × 10−5 , and hence corrected p values at 1.6 × 10−3 . For HEEDB, we present results only for the 10 conditions with the largest number of de novo cases; data for all conditions are reported in Extended Data Fig. 4.
10
b
0
1
0
100 10 1
10 2
N=74,070
N=203,564
N=61,275
N=102,393
N=30,271
103
Diagnostic Performance (%)
N=2
101
104
N=930
N=1,840
102
N=936
N=102
103
N=3,210
104
95.0
105 N=15,179
105
106
102 101 100 0
Patient-level DP Record-level DP
0
Future records with significant change in prediction (count)
106
Historical records with significant change in prediction (count)
a
90.0 87.5 85.0
Patient-level DP Record-level DP
82.5
1
10 3
92.5
10 1
10 2
1
10 3
c
10 2
10 3
10 2
101
10 3
1
10 1
10 2
10 3
Diagnostic Performance (%)
N=129,438
N=29,388
N=6,673
102
0
10 1
0
0
1
0
0
101
103
0
102
104
0
N=27
103
N=3,080
N=210
N=874
104
97.5
105
N=706
105
106
0
Patient-level DP Record-level DP
Historical records with significant change in prediction (count)
d 106
Future records with significant change in prediction (count)
10 1
Privacy Budget (ε)
Privacy Budget (ε)
95.0 92.5 90.0 87.5
Patient-level DP Record-level DP
85.0 1
10 1
10 2
10 3
Privacy Budget (ε)
Privacy Budget (ε)
Fig. 5: Differential privacy reduces memorisation. a-b Results for MIMICECG. a, counts of records with a significant change in prediction due to patient inclusion for models trained with decreasing levels of (ε, δ )-DP privacy protection (increasing ε values). Panels show results for unseen future records (left) and historical training records (right). δ was kept constant at 1/D where D is size of the respective historical training data subset. b, diagnostic performance of the respective models on the unseen test set (M = 200 models for each setting); coloured points, macro average area under the receiver operating characteristic curve scores (AUROC, average across all available classes) of individual models; white markers with error bars denote mean ± s.d. c-d, as in panels (a-b) but for HEEDB. Statistical significance was determined using two-sided, multivariate, nonparametric energy-based hypothesis tests comparing the predicted probabilities of models trained vs. not trained on the respective patient’s historical data (M = 100 models per group for each patient). Multiple comparison correction was performed separately for each experimental setting’s historical/future records using the Benjamini-Hochberg method to control the false discovery rate at α = 0.05. Patient-level DP protection was achieved by discarding all but the most recent historical training record per patient and then applying record-level DP accounting; the future record dataset was not modified. Dashed lines indicate results for non-private baseline models, which were trained with the same data, hyperparameters and training recipe as the private models.
levels of record- and patient-level (ε, δ )-DP privacy protection. Owing to computational resource constraints, the model size had to be reduced from ViT-S-25 (14 million parameters) to ViT-T-25 (1.9 million parameters); we therefore trained separate nonprivate baselines for each accounting regime, matched to the corresponding private models in architecture, data, hyperparameters and training recipe. As expected, we
11
find that memorisation generally decreases with increasing levels of DP privacy protection (i.e. smaller ε values; Fig. 5a&c), at the cost of reduced diagnostic performance on the held-out test set (Fig. 5b&d). However, this expected privacy-utility trade-off [37] conceals a more consequential distinction between the two accounting regimes. Record-level DP, the variant typically used across research and industry for its simplicity, substantially reduced but did not eliminate memorisation bias. Even at the strongest privacy budget tested (ε = 1), we still detected significant memorisation-induced prediction changes in a substantial share of historical and future records (Fig. 5a&c), with effect sizes on future records at the largest budgets reaching over 40 percentage points for both datasets (Extended Data Fig. 6). These findings match results on patient-level membership inference attack susceptibility under record-level DP protection reported by Knolle et al. [7]. By contrast, models trained with patient-level DP protection showed almost no signs of memorisation bias. Although we still detected memorisation on patients’ historical records at weaker privacy budgets for MIMIC-ECG (up to N = 74, 070 records at ε = 103 ), this memorisation essentially did not transfer to patients’ future records, where we detected at most two affected records across all tested budgets (N = 0, 0, 0, 2 for ε = 1, 101 , 102 , 103 ; Fig. 5a). For the larger HEEDB dataset, patient-level DP protection yielded no significant changes in predictions across all tested budgets, in either the historical or future records. Correspondingly, ecCDF plots of energy test statistics and memorisation effect sizes on future records for ε ≤ 103 patient-level DP do not deviate substantially from random partitioning baselines (Extended Data Fig. 6). Together, these findings suggest that DP, implemented at the patient-level, is much more effective at preventing memorisation bias than its record-level counterpart.
Discussion We present data from the first investigation into longitudinal AI memorisation, building on our preliminary, earlier work [39]. Using simulation experiments grounded in real-world longitudinal patient data, we provide early evidence that AI memorisation could translate into downstream diagnostic harm for data contributors. Together, our results indicate that patients whose historical records were used to train an AI model may face an elevated risk of missed diagnoses when they encounter the same model in their future care journey. The number of missed diagnoses attributable to memorisation bias in our simulated deployment experiment is modest. Interpreting this result requires care: de novo cases are rare in routinely collected data, which limits their absolute count in the datasets we investigate, while our temporal split strategy deliberately enriched for such cases. Our estimates should therefore be read as characterising the existence and direction of the effect rather than its expected incidence in deployment. There is reason to expect the number of missed diagnoses attributable to memorisation bias to increase in the future, although we do not test this directly. Prior research has shown that the proportion of training data a model memorises increases with model capacity [6, 7, 9, 10, 40]. This is concerning, given that current AI model development is guided by “scaling laws” [41], which drive rapid growth in both model and dataset sizes in pursuit of improved performance. As larger AI models are trained on historical data sourced
12
from ever-larger patient populations, the absolute number of individuals affected by memorisation bias is likely to rise drastically, exacerbating the risks we identify here. Our findings have immediate and far-reaching implications. Anonymisation has long been assumed to protect data contributors from harm, but our results show that this may no longer hold true in the era of medical AI. Anonymisation cannot and will not protect contributors against future diagnostic errors caused by memorisation bias. Worse, it may even obscure who is at risk. Because AI training datasets are routinely anonymised, neither patients, nor model developers, nor healthcare providers can easily tell who is affected once a model is deployed. Given current deployment practices and infrastructure, it may thus not be straightforward to protect data contributors by excluding them from the AI-assisted interpretation of their own future data. This suggests that the common practice of building medical AI models using anonymised, historical patient data needs a fundamental reassessment and motivates the search for principled risk-mitigation techniques. Effectively mitigating risks from memorisation bias will require concerted efforts by model developers, independent researchers, regulatory authorities and other key stakeholders. This is because the most intuitive mitigation strategies each raise practical difficulties. The simplest option would be to establish new communication channels between training dataset curators and healthcare providers, enabling the exclusion of data contributors from the AI-assisted interpretation of their future data. Such channels, however, could pose privacy risks by revealing membership information [5, 7] and run counter to the purpose of the de-identification procedures currently in use. Exclusion could also be made more targeted by using membership inference attacks as memorisation detectors, flagging patients whose records appear to have been memorised during training. This, too, comes with practical issues: membership inference attacks are computationally expensive to evaluate at patient-level resolution [7], which would be required, as a missed detection leaves a patient unprotected. Underlying all of these approaches is the assumption that patients whose data were not used for model training still exist. Current trends in model development based on scaling laws [41] suggest this may not hold in the future. As a model’s training dataset size approaches population coverage, exclusion becomes self-defeating: withholding a model from a substantial share of a patient population limits its practical usefulness and raises ethical questions about disparities in the provision of care. DP [37] could be one of potentially multiple risk mitigation strategies that do not require excluding data contributors from prospective model deployment, but it requires careful implementation. Our results underscore that the choice of privacy unit for implementing DP is consequential. Record-level DP, the variant most commonly used across research and industry for its simplicity, substantially reduced but did not eliminate memorisation bias in our experiments, with significant memorisation effects still detectable even at ε = 1, the strongest level of protection we tested. Patient-level DP, by contrast, was much more effective at preventing memorisation bias. However, even patient-level DP is not a panacea. Patients with duplicate entries in a database, or familial and hereditary similarities between distinct patients, can violate the assumption that each protected unit is independent, and thereby weaken effective protection. Beyond DP, our results suggest that any effective mitigation will
13
need to protect patients rather than records. Because memorisation bias arises from the self-similarity of a patient’s records over time, safeguards applied at the record level leave the underlying vulnerability intact. Establishing effective patient-level protection at scale, while minimising utility cost, remains an open research problem. Our study has several limitations. First, patients with long follow-up are rare in the datasets we investigate, so our estimates of the persistence of memorisation bias at longer acquisition intervals rest on comparatively few records and are thus likely conservative. Second, our analysis of diagnostic harm is simulated rather than observed. We estimate diagnostic effects from changes in model predictions under post-hoc thresholds, with models assessed in isolation rather than in a clinician-inthe-loop setting, so the true downstream clinical impact of memorisation bias remains to be quantified prospectively. Our results also suggest that prospective trials would require careful design: without a de novo case stratification, a trial would likely reach incorrect conclusions about diagnostic primary endpoints for data contributors, as patients often return with unchanged health states. Third, the utility cost we observe for DP is likely an overestimate. Our patient-level DP results rely on a naive implementation that discards all but one record per patient before applying record-level accounting, so the reduced diagnostic performance on the unseen test data we report reflects this data loss as much as the privacy mechanism itself. On sufficiently large datasets, patient-level DP approaches that retain multiple records per patient and build on the engineering insights of De et al. [42] and Mckenna et al. [43] would likely close some of this gap. Fourth, we did not perform a subgroup analysis. Prior research [7] showed that membership inference risk is not distributed equally, with a disproportionate share of the privacy risk burden falling on patient groups underrepresented in the training data; whether these disparities extend to memorisation bias remains an open question. Fifth, we studied diagnostic models trained for supervised classification tasks. Prior research on memorisation [6, 8, 44, 45] suggests that the phenomenon we report in this study should also arise in other settings (such as e.g., models trained for segmentation/survival prediction, or foundation and large language models), but how it manifests in each of them requires further dedicated research. Together, our findings indicate that including a patient’s personal medical data in an AI model’s training dataset can systematically alter the model’s predictions on the patient’s unseen future data, and that these changes are directionally unfavourable when their health state changes. While the magnitude of the resulting clinical harm remains to be quantified in prospective studies, our findings suggest that memorisation warrants consideration not only as a privacy risk but as a potential source of diagnostic harm concentrated on the very individuals who make medical AI possible.
On the connection to Knolle et al. [7] At a purely technical level, we build closely on the findings of Knolle et al. [7] where the authors demonstrated that the predicted probability an AI model assigns to a given record shifts substantially when that same record is included in the model’s training dataset. For some patients, this shift in predicted probability enables near-perfect success rates for membership inference attacks.
14
With this study, we provide evidence for a closely related phenomenon. Specifically, we show that predicted probability shifts caused by including a patient’s historical records in a model’s training dataset are not limited to those records; they also extend to unseen future records from the same data contributors. Notably, this includes records in which the patient’s health state has changed, suggesting that, in some cases, models rely on patient-specific rather than disease-specific features to make their predictions. This finding carries fundamentally different implications from the ones presented by Knolle et al. [7]. While Knolle et al. [7] focuses on the vulnerability to targeted privacy attacks, our results suggest that using anonymised patient data for AI model training could cause tangible, real-world harm to data contributors in the form of future misdiagnoses. Crucially, this harm arises naturally from the deployment of models trained on data sourced from ever-larger patient populations and does not require interference from adversarial actors.
15
References [1] Zhang, C., Bengio, S., Hardt, M., Recht, B., Vinyals, O.: Understanding deep learning (still) requires rethinking generalization. Communications of the ACM 64(3), 107–115 (2021) [2] Feldman, V., Zhang, C.: What neural networks memorize and why: Discovering the long tail via influence estimation. Advances in Neural Information Processing Systems (2020) [3] Zhang, C., Ippolito, D., Lee, K., Jagielski, M., Tramèr, F., Carlini, N.: Counterfactual memorization in neural language models. Thirty-seventh Conference on Neural Information Processing Systems (2023) [4] Tonekaboni, S., Stempfle, L., Fallahpour, A., Gerych, W., Ghassemi, M.: An investigation of memorization risk in healthcare foundation models. In: The Thirty-ninth Annual Conference on Neural Information Processing Systems (2025). https://openreview.net/forum?id=NMvMYtRjkg [5] Shokri, R., Stronati, M., Song, C., Shmatikov, V.: Membership inference attacks against machine learning models. 2017 IEEE Symposium on Security and Privacy (SP) (2017). IEEE [6] Carlini, N., Chien, S., Nasr, M., Song, S., Terzis, A., Tramer, F.: Membership inference attacks from first principles. 2022 IEEE Symposium on Security and Privacy (SP) (2022). IEEE [7] Knolle, M.A., Menten, M.J., Jungmann, F., Meissen, F., Glocker, B., Rueckert, D., Kaissis, G.: Disparate privacy risks from medical AI. Nature (2026) [8] Carlini, N., Hayes, J., Nasr, M., Jagielski, M., Sehwag, V., Tramer, F., Balle, B., Ippolito, D., Wallace, E.: Extracting training data from diffusion models. 32nd USENIX Security Symposium (USENIX Security 23) (2023) [9] Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., et al.: Extracting training data from large language models. 30th USENIX Security Symposium (USENIX Security 21) (2021) [10] Nasr, M., Rando, J., Carlini, N., Hayase, J., Jagielski, M., Cooper, A.F., Ippolito, D., Choquette-Choo, C.A., Tramèr, F., Lee, K.: Scalable extraction of training data from aligned, production language models. The Thirteenth International Conference on Learning Representations (2025) [11] Esteva, A., Kuprel, B., Novoa, R.A., Ko, J., Swetter, S.M., Blau, H.M., Thrun, S.: Dermatologist-level classification of skin cancer with deep neural networks. Nature 542(7639), 115–118 (2017)
16
[12] De Fauw, J., Ledsam, J.R., Romera-Paredes, B., Nikolov, S., Tomasev, N., Blackwell, S., Askham, H., Glorot, X., O’Donoghue, B., Visentin, D., et al.: Clinically applicable deep learning for diagnosis and referral in retinal disease. Nature Medicine 24(9), 1342–1350 (2018) [13] McKinney, S.M., Sieniek, M., Godbole, V., Godwin, J., Antropova, N., Ashrafian, H., Back, T., Chesus, M., Corrado, G.S., Darzi, A., et al.: International evaluation of an AI system for breast cancer screening. Nature 577(7788), 89–94 (2020) [14] McDuff, D., Schaekermann, M., Tu, T., Palepu, A., Wang, A., Garrison, J., Singhal, K., Sharma, Y., Azizi, S., Kulkarni, K., et al.: Towards accurate differential diagnosis with large language models. Nature 642(8067), 451–457 (2025) [15] Brodeur, P.G., Buckley, T.A., Kanjee, Z., Goh, E., Ling, E.B., Jain, P., Cabral, S., Abdulnour, R.-E., Haimovich, A.D., Freed, J.A., Olson, A., Morgan, D.J., Hom, J., Gallo, R., McCoy, L.G., Mombini, H., Lucas, C., Fotoohi, M., Gwiazdon, M., Restifo, D., Restrepo, D., Horvitz, E., Chen, J., Manrai, A.K., Rodman, A.: Performance of a large language model on the reasoning tasks of a physician. Science 392(6797), 524–527 (2026) https://doi.org/10.1126/science.adz4433 [16] Gommers, J., Hernström, V., Josefsson, V., Sartor, H., Schmidt, D., Hjelmgren, A., Larsson, A.-M., Hofvind, S., Andersson, I., Rosso, A., et al.: Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without AI in the MASAI study: a randomised, controlled, non-inferiority, single-blinded, population-based, screening-accuracy trial. The Lancet 407(10527), 505–514 (2026) [17] Eisemann, N., Bunk, S., Mukama, T., Baltus, H., Elsner, S.A., Gomille, T., Hecht, G., Heywang-Köbrunner, S., Rathmann, R., Siegmann-Luz, K., et al.: Nationwide real-world implementation of AI for cancer detection in population-based mammography screening. Nature Medicine 31(3), 917–924 (2025) [18] Yao, X., Rushlow, D.R., Inselman, J.W., McCoy, R.G., Thacher, T.D., Behnken, E.M., Bernard, M.E., Rosas, S.L., Akfaly, A., Misra, A., et al.: Artificial intelligence–enabled electrocardiograms for identification of patients with low ejection fraction: a pragmatic, randomized clinical trial. Nature Medicine 27(5), 815–819 (2021) [19] Bartlett, P., Freund, Y., Lee, W.S., Schapire, R.E.: Boosting the margin: A new explanation for the effectiveness of voting methods. The annals of statistics 26(5), 1651–1686 (1998) [20] Schapire, R.E.: Explaining AdaBoost. Empirical inference, 37–52 (2013) [21] Wyner, A.J., Olson, M., Bleich, J., Mease, D.: Explaining the success of AdaBoost and random forests as interpolating classifiers. Journal of Machine Learning Research 18(48), 1–33 (2017)
17
[22] Zarifzadeh, S., Liu, P., Shokri, R.: Low-cost high-power membership inference attacks. Proceedings of the 41st International Conference on Machine Learning (2024) [23] Feldman, V.: Does learning require memorization? a short tale about a long tail. Proceedings of the 52nd Annual ACM SIGACT Symposium on Theory of Computing (2020) [24] Leibig, C., Brehmer, M., Bunk, S., Byng, D., Pinker, K., Umutlu, L.: Combining the strengths of radiologists and AI for breast cancer screening: a retrospective analysis. The Lancet Digital Health 4(7), 507–519 (2022) [25] Taylor-Phillips, S., Gilbert, F.J.: Artificial intelligence to help healthcare professionals detect cancer in the UK breast screening programme (EDITH). ISRCTN registry, ISRCTN81384017. Registered 26 November 2025 (2025). https:// doi.org/10.1186/ISRCTN81384017 . https://www.isrctn.com/ISRCTN81384017 [26] European Parliament and Council of the European Union: Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence and amending Regulations (EC) No 300/2008, (EU) No 167/2013, (EU) No 168/2013, (EU) 2018/858, (EU) 2018/1139 and (EU) 2019/2144 and Directives 2014/90/EU, (EU) 2016/797 and (EU) 2020/1828 (Artificial Intelligence Act). Official Journal of the European Union, OJ L, 2024/1689, 12 July 2024. Articles 10(3) and 10(4). Available at https://eur-lex.europa.eu/eli/reg/2024/1689/oj (2024). https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng [27] International Medical Device Regulators Forum: Good machine learning practice for medical device development: Guiding principles. Final Document IMDRF/AIML WG/N88 FINAL:2025, International Medical Device Regulators Forum (2025). https://www.imdrf .org/sites/default/files/2025-02/IMDRF AIML%20WG GMLP N88%20Final.pdf [28] Gow, B., Pollard, T., Nathanson, L.A., Johnson, A., Moody, B., Fernandes, C., Greenbaum, N., Waks, J.W., Eslami, P., Carbonati, T., Chaudhari, A., Herbst, E., Moukheiber, D., Berkowitz, S., Mark, R., Horng, S.: MIMIC-IV-ECG: Diagnostic Electrocardiogram Matched Subset. PhysioNet (2023) https://doi.org/10.13026/ 4nqg-sb35 . Version 1.0 [29] Johnson, A.E., Pollard, T.J., Berkowitz, S.J., Greenbaum, N.R., Lungren, M.P., Deng, C.-y., Mark, R.G., Horng, S.: MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports. Scientific Data (2019) [30] Xie, F., Zhou, J., Lee, J.W., Tan, M., Li, S., Rajnthern, L.S., Chee, M.L., Chakraborty, B., Wong, A.-K.I., Dagan, A., et al.: Benchmarking emergency department prediction models with machine learning and public electronic health
18
records. Scientific Data (2022) [31] Koscova, Z., Li, Q., Robichaux, C., Junior, V.M., Ghanta, M., Gupta, A., Rosand, J., Aguirre, A.D., Reinertsen, E., Song, S., et al.: The Harvard-Emory ECG database. Scientific Data 13(1), 516 (2026) [32] Johnson, A., Bulgarelli, L., Pollard, T., Celi, L., Mark, R., Horng, S.: MIMICIV-ED (version 1.0). PhysioNet (2021) [33] Gorishniy, Y., Rubachev, I., Khrulkov, V., Babenko, A.: Revisiting deep learning models for tabular data. Advances in neural information processing systems (2021) [34] Abdulazeem, H., Whitelaw, S., Schauberger, G., Klug, S.J.: A systematic review of clinical health conditions predicted by machine learning diagnostic and prognostic models trained or validated using real-world primary health care data. PLoS One 18(9), 0274276 (2023) [35] Youden, W.J.: Index for rating diagnostic tests. Cancer 3(1), 32–35 (1950) [36] Peirce, C.S.: The numerical measure of the success of predictions. Science (93), 453–454 (1884) [37] Dwork, C., Roth, A., et al.: The algorithmic foundations of differential privacy. Foundations and Trends in Theoretical Computer Science (2014) [38] Abadi, M., Chu, A., Goodfellow, I., McMahan, H.B., Mironov, I., Talwar, K., Zhang, L.: Deep learning with differential privacy. Proceedings of the 2016 ACM SIGSAC conference on computer and communications security (2016) [39] Knolle, M., Menten, M.J., Rueckert, D., Kaissis, G., Glocker, B.: Memorisation Bias: AI predictions for data contributors are biased towards their health states in the training data. In: International Workshop on Learning with Longitudinal Medical Images and Data, pp. 24–33 (2025). Springer [40] Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramer, F., Zhang, C.: Quantifying memorization across neural language models. International Conference on Learning Representations (2023) [41] Kaplan, J., McCandlish, S., Henighan, T., Brown, T.B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., Amodei, D.: Scaling laws for neural language models. arXiv preprint arXiv:2001.08361 (2020) [42] De, S., Berrada, L., Hayes, J., Smith, S.L., Balle, B.: Unlocking highaccuracy differentially private image classification through scale. arXiv preprint arXiv:2204.13650 (2022) [43] Mckenna, R., Huang, Y., Sinha, A., Balle, B., Charles, Z., Choquette-Choo, C.A.,
19
Ghazi, B., Kaissis, G., Kumar, R., Liu, R., Yu, D., Zhang, C.: Scaling laws for differentially private language models. Proceedings of the 42nd International Conference on Machine Learning (2025) [44] Hayes, J., Shumailov, I., Choquette-Choo, C.A., Jagielski, M., Kaissis, G., Nasr, M., Annamalai, M.S.M.S., Mireshghallah, N., Shilov, I., Meeus, M., Montjoye, Y.-A., Lee, K., Boenisch, F., Dziedzic, A., Cooper, A.F.: Exploring the limits of strong membership inference attacks on large language models. In: The Thirtyninth Annual Conference on Neural Information Processing Systems (2025). https://openreview.net/forum?id=x0i7wvRLHK [45] Wang, W., Kaleem, M.A., Dziedzic, A., Backes, M., Papernot, N., Boenisch, F.: Memorization in self-supervised learning improves downstream generalization. In: International Conference on Learning Representations, vol. 2024, pp. 39935–39959 (2024)
20
Methods MIMIC-ECG processing and model training details. MIMIC-ECG [28, 32] is a large electrocardiography dataset consisting of 799, 981 twelve-lead electrocardiograms (ECGs) collected from 161, 332 patients at the Beth Israel Deaconess Medical Center. Each ECG is 10 seconds long and paired with structured diagnostic labels that we derived by regular-expression matching on machine-generated text reports. Signals were resampled to 250 Hz and pre-processed using a 50 Hz notch filter, a bandpass filter between 0.67 Hz and 40 Hz, and a median filter to reduce baseline wander. Signal data were normalised to zero mean and unit variance. Models for this dataset were trained without data augmentation to perform multi-label classification across the following 15 classes: sinus rhythm, tachycardia, bradycardia, atrial fibrillation, atrial flutter, atrioventricular block, right bundle branch block (RBBB), left bundle branch block (LBBB), pacemaker, infarct, premature ventricular contraction, abnormal ECG, borderline ECG, lateral ST-T changes and fascicular block. For model training, we used a modified vision transformer (ViT-S-25) [46] where we replaced two-dimensional convolution layers with their one-dimensional counterparts. On the unseen test dataset (N = 49, 655 records), a ViT-S-25 with around 14 million parameters, trained on historical records from a randomly selected 50% of patients, achieves a macro-average AUROC of 93.45 ± 0.18% across all classes. MIMIC-CXR processing and model training details. MIMIC-CXR [29] is a large chest radiograph dataset consisting of 357, 522 chest radiographs from 61, 868 patients at the Beth Israel Deaconess Medical Center. We utilise MIMIC-CXR-JPG [47], a subsequent re-release of the original dataset which contains images in “.jpg” format and structured labels derived from free-text radiology reports. The structured labels indicate the presence or absence of 14 common thoracic conditions. For this dataset, we trained Densenet-121 [48] models (pre-trained on ImageNet [49], a large natural image dataset) with data augmentation (random horizontal flipping, random pixel shifts, and random rotations) to detect the presence of the following 14 classes: no finding, enlarged cardiomediastinum, cardiomegaly, lung opacity, lung lesion, oedema, consolidation, pneumonia, atelectasis, pneumothorax, pleural effusion, pleural other, fracture, support devices. On the unseen test dataset (N = 57, 529 records), a DenseNet-121 with around 7 million parameters trained on the historical records from a random 50% patient subset achieves a macro-average AUROC score of 79.68 ± 0.21% across all classes. MIMIC-IV-ED processing and model training details. MIMIC-IV-ED [30, 32] is a large electronic health record dataset comprising 418, 007 records from 201, 213 patients at the emergency department of the Beth Israel Deaconess Medical Center. We followed the pre-processing steps from [30] and trained models without data augmentation to predict hospitalisation and critical outcomes using 64 clinical features. For this dataset, we trained random forest models using the scikit-learn [50] implementation with n estimators=100, min samples leaf=4, and otherwise default parameters. These model hyperparameters were chosen because they yielded the highest diagnostic performance on the validation dataset. On the unseen test dataset (N = 20, 872
21
records), a random forest trained on the historical records from a random 50% patient subset achieves a macro-average AUROC score of 83.39 ± 0.08% across all classes.
HEEDB processing and model training details. The Harvard-Emory ECG Database (HEEDB) [31] is a large electrocardiography dataset from the Massachusetts General Hospital and Emory University Hospital. Because we are interested in patients with long follow-up data, we use only data from the Massachusetts General Hospital, comprising of 10, 457, 413 twelve-lead ECGs from 1, 802, 104 patients. Each ECG is 10 seconds long and paired with structured diagnostic labels derived from a semiautomated diagnosis system. Signals were resampled to a resolution of 250 Hz and pre-processed using a 50 Hz notch filter, a bandpass filter between 0.67 Hz and 40 Hz and a median filter to reduce baseline wander. Signal data were normalised to zero mean and unit variance. Models for this dataset were trained without data augmentation to perform multi-label classification across the following 36 classes: normal ECG, abnormal ECG, normal sinus rhythm, sinus bradycardia, atrial fibrillation, sinus tachycardia, left axis deviation, premature ventricular complexes, borderline ECG, right bundle branch block, septal infarct, non-specific T-wave abnormality, premature atrial complexes, anterior infarct, left bundle branch block, lateral infarct, non-specific ST abnormality, left ventricular hypertrophy, atrial flutter, left anterior fascicular block, right axis deviation, anteroseptal infarct, anterolateral infarct, right atrial enlargement, inferior infarct, bi-fascicular block, left posterior fascicular block, bi-atrial enlargement, 1st degree atrioventricular block, inferior-posterior infarct, supraventricular tachycardia, wide QRS tachycardia, Wolff-Parkinson-White syndrome, acute pericarditis, posterior infarct, bi-ventricular hypertrophy. For model training, we used a modified vision transformer (ViT-S-25) [46] where two-dimensional convolution layers were replaced with their one-dimensional counterparts. On the unseen test dataset (N = 143, 062 records), a ViT-S-25 with around 14 million parameters, trained on historical records from a random 50% subset of patients, achieves a macro-average AUROC score of 95.38 ± 0.22% across all classes. General model training details. All models were trained to perform supervised multi-label classification using the AdamW [51] optimiser, exponential moving parameter average (EMA), weight decay, and a learning rate schedule in form of a cosine decay with linear warm-up. Optimisation hyperparameter values were determined via a random search to maximise diagnostic performance (macro-average AUROC across all available classes) on unseen validation data. Specifically, for each dataset, a random search with T = 50 trials over suitable values for the learning rate, weight decay, momentum and EMA decay was conducted (results are reported in Supplementary Material Table 1). To prevent overfitting, which is known to exacerbate memorisation [52], model checkpointing based on the validation loss was employed, ensuring that only the model weights with the best generalisation performance were retained after training terminated. Unless stated otherwise, model performance figures are reported as the mean ± standard deviation across the M = 200 models trained for the respective configuration.
22
Patient-level temporal split strategy. To obtain the historical and future record datasets, we split the records of each patient in the follow-up cohort at a patientspecific time point. Since patients often present with already documented conditions in routinely collected data, splitting at a fixed calendar date or at a fixed fraction of each patient’s timeline yields future record datasets containing too few de novo cases to reliably study the negative impact of memorisation. We therefore determined the split point adaptively for each patient to enrich the future record datasets for de novo cases. Concretely, we sorted each patient’s records chronologically and, at every record, counted how many of the conditions ever documented for that patient had not yet appeared. As this count is monotonically decreasing over time, we placed the split point just before the last record at which at least C conditions remained undocumented. By construction, a patient’s future records then contain the first documented occurrence of at least C conditions while the historical dataset retains as many of their records as this constraint permits. Patients whose health state does not change across their records, either because they contribute a single record only or because the same conditions are documented throughout, are not split, and all of their records are assigned to the historical dataset. For patients with too few distinct health states to yield C de novo cases (counting the absence of any documented condition as a state), we reduced C for that patient to the largest value that still leaves at least one record in the historical dataset. The split, therefore, guarantees that every patient represented in the future record dataset also contributes at least one record to the historical dataset used for model training. We chose C separately for each investigated dataset to balance the number of historical records available for model training against the number of de novo cases available for our simulated deployment experiment. For MIMIC-ECG, MIMIC-CXR, MIMIC-IV-ED and HEEDB we used C = 1, 1, 1, 2, respectively. Statistical testing of prediction changes. For each record, we tested whether the predictions (vectors of predicted probabilities or all available classes) of models trained on the corresponding patient’s historical data differed significantly from those of models not trained on that data, an approach conceptually related to membership inference. More specifically, we trained M = 200 models on the historical data of randomly selected patient subsets and then partitioned the models by the inclusion/exclusion of each patient in the respective model’s training subset. We then compared the predictions of these two model groups using the energy distance [53], a nonparametric multivariate two-sample statistic that quantifies distributional divergence based on pairwise Euclidean distances between and within samples. For the predictions u1 , . . . , un ∈ Rp of the n models trained on the patient’s data and v1 , . . . , vm ∈ Rp of the m models not trained on it, the energy distance is: En,m (u, v) =
n m n m 2 XX 1 X 1 X ∥ui − vj ∥ − 2 ∥ui − uj ∥ − 2 ∥vi − vj ∥, nm i=1 j=1 n i,j=1 m i,j=1
(1)
where ∥·∥ denotes the Euclidean norm, n + m = M = 200 and p is the number of classes in the respective dataset. The corresponding population quantity is non-negative and zero if and only if the two underlying distributions are identical. Note that subsets were drawn using a balanced random subset design following Carlini et al. [6]: each
23
patient was assigned to a uniformly random half of the M models, independently of all other patients. This guarantees n = m = M/2 for every patient (and thus all of their respective records), which independent Bernoulli sampling would achieve only in expectation. Computation was performed using the Energy test implemented in the hyppo package (v0.5.2) [54]. Energy test p values were calculated using the fast chi-squared approximation [55], which is appropriate at the group sizes used here (n = m = 100), and adjusted for multiple comparisons using the Benjamini–Hochberg procedure to control the false discovery rate at α = 0.05.
Statistical testing of receiver operating point differences. To translate memorisation-induced prediction changes into clinically meaningful quantities, we compared the sensitivity and specificity of the two model groups (those whose training subset included the respective patient’s historical data and those that did not) at a fixed operating point. Class-specific decision thresholds were determined by maximising Youden’s J statistic [35, 36] on the mean predicted probabilities across all M models for the unseen test dataset, and are therefore independent of the subset assignments subsequently tested. Comparisons were restricted to future records for which we detected memorisation and performed separately for de novo conditions, where each positive case was matched to a negative control drawn randomly from the same dataset on sex and 5-year age band, and for records of patients with unchanged health states. Every model’s predictions were binarised at these thresholds individually, retaining the information carried by all M models rather than collapsing each group into a single averaged prediction. For each record j we computed the difference δj in the proportion of correct classifications between the two groups (models trained on that patient’s historical records minus those not trained on them), and took the PN mean (T = N1 j=1 δj ) across the N records as the test statistic. Evaluated over positive cases, T is the difference in sensitivity, and over negative cases, the difference in specificity. Because every patient is assigned to a uniformly random half of the M models independently of all other patients, the assignment matrix is exchangeable in the model index and, under the null hypothesis that inclusion does not alter predictions, independent of the observed classifications. We therefore obtained an exact randomisation test by permuting model labels against subset assignments over B = 100, 000 permutations. Permuting entire assignment vectors preserves the correlation between records of the same patient induced by patient-level subsampling, so no further adjustment for clustering was required. With b+ and b− , the numbers of permuted statistics at least as large and at least as small as the observed T , the one-sided p values are: p± =
b± + 1 . B+1
(2)
Note the addition of 1 to the numerator and denominator, which prevents a p value of zero being reported from a finite number of draws [56]. All reported p values are twosided, obtained by doubling the smaller tail as p = min{1, 2 min(p+ , p− )}. Confidence intervals were obtained by inverting the same permutation test at the corrected level applied to the p values: the interval is the set of effects d for which the shifted statistic
24
T − d is not rejected, which at per-comparison level a, with c = ⌈a(B + 1)/2⌉ − 1, gives ∗ ∗ ∗ [ T − T(B−c+1) , T − T(c) ], where T(i) denotes the i-th smallest permuted statistic. Tabular data model comparison on MIMIC-IV-ED. The tabular setting permits model classes that the imaging and signal datasets do not, so we conducted the analysis described previously on MIMIC-IV-ED with three architectures spanning very different inductive biases: the random forest described above, L2 -regularised logistic regression (C = 3 × 10−5 , lbfgs solver, maximum 5000 iterations), and a fully connected residual network for tabular inputs [33] (“tabular ResNet”, 5 residual blocks of width 500, 3.8 million parameters, trained for 50 epochs with AdamW [51] using a learning rate of 10−2 , weight decay 1.0 and EMA decay 0.99). M = 200 models were trained per architecture under the identical subset protocol previously described. On the unseen test dataset (N = 20, 872 records), the random forest, the logistic regression model, and the tabular ResNet achieve macro-average AUROC scores of 83.39 ± 0.08%, 82.50 ± 0.07% and 83.37 ± 0.17%, respectively. Under random model partitioning, no historical or future records were flagged for memorisation for any of the three architectures. Differential privacy model training details. Models for the differential privacy risk mitigation experiments were trained with DP-Adam [57] [38] using the jax-privacy library [58]. Due to the limited computational resources available to us, model size had to be reduced from ViT-S-25 (14 million parameters) to ViT-T-25 (1.9 million parameters). Hyperparameters for these models were determined separately for MIMIMIC-ECG and HEEDB using a fixed privacy budget of ε = 100 and using a random search with T = 20 trials over suitable values for the learning rate and EMA decay parameter (see Supplementary Material Table 1 for resulting values). Weight decay, dropout or other forms of regularisation were not used. Under record-level protection, the unit of privacy is the individual record. For patient-level protection, we retained only the most recent record per patient, so that one training record corresponds to one patient and the guarantee extends from records to individuals; this reduces the training set from 475, 512 to 146, 332 historical records for MIMIC-ECG and from 6, 147, 511 to 1, 767, 104 for HEEDB. Models were trained at four privacy budgets, ε ∈ {1, 10, 100, 1000}, with δ set to the reciprocal of the number of (historical) training records of a given random patient subset and therefore differing between record- and patient-level DP dataset variants: δ ≃ 2.08 × 10−6 and 1.63 × 10−7 for record-level MIMIC-ECG and HEEDB, and δ ≃ 6.83 × 10−6 and 5.66 × 10−7 for the corresponding patient-level variants. Privacy accounting used the privacy loss distribution accountant implemented in jax-privacy (v2.1.0), with the noise multiplier calibrated to the target ε before training rather than computed post hoc. Note that the diagnostic performance of the models could likely be improved by performing a separate hyperparameter search for each privacy budget and integrating the engineering insights of De et al. [42] and Mckenna et al. [43]. Computational resources. Reproducing all of our experiments requires the training of 5, 200 models and approximately 10, 000 GPU-hours on NVIDIA A100 GPUs. The non-private experiments (200 models per dataset; 200 each for the three MIMICIV-ED architectures) account for 1, 681 GPU-hours, of which HEEDB alone accounts
25
for 1, 069 GPU-hours at 5.3 GPU-hours per model. The differential privacy risk mitigation experiments (200 models for each of the four privacy budgets, at both recordand patient-level protection, for HEEDB and MIMIC-ECG) account for the remaining approximately 8, 200 GPU-hours. Training the logistic regression and random forest baselines for MIMIC-IV-ED requires approximately 4 CPU-hours. These figures exclude hyperparameter search (a further ∼250 GPU-hours) and data preprocessing. Storing the model outputs required for our analysis requires approximately 2 TB of disk space.
References [46] Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., Houlsby, N.: An image is worth 16x16 words: Transformers for image recognition at scale. International Conference on Learning Representations (2021) [47] Johnson, A.E., Pollard, T.J., Greenbaum, N.R., Lungren, M.P., Deng, C.-y., Peng, Y., Lu, Z., Mark, R.G., Berkowitz, S.J., Horng, S.: MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs. arXiv preprint arXiv:1901.07042 (2019) [48] Huang, G., Liu, Z., Van Der Maaten, L., Weinberger, K.Q.: Densely connected convolutional networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4700–4708 (2017) [49] Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., Fei-Fei, L.: Imagenet: A largescale hierarchical image database. In: 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 248–255 (2009). Ieee [50] Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., Duchesnay, E.: Scikit-learn: Machine learning in Python. Journal of Machine Learning Research 12, 2825–2830 (2011) [51] Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. In: International Conference on Learning Representations (2019). https://openreview.net/forum?id=Bkg6RiCqY7 [52] Yeom, S., Giacomelli, I., Fredrikson, M., Jha, S.: Privacy risk in machine learning: Analyzing the connection to overfitting. 2018 IEEE 31st Computer Security Foundations Symposium (CSF) (2018) [53] Székely, G.J., Rizzo, M.L., et al.: Testing for equal distributions in high dimension. InterStat 5(16.10), 1249–1272 (2004) [54] Panda, S., Palaniappan, S., Xiong, J., Bridgeford, E.W., Mehta, R., Shen, C., Vogelstein, J.T.: hyppo: A multivariate hypothesis testing python package. arXiv
26
preprint arXiv:1907.02088 (2019) [55] Shen, C., Panda, S., Vogelstein, J.T.: The chi-square test of distance correlation. Journal of Computational and Graphical Statistics 31(1), 254–262 (2022) https: //doi.org/10.1080/10618600.2021.1938585 [56] Phipson, B., Smyth, G.K.: Permutation P-values should never be zero: Calculating exact P-values when permutations are randomly drawn. Statistical Applications in Genetics and Molecular Biology 9(1) (2010) https://doi.org/ 10.2202/1544-6115.1585 [57] Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization. Proceedings of the 3rd International Conference on Learning Representations (2015) [58] Balle, B., Berrada, L., Charles, Z., Choquette-Choo, C.A., De, S., Doroshenko, V., Dvijotham, D., Galen, A., Ganesh, A., Ghalebikesabi, S., Hayes, J., Kairouz, P., McKenna, R., McMahan, B., Pappu, A., Ponomareva, N., Pravilov, M., Rush, K., Smith, S.L., Stanforth, R.: JAX-Privacy: Algorithms for Privacy-Preserving Machine Learning in JAX. http://github.com/google-deepmind/jax privacy
Data availability. All datasets used in this study are publicly available for research purposes. See GitHub repository (https://github.com/moritzknolle/memorisation bias) for access instructions. Code availability. The code is available on GitHub (https://github.com/ moritzknolle/memorisation bias). Acknowledgments. We thank Georg Schmidt and the members of the AIM Lab for their feedback and support. The authors gratefully acknowledge computational resources provided by the Leibniz Supercomputing Centre of the Bavarian Academy of Sciences and Humanities.
27
Extended Data
28
b
a 93.45±0.18 90
85 83.39±0.08 80
MIMIC-ECG MIMIC-CXR MIMIC-IV-ED HEEDB Random partitioning
10−1 1 - cumulative probability
Test macro AUROC (%)
100
95.38±0.22
95
79.68±0.21
10−2 10−3 10−4 10−5 10−6
0.8
1.0
0.0
0.2
0.4
d
c 100
10−2 10−3 10−4 10−5 10−6
25 50 75 100 125
0
20
40
60
0
20 40 60 80 100
0
105 104 103 102 101 100
100 200 300 400 500 LR
Acquisition interval (months)
e
Training set 95.50 ±0.02
96
83.57 ±0.19
84.0 83.39 ±0.08
94 83.5
92 87.19 ±0.68
90
83.0
88 85.07 ±0.08
82.50 ±0.07
82.5
84
1 - cumulative probability
Diagnostic performance (%)
Future records
Historical records
RF
Future records
DNN
Historical records
100
98
86
f
Test set
Historical records (N = 321,530)
106
−1
0
1.0
Future records (N = 65,402)
Memorised records All records Records with significant change in prediction (count)
1 - cumulative probability
10
0.8
n=300,740
M
0.6
Energy test statistic
n=11,603
0.6
n=2,459
0.4
Energy test statistic
n=83
IM
0.2
n=1,469
IC
H
-IV
0.0
n=25
-E D EE D B
XR
G
-C
-E C
IC IM
IC IM
M
M
10−7
10−1 10−2 10−3 10−4 10−5
0 LR
RF
DNN
LR
RF
DNN
10
0
20
40
0.0
Change in predicted probability (%) LR
RF
0.5
0.0
0.5
1.0
Energy test statistic DNN
Random partitioning
Fig. 1: Additional data on generalisation performance, energy test statistics, acquisition intervals, and memorisation effects for different tabular data architectures. a–c, Results for four datasets: MIMIC-ECG, MIMIC-CXR, MIMIC-IV-ED and HEEDB. a, Diagnostic performance of the random subset models on the unseen test sets (M = 200 models per dataset); coloured points, macro average AUROC (across all available classes) of individual models; white markers with error bars denote mean ± s.d. b, ecCDF of the energy test statistics on unseen future records (left) and historical records used for model training (right). All test statistics were derived from energy-based hypothesis tests comparing the two groups’ predictions (M = 100 models per group for each patient). c, ecCDF of the acquisition intervals of all future records (dashed lines) and the subset of future records for which we detected memorisation (solid lines). d–f, Memorisation across architectures on MIMIC-IV-ED (tabular data): L2 -regularised logistic regression (LR), random forest (RF) and tabular ResNet (DNN), each trained with the identical random subset protocol (M = 200 models each). d, Number of records with significantly different predictions (Benjamini–Hochberg FDR correction, α = 0.05) among future (N = 65,402) and historical (N = 321,530) records. e, Diagnostic performance on the training (left) and unseen test set (right), plotted as in a; note the independent y-axes. f, ecCDF of effect sizes (left pair) and energy test statistics (right pair), for future and historical records. Solid lines in b,f indicate results for models partitioned by patient inclusion/exclusion. Dashed lines in b,f indicate random partitioning baselines.
29
MIMIC-ECG
MIMIC-CXR
MIMIC-IV-ED
HEEDB
17.5 70
Change in predicted probability (%)
40
15.0
60
12.5
50
10.0
40
7.5
30
5.0
20
10
2.5
10
0
0.0
0
60 30
50 40
20
30 20
10
0
99th percentile change in predicted probability (%)
9.0
3.00
8.5 8.0
Random partitioning
2.75
8
7.5
2.50
7
7.0
7.5
2.25 6
7.0 6.5 6.0
6.5
2.00
5
1.75
6.0
1.50
4
5.5
5.5
1.25 20.0
99.9th percentile change in predicted probability (%)
70
Random partitioning
7
18 17.5
16
6
16
14
15.0
5
14 12.5
12
4
12 10.0
10
3
10 7.5 8
2
5.0
8
70 70
Maximum change in predicted probability (%)
40
15.0
Random partitioning
60
60 12.5
50
10.0
40
30
7.5
30
20
5.0
20
10
2.5
10
50
30
40 20
10
0
25
50
75
100
0
10
20
30
40
50
0
20
40
60
80
0
100
200
300
400
Acquisition interval (months)
Fig. 2: Effect sizes and survival-style summary statistics for increasing acquisition intervals. Columns correspond to the four datasets (MIMIC-ECG, MIMIC-CXR, MIMIC-IV-ED, HEEDB). The top row is a scatter plot of the effect size against the acquisition interval (in months) for each future record for which a statistically significant change in predicted probability was detected; the remaining rows show survival-style summary statistics of the effect size computed over all future records: 99th percentile (second row), 99.9th percentile (third row), and maximum (bottom row). For each future record, the effect size is the maximum absolute difference across diagnostic classes between the mean predicted probabilities of models trained with and without the corresponding patient’s historical records (in percentage points). In the lower three rows, at each value X on the x-axis, the summary statistic is computed over all records with an acquisition interval of at least X months. Points and solid lines indicate the partitioning by patient inclusion/exclusion; dashed lines show identical summary statistics for effect sizes computed from models partitioned at random. Within each column, all rows span the same acquisition-interval range, truncated at the largest value beyond which fewer than N = 500 records remain.
30
p
N+
N-
Fascicular block
**
1.6e-03
MIMIC-ECG MIMIC-CXR
p
**
1.6e-03
418
1,696
Lateral ST-T changes
**
1.6e-03
**
1.6e-03
465
961
Borderline ECG
**
1.6e-03
**
1.6e-03
236
1,655
Abnormal ECG
**
1.6e-03
**
1.6e-03
2,148
98
PVC
**
1.6e-03
**
1.6e-03
151
1,333
Infarct
**
1.6e-03
**
1.6e-03
1,035
591
Pacemaker
**
1.6e-03
**
1.6e-03
281
2,176
1.5e-01
28
2,421
Mean Difference (Sensitivity)
Condition Name
LBBB
3.0e-01
RBBB
**
1.6e-03
**
1.6e-03
349
1,800
AV block
**
1.6e-03
**
1.6e-03
268
1,336
Atrial flutter
**
1.6e-03
**
1.6e-03
64
1,901
Atrial fibrillation
**
1.6e-03
**
1.6e-03
306
1,460
Bradycardia
**
1.6e-03
**
1.6e-03
195
1,505
Tachycardia
**
1.6e-03
**
1.6e-03
179
1,622
Sinus
**
1.6e-03
**
1.6e-03
1,698
254
Support devices
**
1.6e-03
**
1.6e-03
996
928
Fracture
**
1.6e-03
**
1.6e-03
27
3,107
Pleural other
**
1.6e-03
*
1.6e-02
30
3,157
Pleural effusion
**
1.6e-03
**
1.6e-03
750
1,342
Pneumothorax
**
1.6e-03
**
1.6e-03
223
2,388
Atelectasis
**
1.6e-03
**
1.6e-03
491
1,079
Pneumonia
**
1.6e-03
**
1.6e-03
129
1,746
Consolidation
**
1.6e-03
**
1.6e-03
121
2,157
Oedema
**
1.6e-03
**
1.6e-03
302
1,778
Lung lesion
**
1.6e-03
**
1.6e-03
107
2,725
Lung opacity
**
1.6e-03
**
1.6e-03
708
847
Cardiomegaly
**
1.6e-03
**
1.6e-03
648
1,004 2,453
8.7e-02
*
1.3e-02
37
**
1.6e-03
**
1.6e-03
592
719
1.0
**
1.6e-03
1,131
1,131
Hospitalisation
**
1.6e-03
**
1.6e-03
753
152
Normal sinus rhythm
**
1.6e-03
**
1.6e-03
9,008
8,269
Sinus tachycardia
**
1.6e-03
**
1.6e-03
3,448
19,911
Premature atrial complexes
**
1.6e-03
**
1.6e-03
860
11,346
Inferior infarct
**
1.6e-03
**
1.6e-03
4,436
10,066
Lateral infarct
**
1.6e-03
**
1.6e-03
2,392
12,059
1st degree AV block
**
1.6e-03
**
1.6e-03
2,888
13,582
Anterior infarct
**
1.6e-03
**
1.6e-03
1,680
12,884
Non-specific T-wave abnormality
**
1.6e-03
**
1.6e-03
1,360
9,909
Septal infarct
**
1.6e-03
**
1.6e-03
1,672
14,153
Normal ECG
**
1.6e-03
**
1.6e-03
740
16,943
MIMIC-IV-ED
Enlarged cardiomediastinum
HEEDB
Mean Difference (Specificity)
No finding Critical Outcome
0
5
10
0
Sensitivity (%)
2
4
6
Specificity (%)
Fig. 3: Memorisation bias artificially inflates diagnostic sensitivity and specificity on future records of patients returning with unchanged health states. Memorisation-induced changes in diagnostic sensitivity and specificity on future records of patients whose health state was unchanged relative to their historical training records. Evaluated cases comprised positive cases of a condition for which at least one positive case was already present in that patient’s historical training records, and negative cases of a condition for which only negative historical cases were present. Each row represents a single condition: the left-hand panel shows the mean difference in sensitivity between models trained vs. not trained on patients’ historical data (M = 100 per group), the right-hand panel the corresponding difference in specificity. Point estimates to the right of zero indicate increased performance on data-contributing patients’ future records, those to the left a reduction. N + and N − denote the number of positive and negative cases for each condition, respectively. Condition-specific decision thresholds maximised Youden’s J statistic on the test set. Statistical significance was determined using exact, two-sided permutation tests that reassign model indices to the corresponding patient subset masks (B = 100,000 permutations per comparison); error bars denote simultaneous 95% confidence intervals obtained by inverting the same test. p values and confidence intervals were corrected across all C = 82 shown comparisons using the Bonferroni method. * and ** denote p ≤ 0.05 and p ≤ 0.01, respectively; the resolution of the permutation test floors raw p values at 2/(B + 1) ≈ 2 × 10−5 , and hence corrected p values at 1.6 × 10−3 . For HEEDB, we present results only for the 10 conditions with the largest number of de novo cases; data for all conditions are reported in Extended Data Fig. 5.
31
p
Mean Difference (Sensitivity)
Condition Name Normal sinus rhythm
**
1.4e-03
Sinus tachycardia
**
1.4e-03
Premature atrial complexes **
Inferior infarct Lateral infarct
Mean Difference (Specificity)
p
N+
N-
1.0
620
489
1.6e-02
589
470
5.5e-02
1.0
319
268
1.4e-03
1.0
294
245
1.0
1.0
292
238
*
1st degree AV block
**
1.4e-03
1.0
289
218
Anterior infarct
**
4.3e-03
1.0
286
239
Non-specific T-wave abnormality
**
1.4e-03
1.0
269
221
Septal infarct
**
1.4e-03
1.0
268
230
Normal ECG
**
1.4e-03
1.0
253
228
Sinus bradycardia
1.0
1.0
237
213
Bi-fascicular block
1.0
1.0
227
170
Premature ventricular complexes
1.0
1.0
220
164
Atrial fibrillation
**
1.4e-03
1.0
217
169
Left axis deviation
**
1.4e-03
1.0
212
174
1.0
1.0
202
163
2.7e-02
1.0
201
144
1.0
1.0
182
137
1.6e-02
1.0
172
166
1.0
1.0
168
139
1.4e-03
1.0
154
146
1.0
1.0
148
130
1.4e-03
1.0
144
122
1.0
1.0
139
117
4.6e-02
1.0
132
118
Right axis deviation
1.0
1.0
127
115
Left bundle branch block
1.0
1.0
101
78
Right atrial enlargement
1.0
1.0
85
74
Wide QRS tachycardia
1.0
1.0
83
73
2.9e-03
1.0
72
61
1.0
1.0
62
47
1.1e-01
1.0
62
52
1.4e-03
1.0
44
41
Wolff-Parkinson-White
1.0
1.0
25
24
Acute pericarditis
1.0
1.0
10
9
Posterior infarct
1.0
1.0
5
5
Anterolateral infarct *
Left anterior fascicular block Left ventricular hypertrophy
*
Borderline ECG Non-specific ST abnormality
**
Abnormal ECG Left posterior fascicular block Right bundle branch block
**
Anteroseptal infarct *
Atrial flutter
**
Inferior-posterior infarct Biventricular hypertrophy Supraventricular tachycardia Bi-atrial enlargement
**
−7.5
−5.0
−2.5
0.0
2.5
5.0
−1.5
Sensitivity (%)
−1.0
−0.5
0.0
0.5
1.0
Specificity (%)
Fig. 4: HEEDB data on diagnostic sensitivity and specificity changes on future de novo cases for all available conditions. Memorisation-induced changes in diagnostic sensitivity and specificity on future records of HEEDB patients returning with a de novo condition, i.e. one for which no positive cases were present in that patient’s historical training record(s). Each row represents a single condition: the lefthand panel shows the mean difference in sensitivity between models trained vs. not trained on patients’ historical data (M = 100 models per group for each patient), the right-hand panel the corresponding difference in specificity. Point estimates to the left of zero indicate reduced performance on data-contributing patients’ future records, those to the right an increase. N + and N − denote the number of positive and negative cases; cases comprised patients returning with a de novo positive diagnosis, each matched to a randomly selected age- and sex-matched control where possible. Condition-specific decision thresholds maximised Youden’s J statistic on the test set. Statistical significance was determined using exact, two-sided permutation tests that reassign model training run indices to the corresponding patient subset masks (B = 100,000 permutations per comparison); error bars denote simultaneous 95% confidence intervals obtained by inverting the same test. p values and confidence intervals were corrected across all C = 72 shown comparisons using the Bonferroni method. * and ** denote p ≤ 0.05 and p ≤ 0.01, respectively; the resolution of the permutation test floors raw p values at 2/(B +1) ≈ 2 × 10−5 , and hence corrected p values at 1.4 × 10−3 .
32
p
N+
N-
Normal sinus rhythm
**
1.4e-03
p
**
1.4e-03
9,008
8,269
Sinus tachycardia
**
1.4e-03
**
1.4e-03
3,448
19,911
Premature atrial complexes
**
1.4e-03
**
1.4e-03
860
11,346
Inferior infarct
**
1.4e-03
**
1.4e-03
4,436
10,066
Lateral infarct
**
1.4e-03
**
1.4e-03
2,392
12,059
1st degree AV block
**
1.4e-03
**
1.4e-03
2,888
13,582
Anterior infarct
**
1.4e-03
**
1.4e-03
1,680
12,884
Non-specific T-wave abnormality
**
1.4e-03
**
1.4e-03
1,360
9,909
Septal infarct
**
1.4e-03
**
1.4e-03
1,672
14,153
Normal ECG
**
1.4e-03
**
1.4e-03
740
16,943
Sinus bradycardia
**
1.4e-03
**
1.4e-03
1,643
11,350
Bi-fascicular block
**
1.4e-03
**
1.4e-03
1,240
19,549
Premature ventricular complexes
**
1.4e-03
**
1.4e-03
1,562
7,602
Atrial fibrillation
**
1.4e-03
**
1.4e-03
2,446
13,433
Left axis deviation
**
1.4e-03
**
1.4e-03
4,656
11,443
Anterolateral infarct
**
1.4e-03
**
1.4e-03
1,878
15,339
Left anterior fascicular block
**
1.4e-03
**
1.4e-03
1,881
16,864
Left ventricular hypertrophy
**
1.4e-03
**
1.4e-03
2,230
16,362
Borderline ECG
**
1.4e-03
**
1.4e-03
808
17,116
Non-specific ST abnormality
**
1.4e-03
**
1.4e-03
303
19,331
Abnormal ECG
**
1.4e-03
**
1.4e-03
22,384
376
Left posterior fascicular block
**
1.4e-03
**
1.4e-03
421
21,404
Right bundle branch block
**
1.4e-03
**
1.4e-03
5,356
14,325
Anteroseptal infarct
**
1.4e-03
**
1.4e-03
1,110
17,847
Atrial flutter
**
1.4e-03
**
1.4e-03
567
18,535
Right axis deviation
**
1.4e-03
**
1.4e-03
580
19,491
Left bundle branch block
**
1.4e-03
**
1.4e-03
640
21,270
Right atrial enlargement
**
1.4e-03
**
1.4e-03
609
20,596
Wide QRS tachycardia
**
1.4e-03
**
1.4e-03
96
22,006
Inferior-posterior infarct
**
1.4e-03
1.0
327
23,024
Biventricular hypertrophy
**
1.4e-03
**
5.8e-03
202
22,728
Supraventricular tachycardia
**
1.4e-03
**
1.4e-03
81
22,412
Bi-atrial enlargement
**
1.4e-03
**
1.4e-03
372
22,477
Wolff-Parkinson-White
**
1.4e-03
1.0
25
24,801
Mean Difference (Sensitivity)
Condition Name
Mean Difference (Specificity)
Acute pericarditis
1.0
**
2.9e-03
12
24,984
Posterior infarct
1.0
**
1.4e-03
39
24,732
0.0
2.5
5.0
7.5
10.0
0
Sensitivity (%)
2
4
6
8
Specificity (%)
Fig. 5: HEEDB data on diagnostic sensitivity and specificity changes on future records of patients returning with unchanged health states for all available conditions. Memorisation-induced changes in diagnostic sensitivity and specificity on future records of HEEDB patients whose health state was unchanged relative to their historical training records; all 36 available conditions are shown. Cases comprised positive future cases of a condition for which at least one positive case was already present in that patient’s historical training records, and negative future cases of a condition for which only negative cases were present. Each row represents a single condition: the left-hand panel shows the mean difference in sensitivity between models trained vs. not trained on that patient’s historical data (M = 100 models per group for each patient), the right-hand panel the corresponding difference in specificity. Point estimates to the right of zero indicate increased performance on data-contributing patients’ future records, those to the left a reduction. N + and N − denote the number of positive and negative cases for each condition, respectively. Condition-specific decision thresholds maximised Youden’s J statistic on the test set. Statistical significance was determined using exact, two-sided permutation tests that reassign model training run indices to the corresponding patient subset masks (B = 100,000 permutations per comparison); error bars denote simultaneous 95% confidence intervals obtained by inverting the same test. p values and confidence intervals were corrected across all C = 72 shown comparisons using the Bonferroni method. * and ** denote p ≤ 0.05 and p ≤ 0.01, respectively; the resolution of the permutation test floors raw p values at 2/(B + 1) ≈ 2 × 10−5 , and hence corrected p values at 1.4 × 10−3 .
33
10−1 10
−2
ε= 101
ε= 102
ε= 102
3
3
ε= 10
ε= ∞
10−3
ε= ∞
10−4 10−5 50
75
100
0
25
50
75
100 10−1 10
−2
10
ε= 1
ε= 101
ε= 101
ε= 10
2
ε= 10
2
ε= 10
3
ε= 10
3
ε= ∞
10−5 0.0
0.5
1.0 0.0
0.5
10−1
ε= 1
ε= 101
ε= 101
ε= 102
ε= 102
3
3
ε= 10
10−3
ε= 10
ε= ∞
ε= ∞
10−5
0
25
50
75
100
0
25
50
75
10−5 25
10−3 10
ε= 1
ε= 101
ε= 101
ε= 102
ε= 102
ε= 103
ε= 103
ε= ∞
ε= ∞
100 10−1
0.00
0.25
0.50
0.75
0.00
0.25
0.50
h
100
0
25
50
75
100
ε= 1
ε= 102
ε= 1
ε= 101
ε= 103
ε= 101
ε= ∞
ε= 102
10
ε= 103 ε= ∞
−3
10−4 10−5 0.5
100
1.0
0.0
0.5
1.0
ε= 1
10−2
ε= 1
ε= 101
ε= 101
ε= 102
ε= 102
3
ε= 103
ε= 10 ε= ∞
10
ε= ∞
−4
10−6 25
50
75
100
0
25
50
75
100
Change in predicted probability (%) ε= 1
100 10−2
ε= 1
ε= 101
ε= 101
ε= 102
ε= 102
3
ε= 103
ε= 10 ε= ∞
ε= ∞
10−4 10−6 0.00
0.75
75
10−2
0
−5
50
Change in predicted probability (%)
100
ε= 1
ε= ∞
10−4
Energy test statistic
1 - cumulative probability
1 - cumulative probability
10−1
ε= 103
10−3
Change in predicted probability (%)
g
ε= 102
3
ε= ∞
0.0
f
ε= 1
ε= 101
ε= 102 ε= 10
1.0
Energy test statistic
ε= 1
ε= 101
10−2
0
−4
e
ε= 1
d
ε= 1
ε= ∞
10−3
100 10−1
100
Change in predicted probability (%)
1 - cumulative probability
25
c 1 - cumulative probability
ε= 1
ε= 101 ε= 10
0
1 - cumulative probability
ε= 1
1 - cumulative probability
b
100
1 - cumulative probability
1 - cumulative probability
a
0.25
0.50
0.75
0.00
0.25
0.50
0.75
Energy test statistic
Energy test statistic
Fig. 6: Effect sizes and energy test statistics for models trained with different levels of record- and patient-level DP protection a-d, results for MIMIC-ECG. a-b, ecCDF analysis of memorisation effect sizes for models trained with decreasing levels of (ε, δ )-DP privacy protection (increasing ε values). Effect size was measured as the absolute change in average predicted probability between models trained vs. not-trained on a given patient’s historical data (maximum across available classes). Panels show results for unseen future records (a) and for historical records used for model training (b). Each panel shows results for models protected by record-level DP (left, blue) and patient-level DP (right, red) separately; colour intensity increases with ε. c-d, as in panels (a-b), but for the energy test statistics from which the corresponding p values were derived. e-h, as in panels (a-d) but for HEEDB. Dashed grey lines indicate the random-partitioning baseline. Effect sizes and test statistics were obtained from multivariate, nonparametric energy-based hypothesis tests comparing the predictions (predicted probabilities for all available classes) of models trained versus not trained on the respective patient’s historical data (M = 100 models per group for each patient). δ was kept constant at 1/D where D is the dataset size. Patient-level DP protection was achieved by discarding all but the most recent historical training record per patient and then applying record-level DP accounting; the future record dataset was not modified.
34
Supplementary Material List of Tables 1
Model training hyperparameters. . . . . . . . . . . . . . . . . . . . . .
35
36
36
ViT-S-25 ViT-T-25 DenseNet-121 ResNet ViT-S-25 ViT-T-25
MIMIC-ECG MIMIC-ECG (DP) MIMIC-CXR MIMIC-IV-ED HEEDB HEEDB (DP)
−3
5 × 10 10−3 5 × 10−4 10−2 10−4 3 × 10−3
α 0.5 0 10−2 100 10−2 0
λ 0.95 0.99 0.9 0.99 0.99 0.99
γ 0.1 0.0 0.0 0.25 0.1 0.0
Dropout
256 (1024) 1024 256 4096 256 (512) 8192
Batch size
50 100 60 50 20 20
Epochs
Table 1: Model training hyperparameters. α, λ and γ denote learning rate, weight decay and EMA decay rate, respectively. Where two batch sizes are given, the second is the effective batch size after gradient accumulation.
Model
Dataset