Conceptio › Archive › arXiv CS
arXiv CSopen access

NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting Tobias Susetzky1,2 , Raphael Rehms1, 5 , Dmitrii Seletkov1,2 , Özgün Turgut1 , Michelle Espranita Liman1 , Lisa Steinhelfer3,4 , Rickmer Braren4 , and Daniel Rueckert1,5,6

arXiv:2609.09140v1 [cs.LG] 8 Sep 2026

1

Chair for AI in Healthcare and Medicine, TUM University Hospital and Technical University of Munich (TUM), Munich, Germany 2 Institute for Diagnostic and Interventional Radiology, TUM University Hospital, Munich, Germany 3 Institute for Diagnostic and Interventional Neuroradiology, TUM University Hospital, Munich, Germany 4 Department of Diagnostic and Interventional Radiology and Nuclear Medicine, University Medical Center Hamburg-Eppendorf, Hamburg, Germany 5 Munich Center for Machine Learning (MCML), Munich, Germany 6 Department of Computing, Imperial College, London, United Kingdom Corresponding author: Tobias Susetzky ([email protected])

Abstract The digitization of healthcare has generated vast, longitudinal, and multimodal patient records over a lifetime, yet fully exploiting these data to represent and predict patient state trajectories remains a critical challenge. Current AI models often struggle to capture the complex, irregular temporal dynamics and inherent stochasticity of real-world multimodal patient data. In the biomedical domain, existing AI approaches for modeling longitudinal patient records are predominantly discriminative, limited to a few modalities, constrained by closed categorical vocabularies, treating time as a monotonic inductive bias, or they are limited in forecasting future patient states. We introduce N OAH, a time-aware, task-agnostic, generative transformer model representing and forecasting the full multimodal patient journey. N OAH features a novel bidirectional time integration and a variational latent space to capture both the continuous evolution of patient states and the stochasticity of clinical trajectories. Built from over 559 million clinical events from 431,000 hospital visits of 299,000 patients across the MIMIC dataset family, N OAH natively processes multiple kinds of medical images, time-series and numeric signals, categorical events, as well as structured and unstructured clinical records. N OAH is the first truly holistic generative model in its field, enabling autoregressive forecasting with optional time control, zero-shot classification, and counterfactual simulation of clinical interventions. Furthermore, its novel approach generates highly informative and predictive patient state representations that demonstrate strong performance in probing for clinical outcomes, 15 ICD chapters, and 29 comorbidities, as well as in time-to-event prediction. Seamlessly handling diverse modalities and complex temporal dynamics, N OAH provides a versatile, task-agnostic, scalable foundation for intelligent predictive systems in personalized clinical care and digital medicine.

1

Introduction Personal health tracking, wearable devices, as well as primary and secondary healthcare systems collect vast amounts of data describing a patient’s medical history over time. This comprises vital measurements, medical imaging, textual reports, signals, e.g. from electrocardiograms (ECG), drug prescriptions and intake, medical procedures, diagnoses, and administrative events. The sheer scale, multimodal complexity, and irregular temporal distribution of these lifelong records exceed human personnel capacities by far. Thus, medicine has a critical need for intelligent systems that enable their semantic analysis, at patient and population level, and also predict future disease trajectories. In particular, these systems must accurately reflect the stochasticity of clinical progression where unforeseeable events are omnipresent. By robustly predicting future events and simulating individualized treatment outcomes under specific clinical interventions, such generative models hold the potential to accelerate and improve personalized clinical care and may even exceed human prognostic abilities. For example, such models could support complex scenarios involving multiple subspecialty disciplines (e.g. an acute medical condition superimposed on a chronic disease state), likely to become a more frequent event in an aging society. With recent AI advances, there is growing interest in exploiting their enormous potential across the medical domain (for instance [1,2]). However, previous approaches largely focus on few isolated modalities (e.g. vision and language [2–6], imaging and tabular [7, 8], or imaging and signals [9–12]). Also, they often specialize in specific types of signals [13], require a fixed sampling rate of signals or events, or restrict to inputs at few specific points in time [5]. In medical representation learning, contrastive approaches have widely proven successful, e.g. [14], albeit subject to the aforementioned restrictions. Simultaneously, transformer-based architectures have been commonly applied e.g. for forecasting [15]. However, like natural language models, these approaches usually use inputs consisting solely of categorical data over a closed vocabulary, for instance sequences of ICD diagnosis codes [15, 16]. Altogether, even though many promising approaches exist, they fail to holistically capture all of a patient’s data and their dynamics in a task-agnostic way that also allows for meaningful, clinical predictions. Among AI architectures, transformer-based models have proven tremendously successful. Adapting them for irregular clinical event sequences necessitates the integration of temporal position information: current methods commonly apply an additive encoding to a transformer’s input tokens or rely on rotary position embedding (RoPE) [17] to inject temporal information directly into the attention module. RoPE has been adopted in both general-purpose and clinical models, e.g. [18]. However, this approach is insensitive to the prediction horizon: a past event may be highly informative for a chronic-disease onset months ahead, yet irrelevant for an acute event in the coming days. To capture the uncertainty resulting from complex longitudinal dynamics, sequential variational models have historically been the architecture of choice. Variational Recurrent Neural Networks (VRNNs) [19] integrate the principles of Variational Autoencoders (VAEs) [20]. By providing randomness at inference time, they nondeterministically generate highly structured sequences such as speech or handwriting. In the clinical domain, a 2

patient’s state at any given time inherently admits multiple plausible future continuations. Robust forecasting models must explicitly represent this stochasticity. Thus, we adapt this sequential variational concept for a transformer. Most recent advances show that including a patient’s long-term history can improve a model’s diagnostic and prognostic abilities (cf. [15, 21]). They have established flexible token representation methods for multimodal patient data: TAMME [21] considers health records a sequence of timestamped events with both continuous and discrete information, and represents them jointly in a shared space. This makes them easily accessible for transformer-based classification models and unsupervised pretraining via random masking and reconstruction. Scaling up TAMME’s token representation methodology and pretraining strategy, A POLLO [22] demonstrates its real-world applicability to representation learning and clinical downstream usage. At inference time, it uses a learned token tied to a specific event type, e.g. diagnosis, to prompt the model to produce one patient representation vector. However, approaches in this family are, at their core, discriminative, lacking the ability to fully reflect temporal dynamics and sequential inter-event dependencies, as well as the capacity for true forecasting and generative application with a source of variability. Neither their architecture nor their training objective directly encourages learning clinical dynamics and progression. To overcome these limitations, we introduce N OAH, a time-aware, task-agnostic, generative, multimodal transformer model designed to learn a latent space of patient states and transitions from lifetime sequences of multimodal patient data. We adapt a context-conditional variational approach to explicitly capture the stochasticity of disease progression, enabling the model to learn predictive patient state representations over time while accounting for unexpected transitions. Notably, N OAH fully consumes the actual content of all modalities, e.g. imaging data or waveforms, while providing variation at inference time when autoregressively forecasting the next events. With a novel bidirectional time integration N OAH can natively handle irregular inter-event time gaps and weight the relevance of elements in a patient’s history for the next prediction based on both recency and prediction horizon. We demonstrate that N OAH serves as a highly versatile task-agnostic foundation for holistic patient trajectory analysis. Pretrained and evaluated on 559 million timestamped clinical events across the MIMIC dataset family [23–28], N OAH successfully captures the temporal dynamics of patient journeys. We validate its performance across a diverse suite of clinical downstream tasks, demonstrating its efficacy in probabilistic zero-shot forecasting, time-to-event prediction, and generating compact patient state representations. Finally, we highlight its generative capabilities for autoregressive rollouts with optional time control and counterfactual simulations of clinical interventions, providing a scalable architecture for hypothesis generation and in silico trial design. Overall, with N OAH, we introduce a ready-to-scale architecture for a foundational holistic patient trajectory model and demonstrate its potential and versatile applicability.

3

Results

20.6 1.5

25

Tx EC t G C X EC R H O

C a N t um

Tx t C a EC t G C X EC R H O

N

um

0

10

g Full forecasting

h Time-controlled forecasting i Simulation

of all event components incl. time of next token at each step

injecting time of next token during rollout steps

1.3

2.5

17.5

37.9

0

21.3

0

20

CXR

20

Portion of stays (%)

f Architecture overview

40

33.9

60

0.7

50

e Modality co-occurrence

16.5 12.1 11.2 10.0 8.0 5.9 5.2 4.9 4.8

Circulatory Injury Digestive Symptoms Mental Neoplasms Respiratory Musculoskeletal Pregnancy Other

6.2

75

d Primary diagnosis

Patients (%)

93.3 53.3

100.0

100

94.9

82M

272M

6

Patients w/ ≥ 1 (%)

10

7

357k

10

8

281k

431K hospital stays 559M events

10

795k

299K patients

Events

10

9

204M

a Merged MIMIC dataset family b Per-modality volume c Per-modality coverage

Txt ECG ECHO

j Patient state embeddings

for zero-shot classification and counterfactuals

Figure 1: N OAH understands, represents, and extends multimodal patient records. (a-e) Patient journey dataset built from MIMIC. (f) N OAH architecture: multimodal health record events embedded through type category, type specifics, and value. A transformer decoder pretrains on these sequences with causal masking and novel time integration. A novel latent module explicitly learns context-conditional patient states and transitions. (g-h) Autoregressive rollout with optional time control. (i) Outcome simulation for Monte Carlo-style classification and counterfactual analysis. (j) Downstream applications of N OAH’s patient state embeddings.

A time-aware model can understand and predict multimodal data over the entire patient journey. N OAH is a transformer model explicitly trained to learn temporal dynamics of patient states across multimodal trajectories (Fig. 1). It introduces a bidirectional temporally enriched attention mechanism to deeply incorporate time features and offer temporal control during generative inference. N OAH is pretrained and evaluated 4

on 559M timestamped events from 299k patients spanning 431k hospital visits (panels a-e) as a versatile foundation for numerous downstream setups in a real-world clinical context. It operates on multimodal event sequences comprising texts, X-ray and echocardiogram (ECHO) imaging, electrocardiogram (ECG) waveforms, numerical and categorical data, in arbitrary order and count (f). Event representations combine high-level type category, free-text type specifics, externally obtained value embedding, and temporal features. Free-text specifics and continuous values overcome limitations of a closed event vocabulary and enable a trained model to consume unseen data such as updated drug names. By integrating timing information for each event such as current age and bidirectional inter-event deltas, N OAH naturally works on temporally irregular sequences. As a key novelty, N OAH is fully multimodal, yet generative. It is a transformer decoder offering variability during autoregressive inference while maintaining a coherent patient state across output components of each timestep. The pretrained model can be directly used for representation, i.e. to obtain lifetime trajectories as sequences of patient state embeddings, and for extending them to forecast future developments (g-j).

N OAH learns patient trajectories, a landscape of states and transitions. We feed N OAH the entire record of each test patient, extract the 768-dimensional transformer output before the variability module as patient state representation for each timestep, and investigate the resulting trajectories (Fig. 2): they unfold and evolve over time in the embedding space (f-g). Qualitatively, patients with similar starting conditions, yet different course and outcome, diverge in this space (a). Unlike static atlases of isolated event types, the learned space is primarily organized by contextualized events, e.g. by event type under care level and clinical progress (b, c). Albeit subject to projection effects, panels d-e indicate a clear separation by care intensity, e.g. long-term ICU vs. general hospital, emergency department (ED), and outside care (e). Lower-intensity care corresponds to the stable, non-critical risk region (d), while a high-risk zone emerges close to the ICU region. For these 2D projections, we use PaCMAP [29] and chain-UMAP, a UMAP customization [30] ensuring a token’s sequence neighbors are in the UMAP nearest neighbors. This preserves within-patient structure without affecting inter-patient and global insights. As a characteristic of N OAH, we can quantify the model’s surprise in patient state transitions, i.e. the difference between prior (”expectation”) and posterior for the next token. We find that this surprise often corresponds to clinical novelty: surprise increases shortly before the formally charted stay beginnings (h). It peaks for ED vitals, followed by drug administration and immediate procedures in other care. This reflects real unpredictability of acute events and first responses, especially in first-time visits. The model reacts to abnormal outpatient measurements that precipitate the stay, before formal transition. Surprise decreases with more available context, stabilizing patient condition, and emerging routines. Among transition events themselves, ED arrival is well predictable due to charting patterns, while death outside clinical settings due to acute events is hardly predictable (i). Overall, N OAH’s surprise coincides with clinical reality. This also holds for ED arrival and acuity (helicopter vs. walk-in, acuity 1 vs. rest), electivity of procedures (planned vs. unplanned), and increasing predictability with more preceding visits and higher age (j-m). Genuinely hard to predict are content-rich modalities such as ECG and X-ray (n, Supplementary Table 1).

5

a Two lifetimes, one diagnosis (chain-UMAP)

b Care level (PaCMAP)

c Event type (PaCMAP)

ward treatment bacteraemia treated lymphoma coded

intubated

home ED: hypotension

discharged alive ED: chest pain

lymphoma coded death ICU admission noradrenaline sedated, lines in

ICU >21d in ICU 7-21d in ICU 3-7d in ICU 0-3d in

1 wk 1d death

t

g State autocorr.

10

6

100

10

4

10

10

colours as in b

2

10

−1

10

1

10

3

10

ED arrival (63k)

10

0

10

hosp. adm. (1st) (27k)

hosp. adm. (rep.) (38k)

1

10

2

sequence lag k

i Surprise of transition tokens

ICU adm. (11k)

ED arrival (1st) (n=30,707) ED arrival (rep.) (n=32,415)

1.2

hosp. adm. (1st) (n=27,185) hosp. adm. (rep.) (n=37,767)

1.1

ICU admission (n=10,799) death in ICU (n=812)

1.0 7 10

time relative to stay onset

j ED arrival surprise

DRUG_PRESCR DRUG_START

DRUG_STOP LAB_TEST

k Triage surprise

l Electivity Surpr.

KL / type median

.032

1.0

.1

.3 .1

.0 walk helicopter ambulance other

1

2

3

4

triage acuity

5

d

0d

+1

+3

h 0 h +1

d

-1

0d

-1

<50 50–65

65–80 80+

all

3

5+

ECG X-ray Echo Note radiology report Death Enter hospitalization Note discharge summary Enter ICU Leave ICU Micro test

1.10 1.08 1.06 1.04

.030

median, 95% CI (patient bootstrap)

SERVICE VITALS

m Surprise by stay rank, age n Surprise by event type (top 10)

3.0

.3

.028 .05 .1 .2 .5 1 all tokens median token KL(q‖p) [nats] (median)

ICU

10.0

1.0 .034

OTHER_EVENT PROCEDURE

1.12

3.0

.036

44,929 test patients · 82.9M tokens

LEAVE_TRANSF MICRO_TEST

hosp.

-3

d

0d

+3

+1

h 0 h +1

d

-1

0d

-1

-3

d

0d

+1

+3

h 0 h +1

-1

d

0d

-1

-3

d

0d

+1

+3

h 0 h +1

-3

d

all tokens (mean)

0d

10

death outside care (n=3,019)

5

-1

10

death in ward (n=500)

6

-1

KL / type median tokens

0.4

record time span (hours)

↑ peak 2.2×

BODY_IN DIAGNOSIS

token KL(q‖p) [nats]

0.5

0.3

1

5

h Surprise (KL) around onsets of stays. Most surprising type and token count per bin. 1.3

admission & transfer complaint notes waveforms imaging

1,000

cos(ht, ht + k)

1 mo

grey: survivors

f Arclength vs record span

time to death

1 yr

e Care level (chain-UMAP)

arclength ∑‖ht + 1 − ht‖

d Time to death (chain-UMAP)

patients / bin

Patient 1 · F, 86 y · lymphoma → intubated, intensive care, died Patient 2 · F, 71→72 y · lymphoma → ward care, discharged alive

vitals & fluids labs & microbiology procedures other medication diagnosis

hospital ED outside care

planned surgical unplanned medical

1

2

4

rank of hospital stay

.01

.1

1

10

KL(q‖p) [nats] — median ± IQR

Figure 2: N OAH learns patient trajectories as states and transitions. (a-c) All patient states in latent space, mainly organized by contextualized event types. (a) Two exemplar trajectories with similar starting conditions, yet diverging course in embedding space. (d-e) Regions of care levels and risk (time until death). (f-g) Trajectories’ exploration of latent space. (h-n) The model’s surprise at state transitions near stay onsets and across different event types correlates with clinical novelty.

6

b Type reliability morning labs hemoglobin 8.3

−1

0

+1

+2

+3

+4

+5

empirical frequency

echos

1 of 50 rollouts shown, 128 tokens each

prompt cut

−2

ecgs

free

vitals every 30 min HR 74–82 SpO2 99–100%

ECE 0.084 · Brier 0.092 ECE 0.086 · Brier 0.089

0.75 0.5 0.25 0

c Modality reliability

texts

time-ctrl.

same routine, own clock — HR ≈74–78 assessment block, 1 h late

1

images

vitals at the true times — HR ≈78–82 imagines a chest X-ray oxycodone + saline

…

−3

numerics

…

tube feeds · heparin nursing assessment

cat

monitoring + assessment vitals continue — HR 77–89

truth

case gastrointestinal bleeding ward · hour 24 prompt M, 56 y · 3,469 events · 1.3 y

empirical frequency

a Exemplar rollout, free and time-controlled vs. truth

1

ECE 0.082 · Brier 0.094 ECE 0.089 · Brier 0.092

0.75 0.5 0.25 0 0

+6

0.25

d Type prediction

e Modality prediction

f Type count error count MAE (events / patient)

0.9 0.8 0.7 0.6 0.5 0.4 4 6

12

24 48

72

1

2

horizon H (h)

4 6

30 classes · H = 24 h

.9 .8 .7 .6 .5 .4

0.4

0.6

0.8

1.0

AUROC gain (TC − free)

free

AUROC, time-ctrl.

2 5 1

24 48

72

0 1

2

horizon H (h)

h Per-type AUROC, free vs. time-ctrl. 1

12

time-controlled

persistence

chance

i Time-control gain .3

12

24 48

72

1

bands: 95% CI

0

.74 .68

.8 .6

AUROC, free

12

24 48

72

horizon H (h)

12

24 48

72

k Zero-shot ROC

balanced acc. .97 .82

1

.61 .57

.4 0

4 6

4 6

horizon H (h)

.2 2

2

band: 95% CI

AUROC

.2

1

1

j Zero-shot outcomes event type modality

.1

4 6

horizon H (h)

score

2

1

10

0 1

0.75

g Modality count error

3

true-positive rate

macro AUROC

1.0

0.5

predicted probability

hours from prompt cut

Prolonged stay

30-d readm.

.8 .6 .4

AUROC 0.74 0.61 0.97

.2 0

72-h mort.

0

0.2

0.4

0.6

0.8

1

false-positive rate

l Exemplar rollouts for zero-shot classification Prolonged stay (>72 h)

30-day readmission

y = 0 · MC P(y=1) = 0.01

y = 1 · MC P(y=1) = 0.86

72-h mortality y = 0 · MC P(y=1) = 0.01

prompt

prompt

prompt

MC 1

MC 1

MC 1

MC 2

MC 2

MC 2

MC 3

MC 3

MC 3

+36

+48

+60

+72

0

hours since admission

+25

+50

+75

days since discharge target hit

true target time

+48

+60

+72

hours since admission

first bin (y = 0)

Figure 3: N OAH can autoregressively roll out future trajectories to make clinical forecasts. (a) Autoregressive next token prediction for an exemplar patient, with the actual time delta to the next token provided (time-ctrl.) and without (free). (b-g) Model calibration and prediction quality for type categories and modalities. (h-i) Effect of time control on performance in terms of occurrence of each type category and modality within a certain temporal horizon. Panel (h) shows results for event type only. (j-l) Zero-shot classification results and examples of N OAH simulating multiple possible trajectory continuations per patient and then using Monte Carlo estimation for probabilities of clinical outcomes.

N OAH can forecast patient trajectories with optional time control. For 13,551 eligible test patients, we prompt N OAH with their lifetime sequences up to a certain point (Fig. 3a). The frozen model then autoregressively rolls out future events to fill the held-out window up to different temporal horizons. In a Monte Carlo fashion, we estimate the probabilities and counts of each event type 7

category and modality to occur within each horizon. We compare against the held-out ground truth (Extended Data Table 1, Supplementary Tables 2, 3). N OAH achieves AUROC scores of 0.83-0.95 with Brier scores 0.03-0.11, outperforming a naive persistence baseline that repeats the pre-anchor time window (b-e). The model is reliable with ECE 0.08 to 0.094, yet with slight overconfidence at high predicted probabilities (b-c). It yields MAEs of ≤ 3.1 and ≤ 10.7 occurrences on average for type and modality classes, naturally increasing within 72 h, as errors propagate during rollout and divergence from ground truth increases (f-g). Within these experiments, providing the ground truth time delta to the next token (time-control), i.e. enforcing the time at which N OAH should predict the next event in each step, improves performance significantly, up to +0.19 AUROC within shorter, more sensitive horizons (h-i). However, unlike the error in the predicted number of modality occurrences, the error in type count actually increases for longer horizons under time control (f-g). We attribute this to errors in other sensitive event components (e.g. the value embedding) where enforced timing moves less accurate predictions even further out of the learned distribution.

Monte Carlo simulation enables zero-shot clinical decisions. N OAH’s autoregressive rollout also enables Monte Carlo estimation of clinical outcome probabilities: we perform binary zero-shot classification for prolonged stay, mortality, and readmission. Prompting again with the full patient history including the beginning of the current stay, we draw M ∈ {25, 50} trajectories, locate the first occurrence of the target token in each (e.g. a discharge event for the length-of-stay task) and bin its time delta to the reference event, which is either admission or discharge depending on the task (Fig. 3l, Extended Data Fig. 1). Under a prompt cutoff at 48h after admission, we achieve an AUROC of 0.97 (Balanced Acc. 0.82) for 72h mortality and AUROC 0.74 (Balanced Acc. 0.68) for length-of-stay ≥ 72 h prediction. Prompting the model with events up to discharge (exclusive), N OAH only slightly separates positives for 30-day readmission with an AUROC of 0.61 (Balanced Acc. 0.57). We attribute this to error propagation and distribution drift over long-term rollouts, since training only predicts one next token instead of multiple autoregressively. The difficulty of this task also matches clinical intuition. For Monte Carlo estimation we discard rollouts without a scorable outcome, resulting in a coverage of 65% for this long-term target, yet 80-82% for the significantly shorter simulations for prolonged stay and mortality. Detailed results in Fig. 3j-k, Extended Data Table 2.

N OAH enables counterfactual simulations of clinical intervention. N OAH can also simulate the effect of counterfactual interventions. We demonstrate this in one high-relevance case that allows direct comparison with a clinical trial: N OAH estimates in-hospital mortality, length of stay, and MAKE-30 (see [32]) for sepsis patients under two arms, the factual and counterfactual choice of 0.9% saline vs. lactated Ringer’s as the first intravenous fluid within six hours after a sepsis marker. We prompt N OAH with the patient’s events truncated after this intervention, for one arm swap the factual fluid for the counterfactual, then for both arms perform autoregressive rollouts and Monte Carlo estimation of outcomes as described above. We observe simulated treatment effects of +8.92 and +8.86 points for MAKE-30 and death for saline vs. Ringer’s (Fig. 4c-f), matching the finding and effect sign in SMART’s sepsis subgroup [31] 8

b Rollout trajectories in latent space

a MLP-retrieved NEWS2 score per arm Patient 1 · received saline

Patient 2 · received saline

Mean NEWS2 |Δ| 1.42

Patient 1 · received saline

Patient 2 · received saline

Mean NEWS2 |Δ| 0.94 chain-UMAP 2

NEWS2 (probe)

8 6 4 2 0

0

1d

2d

3d

4d

5d

0

1d

2d

3d

4d

5d

6d

Patient 4 · received ringers

Patient 3 · received saline Mean NEWS2 |Δ| 0.75

Patient 4 · received ringers

Patient 3 · received saline

Mean NEWS2 |Δ| 0.64 chain-UMAP 2

NEWS2 (probe)

8 6 4 2 0

0

1d

2d

3d

4d

5d

6d

7d

Time since intervention

8d

0

12h

24h

saline ringers

48h

chain-UMAP 1

74.6% 65.8%

NEWS2 medium (5-6) NEWS2 high (>=7)

e Predicted MAKE-30 score

observed 254 h 170 h

73.7%

ringers

±1 s.d. across rollouts NEWS2 low (0-4)

d Predicted length of stay

observed 21.0%

saline

chain-UMAP 1

60h

factual arm (fluid the patient received) counterfactual arm

c Predicted 30-d mortality

matched

36h

Time since intervention start end

f Average treatment effect no effect

observed 27.0% 73.8%

174 h

74.7%

136 h

65.8%

MAKE-30

+8.92

death (30 d)

+8.86

SMART (sepsis)

+4.70 +1.10

SMART (all) 65

70

75

30-day mortality (%)

80

140

160

180

200

65

Median length of stay (h)

70

75

MAKE-30 risk (%)

80

0

5

10

saline − ringers (pp)

Figure 4: N OAH simulates trajectories and effects of counterfactual interventions. For sepsis patients, we prompt N OAH with the patient’s history followed by the administration of either saline or Ringer’s and then estimate outcomes via Monte Carlo simulation before comparing the simulated treatment effect to the sepsis subgroup of the SMART clinical trial [31]. (a) NEWS2 score retrieved from generated patient states during autoregressive rollouts. (b) Projections of those patient state trajectories in latent space via chain-UMAP projection as used and described for Fig. 2. (c-e) Lower predicted mortality, shorter length of stay, and lower MAKE-30 score under the administration of Ringer’s compared to saline. Yet, the model overpredicts mortality and underpredicts stay duration compared to the observed factual outcome. (f) Simulations show an average treatment effect in clear favor of Ringer’s, matching the sign of the effect in SMART’s sepsis subgroup at roughly twice its magnitude.

at roughly twice its magnitude (Extended Data Table 3, Supplementary Tables 4, 5, 6). N OAH overpredicts mortality (and thus MAKE-30) in this setting. This effect grows with increasing prediction horizon. We attribute it to autoregressive rollout drift leaving the learned distribution: a clean in-distribution trajectory of a survivor requires a long-term rollout not emitting significant clinical events. This is unlikely by design. However, we stress that this overprediction does not affect within-patient contrast. We further investigate trajectories qualitatively and observe that arms diverge, consistent with the simulated treatment effect (a-b).

9

a Linear probes, end-of-stay retrieval

c Comorbidity ROC

b ICD-chapter ROC

d NEWS2 band ROC

1

Comorbidities (Quan-Elixhauser) 0.94 0.93 0.91 0.91 0.91 0.91 0.89 0.88 0.88 0.88 0.88 0.88 0.88 0.86 0.86 0.86 0.85 0.85 0.85 0.83 0.83 0.82 0.82 0.82 0.81 0.80 0.78 0.78 0.72

Metastatic cancer Heart failure Psychoses Alcohol abuse Renal failure Diabetes compl. Weight loss Fluid / electrolyte Valvular disease Cardiac arrhythmias Hypertension Paralysis Drug abuse Pulmonary circulation Liver disease AIDS / HIV Coagulopathy Obesity Depression Peripheral vascular Lymphoma Solid tumor (non-met.) Other neurological Diabetes uncompl. Deficiency anemia Blood-loss anemia Hypothyroidism Chronic pulmonary Rheumatoid / collagen 0.5

0.6

0.7

0.8

0.9

1.0

AUROC (end-of-stay retrieval) whiskers: 95% CI over 5-fold × 3-seed CV

AUROC ≥ 5 0.944 ≥ 7 0.905

0 0

0.5

1

0

False-positive rate

0.5

1

0

False-positive rate

0.5

1

False-positive rate

b, c curve colour: AUROC rank as listed in a (best = darkest)

≥5: low vs med+high (prev. 0.47) ≥7: low+med vs high (prev. 0.28)

e NEWS2 score recovery calibration slope 0.98 R² (within-patient) 0.25

15

10

5

identity train mean (4.37) linear regression (L2), 768-d state probe mean ± MAE

0 4 10 10

2

0

1

2

3

4

5

6

7

8

9

10

11

12

13

14

15

16

17

18

true NEWS2 score

g Time-to-event D-calibration

f Time-to-event C-index 0.86

1

mean ± 1 s.d. over 5 outer folds dots: folds · DeepSurv

0.84

observed fraction

0.94 0.89 0.88 0.87 0.86 0.86 0.86 0.85 0.84 0.83 0.81 0.80 0.79 0.76 0.75

Circulatory Neoplasms Endocrine / metabolic Injury / poisoning Genitourinary Factors infl. health (Z/V) Infectious Mental / behavioral Respiratory Digestive / hepatic Symptoms / signs Skin Nervous system Musculoskeletal Eye / ear (sensory)

0.5

predicted NEWS2

ICD chapters

windows

0.87

time-dependent C-index

1.00 0.98

In-hospital mortality Hospital LOS ≥ 72 h ICU LOS ≥ 72 h

True-positive rate

Clinical outcomes

0.82 0.80 0.78

fraction p perfect 0 0.2 0.4 0.6 0.8 1

0.75

0.5

0.25

0.76 0 0

0.2

0.4

0.6

0.8

1

0

within-stay time fraction p

0.25

0.5

0.75

1

probability level

Figure 5: N OAH produces task-agnostic patient state embeddings for multiple downstream applications. (a-c) Linear probing for clinical information retrieval across 3 outcomes, 15 ICD chapters, and 29 comorbidities at the end of a stay. (d-e) Probing for the NEWS2 patient stability score (d) at end of stay and (e) at each complete local window of the vitals required for NEWS2. Bar chart indicates counts of windows. (f-g) The td-C-index and D-calibration for time-to-event survival models trained across the duration of a patient’s stay.

N OAH’s patient state representations carry clinical information and risk over time. We probe N OAH’s patient state embeddings for clinical information, fitting an ℓ2 -regularized logistic regression under stratified five-fold cross-validation over three seeds. For retrieval from a stay’s final patient state, we achieve AUROC scores within [0.72, 0.94] for 15 ICD chapters and 29 comorbidities (Quan-Elixhauser [33]), 10

and 0.87, 0.98, and > 0.99 for ICU length of stay, hospital length of stay, and mortality (Fig. 5a-c). When probing at different points over time, we observe peak performance at the beginning of the stay for most tasks (Extended Data Fig. 2, Supplementary Tables 7, 8, 9). At this point, signals such as complaints and initial treatment are strongest, while in mid- and long-term care, routine procedures dominate (Extended Data Fig. 5). We observe better results on younger patients and shorter stays (Supplementary Fig. 1, Supplementary Table 10) and conclude that the embeddings ”forget” over time, focusing on predictive features more than summarization. This is deliberate given the model’s nature. In addition to these probes, we train a regression MLP to retrieve the NEWS2 score [34] for patient state severity on a scale from 0 to 20 points, covering low (0-4), medium (5-6), and high (7+) risk. We compute ground truth deterministically on patient timelines after aggregating the necessary vitals within 1h windows. On 39,952 windows from 4,414 held-out patients, we achieve an MAE of 1.34 and AUROC values of 0.94 and 0.91 for separating low vs. medium-high and lowmedium vs. high (d, Supplementary Tables 11, 12). The NEWS2 regression is well calibrated on the low and medium levels, only diverging from the ground truth on rarer high-risk scores (e, Extended Data Fig. 3).

N OAH’s patient state representations enable survival analysis. We assess the survival analysis capabilities of N OAH’s patient state representations from the last stay, denoted h<pT , as more within-stay longitudinal context becomes available. We evaluate this by predicting survival from embeddings indexed at increasing within-stay time (T ) fractions (p ∈ {0.0, 0.2, 0.4, 0.6, 0.8, 1.0}, where p = 0.0 is entering the hospital, and p = 1.0 is the last event before death or discharge) using the same patients and endpoint (time-to-death) across all settings. The results demonstrate that N OAH’s embeddings can be effectively used for survival analysis (Fig. 5f-g). Performance improves as embeddings from later phases of the patient’s stay are incorporated, leading to higher discrimination in terms of the time-dependent C-index [35], while maintaining well-calibrated predictions in terms of D-calibration [36, 37]. These findings suggest that progressively richer longitudinal patient representations provide increasingly informative signals for time-to-event prediction.

N OAH relies on time and specializes in clinical channels. We investigate the internals of the pretrained N OAH model: Of 144 attention heads, 141 specialize in clinical relations (Fig. 6a): we measure each query → key relation’s attention share lift over uniform and identify the dominant one. 76 heads relate events of one type to past events of the same type, e.g. X-ray → X-ray for earlier studies. The model maintains per-modality longitudinal context. 65 heads connect observation and reporting, e.g. notes → ECG, leaving ED → vitals. The model has internalized workflows and documentation logic. For explainability, the model allows analyzing which elements in a patient’s record the prediction rests on (c, Extended Data Fig. 4). Attribution decomposition by event types shows N OAH relies on physiology, not just structure (d): patient measurements are crucial for predicting the next token, while clinical structure is most important for the type component yet never outweighs the other inputs. Further, we investigate importance of temporal context: our bidirectional temporal attention decomposes into four terms: content-content (CC), 11

a Head specialization. Dominant query → key relation as lift over uniform attention

2 3

Layer

4 5 6 7 8 9 10 11

services 35× → adm. leave ED 3.3× → vitals X-rays 18× → itself inputs 10× → itself ECHOs 4.3× → itself disch. 3.7× → drug starts prev. drugs 5.6× → itself triage 9.9× → compl. triage 2.9× → vitals outputs 3.8× → other X-rays 13× → itself compl. 12× → itself

outputs 65× → insurance inputs 5.8× → itself diagn. 16× → itself diagn. 10× → itself disch. 4.9× → drug starts input ends 7.3× → inputs adm. 8.3× → proc. enter ED 7.1× → compl. triage 3.4× → vitals leave ICU 2.1× → vitals reports 3.9× → drug starts ECHOs 3.7× → itself

prev. drugs 7.2× → itself disch. 8.4× → prescr. ECHOs 4.8× → itself reports 7.8× → proc. proc. ends 34× → itself disch. 168× → reports prev. drugs 4.8× → itself reports 5.7× → drug starts adm. 11× → itself ECGs 12× → itself other 36× → language

proc. ends 59× → itself disch. 9.1× → prescr. diagn. 12× → itself X-rays 22× → itself diagn. 7.1× → itself prev. drugs 6.5× → itself inputs 8.4× → itself ECHOs 2.9× → itself services 15× → itself compl. 11× → itself proc. ends 3.6× → labs reports 64× → ECGs

0

1

2

3

4

reports 16× → gender reports 87× → services disch. 73× → reports services 25× → adm. proc. ends 4.7× → proc. other 3.4× → itself transf. 9.6× → itself disch. 151× → itself X-rays 13× → itself reports 97× → ECGs ECGs 9.6× → itself disch. 22× → birth

triage 21× → compl. adm. 26× → transf. services 36× → adm. triage 19× → compl. ECHOs 2.9× → itself diagn. 17× → itself reports 5.8× → drug starts reports 92× → ECGs disch. 1.6× → labs leave ICU 4.0× → other triage 9.1× → enter ED leave ICU 4.3× → labs

5

6

content × content

content × key time (incl. Δprev)

disch. 20× → transf. triage 25× → compl. input ends 8.2× → itself leave ED 3.5× → vitals enter ED 8.0× → itself enter ED 19× → compl. proc. ends 19× → outputs prev. drugs 13× → itself ECHOs 5.2× → itself disch. 119× → itself adm. 45× → itself X-rays 14× → itself

compl. 4.3× → itself X-rays 30× → itself ECHOs 5.4× → itself disch. 131× → reports input ends 5.8× → inputs ECHOs 3.8× → itself diagn. 6.1× → prescr. diagn. 7.0× → itself prev. drugs 2.7× → vitals compl. 12× → itself triage 26× → itself ECHOs 5.4× → itself

compl. 6.1× → gender triage 21× → compl. prev. drugs 9.3× → itself proc. ends 64× → itself ECHOs 4.8× → itself reports 78× → ECGs prev. drugs 7.9× → itself ECGs 9.1× → itself adm. 37× → itself disch. 125× → itself prescr. 1.7× → labs

7

8

9

10

11

Attention head

b Term ablation for bidirectional attention

query time (incl. Δnext) × content

L0

1.0

X-rays 9.0× → itself ECGs 5.2× → itself diagn. 15× → itself reports 7.4× → drug starts ECHOs 5.5× → itself diagn. 16× → itself triage 11× → enter ED other 5.0× → itself services 27× → itself triage 5.0× → enter ED diagn. 32× → language

ECGs 5.5× → race ECHOs 4.6× → itself leave ED 3.5× → vitals leave ED 3.5× → vitals outputs 14× → itself adm. 34× → itself reports 30× → itself prev. drugs 3.6× → vitals prev. drugs 86× → compl. diagn. 24× → itself inputs 11× → input ends adm. 12× → itself

time × time

Share of logit variance

0.5

0.127 0.0 L2

L3

0.501 0.5

0.157 0.0

0

3

6

9

0

3

6

9

Attention head

Attribution share over history (%)

Time

50

25

Δt

Admitted from ED Onto the ward, first orders · 3.89 logit Quetiapine started Diphenhydramine overnight · 2.61 logit Lorazepam, second quetiapine Discharge summary · 2.16 logit

0.25

0.50

0.75

0.3 0.5

value

1

2 3

Absolute attribution (logit, log scale)

e Zero-shot AUROC vs visible history

75

modality

Specifics

ED arrival, tox screens Laboratory result · 3.18 logit

-4.7h -4.2h -1.5h -49min 6min 1.9h 2.5h 4.2h 5.7h 6.5h 12.3h 16.6h 17.4h 22.4h

Mean total-variation distance

100

type

Value

prediction point

0.00

d Attribution by input event type category

0

Value × specifics

L1

0.503

1.0

labs vitals drugs X-rays ECGs ECHOs reports diagn. / compl. / triage encounters patient info inputs/outputs microbiology procedures no specialization

c Attribution for exemplar patient (acute mania)

Time since hospital admission

1

transf. end 37× → services ECHOs 5.2× → itself adm. 33× → transf. prev. drugs 12× → itself diagn. 9.6× → itself outputs 27× → itself ECHOs 4.7× → itself reports 62× → itself leave ED 80× → itself adm. 25× → itself transf. 9.1× → itself enter ED 2.1× → vitals

specifics

measurements vitals, lab tests, microbiology, ECG, echo, X-ray, body output clinician actions & orders drugs (prescribed / started / stopped / prior), procedures, body input assessments & diagnoses diagnoses, complaints, triage reports discharge summaries, radiology reports encounters & administration admissions, discharges, transfers, service changes, death other events

Target

0.75

AUROC

0

0.70 0.65

all patients age < 75 age ≥ 75

0.60

48h

72h

1wk

1mo

3mo

1y

full

visible history before anchor

Figure 6: N OAH is transparent: attention distribution and input attribution. (a) The model’s attention heads specialize in clinically relevant relations, capturing within-type longitudinal dynamics and dependencies between observations and charting. The numeric value in each cell is the lift over a uniformly distributed attention. (b) Ablation of single terms in bidirectional temporal attention. Time features on the queries including the forward time delta matter most after the first layer in which the model can merge historic time dynamics into the content channels. (c) Exemplar attribution over a patient timeline for predicting the next token. Obtained using Integrated Gradients with ”event type”-only as baseline. For this case of acute mania, the model attends highly to lab results, admission, and medication. (d) Attribution of event types by target. Patient measurements matter, especially for predicting the next value component. Administrative info is most important for event type and never outweighs other information. This illustrates the model actually learns from physiology, not just structure. (e) Ablation of patient history for zero-shot prediction of prolonged hospital stay. History inclusion matters significantly, yet mostly for elderly patients.

12

time-content (TC, query including time delta to next event), content-time (CT, key including the time delta to previous event), and time-time (TT) attention. We ablate each and measure total variation distance between resulting attention distributions: attention shifts most under omission of CC and TC (b). For CC this is trivially expected. TC being dominant shows the model attends to history differently depending on the prediction horizon. CT is dominant only in the input layer, before the model merges content and time channels (b). The forward time delta can never arise from history under causal masking, so it must be carried explicitly. We also ablate patient history in our zero-shot prolonged hospital stay setting with M = 25 rollouts: with more history AUROC rises from 0.63 to 0.68 for elderly patients. Overall, history improves AUROC from 0.70 to 0.72, while the model performs better on younger subjects for all extents of history (e).

13

Discussion With N OAH, we developed a multimodal, time-aware, generative model that learns longitudinal dynamics of lifetime patient journeys for representation and forecasting. We introduced a novel bidirectional time integration and a variational latent space of context-conditional patient states and transitions. The latter provides variability at inference time while maintaining semantic coherence across components of the generated events. We demonstrated N OAH’s ability to learn and represent patient states over time and verified the usability of these embeddings: lightweight linear models retrieve clinical information and stability scores from highly informative features at any point in time, eliminating the need to directly process complex patient data in downstream systems. Time-to-event models prove these task-independent features highly predictive, underlining their prognostic value. We also established the model’s autoregressive capability for forecasting future clinical events with optional time control. This enables grasping likely developments in the near future easily and may serve as a flexible prediction engine. Here, we estimated probabilities of clinical outcomes in a zero-shot fashion, and simulated treatment effects under counterfactual interventions. N OAH is the first model to holistically capture patient states and their transitions, incorporating any modality acquired in the clinical setting, entire history, and temporal dynamics. Unlike previous approaches, N OAH provides a continuous holistic patient state representation at any point in time, forming semantically meaningful trajectories that the model itself can even extend in a task-agnostic way to simulate any future developments. However, N OAH comes with the exposure bias of autoregressive models: it trains predicting one token ahead under teacher forcing with ground truth predecessors, while at rollout it iteratively consumes re-encoded predictions. This propagates errors for long horizons, yet might be overcome by sufficiently large reliable training data and post-training techniques. Also, we stress that N OAH’s representations should not be considered summaries of the patient history. The model’s objective is to extract features that are predictive. Further, we demonstrate the successful application of N OAH to a broad range of tasks. For feasibility, we focus on selected use cases within them, while each task in general constitutes a challenging problem on its own. Future work may deepen these experiments beyond our scope. Finally, N OAH is not a physiologically grounded deterministic patient simulator, but a probabilistic model capturing statistical patterns. It depends on training data and estimates likely developments, rather than deducing logically. Counterfactual simulations are not causal modeling as known from other fields, but intervention-conditioned stochastic forecasts. In summary, N OAH’s novel approach provides a medical foundational model architecture to serve as both one universal patient state encoder and a prediction engine across all modalities and events in a patient’s lifetime. Its holistic patient view and generative approach surface risks and future developments to clinicians and patients alike. It opens paths to rethink the utilization of large-scale lifetime records currently collected worldwide. A N OAH-based foundation model built from these will impact AI integration in the medical domain and flexibly facilitate clinical care, including simulation, risk prediction, and interaction with patient history. It has the potential to significantly improve personalized treatment and digital medicine, while making the rich information of lifelong multimodal health data accessible to a broad audience. 14

Online Methods Dataset and preprocessing We merge all of the following subsets from the Medical Information Mart for Intensive Care (MIMIC) from PhysioNet [38] on subject IDs to obtain a large-scale dataset of multimodal patient timelines: MIMIC-IV 2.2 [23], MIMIC-ED 2.2 [28], MIMIC-Note 2.2 [26], MIMIC-CXR-JPG 2.0 [27], MIMIC-IV-ECG 1.0 [24], and MIMIC-ECHO 1.0 [25]. Every charted data point of these multimodal health records, i.e. every single piece of information about a patient, is treated and processed as a standalone timestamped event. For ongoing processes such as a surgical procedure covering a certain time span, dedicated end tokens are inserted on the patient timeline at the respective timestamp. In this way, they implicitly provide the duration information to the model without the risk of leaking it at an earlier timestep. Charting frequencies of event types can be highly disproportionate, especially during ICU stays where vital measurements are collected automatically at short, fixed intervals with minimal informational gain. To limit their structural dominance over other events, a sliding window aggregates numeric vital measurements and represents each segment by its minimum and maximum for each type of vitals, timestamped at the window mean. We split our overall dataset into training, validation, and test sets according to a 70/15/15 ratio under the constraint that every patient appears in exactly one split, preventing data leakage. In our resulting dataset, one sequence corresponds to the lifetime records of one patient, starting at birth. It does not necessarily end with death and will usually include large temporal gaps as well as arbitrarily many ED visits, hospitalizations, and ICU stays, but also data from online medical records, and e.g. medication taken outside clinical care. For training and validation, each lifetime sequence is segmented into sliding windows of at most 2048 tokens with an overlap of 256. Every window is prepended with a fixed demographic prefix (date of birth, race, gender, insurance, language, and marital status). The latter four reflect the most recent value available at the window end and fall back to “Unknown” where missing. A mask excludes this metadata prefix and the window overlap region from the training objective. Each event after the prefix contributes to the loss exactly once. Patients with fewer than 64 of these scored tokens are excluded to ensure a minimum of informative history to learn from. In the test set, each patient is represented by a single uncapped sequence spanning their full record, with every token evaluated. Unlike existing work, sequences exceeding the model’s input length are not invalidated by subsampling, but instead processed in consecutive chunks at inference time with the meta prefix prepended for each chunk. For performance reasons, we precompute the value embeddings, i.e. the representations of images, texts, electrocardiography (ECG), and echocardiography (ECHO), as well as type-specifics embeddings obtained from external pretrained models. For images we use RAD-DINO [4], while for textual data we rely on BioLORD [39]. We use ECG data from the MIMIC-IV-ECG corpus containing 800,035 10-second, 12-lead recordings sampled at 500 Hz. We extract one ECG embedding per recording using the frozen OTIS [40] model, a tiny general-purpose time series encoder pretrained on a multi-domain time series corpus. To this end, each ECG recording is partitioned into non-overlapping 1 × 24 patches by a strided convolution, yielding a sequence of 12 × 208 = 2496 tokens of dimension 192. These input tokens are processed by OTIS and the output patch tokens at the final layer are mean-pooled and used as the 192-dimensional global embedding of the ECG recording. From MIMIC-ECHO 1.0, we include 281,059 echocardiogram videos. Embeddings for these videos are generated 15

using the EchoPrime model [41]. The global [CLS] token is used as the embedding representation for each echocardiogram. Prior to embedding, each echocardiogram is preprocessed by first extracting the ultrasound region using the DICOM metadata SequenceOfUltrasoundRegions, followed by removing the ECG overlay and any remaining burned-in annotations (e.g. text). Videos are then resized to 224 × 224 and saved as AVI files. Following EchoPrime, during data loading, a further 10% zoom crop is applied to reduce empty background before resizing back to 224 × 224. Pixels are normalized using the mean and standard deviation calculated from the EchoPrime train set. Additionally, each echocardiogram video is sampled to 16 frames at a stride of 2, with shorter videos zero-padded to a uniform length. For cohort and dataset statistics, see Extended Data Table 4, Supplementary Table 15, and Extended Data Fig. 5.

Model and pretraining

Overall architecture and pretraining. We pretrain a causally masked transformer to learn sequences of patient state representations (ht )t∈{1,...,S} , ht ∈ Rd . For this, every sample in our dataset is a temporally ordered sequence (xt )t∈{1,...,S} of highly heterogeneous events, each with type and time information, as well as an optional continuous value embedding obtained from an external model (see Section Dataset and preprocessing above). We employ trainable projection modules and a composite representation scheme to transform each sample into a sequence of uniform tokens (et )t∈{1,...,S} , et ∈ Rd in a shared space, see Section Multimodal token representation below. To integrate the temporal information deeply into the model, we encode it at two points: directly at input tokens like a conventional position encoding, and in the self-attention modules, where backward-oriented features are added to the keys and forward-oriented features to the queries (cf. Section Bidirectional temporal attention). To reconstruct event components (type, time, value embedding, etc.) from the transformer outputs, we add a dedicated decoding head per component (Section Multimodal output decoding). For variability at inference time without negatively impacting semantic coherence across these output components (i.e. a valid patient state where type, value, etc. in combination are realistic and plausible), we add a latent patient state space between transformer outputs (ht )t∈{1,...,S} and decoding heads: we exploit causal masking and construct a contextual prior pθ (zt | ht−1 ) from the previous timestep’s transformer output ht−1 with parameters θ and an informed posterior qϕ (zt | ht−1 , et ) with access to the current timestep’s input xt via its embedding et , using parameters ϕ. For t = 1, h0 is given by the metadata prefix of the sequence. The model’s objective during self-supervised pretraining is to reconstruct xt from the posterior-sampled zt while simultaneously minimizing the divergence between posterior and prior, i.e. the Kullback-Leibler (KL) divergence:    qϕ (zt | ht−1 , et ) DKL qϕ (zt | ht−1 , et ) pθ (zt | ht−1 ) = Ezt ∼qϕ log pθ (zt | ht−1 ) This directly encourages the model to learn an informative history representation, limited by the surprise of unexpected events, the innovation (see Section Variational autoregression).

16

Multimodal token representation. Like A POLLO [22], we adopt the representation scheme of TAMME [21] to obtain a uniform token sequence (et )t∈{1,...,S} , et ∈ Rd from a given multimodal health record: every item in the multimodal health record is considered an individual timestamped event of a very high-level type category from a fixed vocabulary. It may optionally be further specified by free-text type specifics, and it may be associated with a value of a certain modality, such as the contents of an X-ray image, of an ECG, or with a numeric scalar. While the model dynamically learns an embedding for the type categories during pretraining, textual type specifics and values are embedded by non-trainable external encoders. Numeric scalars are represented by Fourier features [42]. All other modalities use the mentioned dedicated external pretrained AI models as encoders. To map all embeddings to a shared space, projection layers are trained end-to-end along with N OAH. We prepend every input window with a fixed-length metadata prefix holding one token per demographic attribute (date of birth, race, gender, insurance, language, and marital status). Unlike TAMME [21], we also add dedicated tokens to represent the end of events that are not self-contained but actually go on over a period of time. This includes, in particular, clinical procedures and medication intake. We emphasize that those end tokens are inserted at the timestamp of the physically charted termination of the event, not at an earlier position, to prevent longitudinal information leakage.

Bidirectional temporal attention. Our multimodal longitudinal patient records are sequences of clinical events distributed across continuous, irregularly spaced intervals. In such trajectories, a discrete ordinal index of an event within the sequence cannot capture the underlying temporal semantics. Standard positional encodings typically either ignore continuous time entirely by relying on ordinal embeddings, whether sinusoidal or learned, or they compress it into a single relative-time axis, such as rotary position embedding (RoPE) [17]. Neither approach adequately captures a structural property that is fundamental to modeling patient journeys: directional asymmetry of the inter-event temporal gaps. The duration between two adjacent events represents the forward gap of the earlier token and, simultaneously, the backward gap of the later token. When a token acts as an attention query, it effectively queries for reachable patient states and the corresponding time frame. Conversely, when acting as a key, it is expected to implicitly provide information about the temporal proximity to its predecessor. Encoding both directional views symmetrically makes the model conflate functionally distinct temporal signals. To resolve this, we introduce bidirectional temporal attention, an additive temporal encoding that injects a multi-feature temporal position signal directly into the attention mechanism, asymmetric between queries and keys. Let the token at timestep t ∈ {1, . . . , S} in a sequence carry four distinct temporal scalars: an ordinal position index pt , the patient’s age at , the backward time gap since the previous event δtprev , and the forward gap to the next event δtnext . These inter-event gaps in our real-world dataset span multiple orders of magnitude. Therefore, we apply a shifted logarithmic transformation to stabilize the dynamic range:  δet• = log 1 + δt• . We avoid normalizing across the temporal features, e.g. via LayerNorm. All features are mapped to a continuous embedding space using a sinusoidal basis. For an F -dimensional feature vector u = (u1 , ..., uF ) ∈ RF mapped to a dh -dimensional attention head, we define N = dh /(2F ) frequencies per feature requiring 2F | dh , geometrically spaced under a base temperature T = 104 : ωn = T −n/N ,

n = 0, . . . , N − 1. 17

Each scalar feature is encoded by its full sine and cosine bank,   ψ(ui ) = sin(ui ω0 ), . . . , sin(ui ωN −1 ), cos(ui ω0 ), . . . , cos(ui ωN −1 ) ∈ R2N , and the encodings for each feature are concatenated:   ΨF (u) = ψ(u1 ) ∥ · · · ∥ ψ(uF ) ∈ Rdh . The features shared by queries and keys in the well-known attention computation are position and age vtbase = [pt , at ]. In addition, features used for queries include the time delta to the next event vtfwd = [δetnext ] while the features for keys include the backward time delta to the previous event vtbwd = [δetprev ]. During the attention computation, these are combined additively with query (qt ) and key (kt ): qt ← qt + Ψ2 (vtbase ) + Ψ1 (vtfwd ),

kt ← kt + Ψ2 (vtbase ) + Ψ1 (vtbwd ).

This decomposition ensures that queries strictly condition on the forward-looking temporal horizon, while keys independently integrate recency.

Variational autoregression with a context-conditional prior. Standard transformer models in the language domain commonly obtain variability at inference time by sampling from a categorical distribution over a fixed vocabulary. In contrast, our architecture models sequences of composite clinical events, where the output projection comprises multiple component-specific heads predicting continuous embeddings (event value, type specifics) alongside categorical information (event type), timing, etc. Introducing independent variability in each head would decouple the outputs from each other and thus invalidate their overall semantics as parts of one coherent patient state. We therefore route randomness through a single shared continuous latent variable zt ∈ Rk which is sampled once per timestep t and consumed jointly by all decoder heads. Concretely, we model the conditional distribution of the next event xt given the history x<t := (x1 , ..., xt−1 ) via zt : Z pθ (xt | x<t ) =

pθ (xt | zt ) pθ (zt | ht−1 ) dzt

where ht−1 denotes the transformer output for the patient’s record up to t − 1 under a causal attention mask. The decoder pθ (xt | zt ) reconstructs the event components of the current event xt from zt . Crucially, the prior pθ (zt | ht−1 ) is a learned, context-conditional diagonal Gaussian, unlike the fixed standard-normal prior of variational autoencoders. Following the framework of sequential latent-variable models [19], we use a posterior that observes the historical context ht−1 and the input embedding et of the target event xt :  qϕ (zt | ht−1 , et ) = N µq (ht−1 , et ), diag σ 2q (ht−1 , et )

18

This is implemented via an MLP that concatenates ht−1 and et and maps them to (µq , log σ q ) via linear heads. The prior conditions strictly on the history:  pθ (zt | ht−1 ) = N µp (ht−1 ), diag σ 2p (ht−1 ) Because causal masking ensures history ht−1 contains no leakage of xt , any target signal reaches the posterior exclusively through et . Training maximizes the per-timestep evidence lower bound (ELBO, see Pretraining) whose regularization term is the KL divergence between posterior and prior. Since both are diagonal Gaussians, it evaluates in closed form [20], subscripts t, j indicating the j-th component of σq , σp , µq , µp at timestep t: LKL,t =

k  X j=1

2 + (µq,t,j − µp,t,j )2 1 σp,t,j σq,t,j + − log 2 σq,t,j 2 σp,t,j 2



Minimizing the KL divergence between posterior and prior forces the model to learn a prior that is close to the posterior. Simultaneously, a reconstruction loss ensures the posterior path learns a bottleneck representation of the current event xt like an autoencoder. Together, these two objectives therefore maximize the predictive capability of the prior and thus of the contextual history ht−1 . Intuitively, the training objective therefore is to learn an informative prior that closely matches an informed posterior. To minimize the overall loss, in the long run the posterior is compelled to encode only the unpredictable residual, the innovation or surprise of the current event. A predictable continuation of the history results in near-zero KL, whereas a surprising update such as acute injury pays a penalty. In our clinical application, a perfect approximation is unlikely, so the remaining prior-posterior gap after training directly imposes an upper limit on inference performance, the genuine unpredictability in clinical reality. To prevent the degenerate solution where the posterior (and prior jointly) collapses to a deterministic prediction (σ q → 0) while passing et through, we apply a sigmoidal squash to the prior’s uncertainty: log σ p ∈ (log σmin , log σmax ),

where

σmin = 0.05,

σmax = 2

With σ p bounded away from zero, log(σp /σq ) in the KL divergence increases as σ q → 0. This structurally prohibits the posterior from bypassing the history via perfect certainty. Note that during training, we draw a single sample from the posterior via the reparameterization trick [20]: zt = µq + σ q ⊙ ϵ,

ϵ ∼ N (0, I)

Crucially, at inference time with the future target being unavailable, samples are drawn directly from the context-conditional prior:  zt ∼ N µp (ht−1 ), τ 2 diag σ 2p (ht−1 ) The variance of the clinical trajectories generated this way reflects the model’s epistemic uncertainty: a patient trajectory with a definitive, predictable prognosis yields concentrated samples, while a highly ambiguous clinical state naturally results in a broader sampling distribution. Hyperparameter τ ∈ [0, 1] controls the degree of this variability. As in transformer-based language models, we therefore refer to τ as the temperature. 19

Multimodal output decoding. The latent zt must be decoded to reconstruct the multiple components of one clinical event, i.e. type category, value embedding, etc. These components are not independent given zt . Admissible modalities are constrained by the type category ct . For instance, a medication cannot have an ECG value embedding. Optional free-text type specifics st refine the type, and a meaningful value embedding vt requires that a measurement type and modality mt have been recovered, too. Independent per-head sampling could therefore produce clinically impossible outputs that invalidate trajectories. We decompose the distribution of components per event along a clinically motivated hierarchy, with δtnext again denoting the temporal gap to the next event: pθ (xt | zt ) = p(ct | zt ) p(st | ct , zt ) p(mt | ct , st , zt ) p(vt | ct , st , mt , zt ) p(δtnext | zt ) Cascading feature MLPs mirror the structure of the hierarchy formulated above: first, a type category head extracts features ftcat from the latent zt and predicts ct . The subordinate type-specifics head then consumes zt + ftcat , extracts ftspec , and maps to st . The modality and value embedding heads read zt + ftspec , and the time head consumes zt directly. They retrieve mt , vt , and δtnext . The decoded modality mt selects one of the value path’s output heads. The heads for value embedding and type specifics perform regression tasks, aligning reconstructions with the frozen input embedding space. Numeric scalars are emitted as Fourier coefficients in exactly the sinusoidal basis used to encode them in the input sequence. Thus, a sampled numeric token is directly consumable at the next autoregressive step. Inter-event time deltas in our dataset are zero-inflated, since clustered laboratory results, vital sign documentation, and process end markers share similar or identical timestamps. This produces a point mass at δtnext = 0 in a heavy-tailed distribution. A single regressor would be pulled toward zero by this mass and systematically underpredict gaps. We therefore split the time head into a binary decision gate p(δtnext > 0 | zt ) that is supervised at every position and a regression head supervised on log(1 + δtnext ) only where δtnext > 0. The logarithmic transform compresses the tail that spans multiple orders of magnitude. At inference time, we apply a hard gate that, upon a negative decision, clamps the regression output to zero. A soft gate would conflate uncertainty about whether a gap exists with the gap’s magnitude. The presence of type specifics st is gated identically. The heads for category, modality, and type-specifics presence use data-derived class weights to counteract training data imbalance. Further, we enforce plausible pairings of modality and type category p(mt | ct ) during inference by restricting the Softmax distribution of the modality head to modalities ever observed in the dataset given ct . Overall, the decoder’s output scheme is fully compatible with N OAH’s multimodal input token representation. Thus, a generated token can be re-injected into the model at the next timestep, closing the autoregressive generative loop.

Pretraining. The overall pretraining objective is the negative per-timestep evidence lower bound summed across the patient trajectory and averaged over the training corpus D, augmented by a small auxiliary term for prior reconstruction: " L(θ, ϕ) = Ex1:S ∼ D

S X

# α Lrec,t (zt ) + β LKL,t + γ Lrec,t (e zt )

t=1

20

,

where the latent is drawn either from the informed posterior zt ∼ qϕ (· | ht−1 , et ) or from the contextconditional prior e zt ∼ pθ (· | ht−1 ), both reparameterized [20]. The prior path mirrors sampling at inference time. The reconstruction loss Lrec,t is the sum of the negative log-likelihoods for all components of the event xt , i.e. for modality, type category, type specifics, value embedding, and time as described above. Details on the output components are provided above in Section Multimodal output decoding. The first two terms in the above loss formulation implement the per-timestep ELBO derived in Section Variational autoregression: LKL,t is the closed-form Gaussian divergence between posterior and conditional prior, and β is annealed linearly from zero to its target value over the initial fraction of training [43]. The third term is an auxiliary prior-mode reconstruction loss: at every step the decoder is run a second time on a prior-sampled e zt and tasked with reconstructing the same xt as the posterior path. The coefficient γ controls what fraction of the reconstruction budget γ/(α + γ) is spent on this prior-mode pass. Together with the sigmoidal squash bounding the prior’s standard deviation to (σmin , σmax ), these regularizers structurally exclude two degenerate solutions: first, a deterministic pass-through of the input embedding et with σ q → 0. Second, the prior inflating its variance to evade learning actual predictions. Optimization is done end-to-end with the Muon optimizer [44] on the hidden 2D weight matrices and AdamW on the remaining parameters under a cosine learning rate schedule with linear warmup. Hyperparameters are reported in Supplementary Table 16.

Evaluation framework

Time-controlled and free forecasting. We prompt the pretrained N OAH model with a patient’s history and predict multiple next tokens step by step autoregressively, i.e. each generated token is decoded, re-encoded, and appended to the prompt before the model is applied on the new prompt again for the next timestep. In the plain variant, the model predicts a token and decodes all its event components including type category, type specifics, time, etc. freely (free forecasting). In another variant, we enforce a specific time delta δtnext between the current token and the next one (time-controlled) through our temporal position integration. We contrast the two variants against each other and evaluate the forecasting capabilities of N OAH in general: for each test patient, we set an anchor time t⋆ at a fixed offset of 24 h after their most recent hospital admission, prompt N OAH with the patient’s history up to t⋆ , and autoregressively roll out subsequent events for a fixed token budget that aims to cover the held-out window (t⋆ , t⋆ +Hmax ] up to the maximum horizon Hmax . Because each per-timestep latent zt is sampled fresh from the context-conditional prior pθ (zt | h<t ) at every generated event and then decoded by pθ (xt | zt ), our M =50 independent rollouts realize a Monte Carlo (MC) approximation of pθ (x>t⋆ | x≤t⋆ ). The empirical MC frequency then gives the probability of event type category c or modality m occurring within (t⋆ , t⋆ +H]. We evaluate these predicted probabilities against the held-out ground truth across horizons H ∈ {1, 2, 4, 6, 12, 24, 48, 72} hours. Contrasting free and time-controlled performance then isolates the share of forecasting accuracy attributable to correctly anticipating event timing versus event content, per class and per horizon. We report macro AUROC, macro Brier, and a macro count error, i.e. predicted versus actual counts per class, for high-level type category and modality. In the absence of a model directly comparable to N OAH, we benchmark against a prevalence chance floor and a trivial persistence baseline that labels class c or m as recurring in (t⋆ , t⋆ +H] if and only if it occurred in the pre-anchor window (t⋆ −H, t⋆ ]. 21

Zero-shot Monte Carlo classification. The rollout primitive and Monte Carlo estimation described above enable estimating the probability of certain clinical target events within a specific future time frame. By this, N OAH supports training-free classification for standard clinical decision problems, thus zero-shot. We evaluate binary classification for prolonged length of stay (LOS) exceeding 72 h, mortality within 72 h, and 30-day readmission. A reference time tref defines the clock origin. For the length of stay and the mortality tasks, tref is the time of admission, while for the readmission task tref is the time of discharge. The model is prompted with the patient records up to rollout onset t⋆ ≥ tref . We set t⋆ = tref + 48 h for LOS and mortality, while for readmission we set t⋆ = tref , yet truncate the prompt for readmission at one timestep before t⋆ to avoid label leakage by inter-event time deltas. Per patient we draw M trajectories from x≤t⋆ via N OAH’s autoregressive rollout without time control. For LOS and mortality, we use M = 50, while for readmission we reduce to M = 25 due to a longer forecasting horizon necessitating an increased token budget and computation time. For each rollout, we classify by binning the time of the first generated target event relative to tref against a task-specific threshold: we evaluate for a prolonged stay longer than 72 h from admission to discharge, death within 72 h of admission, and readmission within 30 d of discharge. A trajectory that produces no target token but whose simulated horizon reaches past the threshold is assigned to the second bin. Thus, it is counted as a prolonged stay for LOS and negative for mortality and readmission. A simulation that exhausts the generation token budget before either condition is satisfied is declared invalid and excluded. We compute the predicted probability as the frequency over the valid trajectories. To prevent degenerate results under a low number of rollouts and to reduce the influence of patients with few valid rollouts, we apply Laplace smoothing with an additive constant of 0.5. We report coverage, i.e. the mean fraction of valid trajectories per patient, and conversely the invalid-run ratio alongside the AUROC, balanced accuracy, Brier score, and ECE. The first generated event is seeded with a task-specific forward time delta δtnext from the prompt’s last visible token to avoid the rollout stacking at the initial timepoint. To reduce computational cost, we evaluate on 2000 randomly selected test set patients for LOS and mortality, and on 1000 patients for the more resource-intensive task of readmission prediction.

Counterfactuals. To evaluate N OAH’s capability of counterfactual intervention simulation, we choose a problem of high relevance that allows for direct empirical comparison: we formulate N OAH’s counterfactual evaluation as an emulation of the SMART trial [32]: for early fluid resuscitation in critically ill adults, it compared 0.9% saline with balanced crystalloids, predominantly lactated Ringer’s solution as two different possible interventions. Fluid resuscitation is a cornerstone of early sepsis management, yet the optimal crystalloid remains debated since saline has been associated with hyperchloremia, metabolic acidosis, renal vasoconstriction, and acute kidney injury, whereas balanced crystalloids more closely resemble physiological plasma composition. In SMART, balanced crystalloids reduced the incidence of major adverse kidney events within 30 days (MAKE-30: occurrence of at least one of new renal-replacement therapy (RRT), death, or persistent renal dysfunction within 30 d), with particularly pronounced benefit observed among patients with sepsis, making fluid choice a clinically meaningful intervention for counterfactual evaluation [31, 32, 45]. These findings have raised the important question of how alternative fluid strategies may affect outcomes in individual patients. In our setting, we first identify patients with sepsis markers from N OAH’s test set. We then 22

substitute the respective fluid in their factual multimodal event sequences and from the intervention onward perform autoregressive rollout in the Monte Carlo fashion described above to estimate outcome probabilities under different actions. We investigate whether N OAH’s simulation yields results similar to the sepsis cohort in the clinical trial SMART [31]. We identify patients with sepsis markers and administration of either 0.9% NaCl or lactated Ringer’s within six hours of the marker. The first administration time of either fluid within that time window defines t⋆ . We exclude patients whose records end before t⋆ + 24 h without a terminal event, i.e. without death or discharge, as well as patients whose entire record up to t⋆ + 24 h comprises fewer than 256 events in total. For patients with multiple eligible episodes, we include only the chronologically earliest, to avoid information leakage by trends or patterns in a patient’s history. This results in 1,085 samples in total. Per sample, we prompt N OAH with the patient’s lifetime trajectory up to t⋆ and simulate future development via autoregressive rollout under two arms: (1) the factual arm that receives the fluid actually administered and thus uses the prompt without changes, and (2) the counterfactual arm where the type of the fluid actually received is substituted with the type of the other one. We preserve the time of the factual event, so the contrast isolates the effect of fluid composition. For each arm we draw M = 25 trajectories via the same free-mode autoregressive sampling used above. During this rollout, we constrain any generated crystalloid administration to match the type of the fluid associated with the current arm. Further, we again seed the first generated event per rollout with a forward time delta δtnext of 1 h. Both arms share random numbers and receive identical latent noise. We score the outcomes and again perform Monte Carlo estimation to obtain their empirical probability: 30day in-hospital mortality, length of stay, and the composite MAKE-30 score, all anchored at t⋆ . Computation and filtering of invalid rollouts are handled just as above in Section Zero-shot Monte Carlo classification. To quantify the average treatment effect, i.e. the difference in the effects each of the fluids has on patients when used as sepsis interventions, we collect the outcome probabilities for mortality, LOS, and MAKE-30 and measure the difference between the two arms. We then average over all patients. For investigating simulated patient state trajectories qualitatively, we apply chain-UMAP projection to the generated patient states of each arm and also a trained MLP to infer the NEWS2 risk score from each of them (see Section NEWS2 scoring).

Linear probing. Under causal attention, N OAH’s contextualized transformer output ht before the prior network is a representation of the patient state after having observed events for timesteps 0, . . . , t. We treat it as the patient state embedding at time t and write hfinal for its value at the last observed event of a patient’s record. For every test patient, we extract all ht along the whole observed trajectory. We then apply linear probing for different timepoints within and around the last stay in a patient’s records. This can be an ICU or general hospital stay depending on the task we probe for. A probe is an ℓ2 -regularized logistic regression (C=1, L-BFGS) on inputs that are standardized per feature. We estimate performance by stratified five-fold cross-validation repeated over 3 random seeds, resulting in 15 independent fits per target and position. We report mean ± s.d. of AUROC across folds and accuracy as secondary metric. The following families of tasks are evaluated for each patient: first, clinical outcomes including prolonged hospital and ICU stay (≥ 72 h), early in-hospital mortality (≤ 72 h after admission), and overall in-hospital mortality (death at any time during the stay). Second, concept recovery in one-vs-rest probes for each of 15 high-level ICD chapters and 29 Quan-Elixhauser comorbidity flags [33]. A target is probed at a given position only where each class has at 23

least five instances per fold. To trace the inclusion of clinical information in the patient embedding as a stay unfolds, every probe is re-fit at positions defined by absolute time since hospital admission: pre-admission, admission, +12, +24, +36, and +48 h, and discharge (”leave”). A position is dropped for patients whose stay ends before the respective horizon. Thus, the number of samples changes per position and is annotated in our results. The embedding at the ”leave” position has access to all information of the stay, hence probing it provides clean insights about the embeddings’ capabilities for retrieval. For all other positions, the amount and kind of information available in each position vary between stays. Thus, we consider all positions before ”leave” as read-outs that mix prediction and retrieval, and conservatively evaluate them as retrieval.

NEWS2 scoring. To probe for time-varying clinical risk indicators, we train and evaluate an MLP to retrieve the National Early Warning Score 2 (NEWS2) [34] from single N OAH patient state embeddings ht . Ground truth is computed deterministically from the charted vital measurements in the patients’ event streams of N OAH’s test set. From this set, we hold out randomly selected patients for evaluating the MLP. N OAH’s training data never enters this MLP. For ground truth computation, a scoring window opens at the first of the five vitals required for computing NEWS2 (respiration rate, oxygen saturation, temperature, systolic blood pressure, heart rate), keeps the latest value per parameter, and closes once all five are present. It is capped at a total window span of one hour. Windows are non-overlapping and never span across patients. The more sparsely charted oxygen-supplementation and consciousness observations (Level-of-Consciousness text, otherwise mapped from the Glasgow Coma Scale) do not open windows but are carried forward for up to eight hours. We default to ”breathing room air” and ”being alert”. Items are matched by exact identifiers with unit conversion, and implausible values are dropped. Only ICU and emergency department settings chart all five required vitals jointly. Overall, we obtain 1.83 million windows, 272,698 of them over the N OAH test patients. Each window is anchored at the token of its last contributing event. We train and probe with the patient state ht at that anchor associated with the NEWS2 score computed from that window. At this point, the embedding has observed exactly the vitals the label is computed from, making this a retrieval probe of the current state, not a forecast of deterioration. Notably, in other evaluation settings such as counterfactual intervention simulation, we apply the trained MLP not only to embeddings anchored on a ground truth vital window, but to any patient state embedding at any point in time, including the states generated during autoregressive rollout. For training and testing the MLP we split patients by a 70/15/15 ratio, standardize features with training set statistics, and fit an MLP with hidden layers of sizes 512 and 256 (LayerNorm, GELU, dropout 0.1) by minimizing the MSE with AdamW. We use a learning rate of 10−3 , weight decay 0.01, batch size 4096, at most 60 epochs, and early stopping on the validation loss. Three seeds yield near-identical results. We report MAE, RMSE, R2 , Spearman correlation, quadratic-weighted κ over the three commonly established NEWS2 risk bands, and AUROC at the NEWS2 ≥ 5 and ≥ 7 levels, with confidence intervals from resampling patients (1,000 bootstrap draws). We additionally report the within-patient R2 , which scores against each patient’s own label variance. Also, we contrast against trivial constant prediction, linear-regression, and age-only baselines. Finally, as the trained MLP maps any N OAH patient state embedding to a NEWS2 score, we apply the trained model unchanged to the generated states of autoregressive rollouts in counterfactual intervention simulation to track the simulated risk over time (see Fig. 4). 24

Survival analysis. Time-to-event or survival analysis [46] models the data as a set of {xi , ti , ei }, with xi ∈ Rd the d-dimensional embedding of the patient stay, ei ∈ {0, 1} indicates the occurrence of event (ei = 1, death after admission) or right-censoring (ei = 0, discharge without death), ti ∈ R+ the time until event or censoring for the subject i. The prediction targets are the survival function S(t|x), which indicates the probability that an individual survives beyond time t, and the hazard function h(t|x), which describes the instantaneous rate of event conditioned on surviving up to time t: Rt

S(t|x) = P(T > t|X = x) = e− 0 h(u|x)du We choose DeepSurv [47], the deep learning extension of the Cox Proportional Hazard model [48], to effectively incorporate the high-dimensional embeddings [49]. This assumes the hazard function is expressed as a baseline hazard h0 (t) and the risk score h(x), predicted by a neural network: h(t|x) = h0 (t)eh(x) To evaluate predictive performance changes with increasing longitudinal context in embeddings, we build six prediction settings indexed by a within-stay time (T ) fraction (p ∈ {0.0, 0.2, 0.4, 0.6, 0.8, 1.0}, where p = 0.0 is entering the hospital, and p = 1.0 is the last event before death or discharge). For this, we retain only patients with more than five recorded events and at least one hour between hospital admission and discharge (N =26,476). We select the last clinical event with a stay-relative time strictly below the time fraction, resolving concurrent events at the same timestamp by taking the last event. For each of six prediction settings, we employ nested 5-fold cross-validation, with 10% of each training set reserved for hyperparameter tuning and an additional 10% for early stopping, as in [50]. We evaluate the discriminative performance using the time-dependent C-index, C td [35]. For hyperparameter optimization, we use the Tree-Structured Parzen Estimator (Bayesian optimization) in Optuna [51] with 100 trials, setting C td as the optimization target. For the hyperparameter optimization space, see Supplementary Table 14. Additionally, we assess the calibration performance using D-Calibration (D-Cal) [36], visualizing as in [37] and providing p-values for each fold in Supplementary Table 13. The numerical values for C td and D-cal, as well as the IBS as a joint metric of calibration and discrimination, are reported in Supplementary Table 13. To assess marginal calibration, we use Distribution Calibration (D-Cal) [36], which shows how well predicted survival probabilities align with observed outcomes based on a goodness-of-fit test. D-Cal discretizes the predicted survival probabilities at the true event times into n equidistant intervals in [0, 1], and performs a chisquared test for the uniformity of the distribution. As suggested in the original paper [36], we set n = 20 and report the p-values for each fold, with p-values < 0.05 indicating miscalibrated predictions.

25

Computing environment For all model development, training, experiments, and analysis, we use PyTorch 2.10 on Python 3.12 with CUDA 12.8. Everything can be fully replicated using open-source libraries. We train and evaluate on up to 8× NVIDIA A100 GPUs (80 GB).

Data and code availability All data are available from the original providers and distributors of the Medical Information Mart for Intensive Care (MIMIC) dataset family [23–28, 52] via PhysioNet [38] and must not be redistributed by us. Our code will be provided on GitHub for academic research purposes upon publication.

Author contributions T.S. conceived the study and developed the model. T.S. and R.R. prepared the manuscript. T.S., Ö.T., and M.E.L. curated the dataset and performed data preprocessing. Ö.T. contributed to model pretraining. T.S., D.S., and R.R. designed the evaluation framework and carried out the experiments and analysis. R.R., L.S., R.B., and D.R. provided critical insights and conceptual guidance. L.S. and R.B. provided medical expertise and clinical analysis. R.B. and D.R. supervised the research. All authors contributed to the writing. All authors reviewed and approved the final manuscript.

Acknowledgments This work benefited from resources provided by the joint project ”Open Medical Inference”, a Module 3 project of the Medical Informatics Initiative of the Federal Government of Germany, funded by the German Federal Ministry for Research, Technology, and Space (Grant Number 01ZZ2315B).

26

References 1. Tu, T. et al. Towards generalist biomedical AI. NEJM AI 1, AIoa2300138 (2024). 2. Sellergren, A. et al. MedGemma technical report. arXiv preprint arXiv:2507.05201 (2025). 3. Li, C. et al. LLaVA-Med: training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems 36, 28541–28564 (2023). 4. Pérez-Garcı́a, F. et al. Exploring scalable medical image encoders beyond text supervision. Nature Machine Intelligence 7, 119–130 (2025). URL https://doi.org/10.1038/ s42256-024-00965-w. 5. Bannur, S. et al. Learning to exploit temporal structure for biomedical vision-language processing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 15016–15027 (2023). 6. Moor, M. et al. Med-Flamingo: a multimodal medical few-shot learner. In Machine Learning for Health (ML4H), 353–367 (PMLR, 2023). 7. Ebrahimi, S., Arik, S. O., Dong, Y. & Pfister, T. LANISTR: multimodal learning from structured and unstructured data. arXiv preprint arXiv:2305.16556 (2023). 8. Hager, P., Menten, M. J. & Rueckert, D. Best of both worlds: multimodal contrastive learning with tabular and imaging data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 23924–23935 (IEEE, 2023). 9. Turgut, Ö. et al. Unlocking the diagnostic potential of electrocardiograms through information transfer from cardiac magnetic resonance imaging. Medical Image Analysis 101, 103451 (2025). 10. Liman, M. E. et al. Echo2ECG: enhancing ECG representations with cardiac morphology from multi-view echos. arXiv preprint arXiv:2603.08505 (2026). 11. Cenikj, N. et al. Cross-modal contrastive learning of ECG and angiography representations for severe stenosis classification. arXiv preprint arXiv:2606.02605 (2026). 12. Selivanov, A., Müller, P., Turgut, Ö., Stolt-Ansó, N. & Rueckert, D. Global and local contrastive learning for joint representations from cardiac MRI and ECG. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 217–227 (Springer, 2025). 13. Turgut, Ö., Bott, F. S., Ploner, M. & Rueckert, D. Are foundation models useful feature extractors for electroencephalography analysis? arXiv preprint arXiv:2502.21086 (2025). 14. Wang, Z., Wu, Z., Agarwal, D. & Sun, J. MedCLIP: contrastive learning from unpaired medical images and text. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 3876–3887 (2022). 15. Shmatko, A. et al. Learning the natural history of human disease with generative transformers. Nature 647, 248–256 (2025). 16. Steinberg, E., Fries, J., Xu, Y. & Shah, N. MOTOR: a time-to-event foundation model for structured medical records. arXiv preprint arXiv:2301.03150 (2023). 17. Su, J. et al. RoFormer: enhanced transformer with rotary position embedding. Neurocomputing 568, 127063 (2024). 18. Yu, J., Feng, Z., Lu, J., Cai, T. & Zhou, D. Time-aware attention for enhanced electronic health records modeling. arXiv preprint arXiv:2507.14847 (2025). 27

19. Chung, J. et al. A recurrent latent variable model for sequential data. Advances in Neural Information Processing Systems 28 (2015). 20. Kingma, D. P. & Welling, M. Auto-encoding variational Bayes. arXiv preprint arXiv:1312.6114 (2013). 21. Susetzky, T., Qiu, H., Braren, R. & Rueckert, D. A holistic time-aware classification model for multimodal longitudinal patient data. In International Conference on Medical Image Computing and ComputerAssisted Intervention, 24–33 (Springer, 2025). 22. Zhang, A. et al. A multimodal and temporal foundation model for virtual patient representations at healthcare system scale (2026). URL https://arxiv.org/abs/2604.18570. 2604.18570. 23. Johnson, A. et al. 6mm1-ek67.

MIMIC-IV (version 2.2) (2023).

URL https://doi.org/10.13026/

24. Gow, B. et al. MIMIC-IV-ECG: diagnostic electrocardiogram matched subset. PhysioNet (2023). URL https://doi.org/10.13026/4nqg-sb35. Version 1.0. 25. Gow, B. et al. MIMIC-IV-ECHO: echocardiogram matched subset. PhysioNet (2026). URL https: //doi.org/10.13026/nrjh-5r77. Version 1.0. 26. Johnson, A., Pollard, T., Horng, S., Celi, L. A. & Mark, R. MIMIC-IV-Note: deidentified free-text clinical notes (version 2.2) (2023). URL https://doi.org/10.13026/1n74-ne17. 27. Johnson, A. E. W., Pollard, T. J., Mark, R., Berkowitz, S. J. & Horng, S. MIMIC-CXR database (version 2.0.0) (2019). URL https://doi.org/10.13026/C2JT1Q. 28. Johnson, A. et al. MIMIC-IV-ED (version 2.2) (2023). URL https://doi.org/10.13026/ 5ntk-km72. 29. Wang, Y., Huang, H., Rudin, C. & Shaposhnik, Y. Understanding how dimension reduction tools work: an empirical approach to deciphering t-SNE, UMAP, TriMap, and PaCMAP for data visualization. Journal of Machine Learning Research 22, 1–73 (2021). URL http://jmlr.org/papers/v22/20-1061. html. 30. McInnes, L., Healy, J. & Melville, J. UMAP: uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426 (2018). 31. Brown, R. M. et al. Balanced crystalloids versus saline in sepsis: a secondary analysis of the SMART clinical trial. American Journal of Respiratory and Critical Care Medicine 200, 1487–1495 (2019). URL https://doi.org/10.1164/rccm.201903-0557OC. 32. Semler, M. W. et al. Balanced crystalloids versus saline in critically ill adults. New England Journal of Medicine 378, 829–839 (2018). 33. Quan, H. et al. Coding algorithms for defining comorbidities in ICD-9-CM and ICD-10 administrative data. Medical Care 43, 1130–1139 (2005). 34. Smith, G. B. et al. The National Early Warning Score 2 (NEWS2). Clinical Medicine 19, 260 (2019). URL https://doi.org/10.7861/clinmedicine.19-3-260. 35. Antolini, L., Boracchi, P. & Biganzoli, E. A time-dependent discrimination index for survival data. Statistics in Medicine 24, 3927–3944 (2005). 36. Haider, H., Hoehn, B., Davis, S. & Greiner, R. Effective ways to build and evaluate individual survival distributions. Journal of Machine Learning Research 21, 1–63 (2020). URL http://jmlr.org/ papers/v21/18-772.html. 28

37. Bender, A. & Sonabend, R. Machine Learning for Survival Analysis (CRC Press, 2026). URL https: //www.mlsabook.com/. 38. Goldberger, A. L. et al. PhysioBank, PhysioToolkit, and PhysioNet: components of a new research resource for complex physiologic signals. Circulation 101, e215–e220 (2000). URL https://www. ahajournals.org/doi/10.1161/01.CIR.101.23.e215. [Online]. 39. Remy, F., Demuynck, K. & Demeester, T. BioLORD-2023: semantic textual representations fusing large language models and clinical knowledge graph insights. Journal of the American Medical Informatics Association ocae029 (2024). URL https://doi.org/10.1093/jamia/ocae029. 40. Turgut, Ö., Müller, P., Menten, M. J. & Rueckert, D. OTIS: learning high-quality time series features with tiny encoders. arXiv preprint arXiv:2410.07299 (2026). 41. Vukadinovic, M. et al. Comprehensive echocardiogram evaluation with view primed vision language AI. Nature 650, 970–977 (2026). 42. Tancik, M. et al. Fourier features let networks learn high frequency functions in low dimensional domains. Advances in Neural Information Processing Systems 33, 7537–7547 (2020). 43. Bowman, S. et al. Generating sentences from a continuous space. In Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning, 10–21 (2016). 44. Jordan, K. et al. Muon: an optimizer for hidden layers in neural networks (2024). URL https:// kellerjordan.github.io/posts/muon/. 45. Evans, L. et al. Surviving sepsis campaign: international guidelines for management of sepsis and septic shock 2021. Intensive Care Medicine 47, 1181–1247 (2021). 46. Kleinbaum, D. G. & Klein, M. Survival Analysis: A Self-Learning Text. Statistics for Biology and Health (Springer, New York, 2012), 3 edn. URL https://doi.org/10.1007/978-1-4419-6646-9. 47. Katzman, J. L. et al. DeepSurv: personalized treatment recommender system using a Cox proportional hazards deep neural network. BMC Medical Research Methodology 18 (2018). URL http://dx.doi. org/10.1186/s12874-018-0482-1. 48. Cox, D. R. Regression models and life-tables. Journal of the Royal Statistical Society: Series B (Methodological) 34, 187–202 (1972). URL https://doi.org/10.1111/j.2517-6161. 1972.tb00899.x. 49. Huo, Z. F. et al. Time-to-event pretraining for 3D medical imaging. In Yue, Y., Garg, A., Peng, N., Sha, F. & Yu, R. (eds.) International Conference on Learning Representations, vol. 2025, 84072– 84108 (2025). URL https://proceedings.iclr.cc/paper_files/paper/2025/file/ d0de8a3baf58878e6f479f91dfdc672e-Paper-Conference.pdf. 50. Jeanselme, V., Yoon, C. H., Tom, B. & Barrett, J. Neural Fine-Gray: monotonic neural networks for competing risks. In Mortazavi, B. J., Sarker, T., Beam, A. & Ho, J. C. (eds.) Proceedings of the Conference on Health, Inference, and Learning, vol. 209 of Proceedings of Machine Learning Research, 379–392 (PMLR, 2023). URL https://proceedings.mlr.press/v209/jeanselme23a.html. 51. Akiba, T., Sano, S., Yanase, T., Ohta, T. & Koyama, M. Optuna: a next-generation hyperparameter optimization framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2623–2631 (2019). 52. Johnson, A. E. W. et al. MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports. Scientific Data 6, 317 (2019). URL https://doi.org/10.1038/ s41597-019-0322-0. 29

Extended Data Figures cat numerics

images texts

ecgs echos

target (pred) ground-truth target time

Best

first bin (y = 0)

rollout start (prompt end)

Medium

P1 · target 65.8 h

Worst

P2 · no target event

y = 0 · ŷ = 0.01

P3 · target 120.4 h

y = 1 · ŷ = 0.83

y = 1 · ŷ = 0.01

72-h mortality

prompt MC[0] MC[1] MC[2] MC[3] MC[4] +50

+55

+60

+65

+70

+75

+35

+40

+45

hours since admission

+50

+55

+60

+65

+70

+75

+40

+60

hours since admission

P4 · target 63.8 h

P5 · target 239.0 h

y = 0 · ŷ = 0.01

+80

+100

+120

hours since admission

P6 · target 106.5 h

y = 1 · ŷ = 0.73

y = 1 · ŷ = 0.01

Prolonged stay (>72 h)

prompt MC[0] MC[1] MC[2] MC[3] MC[4] +30

+40

+50

+60

+70

+2

+4

hours since admission

+6

+8

+10

+50

+60

days since admission

P7 · target 2194.9 h

P8 · target 172.6 h

y = 1 · ŷ = 0.86

+70

+80

+90

+100

+110

hours since admission

P9 · target 320.1 h

y = 0 · ŷ = 0.26

y = 0 · ŷ = 0.71

30-day readmission

prompt MC[0] MC[1] MC[2] MC[3] MC[4] 0

+17

+34

+51

days since discharge

+68

+85

0

+6

+12

+18

days since discharge

+24

+30

0

+6

+12

+18

+24

+30

days since discharge

Extended Data Figure 1: Zero-shot classification examples. Examples of Monte Carlo rollouts for 72 h mortality, prolonged hospital stay (> 72 h), and 30-day readmission. Each with ground truth binary label y and predicted probability ŷ. The model sees the patient’s history and current events up to a certain cutoff (dotted line). We then roll out the next tokens within a fixed budget until the target event (star) is predicted, the upper edge of the class 0 bin (gray area) is reached, or the token budget is exhausted. The red line marks the position of the ground truth event, either within class 0 or not.

30

Clinical outcomes 1.0

AUROC

ICD chapters

Hospital LOS>=72h

0.9

16k

In-hosp. death<=72h 27k

7.5k

27k

20k

27k

0.8

18k

25k

27k

Comorbidities

ICU LOS>=72h

7.5k

0.7

5.8k

16k

20k

27k

4.4k

7.1k

In-hospital mortality

18k

25k

25k

27k

27k

27k

Factors infl. health (Z/V)

18k

27k

20k

27k

16k

25k 27k

3.5k

7.5k

Endocrine / metabolic

27k 20k

27k

27k

18k

25k

18k

27k

16k

20k

16k

27k 27k

0.8

27k

0.7

27k

25k

20k

27k

25k

27k

20k

16k

25k 27k

16k

ad m en te r +1 2h +2 4h +3 6h +4 8h le av e

Infectious 27k

27k

18k 20k

epr

Genitourinary 27k

18k

27k

ad m en te r +1 2h +2 4h +3 6h +4 8h le av e

pr

Digestive / hepatic 27k

16k

e-

ad m en te r +1 2h +2 4h +3 6h +4 8h le av e

epr

Neoplasms

27k

18k

25k 27k

18k 20k

ad m en te r +1 2h +2 4h +3 6h +4 8h le av e

pr

Circulatory

0.9

e-

ad m en te r +1 2h +2 4h +3 6h +4 8h le av e

e-

Symptoms / signs

1.0

AUROC

pr

pr

e-

ad m en te r +1 2h +2 4h +3 6h +4 8h le av e

0.6

25k 27k

25k

18k

27k

16k

20k

16k

27k 27k

18k

20k

16k

27k

27k 25k 27k

25k 27k

16k

18k 20k

16k

25k

6h

8h av e le av e

le

8h

e

le

av

8h

+4

e

le

av

8h

av

le e

av

le

4h

r

2h

+2

+1

te

en

16k

pr

e-

ad

m

e

le

av

8h

+4

6h

+3

4h

2h

+1

+2

r te

en

18k 20k

8h

27k

pr

e-

ad

m

e

le

av

8h

+4

6h

4h

+3

+2

r

2h

+1

te

en

m ad

+4

27k 16k

20k

16k

pr

Chronic pulmonary 27k

le

av

e

16k

8h

4h

+2

2h

20k

+1

te

en

m ad e-

18k

pr

le

av

8h

+4

6h

+3

4h

2h

r

27k

20k

+1

25k

+4

27k

18k 16k

+2

r te

en

ad

m

e

e-

le

25k 27k

av

8h

+4

6h

4h

27k

16k

+3

+2

r

2h

+1

te

en

27k

18k 20k

e

25k

pr

pr

e-

ad

m

e

le

4h

r

2h

+2

+1

te

en

m ad pr

e-

le

av

8h

+4

6h

+3

4h

2h

+1

+2

r te

e

Depression

27k

18k

Hypothyroidism

27k

av

8h

+4

6h

4h

en

pr

e-

ad

m

e

le

av

8h

+4

6h

4h

+3

+2

r

2h

+1

te

en

m ad pr

e-

le

av

8h

+4

e

27k

16k

+3

r

2h

+2

+1

te

en

e-

ad

m

16k

16k

27k

18k 20k

pr

pr

e-

ad m en te r +1 2h +2 4h +3 6h +4 8h le av e

18k 16k

25k 27k

ad m en te r +1 2h +2 4h +3 6h +4 8h le av e

27k

0.6

25k 20k

20k 18k

25k 27k

Paralysis 27k

27k

25k 27k

e-

0.7

e-

4h

+3

r

2h

+2

+1

te

en

m ad pr

e-

ad m en te r +1 2h +2 4h +3 6h +4 8h le av e

pr 27k 27k

0.8

27k

18k 20k

6h

4h 20k

27k

Obesity 27k

pr

AUROC

AIDS / HIV

0.9

Metastatic cancer

25k

6h

4h

18k

27k

8h

16k

25k

+4

27k

16k

18k 16k

e

20k

av

25k 27k

e-

ad m en te r +1 2h +2 4h +3 6h +4 8h le av e

epr

Rheumatoid / collagen

6h

4h

+3

r

2h

+2

+1

te

en

m

16k

27k 27k

le

20k

27k

8h

16k

18k

25k 27k

+3

r

2h

+2

+1

te

en

27k

18k 20k

e

25k 27k

27k

+3

r

2h

+2

+1

te

en

ad epr

Deficiency anemia 27k

27k

Weight loss 27k

+4

27k

18k 20k

ad pr

Blood-loss anemia 27k

6h

25k

e-

ad m en te r +1 2h +2 4h +3 6h +4 8h le av e

pr

e-

ad m en te r +1 2h +2 4h +3 6h +4 8h le av e

epr

AUROC

Liver disease

16k

pr

e-

ad

m

e

le

av

8h

+4

6h

+3

4h

2h

+1

+2

r te

en

Valvular disease

18k 20k

20k

+4

16k

18k

27k

16k

27k

25k

6h

25k 27k

27k

18k 20k

pr

e-

ad

m

e

le

av

8h

+4

6h

4h

+3

+2

r

2h

+1

te

en

27k

m

e

le

av

8h

+4

6h

+3

4h

2h

+1

+2

r te

en

27k

27k

18k 20k

16k

ad epr

16k

25k

+3

25k 27k

m

e

le

av

8h

+4

6h

4h

+3

+2

r

2h

+1

te

en

ad epr e-

ad

m

e

le

av

8h

+4

6h

4h

20k

27k

16k

+3

r

2h

+2

+1

te

en

18k

6h

20k

m

e

le

av

8h

+4

6h

4h

+3

r

2h

+2

+1

te

en

m m

27k

18k

27k

27k

16k

Psychoses 27k

Pulmonary circulation

0.6

1.0

+4

20k

+3

27k 16k

27k

0.7

+4

4h

18k

6h

20k

25k

25k 27k

18k 20k

Coagulopathy

27k

27k 27k

Alcohol abuse

27k

6h

4h

25k 27k

+3

27k

0.8

+3

27k

27k

0.6

0.9

16k

Renal failure 27k

27k

27k

Drug abuse

27k

1.0

+3

2h

+2

r r

2h

+2

te

en

m

20k

pr

e-

ad

av e

le

8h

+4

6h

+3

4h

2h

+1

+2

r te

en

m

27k

18k

18k

27k

Hypertension

20k

+1

te

en

16k

25k

+1

20k

27k

pr

e-

ad

av e

le

8h

+4

6h

4h

+3

+2

r

2h

+1

te

en

m pr

e-

ad

av e

le

8h

+4

6h

4h

+3

r

2h

+2

27k

pr

pr

Other neurological

18k

25k 27k

16k

pr

ad m en te r +1 2h +2 4h +3 6h +4 8h le av e

epr

27k

18k 20k

16k

0.9

0.7

ad pr

25k 27k

Diabetes uncompl.

25k

e-

ad m en te r +1 2h +2 4h +3 6h +4 8h le av e

e-

18k

0.6

27k

27k

18k

16k

Cardiac arrhythmias 27k

ad

25k 20k

0.8

25k 27k

16k

27k

27k

ad m en te r +1 2h +2 4h +3 6h +4 8h le av e

27k

0.7

18k 20k

25k

27k 27k

Lymphoma

e-

0.8

25k 27k

Peripheral vascular

27k

27k

1.0

27k

pr

epr 0.9

18k 20k

Diabetes compl. 27k

16k

27k

20k

Solid tumor (non-met.)

1.0

20k

Fluid / electrolyte

27k

18k

m

Mental / behavioral 27k

16k

ad m en te r +1 2h +2 4h +3 6h +4 8h le av e

0.6

18k

25k

25k 27k

16k

pr

0.9 27k

te

ad m en te r +1 2h +2 4h +3 6h +4 8h le av e

pr

Heart failure 27k

0.7

20k

epr

Respiratory 27k

27k

18k

27k 16k

e-

ad m en te r +1 2h +2 4h +3 6h +4 8h le av e

e-

Eye / ear (sensory)

0.8

18k 20k

+1

25k 27k

25k

en

16k

m

20k

27k

ad

8h av e le

6h

+4

4h

+3

2h

+1

+2

r te

en

m pr

e-

ad

8h av e le

6h

+4

4h

+3

+2

r

2h

+1

te

en

m pr

e-

ad

8h av e le

6h

+4

4h

+3

r

2h

+2

+1

te

en

Nervous system 27k

27k

27k

ad

27k

0.7

18k

e-

25k

pr

AUROC

Skin

e-

27k

0.8

1.0

AUROC

m

Musculoskeletal 27k

0.9

0.6

AUROC

ad epr

ad m en te r +1 2h +2 4h +3 6h +4 8h le av e

e-

Injury / poisoning

1.0

AUROC

pr

pr

e-

ad m en te r +1 2h +2 4h +3 6h +4 8h le av e

0.6

Extended Data Figure 2: Linear probing of N OAH’s patient states over time. Cross-validated AUROC of linear probes fit on the frozen patient state h at successive points of each test patient’s last stay. Each ICD and comorbidity panel is a binary classification with ground truth collected at discharge. At the pre-admission probe, the model has seen only the patient’s history, but no data after the formally charted beginning of the current stay. The probe at ”leave” is a clean retrieval. Earlier positions are read-outs mixing prediction and retrieval, since they may carry information on ICD chapters and comorbidities after some timepoint. Probes are ℓ2 -regularized logistic regression, one fit per position. Plots show the mean and 2.5-97.5th percentile AUROC across all folds. Bar annotations show the number of samples per timepoint. 31

P1 · 202 windows over 851 h

NEWS2

15

NEWS2 (truth)

median score 7 · probe MAE 1.77

prediction

10 7 5 0 0

3d

6d

9d

12d

15d

18d

21d

24d

27d

30d

33d

36d

9d

12d

15d

18d

21d

24d

27d

30d

33d

36d

P2 · 194 windows over 851 h

NEWS2

15

median score 7 · probe MAE 1.74

10 7 5 0 0

3d

6d

P3 · 193 windows over 827 h

NEWS2

15

median score 8 · probe MAE 1.62

10 7 5 0 0

3d

6d

9d

12d

15d

18d

21d

24d

27d

30d

33d

P4 · 190 windows over 860 h

NEWS2

15

median score 8 · probe MAE 1.62

10 7 5 0 0

3d

6d

9d

12d

15d

18d

21d

24d

27d

30d

33d

36d

P5 · 170 windows over 735 h

NEWS2

15

median score 7 · probe MAE 1.51

10 7 5 0 0

3d

6d

9d

12d

15d

18d

21d

24d

27d

30d

P6 · 165 windows over 704 h

NEWS2

15

median score 8 · probe MAE 1.71

10 7 5 0 0

3d

6d

9d

12d

15d

18d

21d

24d

27d

30d

time since the stay's first scored window

Extended Data Figure 3: Read-out of NEWS2 risk score from N OAH’s patient states over time. On our test set, we build an MLP to map patient state embeddings to the NEWS2 risk score. For this, we compute ground truth deterministically on windows where the necessary vitals are all available. We then evaluate on a held-out set and obtain the exemplar predictions above, one time series per patient. The model overall performs very well in distinguishing the widely established risk zones of NEWS2 (green to red shading). It only deviates significantly within the high-risk area (red zone).

32

Value

Value × specifics

P1 · Hospital ward — acute mania, admitted to psychiatry

Time (lower bound)

P3 · Emergency department — fall, head bleed, tachycardia

Trauma CT reported

-2.6h

ED arrival, tox screens · 4.95 logit

-4.4h

Specifics

P2 · Intensive care — ventilated, extubated, stepped down -9.4h

-1.6h Admitted from ED

-57min

Admitted from ED

6min

6.1h

Diphenhydramine overnight

12.3h

Lorazepam, second quetiapine · 2.82 logit

16.9h

Discharge summary 22.4h

prediction point

0

1

2

3

4

absolute attribution (logit)

6.6h

Midazolam sedation

9.3h

Extubated

11.6h 12.2h

Waking: GCS eye 3 · 2.77 logit

13.2h

time since ED arrival

Quetiapine started

4.2h

time since hospital admission

Onto the ward, first orders · 3.89 logit

1.9h 2.5h

time since hospital admission

Into the trauma ICU

1.7h

2.1h 3h

15.4h

4h 4.3h

17.6h 18.2h

Left the ICU

19.1h

First oral oxycodone · 7.24 logit Oxycodone dose repeated · 7.96 logit

21.4h 22.1h

prediction point

0

2

4

6

absolute attribution (logit)

Ambulance: fall, head bleed · 3.62 logit

51min

14.2h

16.5h

ED arrival

-2min

First ECG

5.6h

ECGs, metoprolol, HR 170

7h 7.8h

Admitted to hospital · 9.96 logit

8.6h

prediction point

0

2

4

6

8

absolute attribution (logit)

Extended Data Figure 4: N OAH’s attribution over exemplar patient timelines. We visualize how much each input token and its components (type category, specifics, time, and value) contribute to the prediction of the event type of the next token following these visualized prompts. Overall, time plays the most crucial role. Value is most important in events where the type alone is actually not meaningful in a clinical sense, e.g. lab results for tox screening, discharge summary, and ECG, while time and specifics usually dominate. Attribution values are computed using Integrated Gradients with an ”event type”-only baseline.

33

100

10

1

0.1 0 admit

b Inter-event interval

0.25

0.5

0.75

1 discharge

2h 1h 30 m 12 m 6m 3m 1.2 m

Q2 (26–66h)

Q3 (66–126h)

Fraction of tokens per bin

Shannon entropy of type-category dist. (bits)

2.0 1.8 0.75

Relative time within stay

0.01

0.25

0.5

0.75

1 discharge

0 admit

1 discharge

0.25

0.5

0.75

1 discharge

Relative time within stay Complaint Radiology report Vitals

e

2.2

0.5

0.1

Q4 (>126h)

2.4

0.25

1

Relative time within stay

d Category entropy

0 admit

Key event-type rates

10

0.001

0 admit

Relative time within stay Q1 (≤26h)

c

IQR (P25–P75) median

Events per stay per bin

Token rate (by stay-duration quartile) Time to next distinct event

Events per stay (in each 2.5 %-of-stay bin)

a

Lab test Drug start

Category composition

100 %

Vitals & labs Procedures

75 %

Meta / other Body i/o

50 %

Medications Diagnoses & complaints

25 %

Enter/transfer/leave Microbiology

0 0 admit

Discharge / radiology notes

0.25

0.5

0.75

1 discharge

Imaging / ECG / ECHO

Relative time within stay

Extended Data Figure 5: Within-stay distribution of events in our dataset. Events are binned by relative temporal progress within a hospital stay (0 = admission, 1 = discharge). Each stay contributes once (all stays ≥ 1 h). (a) Mean number of events per stay per bin by stay-duration quartile. The rate of collected events is highest at the beginning of a stay, decreases steadily over the progression of the stay, and rebounds sharply in the final bin. (b) Median time to the next distinct event. Gaps lengthen from ∼7 to ∼30 minutes over the first three quarters of the stay and plateau toward discharge. (c) Rates of key event types. (d) Shannon entropy of the type category distribution. Informational variety peaks at admission, plateaus mid-stay, and drops toward the end. (e) Average composition of event type categories over time.

34

Extended Data Tables Extended Data Table 1: Forecasting. Generative forecasting of future clinical events. Given each patient record up to a fixed cutoff (stay-anchored, prompt cut at enter time +24 h), N OAH draws 50 Monte Carlo rollouts and we score whether each event-type (C=31) or modality (C=6) class actually occurs within a horizon H. Macro-averaged AUROC and Brier score at four horizons, N =13,551 patients. N OAH (free) predicts both event content and timing while the time-controlled variant is provided the ground truth forward time delta at each next token prediction. Persistence repeats the previous H-window looking in the opposite temporal direction, i.e. on the prompt. Marginal is the prevalence floor (chance). Best model per column is in bold. Parentheses give 95% patient-bootstrap confidence intervals (B=1,000 resamples). Predictor Macro AUROC

Granularity

1h

6h

24 h

72 h

↑ better

Marginal event-type Persistence event-type N OAH (free) event-type N OAH (time-controlled) event-type Marginal modality Persistence modality N OAH (free) modality N OAH (time-controlled) modality Macro Brier score

0.500 (0.500, 0.500) 0.500 (0.500, 0.500) 0.500 (0.500, 0.500) 0.500 (0.500, 0.500) 0.582 (0.580, 0.593) 0.627 (0.624, 0.630) 0.610 (0.604, 0.618) 0.592 (0.585, 0.598) 0.747 (0.714, 0.756) 0.826 (0.810, 0.847) 0.829 (0.818, 0.840) 0.820 (0.813, 0.827) 0.866 (0.853, 0.903) 0.854 (0.847, 0.861) 0.856 (0.846, 0.866) 0.842 (0.835, 0.849) 0.500 (0.500, 0.500) 0.500 (0.500, 0.500) 0.500 (0.500, 0.500) 0.500 (0.500, 0.500) 0.657 (0.652, 0.693) 0.604 (0.597, 0.611) 0.583 (0.576, 0.592) 0.571 (0.565, 0.579) 0.764 (0.702, 0.779) 0.818 (0.790, 0.838) 0.811 (0.799, 0.822) 0.813 (0.802, 0.824) 0.952 (0.934, 0.960) 0.894 (0.868, 0.919) 0.839 (0.827, 0.852) 0.828 (0.816, 0.839)

↓ better

Marginal event-type Persistence event-type N OAH (free) event-type N OAH (time-controlled) event-type Marginal modality Persistence modality N OAH (free) modality N OAH (time-controlled) modality

0.050 (0.050, 0.050) 0.088 (0.088, 0.088) 0.106 (0.106, 0.106) 0.120 (0.120, 0.120) 0.059 (0.058, 0.061) 0.119 (0.117, 0.120) 0.274 (0.272, 0.276) 0.377 (0.375, 0.379) 0.045 (0.044, 0.046) 0.073 (0.072, 0.074) 0.092 (0.090, 0.093) 0.105 (0.104, 0.107) 0.032 (0.032, 0.033) 0.061 (0.060, 0.063) 0.089 (0.088, 0.090) 0.110 (0.109, 0.111) 0.111 (0.111, 0.111) 0.116 (0.116, 0.116) 0.102 (0.102, 0.102) 0.097 (0.097, 0.097) 0.091 (0.089, 0.093) 0.147 (0.145, 0.150) 0.147 (0.145, 0.149) 0.174 (0.171, 0.176) 0.097 (0.095, 0.098) 0.101 (0.100, 0.103) 0.094 (0.093, 0.096) 0.087 (0.086, 0.089) 0.038 (0.037, 0.040) 0.075 (0.073, 0.077) 0.092 (0.091, 0.094) 0.084 (0.082, 0.085)

Extended Data Table 2: Zero-shot Monte Carlo estimation. N OAH estimates the empirical probability of each clinical outcome via occurrences of the target event across Monte Carlo rollouts without task-specific training (zero-shot). Cells are point estimates with 95% patient-level bootstrap confidence intervals. Coverage is the fraction of autoregressive rollouts producing a scorable outcome. The lower half provides details on the configuration for each task.

AUROC ↑ Balanced accuracy ↑ Brier ↓ ECE ↓ Coverage ↑ Prediction time MC rollouts M Rollout length Temperature

Prolonged stay (> 72 h)

30-day readmission

72 h mortality

0.741 (0.710, 0.771) 0.683 (0.652, 0.714) 0.203 (0.191, 0.216) 0.201

0.609 (0.560, 0.655) 0.569 (0.534, 0.603) 0.232 (0.211, 0.252) 0.146

0.972 (0.943, 0.996) 0.821 (0.588, 0.990) 0.035 (0.028, 0.042) 0.085

0.804

0.645

0.822

admission +48 h 50 128 tokens 1

discharge 25 1,024 tokens 1

admission +48 h 50 128 tokens 1

35

Extended Data Table 3: Counterfactual simulation. Simulated effect of saline versus balanced crystalloid (Ringer’s) on 30-day in-hospital mortality and length of stay (N =1,085 patients). For each patient we run simulations under two arms, the runs sharing the same prompt up to the factual versus the counterfactual fluid given at the real treatment time. Matched is the result for the fluid actually received. Parentheses give patientbootstrap 95% intervals (B=1,000). Paired contrasts use only patients with a defined outcome in both arms, while each scenario row uses all patients defined in its arm. In particular, Effect is not the difference of the rows above for LOS, where a simulation without discharge leads to an undefined outcome. The sepsis subgroup of the SMART trial serves as reference: 30-day in-hospital mortality 31.2 vs. 26.3% for saline vs. balanced crystalloid, adjusted OR 0.74 (95% CI 0.59-0.93) [31]. The interval shown is derived from the published arm rates (unadjusted). 30-day in-hospital mortality (%) Scenario Observed Matched (calibration) Saline Ringer’s Effect (Saline − Ringer’s) SMART, sepsis [31] SMART, all [32]

over runs with explicit outcome

over all runs

21.0 (18.9, 23.5) 73.7 (72.0, 75.4) 48.4 (46.9, 50.2) 74.6 (72.9, 76.4) 49.2 (47.6, 50.9) 65.8 (63.8, 67.7) 44.9 (43.2, 46.7) +8.86 pp (7.83, 9.89) +4.9 pp (+0.4, +9.2) +0.8 pp (n.s.)

36

Hospital LOS (h)

+4.30 pp (3.29, 5.32) -

253.7 (238.8, 270.4) 170.1 (163.5, 176.9) 174.3 (168.2, 181.4) 135.7 (131.4, 140.1) +40.5 (35.3, 46.0) -

Extended Data Table 4: Dataset. Cohort demographic and clinical characteristics over all patients. Admission-derived attributes (age at first admission, length of stay, in-hospital mortality) are defined only for hospitalized patients. All percentages are of the full cohort. Attribute

Value

Count

Hospitalized

Yes No

180,733 (60.3%) 118,979 (39.7%)

Sex

Female Male

158,525 (52.9%) 141,187 (47.1%)

Age at first admission (years)

18-29 30-44 45-64 65-79 80+

24,596 (8.2%) 31,277 (10.4%) 58,410 (19.5%) 41,010 (13.7%) 25,440 (8.5%)

Race or ethnicity

White 166,855 (55.7%) Black or African American 36,941 (12.3%) Hispanic or Latino 16,010 (5.3%) Asian 13,370 (4.5%) American Indian or Alaska Native 558 (0.2%) Native Hawaiian or other Pacific Islander 341 (0.1%) Other 15,202 (5.1%) Unknown 50,435 (16.8%)

Insurance

Unknown Other Medicare Medicaid

118,979 (39.7%) 108,817 (36.3%) 57,735 (19.3%) 14,181 (4.7%)

Hospitalizations per patient

0 1 2 3-5 6+

118,979 (39.7%) 101,198 (33.8%) 35,712 (11.9%) 29,866 (10.0%) 13,957 (4.7%)

Length of first hospital stay

<1 day 1-3 days 3-7 days 1-2 weeks >2 weeks

47,804 (15.9%) 57,392 (19.1%) 48,229 (16.1%) 18,594 (6.2%) 8,647 (2.9%)

In-hospital mortality (any stay) Died in hospital Survived hospitalization

8,493 (2.8%) 172,240 (57.5%)

Mortality (any cause, in record) Deceased Alive / censored

29,076 (9.7%) 270,636 (90.3%)

Total

Patients

299,712

37

Supplementary Information Supplementary Table 1: Per-event-type latent surprise. Median, upper-tail (q95 ), and mean KL(q∥p) per event describing N OAH’s latent innovation, i.e. the surprise of the trained model on the test set. By event type, sorted by median descending. A larger value means the event type is less predictable from the patient history the model has seen up to this point. Routine measurements sit near zero, while imaging and waveform events with their rather unpredictable contents, as well as terminal events, carry the most surprise. Event type

Median KL

q95 KL

Mean KL

38 37 22.5 2.86 0.961 0.558 0.452 0.16 0.105 0.0605 0.0571 0.0553 0.0545 0.0506 0.0436 0.0401 0.04 0.0381 0.0365 0.0355 0.0347 0.031 0.0309 0.0306 0.0305 0.0286 0.0276 0.0258 0.0256 0.0242 0.0222

49.9 57.3 41.9 6.45 21.5 6.39 5.96 6.65 7.95 3.4 4.18 5.63 2.66 6.81 0.394 0.235 2.5 0.574 0.964 0.0588 0.548 0.932 0.0376 0.51 0.0497 0.0757 0.0457 0.0339 0.0526 0.0593 0.0544

38.7 37.2 24 3 4.24 1.39 1.33 1.1 1.06 0.75 0.644 1.12 0.504 1.34 0.179 0.0667 0.374 0.135 0.198 0.0384 0.097 0.18 0.0316 0.0946 0.038 0.0344 0.0302 0.037 0.0584 0.0533 0.0281

ECG X-ray ECHO Note Radiologyreport Death Enter Hospitalization Note Dischargesummary Enter ICU Leave ICU Microbiology Test Service Leave Transfer Leave Hospitalization Procedure End Leave ED Body Output Transfer Drug Prescription Drug Start Body Input Complaint Drug Stop Enter ED Drug Previously Procedure Diagnosis Other Event Triage Acuity Vitals Lab Test Body Input End

38

Supplementary Table 2: Forecasting. The same forecasts as Extended Data Table 1, evaluated at a single horizon H=24 h, i.e. scoring events that occur within 24 h after the prompt cutoff. Patients are split into quartiles over length of stay. Each cell is the macro-averaged event-type AUROC (↑ better) or Brier score (↓ better) for that subgroup. Computed over N =13,551 patients. Best model per column is in bold, and parentheses give 95% patient-bootstrap CIs (B=1,000 resamples). Predictor

Metric

Q1 (≤ 48 h)

Q2 (48-91 h)

Q3 (91-164 h)

Q4 (> 164 h)

Persistence AUROC 0.597 (0.588, 0.609) 0.647 (0.636, 0.656) 0.639 (0.631, 0.647) 0.610 (0.605, 0.631) Persistence Brier 0.287 (0.283, 0.290) 0.293 (0.288, 0.297) 0.317 (0.312, 0.322) 0.342 (0.338, 0.348) N OAH (free) AUROC 0.828 (0.807, 0.847) 0.823 (0.809, 0.836) 0.833 (0.822, 0.842) 0.835 (0.821, 0.844) N OAH (free) Brier 0.084 (0.082, 0.086) 0.099 (0.096, 0.101) 0.103 (0.100, 0.106) 0.110 (0.107, 0.114) N OAH (time-controlled) AUROC 0.868 (0.851, 0.882) 0.872 (0.859, 0.885) 0.864 (0.855, 0.874) 0.865 (0.853, 0.873) N OAH (time-controlled) Brier 0.096 (0.093, 0.098) 0.085 (0.082, 0.087) 0.095 (0.093, 0.098) 0.106 (0.104, 0.109)

Supplementary Table 3: Forecasting. The 10 best- and 10 worst-predicted event types in terms of AUROC at the H=24 h horizon over N =13,551 patients. Prev. is the cohort prevalence. AUROC for free and timecontrolled (TC) forecasting. Count MAE (free) is the mean absolute error between predicted and actual perpatient event counts in the window. Best rollout variant is in bold. Class

Prev. AUROC (free) AUROC (TC) Count MAE (free)

Top-10 by free-mode AUROC Leave ED Leave ICU Body Output Other Event Complaint Body Input End Body Input Vitals Procedure End X-ray

1.2% 1.8% 20.4% 25.9% 7.2% 19.1% 19.6% 31.8% 5.9% 1.9%

0.982 0.967 0.964 0.953 0.941 0.930 0.930 0.930 0.925 0.901

0.914 0.982 0.974 0.979 0.960 0.960 0.961 0.954 0.938 0.902

0.008 0.045 2.305 3.245 1.049 1.263 1.474 7.949 0.401 0.139

0.687 0.696 0.696 0.713 0.723 0.732 0.732 0.760 0.773 0.778

0.735 0.806 0.777 0.731 0.877 0.689 0.757 0.734 0.810 0.808

0.361 0.058 0.175 5.392 0.367 0.010 0.012 0.464 13.171 0.240

Bottom-10 by free-mode AUROC Note Radiologyreport Service Transfer Diagnosis Note Dischargesummary Triage Acuity Enter ED Leave Transfer Lab Test ECG

11.7% 1.9% 9.5% 31.0% 12.1% 0.1% 0.1% 24.9% 63.6% 6.2%

39

Supplementary Table 4: Counterfactual simulation. Calibration of 30-day in-hospital mortality prediction under the factual intervention against the observed outcome (n=1,085 uncensored patients), by decile of predicted risk. The final column gives each decile’s contribution to the ECE. N OAH significantly tends toward overpredicting mortality for this subcohort. n Mean predicted Observed risk

Decile 1 2 3 4 5 6 7 8 9 10

108 109 108 109 108 109 108 109 108 109

3.04% 13.69% 25.56% 35.38% 43.93% 51.96% 60.59% 70.72% 84.30% 96.26%

Total 1,085

-

|gap| Bin-ECE contribution

3.70% 0.67 pp 10.09% 3.60 pp 20.37% 5.19 pp 15.60% 19.78 pp 25.00% 18.93 pp 15.60% 36.37 pp 25.93% 34.67 pp 22.02% 48.70 pp 25.93% 58.37 pp 45.87% 50.39 pp -

-

0.066 pp 0.361 pp 0.516 pp 1.987 pp 1.884 pp 3.653 pp 3.451 pp 4.892 pp 5.810 pp 5.062 pp 27.683 pp (ECE)

Supplementary Table 5: Counterfactual simulation. Fraction of Monte Carlo simulations that reach a determined outcome within the 30-day window. Per arm (intervention) with the associated invalid-run ratio. Arm

Mean coverage Invalid-run ratio 95% CI on coverage

Matched Saline Ringer’s

68.57% 68.48% 71.78%

31.43% 31.52% 28.22%

(66.89%, 70.22%) (66.84%, 70.21%) (70.21%, 73.34%)

Supplementary Table 6: Counterfactual simulation. Simulated effect of saline versus Ringer’s on MAKE30, the primary endpoint of the SMART trial (i.e. occurrence of death, new renal-replacement therapy, or a final serum creatinine ≥2× baseline within 30 days). Rows as in Extended Data Table 3. Observed mortality counts any death within 30 d, following the MAKE-30 definition. The renal component is evaluated only where the simulation emits a creatinine and a baseline exists. Reference: all SMART patients [32] and sepsis subgroup [31] with MAKE-30 at 40.1 vs. 35.4% for saline vs. balanced crystalloid, adjusted OR 0.78 (95% CI 0.63-0.97), interval derived from the published arm rates (unadjusted). MAKE-30

Components (%)

Scenario

composite (%)

death

new RRT

renal dysf.

Observed Matched (calibration) Saline Ringer’s

27.0 (24.4, 29.6) 73.8 (71.9, 75.4) 74.7 (72.8, 76.6) 65.8 (63.9, 67.8)

23.4 (20.8, 25.7) 73.7 (71.9, 75.4) 74.6 (72.9, 76.3) 65.8 (63.9, 67.9)

5.6 (4.3, 7.0) 0.2 (0.1, 0.3) 0.2 (0.1, 0.3) 0.1 (0.0, 0.1)

3.7 (2.6, 4.9) 0.2 (0.1, 0.3) 0.2 (0.1, 0.3) 0.0 (0.0, 0.1)

Effect (Saline − Ringer’s) +8.92 pp (7.85, 9.93) +8.86 pp (7.83, 9.93) +.13 pp (.04, .23) +.13 pp (.05, .22) SMART, sepsis [31] +4.7 pp (+0.1, +9.4) SMART, all [32] +1.1 pp (+0.2, +1.9) -

40

Supplementary Table 7: Linear probing. AUROC of ℓ2 -regularized logistic probes that decode clinical outcomes from the frozen patient state embeddings at fixed points along the last stay (mean ± s.d., 5-fold × 3seed stratified cross-validation, one probe per position). The probe at discharge is a clean retrieval probe, after the outcome is already known to the model and thus potentially to the patient embedding used for this probe. For the other timepoints, the kind and degree of information available to the model at each timestep when producing the embeddings vary between patients and stays. Thus, the embeddings up to a certain timepoint do not include explicit or implicit information about the label and their probes are therefore predictive rather than retrieving. However, this point is individual for each patient and stay, so we consider all probes during the stay as read-outs that mix prediction and retrieval. For each of the temporal positions within the stay, we only score stays long enough to reach them. Thus, the number of samples naturally shrinks over rows (total number of samples: hospital N = 27,185, ICU N = 7,479). Length of stay ≥ 72 h Position

Hospital

ICU

In-hospital mortality ≤ 72 h

Whole stay

Read-out, prediction/retrieval mixed (State during the stay) Pre-admission 0.836 ± 0.003 0.688 ± 0.012 0.907 ± 0.022 At admission 0.860 ± 0.005 0.712 ± 0.011 0.919 ± 0.020 +12 h 0.867 ± 0.005 0.754 ± 0.010 0.914 ± 0.016 +24 h 0.861 ± 0.005 0.784 ± 0.012 0.919 ± 0.014 +36 h 0.888 ± 0.004 0.764 ± 0.014 0.926 ± 0.025 +48 h 0.925 ± 0.004 0.770 ± 0.020 0.932 ± 0.025

0.900 ± 0.007 0.913 ± 0.007 0.914 ± 0.007 0.896 ± 0.010 0.897 ± 0.011 0.881 ± 0.012

Clean retrieval (State at discharge) At discharge 0.977 ± 0.002

0.998 ± 0.003

0.871 ± 0.009

41

0.993 ± 0.002

Supplementary Table 8: Linear probing. Read-out and retrieval of the ICD chapters of a stay from the patient state embedding at the beginning and end of the stay. AUROC (mean ± s.d. over 5-fold × 3-seed stratified cross-validation) of a logistic probe on the states at admission, at +48 h, and at discharge (clean retrieval). The +48 h column is restricted to the subcohort of stays that reach it. Rows sorted by discharge AUROC descending. Read-out Target Circulatory Neoplasms Endocrine / metabolic Injury / poisoning Genitourinary Factors infl. health (Z/V) Infectious Mental / behavioral Respiratory Digestive / hepatic Symptoms / signs Skin Nervous system Musculoskeletal Eye / ear (sensory)

Retrieval

Prev. (%)

Admission

+48 h

Discharge

58.2 14.1 55.8 27.4 29.9 65.4 17.1 40.9 27.3 37.3 52.7 9.7 28.9 26.7 7.5

0.913 ± 0.004 0.828 ± 0.009 0.837 ± 0.005 0.791 ± 0.004 0.817 ± 0.005 0.820 ± 0.005 0.803 ± 0.006 0.745 ± 0.007 0.769 ± 0.005 0.777 ± 0.005 0.783 ± 0.006 0.755 ± 0.011 0.738 ± 0.006 0.726 ± 0.008 0.696 ± 0.010

0.907 ± 0.004 0.768 ± 0.010 0.815 ± 0.007 0.729 ± 0.010 0.793 ± 0.005 0.783 ± 0.006 0.778 ± 0.009 0.706 ± 0.005 0.766 ± 0.007 0.735 ± 0.008 0.762 ± 0.008 0.709 ± 0.012 0.704 ± 0.010 0.691 ± 0.009 0.643 ± 0.016

0.937 ± 0.003 0.887 ± 0.007 0.878 ± 0.005 0.868 ± 0.007 0.865 ± 0.004 0.864 ± 0.004 0.857 ± 0.007 0.850 ± 0.006 0.838 ± 0.007 0.831 ± 0.005 0.814 ± 0.006 0.800 ± 0.007 0.792 ± 0.004 0.764 ± 0.006 0.747 ± 0.010

42

Supplementary Table 9: Linear probing. Read-out and retrieval of Quan-Elixhauser comorbidities of a stay from the patient state embedding at the beginning and end of the stay. AUROC (mean ± s.d. over 5-fold × 3seed stratified cross-validation) of a logistic probe on the states at admission, at +48 h, and at discharge (clean retrieval). The +48 h column is restricted to the subcohort of stays that reach it. Rows sorted by discharge AUROC descending. Read-out Target Metastatic cancer Heart failure Psychoses Alcohol abuse Renal failure Diabetes compl. Weight loss Fluid / electrolyte Valvular disease Cardiac arrhythmias Hypertension Paralysis Drug abuse Pulmonary circulation Liver disease AIDS / HIV Coagulopathy Obesity Depression Peripheral vascular Lymphoma Solid tumor (non-met.) Other neurological Diabetes uncompl. Deficiency anemia Blood-loss anemia Hypothyroidism Chronic pulmonary Rheumatoid / collagen

Retrieval

Prev. (%)

Admission

+48 h

Discharge

4.4 11.9 2.7 8.4 11.0 6.9 4.4 17.5 7.1 20.2 44.6 2.1 5.2 4.0 6.1 0.5 7.3 8.3 16.2 5.7 1.3 3.9 8.8 11.9 7.3 0.6 10.4 15.9 2.7

0.880 ± 0.010 0.907 ± 0.004 0.812 ± 0.020 0.861 ± 0.007 0.879 ± 0.006 0.861 ± 0.008 0.811 ± 0.012 0.827 ± 0.006 0.847 ± 0.007 0.838 ± 0.007 0.858 ± 0.004 0.808 ± 0.013 0.815 ± 0.011 0.801 ± 0.011 0.809 ± 0.010 0.787 ± 0.037 0.791 ± 0.009 0.808 ± 0.008 0.711 ± 0.009 0.801 ± 0.010 0.778 ± 0.030 0.797 ± 0.016 0.784 ± 0.011 0.755 ± 0.008 0.731 ± 0.011 0.725 ± 0.029 0.723 ± 0.009 0.709 ± 0.008 0.680 ± 0.013

0.840 ± 0.012 0.881 ± 0.006 0.764 ± 0.036 0.766 ± 0.013 0.838 ± 0.009 0.827 ± 0.010 0.776 ± 0.015 0.812 ± 0.008 0.821 ± 0.012 0.821 ± 0.005 0.825 ± 0.007 0.777 ± 0.030 0.763 ± 0.019 0.751 ± 0.017 0.782 ± 0.014 0.717 ± 0.056 0.762 ± 0.011 0.755 ± 0.012 0.660 ± 0.014 0.733 ± 0.010 0.735 ± 0.037 0.700 ± 0.028 0.746 ± 0.011 0.721 ± 0.008 0.678 ± 0.015 0.606 ± 0.036 0.654 ± 0.014 0.673 ± 0.014 0.620 ± 0.021

0.945 ± 0.008 0.934 ± 0.003 0.914 ± 0.012 0.910 ± 0.007 0.910 ± 0.005 0.905 ± 0.008 0.887 ± 0.010 0.885 ± 0.004 0.883 ± 0.006 0.882 ± 0.006 0.882 ± 0.004 0.879 ± 0.021 0.877 ± 0.009 0.862 ± 0.009 0.861 ± 0.009 0.859 ± 0.036 0.855 ± 0.009 0.849 ± 0.008 0.846 ± 0.006 0.830 ± 0.010 0.828 ± 0.029 0.822 ± 0.014 0.821 ± 0.009 0.819 ± 0.006 0.810 ± 0.010 0.801 ± 0.037 0.778 ± 0.011 0.777 ± 0.008 0.717 ± 0.028

43

Supplementary Table 10: Linear probing. Result of the retrieval at the end of the stay split by subpopulation. For every target, the last patient state embedding in the stay is probed separately within subgroups of age and length of stay. This investigates whether N OAH carries clinical information well over time and across the cohort. Each value is the AUROC of a single retrieval probe per target. A dash marks subgroups that are not evaluable, and the best subgroup per row is in bold. Rows are sorted by mean AUROC. Age (yr)

LOS quartile

<50

50-65

65-80

80+

Q1 Q2 0.0-1.0 d 1.0-2.6 d

0.999 0.978 0.898

0.999 0.981 0.866

0.998 0.977 0.875

0.999 0.963 0.867

0.996 0.916

0.999 0.871

0.997 0.825 0.859

0.999 0.859

Circulatory 0.895 Neoplasms 0.883 Injury / poisoning 0.912 Genitourinary 0.846 Factors infl. health (Z/V) 0.867 Mental / behavioral 0.911 Infectious 0.848 Endocrine / metabolic 0.861 Respiratory 0.834 Digestive / hepatic 0.879 Symptoms / signs 0.836 Skin 0.849 Nervous system 0.832 Musculoskeletal 0.803 Eye / ear (sensory) 0.780

0.873 0.907 0.864 0.846 0.830 0.845 0.856 0.835 0.839 0.820 0.813 0.810 0.777 0.749 0.719

0.884 0.881 0.835 0.828 0.834 0.809 0.868 0.798 0.818 0.780 0.782 0.777 0.748 0.697 0.667

0.891 0.812 0.804 0.805 0.806 0.749 0.843 0.744 0.810 0.743 0.761 0.709 0.727 0.657 0.636

0.925 0.890 0.887 0.841 0.830 0.902 0.788 0.876 0.813 0.846 0.822 0.833 0.804 0.768 0.769

0.932 0.888 0.888 0.862 0.850 0.857 0.831 0.863 0.828 0.836 0.817 0.811 0.768 0.778 0.766

0.940 0.882 0.863 0.848 0.862 0.839 0.836 0.865 0.828 0.820 0.813 0.805 0.776 0.790 0.749

0.924 0.861 0.801 0.839 0.833 0.786 0.825 0.842 0.810 0.776 0.784 0.720 0.744 0.688 0.671

0.949 0.924 0.902 0.899 0.905 0.882 0.899 0.881 0.772

0.940 0.894 0.873 0.871 0.866 0.822 0.870 0.864 0.875

0.885 0.873 0.849 0.778 0.802 0.750 0.804 0.868 -

0.966 0.931 0.897 0.950 0.920 0.958 0.902 0.849 0.811

0.944 0.937 0.919 0.915 0.923 0.902 0.879 0.847 0.934

0.952 0.935 0.930 0.909 0.915 0.916 0.881 0.885 0.914

0.917 0.903 0.868 0.888 0.863 0.832 0.813 0.845 0.820

Target

Q3 Q4 2.6-5.1 d 5.1-220.0 d

Clinical outcomes In-hospital mortality Hospital LOS ≥ 72 h ICU LOS ≥ 72 h ICD chapters

Elixhauser comorbidities Metastatic cancer Heart failure Diabetes compl. Psychoses Renal failure Alcohol abuse Weight loss Paralysis AIDS / HIV

0.965 0.955 0.957 0.958 0.947 0.948 0.930 0.890 0.923

continued on next page

44

Supplementary Table 10 (continued) Age (yr)

LOS quartile

Target

<50

50-65

65-80

80+

Q1 Q2 0.0-1.0 d 1.0-2.6 d

Fluid / electrolyte Valvular disease Cardiac arrhythmias Pulmonary circulation Liver disease Obesity Depression Drug abuse Hypertension Coagulopathy Lymphoma Blood-loss anemia Solid tumor (non-met.) Peripheral vascular Diabetes uncompl. Other neurological Deficiency anemia Chronic pulmonary Hypothyroidism Rheumatoid / collagen

0.906 0.922 0.833 0.940 0.905 0.896 0.901 0.874 0.900 0.897 0.868 0.900 0.923 0.885 0.885 0.863 0.858 0.804 0.844 0.822

0.894 0.857 0.853 0.848 0.861 0.852 0.839 0.822 0.809 0.870 0.843 0.853 0.823 0.823 0.792 0.823 0.818 0.773 0.730 0.707

0.860 0.833 0.845 0.799 0.810 0.791 0.813 0.798 0.757 0.827 0.815 0.728 0.772 0.763 0.753 0.803 0.771 0.741 0.707 0.643

0.808 0.786 0.814 0.792 0.773 0.806 0.756 0.679 0.696 0.773 0.734 0.716 0.684 0.700 0.742 0.730 0.738 0.719 0.689 0.592

0.847 0.862 0.867 0.903 0.843 0.870 0.902 0.908 0.904 0.842 0.847 0.920 0.861 0.883 0.856 0.800 0.859 0.800 0.809 0.737

45

0.888 0.874 0.885 0.868 0.853 0.879 0.867 0.904 0.896 0.840 0.815 0.886 0.853 0.850 0.858 0.831 0.829 0.798 0.806 0.722

Q3 Q4 2.6-5.1 d 5.1-220.0 d 0.866 0.879 0.883 0.849 0.858 0.856 0.841 0.871 0.886 0.812 0.829 0.805 0.815 0.812 0.828 0.828 0.794 0.773 0.782 0.743

0.830 0.857 0.846 0.783 0.844 0.793 0.778 0.834 0.811 0.793 0.805 0.679 0.752 0.758 0.741 0.763 0.730 0.715 0.710 0.634

a Age at admission (yr) <50 (n=9.7k) 50–65 (n=6.7k)

b Last-stay LOS quartile

65–80 (n=6.4k) 80+ (n=4.3k)

Q1 (0.0–1.0 d) (n=6.8k) Q2 (1.0–2.6 d) (n=6.8k)

Clinical outcomes

Clinical outcomes

ICD chapters

ICD chapters

Comorbidities

Comorbidities

Q3 (2.6–5.1 d) (n=6.8k) Q4 (5.1–220.0 d) (n=6.8k)

In-hospital mortality Hospital LOS>=72h ICU LOS>=72h

Circulatory Neoplasms Injury / poisoning Infectious Factors infl. health (Z/V) Genitourinary Mental / behavioral Respiratory Endocrine / metabolic Digestive / hepatic Symptoms / signs Skin Nervous system Musculoskeletal Eye / ear (sensory) Metastatic cancer Heart failure Diabetes compl. Renal failure Psychoses Weight loss Paralysis Fluid / electrolyte AIDS / HIV Alcohol abuse Valvular disease Pulmonary circulation Coagulopathy Liver disease Cardiac arrhythmias Obesity Depression Lymphoma Other neurological Solid tumor (non-met.) Blood-loss anemia Deficiency anemia Drug abuse Peripheral vascular Diabetes uncompl. Hypertension Chronic pulmonary Hypothyroidism Rheumatoid / collagen 0.6

0.7

0.8

0.9

1.0

AUROC (end-of-stay retrieval)

0.6

0.7

0.8

0.9

1.0

AUROC (end-of-stay retrieval)

Supplementary Figure 1: Linear probing of the final patient state for retrieval. Results split by age at admission and by length of hospital stay. For details on the probing setup, see Section Online Methods above and the caption of Extended Data Fig. 2. We observe that the probe performs better on younger patients and shorter stays across most tasks.

46

Supplementary Table 11: NEWS2 scoring. Recovery of the NEWS2 score from N OAH’s patient state embeddings on the held-out test patients (39,952 windows, 4,414 patients, mean ground truth NEWS2 is 4.25). 2 Linear regression is done using ordinary least squares with an ℓ2 penalty. Rwithin normalizes the error by the within-patient label variance, so 0 means no better than that patient’s own mean score. ρ is Spearman over windows for the window-level results and it is the mean within-patient Pearson for the patient macro results. κw is quadratic-weighted over the three clinical bands. Square brackets are 95% percentile intervals from resampling patients with replacement. Predictor N OAH MLP 512-256

constant (train mean)

constant (train median)

linear regression (ℓ2 ), age only

linear regression (ℓ2 ), 768-dimensional state

Window level MAE

1.336

2.958

2.939

2.841

1.509

[1.30, 1.36]

[2.90, 3.02]

[2.87, 3.00]

[2.78, 2.91]

[1.48, 1.54]

1.784

3.436

3.443

3.355

1.958

[1.75, 1.82]

[3.37, 3.50]

[3.37, 3.52]

[3.28, 3.43]

[1.92, 1.99]

0.730

-0.001

-0.005

0.045

0.675

[0.72, 0.74]

[-0.01, -0.00]

[-0.02, -0.00]

[0.02, 0.06]

[0.66, 0.69]

0.252

-1.669

-1.702

-1.565

0.103

[0.22, 0.29]

[-1.85, -1.52]

[-1.88, -1.54]

[-1.75, -1.41]

[0.06, 0.14]

ρ macro F1 κw AUROC≥7

0.862 0.658 0.757 0.905

0.230 0.000 0.500

0.230 0.000 0.500

0.220 0.315 0.145 0.589

0.836 0.637 0.733 0.891

Patient macro MAE RMSE ρ

0.814 0.940 0.483

3.423 3.545 0.000

3.148 3.278 -

3.111 3.240 -0.066

1.062 1.210 0.335

RMSE R2 R2within

47

Supplementary Table 12: NEWS2 scoring. Top: recovery of the three clinical response bands of the NEWS2 score from the rounded, clipped prediction by the MLP. Bottom: error per true integer score. Risk band

Windows

Share

Precision

Recall

low (0-4) medium (5-6) high (≥7)

21,004 7,740 11,208

52.6% 19.4% 28.1%

0.889 0.387 0.693

0.852 0.411 0.718

True score

Windows

Mean ŷ

MAE

Bias

0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18

7,657 4,394 3,414 3,019 2,520 4,107 3,633 3,261 2,955 2,116 1,298 793 443 182 95 46 11 6 2

0.80 1.42 2.87 3.74 4.51 5.63 6.11 6.64 7.06 7.53 8.01 8.41 9.00 9.41 9.98 10.40 10.28 11.54 13.73

0.81 0.79 1.48 1.55 1.45 1.46 1.32 1.24 1.42 1.75 2.12 2.65 3.07 3.60 4.02 4.60 5.72 5.46 4.27

0.80 0.42 0.87 0.74 0.51 0.63 0.11 -0.36 -0.94 -1.47 -1.99 -2.59 -3.00 -3.59 -4.02 -4.60 -5.72 -5.46 -4.27

Supplementary Table 13: Time-to-event prediction. Survival analysis for within-stay time fractions evaluated using time-dependent C-index (td-CI), integrated Brier score (IBS), and D-calibration (D-cal) p-values. For td-CI and IBS, the mean ± s.d. across the five folds is reported. D-calibration is reported as the p-value from each outer fold. td-CI ↑

IBS ↓

D-cal p-values (< 0.05 → miscalibration)

0.0

0.760 ± 0.005

0.023 ± 0.005

{0.860, 0.994, 0.998, 0.994, 0.994}

0.2

0.784 ± 0.018

0.022 ± 0.006

{0.920, 0.868, 0.882, 0.963, 0.594}

0.4

0.793 ± 0.014

0.021 ± 0.005

{0.864, 0.945, 0.974, 0.854, 0.983}

0.6

0.813 ± 0.014

0.023 ± 0.005

{0.301, 0.561, 0.993, 0.925, 0.943}

0.8

0.830 ± 0.011

0.023 ± 0.009

{0.337, 0.310, 0.958, 0.200, 0.635}

1.0

0.849 ± 0.013

0.024 ± 0.026

{0.603, 0.183, 0.383, 0.417, 0.054}

Within-stay time fraction p

48

Supplementary Table 14: Survival analysis. Summary of the hyperparameter search space. Model

Hyperparameter

Domain {Adam, SGD} [1e-6, 1e-2] [0.0, 0.9] [1e-5, 1e-2] [0.0, 0.5] layers = [ [width] * depth for width in {32, 64, 128, 256} for depth in [1, 6] ]

DeepSurv [47] Optimizer Weight decay Momentum Learning rate Dropout Layers

Supplementary Table 15: Dataset. Event volume statistics per modality along with the value embedding dimensionality where applicable. Pt. cov. is the fraction of cohort patients with at least one event of that modality. Modality

Total Train

Val Test Dim

Categorical 82M 57M 12M Numeric 272M 190M 41M Imaging (CXR) 357k 250k 53k Text 204M 143M 31M ECG 795k 556k 118k ECHO 281k 198k 39k Total

12M 41M 53k 30M 119k 43k

559M 391M 84M 83M

49

768 768 192 512

Pt. cov. 100.0% 94.9% 20.6% 93.3% 53.3% 1.5%

- 299k pts

Supplementary Table 16: Hyperparameter configuration for pretraining N OAH. Training was conducted on 8× NVIDIA A100 80GB GPUs. Hyperparameter

Value

Architecture Trainable parameters Hidden dimension Transformer layers Attention heads FFN expansion ratio Latent bottleneck dimension Context window (tokens) Attention mask Position encoding Temporal encoding layers Numerics Fourier basis LayerScale init Optimization Optimizer Peak learning rate Learning rate schedule Warmup epochs Weight decay Gradient clip norm Dropout Drop path Attention dropout

150,322,970 768 12 12 8 384 2,048 Causal (autoregressive) Temporal (bidirectional time-delta attention) 0-3 22 dyadic scales (2−7 -214 ) 1 × 10−4 Muon (hidden matrices) + AdamW (embeddings, heads, non-2D) 6 × 10−4 Cosine with linear warmup 6 0.01 10 0.1 0.1 0.1

Training Objective Prior-decoded reconstruction weight Posterior-decoded reconstruction weight

0.1 0.3

Variational Latent KL weight βmax KL warmup epochs (linear anneal) Prior standard-deviation floor Standard-deviation cap (smooth squash)

1 10 0.05 2

50

Record · ID 668022 · SHA-256 0f3a7bb049d6f7eb
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.