Conceptio › Archive › arXiv CS
arXiv CSopen access

Intervention Granularity Matters: Coherent Treatment Bundles in Counterfactual Simulation with Clinical World Models

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Intervention Granularity Matters: Coherent Treatment Bundles in Counterfactual Simulation with Clinical World Models

arXiv:2609.21906v1 [cs.LG] 18 Sep 2026

Fangzhou Wang∗† Duke University

Yixuan Yang∗ Duke University

Camilla Balzarotti Duke University

Rishikesan Kamaleswaran Duke University

Abstract Counterfactual simulation with a clinical world model means fixing a patient’s history, changing the treatment, and reading off the predicted response. Doing so requires deciding what counts as one intervention. In clinical settings, interventions are documented as bundles: a co-occurrence audit of 945,707 patienthours from MIMIC-IV shows groups of components, such as every parameter of a dialysis circuit, that never appear apart, so an edit that changes one component on its own describes an hour that never occurs in the data. We hypothesize that the granularity at which an intervention is edited changes how a world model responds, and test this with Clin-JEPA, a latent world model of patient trajectories conditioned on hourly treatment text. At 1,019 documented onsets of invasive ventilation, we keep the patient’s history and other treatments fixed and compare editing one ventilator setting with editing the complete configuration recorded for a real patient with the most similar recent trajectory. The complete bundle moves the predicted next state further than any single setting, consistently across all five settings, and the difference remains after accounting for how much each edit changes the model’s input. Intervention granularity therefore materially affects the response of a clinical world model: single-component edits may understate treatment sensitivity, and bundle-aware editing may offer a better-supported basis for counterfactual treatment simulation.

1

Introduction

Simulating a treatment arm with a patient world model requires a model that responds to the treatment, and a decision about what counts as one treatment. Action-conditioned world models come from reinforcement learning [1, 5, 6]; clinical models now generate patient trajectories from electronic health records [14, 17, 20] and sometimes condition the learned transition on interventions [18, 24, 25, 27]. Clin-JEPA [27] does this for clinical patients: each hour’s observations and treatments are written as text, embedded by a language model, and a predictor maps the recent history of state and treatment embeddings to the next state embedding. In principle one can fix a patient’s history, change the treatment at one hour, and read off the change in the prediction. The question is what to change. ICU treatments are not delivered one item at a time. A co-occurrence audit of 945,707 ICU patienthours (Section 2) shows groups of components, such as every parameter of a dialysis circuit, or every vasopressor together with the norepinephrine-equivalent dose that summarizes them, that appear only together: the probability of seeing one without the others is zero. An edit that adds or removes such a component on its own therefore describes an hour that never occurs in the training data, a violation of positivity [4, 8, 19], and a model that fits the data is free to treat such an input as noise [3]. Action-conditioned world models in other domains are known to disregard actions when the ∗ Equal contribution. † Corresponding author: [email protected]

Preprint.

history fixed

predicted next state

onset hour t+1

24-hour history vital signs, labs

drugs, ventilator settings

t−23

recorded treatment

single setting

coherent bundle

invasive ventilation FiO2

invasive ventilation

invasive ventilation ventilator mode FiO2 PEEP set respiratory rate set tidal volume

PEEP

coherent bundle

world model (frozen)

single setting dz

dz

recorded other treatments kept other treatments kept other treatments kept next-hour response dz observed care reference patient isolated component (most similar trajectory) coherent bundle (never documented (observed in real care) copied from the most similar real patient on its own) hour t

Figure 1: Bundle-consistent editing at a ventilation onset. The 24-hour history is fixed. At the onset hour, the ventilation description is replaced by one setting or by the complete configuration recorded for the reference patient, the real patient with the most similar trajectory; other treatments are kept. The frozen world model predicts the next state under each text, and the response dz is the relative distance from the prediction under the recorded text. Inset: a coherent bundle is a configuration observed in real care; a component on its own is not.

future is predictable from the past [22, 26]. We therefore hypothesize that the granularity at which an intervention is edited matters: a clinically meaningful treatment may be better represented by the full set of settings that are recorded together than by any one of them alone. We test this hypothesis on invasive mechanical ventilation, whose bundle (the ventilation status plus the ventilator settings) is explicitly documented. We propose bundle-consistent editing: at a documented onset of ventilation, the ventilation description is replaced by the complete configuration recorded for the real patient with the most similar 24-hour trajectory (Section 3). On 1,019 onsets, the complete bundle produces a larger next-hour response than any single setting, and the difference remains after adjusting for how far each edit moves the treatment embedding (Section 4).

2

Interventions in ICU records come as bundles

In the data used by Clin-JEPA [10, 27], every ICU hour has a treatment text that lists that hour’s interventions (drugs with doses and rates, ventilator and dialysis settings, fluids, procedures); hours without interventions receive the sentence “No active interventions.” We audited how the components of these texts co-occur in the test split (945,707 unique patient-hours from 12,631 stays). For six common ICU interventions we counted every component present in the same hour as the intervention and computed its co-occurrence ratio ρ(B | I) = P (B | I)/P (B): how many times more often component B appears in hours with intervention I than in hours overall (Appendix C). For blood-pressure management (146,829 hours with an arterial line or a vasopressor), the seven vasoactive agents and the norepinephrine-equivalent dose, a summary of all vasopressors in one unit [13], all have exactly the same ratio, 6.44 = 945,707/146,829; for dialysis (40,604 hours), twenty parameters of the renal-replacement circuit all have exactly the same ratio, 23.29 = 945,707/40,604. A ratio equal to N/NI (N hours in total, NI with the intervention) arises only when every hour containing the component contains the intervention, P (B | ¬I) = 0: these components are never documented apart from the intervention they belong to, and the ventilator settings used in Section 3 behave the same way inside the ventilation bundle. This is how protocol-driven ICU care [9, 12, 21] appears in the record: as bundles of co-interventions. Two design rules for counterfactual edits follow. First, an edit that adds a ventilator setting to an unventilated patient, changes one dialysis parameter on its own, or removes one vasopressor while leaving the summary dose unchanged describes an hour with zero support in the data; components that only occur together should be edited together. Second, derived summaries such as the norepinephrine-equivalent dose (present in 88.3% of blood-pressure-management hours and never outside them) are computed from the individual doses and must be recomputed after an edit; changing a dose without updating the summary produces a text that contradicts itself. 2

3

Bundle-consistent editing at invasive ventilation onset

Model and measurements. We use the released Clin-JEPA checkpoint [27] without retraining. Each hour t of a stay is written as a state text st (vital signs, laboratory results, scores) and the treatment text at described above. A Qwen3-8B encoder with LoRA adapters maps each text to a 4096-dimensional embedding, zt = E(st ) and ut = E(at ), and a 92M-parameter transformer predictor reads the interleaved embeddings of the previous 24 hours and predicts the next state embedding ẑt+1 . All quantities below compare the model’s own predictions; no observed future serves as counterfactual ground truth. When the history is fixed and one hour’s treatment text is edited, the next-hour response is the relative distance between the prediction under the edited text and the prediction under the recorded text, dz = ∥ẑ edit − ẑ rec ∥2 /∥ẑ rec ∥2 , and the edit magnitude is the distance between the two treatment embeddings, da = ∥uedit − urec ∥2 . Bundle and configuration bank. We test the granularity hypothesis on one intervention with a clear bundle, invasive mechanical ventilation, at its onset. The index component is the ventilation status InvasiveVent; the clinician-set components are ventilator mode, FiO2 , PEEP, set respiratory rate, and set tidal volume. Measured consequences of ventilation, such as plateau pressure and minute volume, are excluded: they are readouts, not decisions. From the test split we built a bank of every invasively ventilated hour (229,037 hours, 4,031 stays), keeping for each hour the subset of the five settings actually documented (66.4% record only the status, 16.5% all five). Decision points. A target is an hour t without invasive ventilation followed by an hour t + 1 with it, with no invasive ventilation in the preceding 12 hours, a complete 24-hour history, at least 12 further hours of the stay, and ventilation persisting for at least two hours. Of 3,553 raw transitions in the test split, 1,089 (788 stays) qualify; only 1.5% had six or more hours of non-invasive ventilation in the preceding day, so late rescue after failed non-invasive support is rare (Appendix A). Reference bundles from trajectory-matched patients. Rather than writing ventilator settings by hand, we take them from a real patient: for each target we compared its 24 state embeddings zt−23 , . . . , zt with the 24 pre-onset state embeddings of every other documented onset in the test split (1,130 eligible onsets), using the mean of the hourly cosine similarities, and took the most similar onset from a different stay as the reference; treatment embeddings and post-onset states play no role. Every target found a reference (mean similarity 0.929), 88.2% of reference bundles contain at least three settings and 52.6% all five, and masking the six hours before onset from the matching changes none of this (Appendix B). Conditions and endpoint. For every target we build three versions of the onset-hour treatment text, keeping every non-ventilation fragment as recorded and changing only the ventilation description (Figure 1): the recorded text; the single-setting text, which keeps the ventilation status and copies one setting from the reference bundle, one condition per setting the reference documents; and the bundle text, which keeps the status and copies every setting the reference documents. Each version is re-encoded and passed to the frozen predictor with the unchanged history, and the endpoint is the next-hour response dz . We compare the bundle with each single-setting condition on the same targets by the paired difference in dz with a bootstrap 95% confidence interval (5,000 resamples). To check that the difference is not a matter of how far each edit moves the treatment embedding, we also fit the linear model dz = β0 + βb 1[bundle] + β1 da + ε on the single-setting and bundle rows of the same targets (with a d2a term as a sensitivity analysis); the adjusted bundle effect is βb , the extra response attributed to the bundle at equal edit magnitude. We stop at the next hour on purpose: rolling further would require deciding what treatments follow under each condition.

4

Results

1,019 of the 1,089 targets have a usable reference bundle and enter the comparison; the remaining 70 matched reference bundles contain no eligible clinician-controlled ventilator setting from which to construct a bundle edit. Table 1 and Figure 2 give the mean next-hour response under each condition. On the same patients, the bundle moves the prediction further from the prediction under the recorded text than any single ventilator setting: the paired difference is 0.018 for FiO2 , 0.018 for PEEP, 0.025 for set respiratory rate, 0.014 for set tidal volume, and 0.024 for ventilator mode, 3

Table 1: The complete bundle produces a larger next-hour response than any single setting. Mean next-hour response dz under the single-setting and the bundle condition, on the targets whose reference documents the setting in the row. Bundle − single: mean paired difference with its bootstrap 95% confidence interval. Adjusted: the bundle effect βb at equal edit magnitude. Adj. / single: the adjusted effect as a fraction of the single-setting mean, for scale only. Setting FiO2 PEEP Set respiratory rate Set tidal volume Ventilator mode

n Single setting Coherent bundle Bundle − single [95% CI]

Adjusted

Adj. / single

+0.0178 [0.0157, 0.0201] +0.0178 [0.0157, 0.0199] +0.0246 [0.0220, 0.0274] +0.0140 [0.0112, 0.0167] +0.0241 [0.0202, 0.0281]

+0.0083 +0.0088 +0.0121 +0.0057 +0.0253

27% 29% 39% 14% 61%

918 967 685 677 338

0.0309 0.0303 0.0307 0.0417 0.0415

0.0487 0.0481 0.0553 0.0556 0.0656

(a) Immediate response PEEP Resp. rate

0.031

0.030

0.048

0.031

n=967 0.055

0.042

0.056

0.041

0.03

0.04

n=685 n=677 0.066

0.05

0.06

0.07

Immediate latent response dz

n=338

(c) Action-adjusted difference

0.6

n=918

Single Bundle

Tidal volume Vent. mode 0.02

(b) Action--response relation

0.049

Latent response dz

FiO2

0.5 0.4 0.3

+0.008 (27%)

PEEP

+0.009 (29%)

Resp. rate

0.2

+0.012 (39%)

Tidal volume

0.1

+0.006 (14%)

Vent. mode

0.0

0.08

FiO2

20

40

60

80

Action perturbation da

100

+0.025 (61%)

0.00

0.01

0.02

0.03

Adjusted bundle single dz

0.04

Figure 2: Immediate representation responses to ventilation intervention construction. (a) Mean factual-relative next-state latent response dz for single-component (blue) and coherent-bundle (green) edits on paired targets; n is the number of paired target instances for each setting. (b) Relationship between action-embedding perturbation magnitude da and dz ; curves are descriptive quadratic fits (overall Pearson r = 0.21). (c) Action-distance-adjusted bundle-minus-single differences in dz with 95% target-level bootstrap confidence intervals. Percentages give the adjusted difference relative to the corresponding mean single-component dz . (Appendix D) and every bootstrap confidence interval lies above zero. Adjusting for the edit magnitude changes the picture little (Appendix D): at equal edit magnitude the bundle advantage remains positive for each setting, between 0.006 and 0.025, which is 14% to 61% of the corresponding single-setting response, and the pooled effect is +0.0096 with a linear and +0.0080 (95% CI 0.005 to 0.011) with a quadratic adjustment. The bundle produces the larger response in 70.6–89.5% of paired targets across settings. Mean da ranges from 45.1 to 68.3 across intervention constructions and is not uniformly larger for bundle edits. Excluding the six hours immediately preceding onset also leaves the reference-matching neighborhood largely stable, with 98.2% of masked-window Top-1 references remaining within the original Top-20 neighborhood. The larger immediate response to bundle edits is therefore not fully explained by action-embedding perturbation magnitude alone.

5

Discussion and limitations

Intervention granularity changes what a clinical world model predicts. Under identical histories, a coherent configuration taken from a similar real patient produces a clearly larger immediate response than any of its components alone, and the gap is not a matter of edit size. With the audit, this suggests how to construct counterfactual treatment arms: define the unit of intervention as a bundle found in the data and reviewed clinically; take its values from observed care for similar patients, and report the reference similarity with the effect; and recompute derived summary variables after every edit. Target-trial emulation asks the same of observational analyses, a well-defined and sustained intervention [7], and counterfactual models over time estimate treatment regimes rather than single actions [2, 15, 16]. For treatment-arm simulation [23], single-component edits may understate a model’s treatment sensitivity, and bundle-aware editing may offer a better-supported basis. The claims are limited to representation: a larger response shows that the model distinguishes the bundle from its components, not that the unobserved counterfactual is correct. Because singlesetting configurations also occur in the data, this experiment tests edit granularity rather than empirical support; unsupported edits require direct evaluation.

4

References [1] Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Mojtaba Komeili, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, Sergio Arnaud, Abha Gejji, Ada Martin, Francois Robert Hogan, Daniel Dugas, Piotr Bojanowski, Vasil Khalidov, Patrick Labatut, Francisco Massa, Marc Szafraniec, Kapil Krishnakumar, Yong Li, Xiaodong Ma, Sarath Chandar, Franziska Meier, Yann LeCun, Michael Rabbat, and Nicolas Ballas. V-jepa 2: Self-supervised video models enable understanding, prediction and planning. arXiv preprint arXiv:2506.09985, 2025. URL https://arxiv.org/abs/2506.09985. [2] Ioana Bica, Ahmed M. Alaa, James Jordon, and Mihaela van der Schaar. Estimating counterfactual treatment outcomes over time through adversarially balanced representations. In International Conference on Learning Representations (ICLR), 2020. URL https://arxiv. org/abs/2002.04083. [3] Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann. Shortcut learning in deep neural networks. Nature Machine Intelligence, 2(11):665–673, 2020. doi: 10.1038/s42256-020-00257-z. [4] Omer Gottesman, Fredrik Johansson, Matthieu Komorowski, Aldo Faisal, David Sontag, Finale Doshi-Velez, and Leo Anthony Celi. Guidelines for reinforcement learning in healthcare. Nature Medicine, 25(1):16–18, 2019. doi: 10.1038/s41591-018-0310-5. [5] David Ha and Jürgen Schmidhuber. Recurrent world models facilitate policy evolution. In Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018. URL https://proceedings.neurips.cc/paper_files/paper/ 2018/file/2de5d16682c3c35007e4e92982f1a2ba-Paper.pdf. [6] Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. Mastering diverse control tasks through world models. Nature, 640(8059):647–653, 2025. doi: 10.1038/ s41586-025-08744-2. [7] Miguel A. Hernán and James M. Robins. Using big data to emulate a target trial when a randomized trial is not available. American Journal of Epidemiology, 183(8):758–764, 2016. doi: 10.1093/aje/kwv254. [8] Miguel A. Hernán and James M. Robins. Causal Inference: What If. Chapman & Hall/CRC, Boca Raton, 2020. URL https://miguelhernan.org/whatifbook. [9] Samir Jaber, Boris Jung, Philippe Corne, Mustapha Sebbane, Laurent Muller, Gerald Chanques, Daniel Verzilli, Olivier Jonquet, Jean-Jacques Eledjam, and Jean-Yves Lefrant. An intervention to decrease complications related to endotracheal intubation in the intensive care unit: a prospective, multiple-center study. Intensive Care Medicine, 36(2):248–255, 2010. doi: 10.1007/s00134-009-1717-8. [10] Alistair E. W. Johnson, Lucas Bulgarelli, Lu Shen, Alvin Gayles, Ayad Shammout, Steven Horng, Tom J. Pollard, Sicheng Hao, Benjamin Moody, Brian Gow, Li-wei H. Lehman, Leo A. Celi, and Roger G. Mark. MIMIC-IV, a freely accessible electronic health record dataset. Scientific Data, 10(1):1, 2023. doi: 10.1038/s41597-022-01899-x. [11] Kirsten Neudoerffer Kangelaris, Lorraine B. Ware, Chen Yu Wang, David R. Janz, Hanjing Zhuo, Michael A. Matthay, and Carolyn S. Calfee. Timing of intubation and clinical outcomes in adults with acute respiratory distress syndrome. Critical Care Medicine, 44(1):120–129, 2016. doi: 10.1097/CCM.0000000000001359. [12] Michael Klompas, Richard Branson, Kelly Cawcutt, Matthew Crist, Eric C. Eichenwald, Linda R. Greene, Grace Lee, Lisa L. Maragakis, Krista Powell, Gregory P. Priebe, Kathleen Speck, Deborah S. Yokoe, and Sean M. Berenholtz. Strategies to prevent ventilator-associated pneumonia, ventilator-associated events, and nonventilator hospital-acquired pneumonia in acute-care hospitals: 2022 update. Infection Control & Hospital Epidemiology, 43(6):687– 713, 2022. doi: 10.1017/ice.2022.88. 5

[13] Yuki Kotani, Annamaria Di Gioia, Giovanni Landoni, Alessandro Belletti, and Ashish K. Khanna. An updated “norepinephrine equivalent” score in intensive care as a marker of shock severity. Critical Care, 27(1):29, 2023. doi: 10.1186/s13054-023-04322-y. [14] Zeljko Kraljevic, Dan Bean, Anthony Shek, Rebecca Bendayan, Harry Hemingway, Joshua Au Yeung, Alexander Deng, Alfred Balston, Jack Ross, Esther Idowu, James T. Teo, and Richard J. B. Dobson. Foresight—a generative pretrained transformer for modelling of patient timelines using electronic health records: a retrospective modelling study. The Lancet Digital Health, 6 (4):e281–e290, 2024. doi: 10.1016/S2589-7500(24)00025-6. [15] Rui Li, Stephanie Hu, Mingyu Lu, Yuria Utsumi, Prithwish Chakraborty, Daby M. Sow, Piyush Madan, Jun Li, Mohamed Ghalwash, Zach Shahn, and Li-wei Lehman. G-net: A recurrent network approach to g-computation for counterfactual prediction under a dynamic treatment regime. In Proceedings of Machine Learning for Health (ML4H), volume 158 of Proceedings of Machine Learning Research, pages 282–299. PMLR, 2021. URL https://proceedings. mlr.press/v158/li21a.html. [16] Bryan Lim, Ahmed Alaa, and Mihaela van der Schaar. Forecasting treatment responses over time using recurrent marginal structural networks. In Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018. URL https://papers.nips. cc/paper/2018/hash/56e6a93212e4482d99c84a639d254b67-Abstract.html. [17] Nikita Makarov, Maria Bordukova, Papichaya Quengdaeng, Daniel Garger, Raul RodriguezEsteban, Fabian Schmich, and Michael P. Menden. Large language models forecast patient health trajectories enabling digital twins. npj Digital Medicine, 8(1):588, 2025. doi: 10.1038/ s41746-025-02004-3. [18] Linjie Mu, Zhongzhen Huang, Yannian Gu, Shengqian Qin, Shaoting Zhang, and Xiaofan Zhang. Ehrworld: A patient-centric medical world model for long-horizon clinical trajectories. arXiv preprint arXiv:2602.03569, 2026. URL https://arxiv.org/abs/2602.03569. [19] Maya L. Petersen, Kristin E. Porter, Susan Gruber, Yue Wang, and Mark J. van der Laan. Diagnosing and responding to violations in the positivity assumption. Statistical Methods in Medical Research, 21(1):31–54, 2012. doi: 10.1177/0962280210386207. [20] Pawel Renc, Yugang Jia, Anthony E. Samir, Jaroslaw Was, Quanzheng Li, David W. Bates, and Arkadiusz Sitek. Zero shot health trajectory prediction using transformer. npj Digital Medicine, 7(1):256, 2024. doi: 10.1038/s41746-024-01235-0. [21] Vincenzo Russotto, Sheila Nainan Myatra, John G. Laffey, et al. Intubation practices and adverse peri-intubation events in critically ill patients from 29 countries. JAMA, 325(12): 1164–1172, 2021. doi: 10.1001/jama.2021.1727. [22] Yuhong Shi, Zhenhao Chu, Jie Wei, Jun Hao, Jianyi Liu, and Jingwen Fu. Overcoming statistical bias in action-controllable world models. arXiv preprint arXiv:2608.04653, 2026. URL https://arxiv.org/abs/2608.04653. [23] Kristian Thorlund, Louis Dron, Jay J. H. Park, and Edward J. Mills. Synthetic and external controls in clinical trials – a primer for researchers. Clinical Epidemiology, 12:457–467, 2020. doi: 10.2147/CLEP.S242097. [24] Jiangyuan Wang, Xuyong Chen, Junwei He, Xu Xu, Shasha Xie, and Fuman Han. Chronomedicalworld: A medical world model for learning patient trajectories from longitudinal care data. arXiv preprint arXiv:2605.21963, 2026. URL https://arxiv.org/abs/2605.21963. [25] Qianyi Xu, Gousia Habib, Feng Wu, Dilruk Perera, and Mengling Feng. meddreamer: Modelbased reinforcement learning with latent imagination on complex ehrs for clinical decision support. arXiv preprint arXiv:2505.19785, 2025. URL https://arxiv.org/abs/2505. 19785. [26] Tianzhuo Yang, Zihan Shen, Zirui Mi, Zhaoyi Zhang, Jiayi Zhou, Jiaming Ji, Juntao Dai, Jiawei Chen, Boyuan Chen, and Yaodong Yang. MiraBench: Evaluating action-conditioned reliability in robotic world models. arXiv preprint arXiv:2605.29360, 2026. URL https: //arxiv.org/abs/2605.29360. 6

[27] Yixuan Yang, Mehak Arora, Ryan Zhang, Baraa Abed, Junseob Kim, Tilendra Choudhary, Md Hassanuzzaman, Kevin Zhu, Ayman Ali, Chengkun Yang, Alasdair Edward Gent, Victor Moas, and Rishikesan Kamaleswaran. Clin-jepa: A multi-phase co-training framework for joint-embedding predictive pretraining on ehr patient trajectories. arXiv preprint arXiv:2605.10840, 2026. URL https://arxiv.org/abs/2605.10840.

7

Responsible-use statement This work uses the de-identified MIMIC-IV database under its data use agreement. The world model and the editing protocol are research tools for evaluating models; neither is validated for clinical use, and no output is a basis for an individual treatment decision. One implication is cautionary: a model that appears to support counterfactual simulation may respond to interventions for reasons unrelated to their clinical effect, and the effect sizes read from it depend on how the intervention is specified. We make no claim about the benefit or harm of intubation. Any use of simulated treatment arms as evidence should report the data support of the edited configurations, the similarity of the reference patients, and the sensitivity of the conclusion to the choice of bundle. Responses may also differ across patient groups unevenly represented in ICU data, which we have not examined.

8

A

Decision-point cohort

Table 2: Cohort flow for invasive-ventilation onset targets in the test split. The primary analysis uses the 12-hour washout (bold). Criterion

Target hours

Stays

Raw transition (no InvasiveVent at t, InvasiveVent at t + 1)

3,553

2,903

6-h washout + 24-h history + 12-h future + persistent ≥ 2 h

3,442 1,127 1,125

2,877 805 803

12-h washout + 24-h history + 12-h future + persistent ≥ 2 h (primary cohort)

3,404 1,125 1,091 1,089

2,869 809 790 788

24-h washout + 24-h history + 12-h future + persistent ≥ 2 h

3,108 816

2,813 675

Respiratory support before onset. Late intubation after prolonged non-invasive support is a clinically distinct path with worse outcomes [11], so we checked how common it is among the 1,089 targets. Non-invasive ventilation in the preceding 24 hours: 0 h for 1,069 targets (98.2%), under 6 h for 4, 6 to 12 h for 6, and 12 to 24 h for 10. The documented ventilation status at hour t is: none, 749 (68.8%); supplemental oxygen, 24.5%; high-flow nasal cannula, 4.0%; tracheostomy, 1.7%; non-invasive ventilation, 1.0%. Of the 749 targets with no documented status at t, 449 have no documented status anywhere in the preceding 24 hours; for the other 300 the last documented status was invasive ventilation (225, all 12 to 24 hours before t, consistent with re-initiation rather than a first decision), supplemental oxygen (61), tracheostomy (6), high-flow nasal cannula (5), or non-invasive ventilation (3). SpO2 is available at t for 87.2% of the 749 but an oxygen flow rate for only 2.3%, so we did not impute a status for this heterogeneous group. The targets occur a median of 95 hours after ICU admission (interquartile range 49 to 171, range 23 to 323).

B

Reference matching and masked-window sensitivity

To reduce sensitivity to documentation immediately preceding ventilation onset, masked-window matching excludes the final six pre-onset hours from the similarity calculation, using only state embeddings from t−23 through t−6; the target onset time and subsequent intervention construction are unchanged. Table 3: Reference matching with the full 24-hour window versus the window with the last six hours before onset masked. Similarity: mean hourly cosine similarity of the best match.

Targets matched Mean similarity of best match 5th percentile of similarity Distinct references selected Reuse per reference: median / 95th pct. / max Bundles with ≥ 3 of 5 settings Bundles with all 5 settings Same best match under both windows Masked best match within full-window top 20

C

Full window (t − 23 to t)

Masked window (t − 23 to t − 6)

1,089 / 1,089 0.929 0.894 518 1 / 5 / 12 88.2% 52.6%

1,089 / 1,089 0.934 0.899 541 1 / 5 / 15 88.3% 55.2% 59.2% 98.2%

Co-occurrence audit

The audit uses one row per unique (stay, hour) pair of the test split (N = 945,707) and treats every fragment of the treatment text as a component (repeated mentions in one hour count once). 9

The six interventions are invasive ventilation, blood-pressure management, oxygen escalation, dialysis, neuromuscular blockade, and antipsychotic treatment. Intervention hours are defined from explicit components: invasive ventilation from the status field; blood-pressure management from an arterial-line record or any explicit vasoactive agent, with the derived norepinephrine-equivalent dose deliberately excluded from the definition; dialysis from the renal-replacement status; and likewise for the other three. For each intervention, components are ranked by the co-occurrence ratio after support gates (at least 10 co-occurring hours, at least 20 hours overall, presence in at least 0.5% of intervention hours). Identical ratios within a group identify components that occur only inside the intervention. For blood-pressure management, eight components had the same ratio of 6.44, including the seven intervention-defining vasoactive agents and the norepinephrine-equivalent dose. For dialysis, twenty dialysis/CRRT components were similarly fully contained within dialysis hours (ratio 23.29). Notably, although the norepinephrine-equivalent dose was excluded from the blood-pressure-management definition, it appeared in 88.3% of blood-pressure-management hours and never outside them.

D

Response versus edit magnitude

This figure plots the next-hour response dz against edit magnitude da for every single-setting (blue) and bundle (green) edit. The response rises only weakly with da (Pearson correlation 0.21), but the main comparison still adjusts for it for the sake of fairness; the trend lines are descriptive.

10

Record · ID 1006850 · SHA-256 44a7cc6c3bc7d099
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.