ConceptioArchivearXiv CS
arXiv CSopen access

Learning Normal Representations for Blood Biomarkers

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Learning Normal Representations for Blood Biomarkers Aashna P. Shah1,2,6 , Michelle M. Li1,6 , Yash Lal4 , Seffi Cohen1,6,7 , Liat F. Antwarg1,6,7 , Morgan Sanchez1,6 , James A. Diao1,3,6 , Chirag J. Patel1 , Ben Y. Reis1,5,6,7 , Ran D. Balicer6,7 *, Noa Dagan6,7,8 *, and Arjun K. Manrai1,6 * 1. Department of Biomedical Informatics, Harvard Medical School, Boston, MA, USA 2. Department of Systems Biology, Harvard Medical School, Boston, MA, USA 3. Department of Medicine, Brigham and Women’s Hospital, Boston, MA, USA 4. Department of Mathematics, Johns Hopkins University, Baltimore, MD, USA 5. Computational Health Informatics Program (CHIP), Boston Children’s Hospital, Boston, MA, USA 6. The Ivan and Francesca Berkowitz Family Living Laboratory Collaboration at Harvard Medical School and Clalit Research Institute,

arXiv:2605.18701v1 [cs.LG] 18 May 2026

USA and Israel 7. Clalit Research Institute, Innovation Division, Clalit Health Services, Ramat-Gan, Israel 8. Faculty of Computer and Information Science, Ben Gurion University, Be’er Sheva, Israel

* Co-senior authors Correspondence: [email protected]

Abstract Blood-based biomarkers underpin clinical diagnosis and management, yet their interpretation relies largely on fixed population reference intervals that ignore stable, intra-patient variability. As such, population-based interpretation can mask meaningful deviation from an individual’s baseline, risking delayed disease detection. To remedy this, there have been increasing efforts to personalize blood biomarker interpretation using individual testing histories. However, these methods may overfit to sparse data, inflating false-positive rates and unnecessary follow-up, and can also unwittingly include unrecognized or subclinical disease. Here, we leverage nearly 2 billion longitudinal laboratory measurements from over 1.6 million individuals across North America, the Middle East, and East Asia, to show that while laboratory values are highly individual, purely personalized intervals routinely overfit, classifying up to 68% of measurements as abnormal, without corresponding associations with adverse clinical outcomes. We then introduce NORMA, a conditional transformer-based framework that generates reference intervals by conditioning on both a patient’s history and population-level data about “normal” variation. NORMA-derived intervals achieve higher precision for predicting outcomes, including mortality, acute kidney injury, and chronic disease. These findings caution against over-personalization in laboratory medicine and demonstrate that anchoring individual trajectories to population-level priors outperforms either approach alone. To promote transparency, we publicly release the model, code, and an interactive user interface for accessible, individualized laboratory interpretation.

Introduction Laboratory testing is among the most frequently performed medical tests, with more than 14 billion tests ordered annually in the United States alone 1 . Yet interpretation has changed remarkably little since the 1960s; values are classified as “low,” “normal,” or “high” against fixed population reference intervals (PopRI ) that represent the central 95% of measurements from ostensibly healthy individuals 2–8 . Prior studies have recognized that many routine biomarkers fluctuate around narrow individual “setpoints” and that within-person change carries prognostic information that population intervals miss 9–20 . Despite this, clinical practice often still relies on universal thresholds, partly because of the ease of a one-size-fits-all approach, and because operationalizing individualized interpretation requires modeling patient-specific trajectories at scale.

1

Efforts to move beyond universal thresholds have taken two directions. The first refines reference intervals for predefined subgroups, for example, sex-specific hematologic ranges or adjusted HbA1c thresholds for patients with hemoglobin variants 12,14,21–24 . Defining population subgroups introduces several challenges— such groups are often defined coarsely or arbitrarily, may encode existing biases, and cannot capture the full spectrum of individual variation embedded in longitudinal trajectories. The practice of race adjustment, in particular, has been increasingly challenged with major clinical and non-clinical implications for patients 25–27 . The second approach derives personalized reference intervals (PerRI ) directly from each patient’s laboratory testing history. Foy et al., for example, showed that deviations from individually derived hematologic setpoints associate more strongly with mortality than deviations from population thresholds, underscoring the prognostic value of within-person change 16 . However, purely individualized approaches that derive baselines entirely from a patient’s own data risk incorporating unrecognized chronic disease into the estimated setpoint and overfitting to sparse histories, producing overly narrow intervals that label benign physiological variation as pathological or vice versa 16,18,28,29 . Here, we ask whether these competing approaches can be combined using nearly 2 billion measurements across 30 analytes, spanning two decades of national longitudinal care, dense ICU monitoring, and perioperative data across three countries (Fig. 1). We first quantify the individuality of routine blood tests and then systematically compare population and personalized reference intervals. We then introduce NORMA (Normal Outcome Range Modeling with Attention), a new conditional transformer framework that combines both approaches to generate individualized reference intervals by anchoring each patient’s trajectory to populationlevel expectations for a healthy state. Rather than choosing between the false dichotomy of purely population and personalized approaches, NORMA learns where each patient’s interval is expected to fall between the two extremes. We show that purely personalized intervals overcall abnormalities and lose prognostic signal, that NORMA resolves this tradeoff, and test whether the resulting intervals detect clinically meaningful change months to years before population intervals.

Results Study populations We assessed population-based (PopRI ), personalized (PerRI ), and NORMA-derived (NORMARI ) reference intervals across 30 routine laboratory analytes spanning four clinical panels: complete blood count, comprehensive metabolic panel, lipid panel, and hemoglobin A1c (HbA1c). These panels are used in routine health evaluation, cardiometabolic risk assessment, and diabetes screening or monitoring, as detailed in the Methods. Within this analyte set, NORMA was trained on 3.4 million longitudinal laboratory sequences from two development cohorts: MIMIC-IV (179,601 patients) and EHRSHOT (5,676 patients); no development data were used in external validation (Supplementary Tables 1–3). We validated NORMA in three independent cohorts that differed in temporal scale and clinical setting. The Clalit Health Services (CHS) cohort included 1,450,862 adults (63% female; mean age 55.1 ± 18.6 years) contributing approximately 1.9 billion laboratory measurements over more than two decades of outpatient follow-up (Fig. 2a; Supplementary Table 3). To ensure stable baseline estimation, we required at least five outpatient measurements per analyte spaced at least 90 days apart, yielding trajectories spanning a median of more than four years. The eICU Collaborative Research Database (eICU-CRD) cohort included 98,432 ICU patients (46% female; mean age 63.3 ± 16.1 years) across 208 U.S. hospitals, contributing over 20 million measurements with dense sampling (median number of measurements per patient ranged between 8-35 for high-frequency analytes) over a median stay of approximately 8 days.

2

The INSPIRE cohort included 51,159 surgical patients (52% female; mean age 57.6 ± 14.9 years) undergoing perioperative care in South Korea, contributing over 10 million laboratory measurements collected across preoperative, intraoperative, and postoperative timepoints, with a median follow-up of approximately 175 days. Laboratory distributions in CHS, EHRSHOT, and INSPIRE largely overlapped population reference intervals, whereas eICU-CRD and MIMIC-IV exhibited shifts consistent with greater rates of critical illness, including lower albumin (ALB) and higher glucose (GLU) variability (Supplementary Table 3).

Lab values are individualized Across all three cohorts, within-person variability was substantially smaller than between-person variability for most analytes, confirming that laboratory values fluctuate around narrow individual setpoints (Fig. 2b; Supplementary Table 4). We quantified this finding using the individuality index (II), the ratio of within-person to between-person coefficient of variation, where values below 0.6 indicate high biological individuality. Hematologic indices showed the lowest intra-individual variability relative to inter-individual variability. Mean corpuscular hemoglobin (MCH), mean corpuscular volume (MCV), and red blood cell count (RBC) had individuality indices ranging from approximately 0.2 to 0.6 across CHS, eICU-CRD, and INSPIRE, with values in eICU-CRD generally equal to or lower than those observed in CHS and INSPIRE (Supplementary Table 4). In contrast, electrolytes and metabolic analytes exhibited lower biological individuality; potassium (K), glucose (GLU), sodium (NA), and calcium (CA) had individuality indices ranging from approximately 0.7 to 1.2 across cohorts, with within-person variability approaching the population spread (Fig. 2b; Supplementary Table 4). INSPIRE showed patterns consistent with an intermediate regime between outpatient and ICU settings. Within-person stability carried prognostic signal. Greater deviation from a patient’s baseline, quantified as the absolute z-score of each index measurement relative to that patient’s baseline average, was monotonically associated with mortality across all cohorts (Fig. 2c; Supplementary Fig. 3). When stratified by raw analyte values, several biomarkers, including white blood cell count (WBC), glucose (GLU), aspartate aminotransferase (AST), and hemoglobin A1c (HbA1c), exhibited U-shaped mortality curves, in which both abnormally high and low values were associated with elevated risk.

NORMA generates calibrated individualized intervals NORMA is a conditional, autoregressive transformer framework with multiple configurations that models the distribution of a patient’s next laboratory value given their longitudinal measurement history (Fig. 3a; Methods) and population data on “healthy” variation. The model is trained to forecast the next observed value. To derive a personalized reference interval, we condition the query token on values within the population reference range, so that the resulting 95% prediction interval represents the expected range for that patient given their prior trajectory and population context, rather than a strictly disease-free state. We trained NORMA on 3.4 million longitudinal sequences from MIMIC-IV and EHRSHOT under two output parameterizations (Gaussian and quantile), which share the same architecture but differ in distributional assumptions (Methods and Supplementary Table 2). On a held-out test set, NORMA outperformed all baselines in next-step forecasting accuracy across both output parameterizations. Baseline models included last value carried forward, autoregressive integrated moving average (ARIMA), and patient-specific mean. Mean absolute error was lowest for NORMA quantile (5.9 [IQR 3.6–9.1]) and NORMA Gaussian (6.0 [2.6–10.2]), compared with ARIMA (6.7 [4.3–11.9]), last value carried forward (7.3 [3.9–10.7]), and patient-specific mean (8.4 [4.5–13.6]; Supplementary Table 5; Fig. 3b). Forecasting accuracy varied by analyte, with mean platelet volume (MPV) showing the weakest performance, consistent with its having the smallest training sample size (Fig. 3c; Supplementary Fig. 2).

3

NORMA’s prediction intervals responded appropriately to the factors that govern uncertainty (Fig. 3d; Supplementary Table 6). More prior measurements narrowed intervals, longer prediction horizons widened them, and greater within-person variability produced substantially wider intervals. The quantile parameterization was more responsive than the Gaussian to changes in input features; for example, doubling within-person variability roughly doubled the predicted interval width under the quantile model, compared with minimal change under the Gaussian. Importantly, interval width stabilized after approximately 30 prior measurements and did not continue to narrow with additional data, indicating that the model does not overweight longer histories.

NORMARI detects abnormalities earlier than population intervals In all three external validation cohorts, PerRI flagged the most tests as abnormal, PopRI the fewest, and NORMARI was intermediate. In CHS, abnormality rates were 29.6% (PopRI ), 39.1% (NORMARI ), and 46.8% (PerRI ). In eICU-CRD, where baseline abnormality rates were higher due to acute illness, the same ordering held: 50.2% (PopRI ), 55.8% (NORMARI ), and 68.1% (PerRI ). In INSPIRE, abnormality rates followed the same pattern, with values of 35.1% (PopRI ), 42.5% (NORMARI ), and 57.0% (PerRI ), consistent with its intermediate clinical setting (Fig. 4a; Fig. 5a; Fig. 6a; Supplementary Table 7). Among tests classified as normal by population intervals, NORMARI reclassified 12, 9, and 14 per 100 as abnormal in CHS, eICU-CRD, and INSPIRE, respectively, compared with 27, 37, and 34 per 100 for PerRI (Supplementary Tables 7–8). NORMARI flagged abnormalities earlier than population intervals across all cohorts. In CHS, 49% of tests NORMARI flagged as abnormal were later also flagged by PopRI , with a median lead time of 8.7 months (IQR 1.7–33.8; Supplementary Table 20; Fig. 4e). In eICU-CRD, 25% were later confirmed by PopRI , with a median lead time of 34.8 hours (IQR 23.1–72.6; Supplementary Table 20; Fig. 5e). In INSPIRE, 30% were later confirmed by PopRI , with a median lead time of 23.7 hours (IQR 11.4–54.7), consistent with its perioperative timescale (Fig. 6e; Supplementary Table 20).

NORMARI detects clinically meaningful abnormalities missed by population intervals We next asked whether the additional abnormalities detected by NORMARI carried clinical meaning. Restricting to measurements classified as normal by population intervals, where personalization could add signal beyond standard practice, NORMARI reclassifications carried substantially higher positive predictive value than PerRI . For every 100 patients PopRI classified as normal but NORMARI flagged as abnormal in eICU-CRD, 13 died in hospital versus 10 under PerRI (∆ = +3), 16 developed AKI versus 15 (∆ = +1), and 27 had prolonged ICU stays versus 23 (∆ = +4). In CHS, 30 died versus 24 under PerRI (∆ = +6), 46 developed CKD versus 39 (∆ = +7), and 80 had type 2 diabetes versus 75 (∆ = +5). In INSPIRE, NORMARI identified 6 more unplanned ICU admissions, 1 more in-hospital death, and 1 more perioperative infection per 100 reclassified patients than PerRI (Supplementary Tables 17–19). Across all assessed outcomes – mortality, type 2 diabetes, and CKD over 10 years in CHS; in-hospital mortality, AKI, sepsis, and prolonged LOS in eICU-CRD; and perioperative mortality, prolonged LOS, unplanned ICU admission, and postoperative infection in INSPIRE – PerRI achieved higher sensitivity but NORMARI achieved substantially higher specificity and positive predictive value (Figs. 4c,d, 5c,d, 6c,d and Supplementary Figs. 9–11).

NORMARI improves clinical prediction In time-to-event analyses across all cohorts, NORMARI and PopRI abnormality flags showed comparable prognostic associations with clinical outcomes, while PerRI flags were substantially weaker (Figs. 4b, 5b, 6b and Supplementary Tables 21–29). This was most pronounced in CHS, where NORMARI abnormality flags

4

were strongly associated with all-cause mortality for hematologic and metabolic analytes, with several showing 2- to 4-fold elevated risk after adjusting for age and sex – including creatinine (HR 3.30), hemoglobin (HR 3.72), and hematocrit (HR 3.36). Effect sizes were comparable to or slightly exceeded those of PopRI for most analytes, while PerRI flags carried essentially no prognostic signal. Mean concordance indices for mortality were similar between NORMARI and PopRI across all three cohorts (CHS: 0.752 vs. 0.758; eICU-CRD: 0.567 vs. 0.571; INSPIRE: 0.730 vs. 0.731), while PerRI was consistently lowest (CHS: 0.748; eICU-CRD: 0.562; INSPIRE: 0.695). The gap was most pronounced in INSPIRE, where PerRI concordance (median 0.718, IQR [0.673, 0.742]) fell well below both NORMARI (0.754 [0.712, 0.780]) and PopRI (0.742 [0.722, 0.778]; Supplementary Tables 21, 22, 26).

Discussion Routine laboratory testing is becoming more frequent, more accessible, and more central to medical decisionmaking 30,31 . Beyond tests that clinicians order within the healthcare system, patients are increasingly seeking out repeated panel-based screening through direct-to-consumer companies. In parallel, interpretation is shifting away from static cutoffs and monolithic “normal ranges” towards “personalized” ranges that leverage trajectories of variation and individual baselines, yet the implications and risks of this shift remain largely unclear 17–19,28 . Across nearly two billion laboratory measurements spanning outpatient and intensive care settings, we find that purely personalized reference intervals consistently overcall abnormalities and are poorly associated with adverse outcomes. NORMA balances purely personalized and population intervals by training a transformer to predict what a patient’s next laboratory value is expected to fall if they remain physiologically stable, and defining the reference interval from this conditional distribution. In doing so, NORMARI inherits the sensitivity of personalized interpretation while maintaining the specificity of population-level definitions of normal variation. In eICU, PerRI flagged the majority of laboratory measurements as abnormal, and similarly elevated abnormality rates were observed in CHS and INSPIRE compared to population-based intervals, indicating that a purely personalized framework would likely generate alerts for a substantial fraction of tests across critical care, outpatient, and perioperative settings. At the scale of a national health system processing millions of tests annually, this could translate to an enormous burden of false alarms, unnecessary follow-up testing, and potential patient anxiety 29,32,33 . By contrast, NORMARI moderated this alert burden, flagging fewer measurements than PerRI while yielding abnormalities that were more strongly enriched for adverse clinical outcomes. This benefit was not apparent from aggregate discrimination alone. Across individual analytes and outcomes, NORMARI and PopRI achieved broadly similar concordance indices, whereas PerRI consistently performed worst. This similarity between NORMARI and PopRI might suggest that personalization adds little value; however, the value of NORMARI is seen in individuals that PopRI labels normal but NORMARI flags as abnormal. In this reclassified subset, NORMARI achieved substantially higher precision and specificity than PerRI . By anchoring individualized intervals to population-level expectations, NORMA reduces the risk that chronic or subclinical disease is absorbed into a patient’s estimated baseline, a limitation of purely personalized approaches such as PerRI . This limitation is not unique to the Gaussian mixture approach used for PerRI . Even with Bayesian updating or hierarchical extensions, a population prior converges toward the patient’s observed distribution as measurements accumulate, eventually washing out the prior and reproducing the same overcalling behavior 28,34 . However, such methods are also nondeterministic, computationally expensive, introduce prior-strength and group-structure hyperparameters, and lack native handling of time or irregular

5

sampling. NORMA avoids this by conditioning on a healthy state at every prediction, regardless of how many measurements are available, which is why its interval width stabilizes rather than continuing to narrow. Furthermore, it does not require explicit pre-filtering of trajectories for stability; it handles irregular measurement spacing natively and its inputs are limited to data already present in most electronic health records. In a clinical deployment, NORMA could provide more "personalized" assessment for patients even without sufficient measurement history. NORMA builds on a parallel line of work that scales sequence modelling over longitudinal health records into clinical foundation models. Early efforts demonstrated that deep learning over raw EHR streams could match or exceed task-specific models for mortality, readmission, and laboratory forecasting 35–38 . More recent transformer-based foundation models, trained on event streams from millions of patient timelines, have demonstrated that longitudinal records can support zero- or few- shot forecasting of diagnoses, procedures, and disease progression at increasing scale, with prediction horizons extending years into the future and architectures expanding to incorporate multimodal clinical data 39–42 . The same paradigm is now extending beyond structured EHRs to continuous physiological streams such as continuous glucose monitoring and other wearables 43,44 . These models demonstrate that longitudinal patient representations can support broad outcome prediction across clinical domains. While NORMA learns from patient trajectories, it is distinct in the prediction task. Rather than learning a general representation of trajectories, such as APOLLO, to predict diagnoses, procedures, or downstream clinical outcomes, it models continuous biomarker distributions and estimates the range of values expected for a given patient under a specified future laboratory state 42 . NORMA therefore draws on counterfactual prediction and sequence-editing frameworks, such as CLEF, which ask how a trajectory would change under an imposed condition or intervention 45–47 . Here, the imposed condition is a future normal laboratory state, allowing NORMA to estimate an individualized healthy-state reference interval rather than simply forecast the next observed measurement. The clinical outcomes evaluated here, including all-cause mortality, acute kidney injury, type 2 diabetes, and chronic kidney disease, were used only for downstream validation rather than model training. This disease-agnostic design suggests that deviations from an individual’s expected healthy trajectory may carry prognostic relevance across diverse disease states without requiring outcome-specific retraining. As longitudinal biomarker monitoring expands further into precision medicine, the challenge of reconciling population-derived and individualized reference standards will broaden, and conditional prediction anchored to a healthy reference state may generalize across these settings 43,44,47,48 . Our study had several limitations. First, NORMA’s current scope is limited by the analytes and clinical inputs used for training. NORMA was trained and evaluated on 30 common blood tests and conditioned only on age, sex, and laboratory trajectories; it does not currently incorporate comorbidities, medications, or other clinical context, which could further refine expected trajectories and abnormality thresholds 49 . Performance was also not uniform across analytes. MPV showed the lowest forecasting accuracy, consistent with its having the smallest training sample size and weakest forecasting accuracy among the 30 analytes. Second, some modeling assumptions may affect calibration. In CHS, we modeled predictive distributions as Gaussian, which may not fully capture skewed or heavy-tailed analytes. However, we implemented quantile regression in eICU, and results were consistent across both parameterizations, suggesting that NORMA’s clinical value does not depend on a strict Gaussian assumption. Broader exploration of output distributions may further improve calibration for skewed or heavy-tailed analytes. Third, generalizability and clinical impact remain to be established prospectively. The CHS analysis was conducted within a single national health system; eICU, although geographically diverse, represents only the ICU segment of hospital care; and INSPIRE reflects perioperative care at a single academic medical center. Performance may differ in emergency departments, primary care clinics, non-surgical inpatient settings, or populations with different demographics and laboratory utilization patterns. Finally, this study evaluated retrospective associations between abnormality flags and

6

clinical outcomes. Prospective trials are needed to determine whether NORMARI -augmented interpretation accelerates time to diagnosis, alters downstream testing, and influences clinician behavior. Such evaluation may further benefit from cohort-specific fine-tuning, for example on Clalit data before deployment, which could improve local performance at the cost of generalizability. Individualized laboratory interpretation can be made more precise and more clinically useful by combining patient-specific trajectories with population-level expectations and uncertainty-aware prediction. By validating across longitudinal outpatient and acute settings, outcome horizons, and patient populations, we show that this framework generalizes beyond a single clinical context. NORMA provides a framework for integrating personalized laboratory interpretation into routine practice, balancing the sensitivity of personalized intervals with the specificity of population-anchored prediction.

Methods Study Populations We used five cohorts spanning model development and external validation. Model development used MIMIC-IV and EHRSHOT. External validation used three independent cohorts representing distinct clinical settings and time horizons: Clalit Health Services (CHS), a national longitudinal health system cohort from Israel; the eICU Collaborative Research Database (eICU-CRD), a multicenter critical care cohort from the United States; and INSPIRE, a perioperative cohort from Seoul National University Hospital in South Korea. Across these cohorts, we evaluated three reference interval frameworks: population-based reference intervals (PopRI), personalized reference intervals (PerRI ), and NORMA-derived reference intervals (NORMARI ). No development data were used in external validation. MIMIC-IV and EHRSHOT. MIMIC-IV contributed 179,601 patients from a tertiary-care hospital system and was used for acute-care, short-horizon model development. EHRSHOT contributed 5,676 patients from longitudinal electronic health records and was used for intermediate-horizon model development. Together, these cohorts provided the development data used for NORMA training, validation, and held-out testing. Clalit Health Services. The CHS cohort comprises nationwide longitudinal electronic healthcare data from Israel’s largest health maintenance organization 50 . We included adults aged 18–99 years with repeated outpatient laboratory testing between 2000 and 2024. For each analyte, eligibility required at least five outpatient measurements spaced at least 90 days apart prior to January 1, 2015 (baseline period) and at least one measurement between January 1, 2015 and January 1, 2016 (classification period). We excluded hospitalizations, emergency department visits, and urgent encounters during baseline ascertainment. The first measurement during the classification period served as the index value and defined the index date. We assessed eligibility independently per analyte, allowing individuals to contribute to multiple analyses. We followed individuals for up to ten years from each analyte-specific index date. For those without events, we censored follow-up at the earliest of health plan disenrollment or death. Patients remaining enrolled without a recorded death were censored at the last available date in the dataset (January 1, 2024). Primary outcomes included all-cause mortality and incident diagnoses of type 2 diabetes (T2D) and chronic kidney disease (CKD). We ascertained outcomes using ICD-9/10 codes supplemented by laboratory criteria. We defined T2D as HbA1c ≥6.5% or fasting glucose ≥126 mg/dL on two occasions; CKD as an estimated glomerular filtration rate <60 mL/min/1.73 m² on two measurements at least 90 days apart. eICU Collaborative Research Database. The eICU Collaborative Research Database (eICU-CRD v2.0) is a multicenter critical care cohort comprising more than 200,859 ICU admissions from 208 hospitals across

7

the United States 51 . We included adults aged 18–99 years with repeated laboratory measurements during their ICU stay. We applied a baseline filter requiring at least five measurements. After filtering, 98,432 unique patients contributed over 20 million measurements across 29 analytes. For each patient–analyte sequence, we used the first 75% of measurements as the baseline to estimate PerRI and NORMARI , and we used the remaining measurements as index values for classification. Primary outcomes included in-hospital mortality, acute kidney injury (AKI), sepsis, and prolonged ICU length of stay (>7 days). Outcomes were ascertained using ICD-9 diagnosis codes and structured clinical documentation. INSPIRE. The INSPIRE dataset is a publicly available perioperative research cohort comprising approximately 130,000 surgical cases from a single academic medical center (Seoul National University Hospital) in South Korea between 2011 and 2020 52 . We included adults aged 20–90 years with repeated laboratory measurements during the perioperative period. We applied a baseline filter requiring at least five measurements. After filtering, 51,159 unique patients contributed over 10 million measurements across 19 analytes. For each patient–analyte sequence, we used the first 75% of measurements as the baseline to estimate PerRI and NORMARI , and we used the remaining measurements as index values for classification. Primary outcomes included in-hospital mortality, prolonged length of stay, unplanned ICU admission, and perioperative infection. Outcomes were ascertained using ICD-10-CM diagnosis codes and structured clinical documentation.

Laboratory Measurements We selected 30 routinely measured blood analytes from common clinical panels based on clinical ubiquity and sufficient repeated measurement across development and validation datasets. We grouped analytes into four standard clinical panels: complete blood count (hematocrit, hemoglobin, mean corpuscular hemoglobin, mean corpuscular hemoglobin concentration, mean corpuscular volume, mean platelet volume, platelet count, red blood cell count, red cell distribution width, white blood cell count); comprehensive metabolic panel (sodium, potassium, chloride, calcium, bicarbonate, blood urea nitrogen, creatinine, glucose, aspartate aminotransferase, alanine aminotransferase, alkaline phosphatase, total bilirubin, direct bilirubin, total protein, albumin); lipid panel (total cholesterol, low-density lipoprotein cholesterol, high-density lipoprotein cholesterol, triglycerides); glycemic control (glycated hemoglobin). We mapped all laboratory results to Logical Observation Identifiers Names and Codes (LOINC) 53 . We retained only positive values, standardized units to canonical formats, and excluded analyte-specific outliers exceeding ±3 standard deviations from the analyte median. When duplicate results shared identical timestamps, we averaged values within each analyte.

Reference Interval Frameworks Population Reference Intervals. We defined PopRI using externally specified laboratory reference ranges and clinical decision thresholds rather than estimating cohort-specific intervals. For each analyte, we used the adult reference range reported by the American Board of Internal Medicine Laboratory Test Reference Ranges 54 . Where sex-specific reference intervals were provided, classifications were made using the patient’s sex-specific range. Supplementary Table 1 provides the reference interval, sex stratification, and canonical unit used for each analyte. Personalized Reference Intervals. We estimated PerRI following Foy et al. For each patient-analyte pair, we fit Gaussian mixture models with one to three components to baseline measurements and selected model order using the Akaike information criterion 16 . From the selected model, we chose the mixture component that explained the largest number of that patient’s baseline measurements and treated it as the individual’s 8

physiological setpoint. We defined the PerRI as the mean of this component ± 2 standard deviations.

NORMA (Normal Outcome Range Modeling with Attention) NORMA is a conditional, decoder-only transformer that models the distribution of a patient’s next laboratory value given their longitudinal measurement history and a query specifying a future health state and prediction horizon. We trained the model on 3.4 million longitudinal sequences from MIMIC-IV and EHRSHOT, requiring at least three repeated measurements per patient without constraints on sampling intervals or clinical setting 55,56 . Data were split at the patient level into training (70%), validation (10%), and test (20%) sets to prevent leakage across partitions. Each sequence was tokenized, padded, and masked to handle variable-length histories. A weighted sampling scheme was used during training to balance representation across 30 laboratory analytes. Alternative input encodings with justifications are summarized (Supplementary Table 3). Input representation. For a given patient-biomarker pair, we denote the ordered sequence of observed laboratory values as x = (x1 , . . . , xT ), with corresponding measurement times t = (t1 , . . . , tT ) and clinical states s = (s1 , . . . , sT ), where each si indicates the laboratory state relative to the population reference interval (low, normal, or high). The model processes these inputs as a sequence of three token types: a context token, history tokens, and a query token. The context token zc encodes static patient covariates: sex g, age a, and laboratory test code c, as a sum of learned embeddings:

zc = eg (g) + ea (a) + ec (c) Each history token zihist represents one prior measurement and combines three components: zihist = fv (x̃i ) + es (si ) + fτ (∆ti ) where fv is a linear projection of the input value x̃i , es is a learned state embedding, and fτ encodes the inter-measurement interval ∆ti = ti − ti −1 . The query token, zquery , specifies the context at the time of prediction: zquery = es (sT +1 ) + fh (tT +1 − tT ) where sT +1 denotes the future health state and tT +1 − tT the prediction horizon. The full sequence is processed using a causally masked Transformer decoder to obtain the predictive distribution of the next observed laboratory value by P(xT +1 |x, t, s, g, a, c, s∗, t ∗). Training Objective. Given the input sequence, the model is optimized to predict the conditional distribution of the next observed value xT +1 . We trained NORMA under two output parameterizations. Under the Gaussian parameterization, the model jointly estimates a predicted mean µ and log-variance log

σ ²: L = 12 [log σ 2 + (xT +1 − µ)2 /σ 2 ] Under the quantile parameterization, the model directly predicts fixed quantiles τ ∈ {0.025, 0.25, 0.50, 0.75, 0.975} using a pinball loss without distributional assumptions:

9

Lτ = τ × max(0, xT +1 − qτ ) + (1 − τ ) × max(0, qτ − xT +1 ) where qτ is the predicted τ -th quantile. Both parameterizations share the same core architecture and were trained on the same development data. NORMA reference intervals. We define the NORMA reference interval (NORMARI ) as the 95% prediction interval obtained by conditioning on a future normal laboratory state. Under the Gaussian parameterization, this corresponds to [µ− 1.96σ , µ + 1.96σ ]; under the quantile parameterization, it corresponds to [q̂0.025 , q̂0.975 ]. Evaluation. We first evaluated forecasting performance for next-step prediction. For point accuracy, we used the 50th percentile (median) prediction and compared it against three baselines: a person-specific historical mean model, last observation carried forward ("Last"), and an autoregressive integrated moving average model (ARIMA). Performance was quantified using mean absolute error (MAE), mean absolute percentage error (MAPE), and coefficient of determination (R²). Sensitivity Analysis. To assess how NORMA captures uncertainty across clinical contexts, we conducted one-at-a-time sensitivity sweeps over five input features: patient age, sex, number of prior measurements, prediction horizon, and within-sequence variability. For each of 30 laboratory tests, we constructed a synthetic baseline sequence: 10 measurements spaced 90 days apart, each set to the midpoint of the sex-specific population reference interval, for a 50-year-old male with a 30-day prediction horizon. We then varied one feature at a time while holding all others at baseline values. Age was swept from 20 to 80 years in 5-year increments; history length from 2 to 300 measurements; and prediction horizon from 7 days to 10 years. For within-sequence variability, we parameterized the standard deviation as a multiplier of one-tenth of the reference range width (0.0 to 3.0) and, at each noise level, drew 30 independent histories from N(midpoint, σ ) to capture sampling variability. For each configuration, we recorded the predicted median (50th quantile), the 90% prediction interval width, and the percent change in interval width relative to baseline, enabling direct comparison of sensitivity magnitude across biomarkers. NORMA configuration. For CHS validation, we applied NORMA with the default configuration: Time2Vec temporal encoding, ternary health-state indicators, raw-value laboratory inputs, raw age projected linearly, and the Gaussian output parameterization. For eICU validation, we used log-delta-t temporal encoding with periodic components, ternary health-state indicators, within-sequence normalization of laboratory values, binned age embeddings in decade-wide bins, a dedicated context token for patient covariates, and the quantile output parameterization. INSPIRE used the same configuration as eICU. All cohort-specific NORMA configurations are summarized in Supplementary Table 3.

Statistical Analysis Mortality association analysis. We assessed the association between laboratory values and mortality using two complementary approaches. First, we grouped index-period measurements into quintiles of raw analyte values within each analyte and estimated mortality rates per quintile. Second, to evaluate whether deviation from a patient’s personal baseline predicts mortality, we computed a deviation score for each index measurement as the absolute z-score from the patient’s baseline mean (|value −µbaseline | / σbaseline ), where

µbaseline and σbaseline are the mean and standard deviation of the patient’s baseline measurements. We grouped deviation scores into deciles and estimated mortality rates within each decile. For both analyses, we computed 95% confidence intervals using Wilson score intervals. Analytes with fewer than 20 observations were excluded.

10

Abnormality classification. We classified each index measurement as normal or abnormal under PopRI , PerRI , and NORMARI . To avoid incorporating pre-existing pathology into PerRI thresholds, we excluded patients whose PerRI mean fell outside the PopRI . Measurements classified as PopRI -abnormal were also classified as abnormal using PerRI and NORMARI . Clinical Risk Stratification and Prediction. We evaluated whether abnormal classifications under each framework identified individuals at elevated risk of future events. For each analyte and framework (PopRI , PerRI , NORMARI ), we defined a binary indicator of abnormality for values outside the corresponding interval. For each analyte–outcome pair, we calculated positive predictive value (probability of the event given abnormal classification), sensitivity (probability of abnormal classification among those with the event), and specificity (probability of normal classification among those without the event). To isolate differences between interval paradigms, we restricted analyses to measurements within the PopRI , because values outside the PopRI necessarily fall outside PerRI and NORMARI . We estimated associations using Cox proportional hazards models with a single binary abnormality indicator adjusted for age and sex. We split data into training (60 percent) and test (40 percent) sets stratified by event status and approximate event time in one-year bins. We evaluated performance at 1, 3, 5, and 10 years using time-dependent receiver operating characteristic curves and quantified uncertainty with 95% confidence intervals from 1,000 bootstrap resamples. We applied Benjamini-Hochberg false discovery rate correction to account for multiple comparisons across outcomes and analytes.

Ethics Approval The use of Clalit Health Services data in this study was approved by the Clalit Health Services Institutional Review Board (Helsinki Committee). MIMIC-IV, eICU-CRD, and INSPIRE were accessed under the PhysioNet Credentialed Health Data Use Agreement; EHRSHOT was accessed under the Stanford Research Use Agreement. All datasets were de-identified prior to release.

Data Availability This study utilizes two publicly available datasets for the development of NORMA: EHRSHOT and MIMIC-IV, both accessible to qualified researchers under their respective data-use agreements. The eICU Collaborative Research Database is also publicly available via PhysioNet. Clalit Health Services data are not publicly available.

Code Availability The full source code for the NORMA model, including data processing, training, and evaluation pipelines, is publicly available at https://github.com/aashnapshah/NORMA. The interactive web-based user interface for individualized laboratory interpretation is available at https://norma-tpy0.onrender.com/.

Acknowledgments We gratefully acknowledge support from the Ivan and Francesca Berkowitz Family Living Laboratory Collaboration at Harvard Medical School and Clalit Research Institute, NIEHS R01ES032470, and NIDDK R01DK137993.

11

References 1. Danielle B Freedman. Towards better test utilization - strategies to improve physician ordering and their impact on patient outcomes. EJIFCC, 26(1):15–30, January 2015. 2. G H Guyatt, A D Oxman, M Ali, A Willan, W McIlroy, and C Patterson. Laboratory diagnosis of irondeficiency anemia: an overview. J. Gen. Intern. Med., 7(2):145–153, March 1992. 3. P Newsome, R Cramb, S Davison, J Dillon, M Foulerton, E Godfrey, Richard Hall, Ulrike Harrower, M Hudson, A Langford, A Mackie, R Mitchell-Thain, K Sennett, N Sheron, J Verne, Martine Walmsley, and A Yeoman. Guidelines on the management of abnormal liver blood tests. Gut, 67:6–19, November 2017. 4. Silvio E Inzucchi. Clinical practice. diagnosis of diabetes. N. Engl. J. Med., 367(6):542–550, August 2012. 5. Kenneth A Sikaris. Enhancing the clinical value of medical laboratory testing. Clin. Biochem. Rev., 38(3): 107–114, November 2017. 6. Ferruccio Ceriotti, Rolf Hinzmann, and Mauro Panteghini. Reference intervals: the way forward. Ann. Clin. Biochem., 46(Pt 1):8–17, January 2009. 7. Richard C Friedberg, Rhona Souers, Elizabeth A Wagar, Ana K Stankovic, Paul N Valenstein, and College of American Pathologists. The origin of reference intervals: A college of american pathologists Q-probes study of “normal ranges” used in 163 clinical laboratories. Arch. Pathol. Lab. Med., 131(3):348–357, March 2007. 8. Joris L J M Müskens, Rudolf Bertijn Kool, Simone A van Dulmen, and Gert P Westert. Overuse of diagnostic testing in healthcare: a systematic review. BMJ Qual. Saf., 31(1):54–63, January 2022. 9. Brian A Ference, Henry N Ginsberg, Ian Graham, Kausik K Ray, Chris J Packard, Eric Bruckert, Robert A Hegele, Ronald M Krauss, Frederick J Raal, Heribert Schunkert, Gerald F Watts, Jan Borén, Sergio Fazio, Jay D Horton, Luis Masana, Stephen J Nicholls, Børge G Nordestgaard, Bart van de Sluis, Marja-Riitta Taskinen, Lale Tokgözoglu, Ulf Landmesser, Ulrich Laufs, Olov Wiklund, Jane K Stock, M John Chapman, and Alberico L Catapano. Low-density lipoproteins cause atherosclerotic cardiovascular disease. 1. evidence from genetic, epidemiologic, and clinical studies. a consensus statement from the european atherosclerosis society consensus panel. Eur. Heart J., 38(32):2459–2472, August 2017. 10. Thore Buergel, Jakob Steinfeldt, Greg Ruyoga, Maik Pietzner, Daniele Bizzarri, Dina Vojinovic, Julius Upmeier Zu Belzen, Lukas Loock, Paul Kittner, Lara Christmann, Noah Hollmann, Henrik Strangalies, Jana M Braunger, Benjamin Wild, Scott T Chiesa, Joachim Spranger, Fabian Klostermann, Erik B van den Akker, Stella Trompet, Simon P Mooijaart, Naveed Sattar, J Wouter Jukema, Birgit Lavrijssen, Maryam Kavousi, Mohsen Ghanbari, Mohammad A Ikram, Eline Slagboom, Mika Kivimaki, Claudia Langenberg, John Deanfield, Roland Eils, and Ulf Landmesser. Metabolomic profiles predict individual multidisease outcomes. Nat. Med., 28(11):2309–2320, November 2022. 11. Edoardo G Giannini, Roberto Testa, and Vincenzo Savarino. Liver enzyme alteration: a guide for clinicians. CMAJ, 172(3):367–379, February 2005. 12. M C Walters and H T Abelson. Interpretation of the complete blood count. Pediatr. Clin. North Am., 43(3): 599–622, June 1996. 13. James C Boyd. Defining laboratory reference values and decision limits: populations, intervals, and interpretations. Asian J. Androl., 12(1):83–90, January 2010. 14. Arjun K Manrai, Chirag J Patel, and John P A Ioannidis. In the era of precision medicine and big data, who is normal? JAMA, 319(19):1981–1982, May 2018. 15. D J Nazir, R S Roberts, S A Hill, and M J McQueen. Monthly intra-individual variation in lipids over a

12

1-year period in 22 normal subjects. Clin. Biochem., 32(5):381–389, July 1999. 16. Brody H Foy, Rachel Petherbridge, Maxwell T Roth, Cindy Zhang, Daniel C De Souza, Christopher Mow, Hasmukh R Patel, Chhaya H Patel, Samantha N Ho, Evie Lam, Camille E Powe, Robert P Hasserjian, Konrad J Karczewski, Veronica Tozzo, and John M Higgins. Haematological setpoints are a stable and patient-specific deep phenotype. Nature, 637(8045):430–438, January 2025. 17. Shuo Wang, Min Zhao, Zihan Su, and Runqing Mu. Annual biological variation and personalized reference intervals of clinical chemistry and hematology analytes. Clin. Chem. Lab. Med., 60(4):606–617, March 2022. 18. Abdurrahman Coskun, Sverre Sandberg, Ibrahim Unsal, Fulya G Yavuz, Coskun Cavusoglu, Mustafa Serteser, Meltem Kilercik, and Aasne K Aarsand. Personalized reference intervals - statistical approaches and considerations. Clin. Chem. Lab. Med., 60(4):629–635, March 2022. 19. Abdurrahman Coşkun, Sverre Sandberg, Ibrahim Unsal, Coskun Cavusoglu, Mustafa Serteser, Meltem Kilercik, and Aasne K Aarsand. Personalized reference intervals in laboratory medicine: A new model based on within-subject biological variation. Clin. Chem., 67(2):374–384, January 2021. 20. A E Obstfeld, K Patel, J C Boyd, J Drees, D T Holmes, J P Ioannidis, and A K Manrai. Data mining approaches to reference interval studies. Clinical Chemistry, 67(9):1175–1181, 2021. 21. Mary E Lacy, Gregory A Wellenius, Anne E Sumner, Adolfo Correa, Mercedes R Carnethon, Robert I Liem, James G Wilson, David B Sacks, David R Jacobs, Jr, April P Carson, Xi Luo, Annie Gjelsvik, Alexander P Reiner, Rakhi P Naik, Simin Liu, Solomon K Musani, Charles B Eaton, and Wen-Chih Wu. Association of sickle cell trait with hemoglobin A1c in african americans. JAMA, 317(5):507–515, February 2017. 22. Kenneth R Feingold. Guidelines for the management of high blood cholesterol. In Endotext [Internet]. MDText. com, Inc., 2025. 23. O Yaw Addo, Emma X Yu, Anne M Williams, Melissa Fox Young, Andrea J Sharma, Zuguo Mei, Nicholas J Kassebaum, Maria Elena D Jefferds, and Parminder S Suchdev. Evaluation of hemoglobin cutoff levels to define anemia among healthy individuals. JAMA Netw. Open, 4(8):e2119123, August 2021. 24. D H Rushton, R Dover, A W Sainsbury, M J Norris, J J Gilkes, and I D Ramsay. Why should women have lower reference limits for haemoglobin and ferritin concentrations than men? BMJ, 322(7298):1355–1357, June 2001. 25. James A Diao, Yixuan He, Rohan Khazanchi, Max Jordan Nguemeni Tiako, Jonathan I Witonsky, Emma Pierson, Pranav Rajpurkar, Jennifer R Elhawary, Luke Melas-Kyriazi, Albert Yen, Alicia R Martin, Sean Levy, Chirag J Patel, Maha Farhat, Luisa N Borrell, Michael H Cho, Edwin K Silverman, Esteban G Burchard, and Arjun K Manrai. Implications of race adjustment in lung-function equations. N. Engl. J. Med., 390(22):2083–2097, June 2024. 26. Darshali A Vyas, Leo G Eisenstein, and David S Jones. Hidden in plain sight—reconsidering the use of race correction in clinical algorithms. New England Journal of Medicine, 383(9):874–882, 2020. 27. Aashna P Shah, James A Diao, Emma Pierson, Chirag J Patel, and Arjun K Manrai. Disentangling proxies of demographic adjustments in clinical equations. arXiv [q-bio.QM], November 2025. 28. Nabihah Tayob and Ziding Feng. Personalized statistical learning algorithms to improve the early detection of cancer using longitudinal biomarkers. Cancer Biomark., 33(2):199–210, 2022. 29. Isaac S Kohane, Daniel R Masys, and Russ B Altman. The incidentalome: a threat to genomic medicine. JAMA, 296(2):212–215, July 2006. 30. Christina Koch, Katherine Roberts, Christopher Petruccelli, and Daniel J Morgan. The frequency of unnecessary testing in hospitalized patients. Am. J. Med., 131(5):500–503, May 2018. 13

31. Henrik L Jørgensen and Bent S Lind. Blood tests - too much of a good thing. Scand. J. Prim. Health Care, 40(2):165–166, June 2022. 32. Christopher Naugler and Irene Ma. More than half of abnormal results from laboratory tests ordered by family physicians could be false-positive. Can. Fam. Physician, 64(3):202–203, March 2018. 33. Tony Badrick, Joe M El-Khoury, and Elvar Theodorsson. Laboratory reference intervals - history and modern approaches for improved utility. Scand. J. Clin. Lab. Invest., 85(4):229–241, June 2025. 34. Davood Roshan, John Ferguson, Charles R Pedlar, Andrew Simpkin, William Wyns, Frank Sullivan, and John Newell. A comparison of methods to generate adaptive reference ranges in longitudinal monitoring. PLoS One, 16(2):e0247338, February 2021. 35. Alvin Rajkomar, Eyal Oren, Kai Chen, Andrew M Dai, Nissan Hajaj, Michaela Hardt, Peter J Liu, Xiaobing Liu, Jake Marcus, Mimi Sun, Patrik Sundberg, Hector Yee, Kun Zhang, Yi Zhang, Gerardo Flores, Gavin E Duggan, Jamie Irvine, Quoc Le, Kurt Litsch, Alexander Mossin, Justin Tansuwan, De Wang, James Wexler, Jimbo Wilson, Dana Ludwig, Samuel L Volchenboum, Katherine Chou, Michael Pearson, Srinivasan Madabushi, Nigam H Shah, Atul J Butte, Michael D Howell, Claire Cui, Greg S Corrado, and Jeffrey Dean. Scalable and accurate deep learning with electronic health records. NPJ Digit. Med., 1(1):18, May 2018. 36. Matthew B A McDermott, Bret Nestor, Peniel Argaw, and Isaac Kohane. Event stream GPT: A data preprocessing and modeling library for generative, pre-trained transformers over continuous-time sequences of complex events. arXiv [cs.LG], June 2023. 37. Zhichao Yang, Avijit Mitra, Weisong Liu, Dan Berlowitz, and Hong Yu. TransformEHR: transformer-based encoder-decoder generative model to enhance prediction of disease outcomes using electronic health records. Nat. Commun., 14(1):7857, November 2023. 38. Lavender Yao Jiang, Xujin Chris Liu, Nima Pour Nejatian, Mustafa Nasir-Moin, Duo Wang, Anas Abidin, Kevin Eaton, Howard Antony Riina, Ilya Laufer, Paawan Punjabi, Madeline Miceli, Nora C Kim, Cordelia Orillac, Zane Schnurman, Christopher Livia, Hannah Weiss, David Kurland, Sean Neifert, Yosef Dastagirzada, Douglas Kondziolka, Alexander T M Cheung, Grace Yang, Ming Cao, Mona Flores, Anthony B Costa, Yindalon Aphinyanaphongs, Kyunghyun Cho, and Eric Karl Oermann. Health system-scale language models are all-purpose prediction engines. Nature, 619(7969):357–362, July 2023. 39. Artem Shmatko, Alexander Wolfgang Jung, Kumar Gaurav, Søren Brunak, Laust Hvas Mortensen, Ewan Birney, Tom Fitzgerald, and Moritz Gerstung. Learning the natural history of human disease with generative transformers. Nature, 647(8088):248–256, November 2025. 40. Shane Waxler, Paul Blazek, Davis White, Daniel Sneider, Kevin Chung, Mani Nagarathnam, Patrick Williams, Hank Voeller, Karen Wong, Matthew Swanhorst, Sheng Zhang, Naoto Usuyama, Cliff Wong, Tristan Naumann, Hoifung Poon, Andrew Loza, Daniella Meeker, Seth Hain, and Rahul Shah. Generative medical event models improve with scale. arXiv [cs.LG], November 2025. 41. Pawel Renc, Yugang Jia, Anthony E Samir, Jaroslaw Was, Quanzheng Li, David W Bates, and Arkadiusz Sitek. Zero shot health trajectory prediction using transformer. NPJ Digit. Med., 7(1):256, September 2024. 42. Andrew Zhang, Tong Ding, Sophia J Wagner, Caiwei Tian, Ming Y Lu, Rowland Pettit, Joshua E Lewis, Alexandre Misrahi, Dandan Mo, Long Phi Le, and Faisal Mahmood. A multimodal and temporal foundation model for virtual patient representations at healthcare system scale. arXiv [cs.LG], April 2026. 43. Guy Lutsker, Gal Sapir, Smadar Shilo, Jordi Merino, Anastasia Godneva, Jerry R Greenfield, Dorit Samocha-Bonet, Raja Dhir, Francisco Gude, Shie Mannor, Eli Meirom, Eric P Xing, Gal Chechik, Hagai Rossman, and Eran Segal. A foundation model for continuous glucose monitoring data. Nature, 650

14

(8103):978–986, February 2026. 44. Ahmed A Metwally, A Ali Heydari, Daniel McDuff, Alexandru Solot, Zeinab Esmaeilpour, Anthony Z Faranesh, Menglian Zhou, Girish Narayanswamy, Maxwell A Xu, Xin Liu, Yuzhe Yang, David B Savage, Mark Malhotra, Conor Heneghan, Shwetak Patel, Cathy Speed, and Javier L Prieto. Insulin resistance prediction from wearables and routine blood biomarkers. Nature, March 2026. 45. Valentyn Melnychuk, Dennis Frauen, and Stefan Feuerriegel. Causal transformer for estimating counterfactual outcomes. arXiv [cs.LG], April 2022. 46. Michelle M Li, Kevin Li, Yasha Ektefaie, Ying Jin, Yepeng Huang, Shvat Messica, Tianxi Cai, and Marinka Zitnik. Controllable sequence editing for biological and clinical trajectories. arXiv [cs.LG], February 2025. 47. Shakson Isaac, Yentl Collin, and Chirag Patel. SSM-CGM: Interpretable state-space forecasting model of continuous glucose monitoring for personalized diabetes management. arXiv [cs.LG], October 2025. 48. Martin W McIntosh, Nicole Urban, and Beth Karlan. Generating longitudinal screening algorithms using novel biomarkers for disease. Cancer Epidemiol. Biomarkers Prev., 11(2):159–166, February 2002. 49. Ruth Johnson, Uri Gottlieb, Galit Shaham, Lihi Eisen, Jacob Waxman, Stav Devons-Sberro, Curtis R Ginder, Peter Hong, Raheel Sayeed, Xiaorui Su, Ben Y Reis, Ran D Balicer, Noa Dagan, and Marinka Zitnik. ClinVec: Unified embeddings of clinical codes enable knowledge-grounded AI in medicine. medRxiv, May 2025. 50. Ran D Balicer, Efrat Shadmi, Nicky Lieberman, Sari Greenberg-Dotan, Margalit Goldfracht, Liora Jana, Arnon D Cohen, Sigal Regev-Rosenberg, and Orit Jacobson. Reducing health disparities: strategy planning and implementation in israel’s largest health care organization. Health Serv. Res., 46(4): 1281–1299, August 2011. 51. Tom J Pollard, Alistair E W Johnson, Jesse D Raffa, Leo A Celi, Roger G Mark, and Omar Badawi. The eICU collaborative research database, a freely available multi-center database for critical care research. Sci. Data, 5(1):180178, September 2018. 52. Leerang Lim, Hyeonhoon Lee, Chul-Woo Jung, Dayeon Sim, Xavier Borrat, Tom J Pollard, Leo A Celi, Roger G Mark, Simon T Vistisen, and Hyung-Chul Lee. INSPIRE, a publicly available research dataset for perioperative medicine. Sci. Data, 11(1):655, June 2024. 53. Clement J McDonald, Stanley M Huff, Jeffrey G Suico, Gilbert Hill, Dennis Leavelle, Raymond Aller, Arden Forrey, Kathy Mercer, Georges DeMoor, John Hook, Warren Williams, James Case, and Pat Maloney. LOINC, a universal standard for identifying laboratory observations: a 5-year update. Clin. Chem., 49(4): 624–633, April 2003. 54. American Board of Internal Medicine. ABIM laboratory test reference ranges. Technical report, January 2025. 55. Alistair Johnson, Lucas Bulgarelli, Tom Pollard, Brian Gow, Benjamin Moody, Steven Horng, Leo Anthony Celi, and Roger Mark. MIMIC-IV, October 2024. 56. Michael Wornow, Rahul Thapa, Ethan Steinberg, Jason A Fries, and Nigam H Shah. EHRSHOT: An EHR benchmark for few-shot evaluation of foundation models. arXiv [cs.LG], July 2023.

15

Figures a Contextual Sequence Modeling

b Training and Validation Cohorts

Labs Time

Sex

NORMA

TRAINING

Autoregressive Transformer

PopRI State

Age

VALIDATION

EHRSHOT

MIMIC-IV

CLALIT

eICU

INSPIRE

5.7k patients 122k sequences

180k patients 3.3M sequences

1.4M patients 32M sequences

98k patients 1M sequences

51K patients 550k sequences

Abn

Predicts distribution for next “healthy” lab

>20 years | 37M sequences | 30 labs

c Reference Interval Paradigms

PopRI

PerRI

NORMARI

Norm

Complete Blood Count HGB (g/dL)

HGB (g/dL)

Population

Hepatic Function

Metabolic Function

Lipid Panel

HGB (g/dL)

Individualized

Contextualized

✗ Personalized

✗ Robust to noise

✓ Personalized

✗ Repeat Labs

✗ Time-Sensitive

✓ Population-Aware

✗ Handles sparsity

✓ Time-Sensitive

d Clinical Utility of NORMARI Glucose → Type 2 Diabetes

✓ Models Uncertainty

PopRI

PerRI

NORMARI

Normal

High

High

e Early Risk Stratification + 12%

+ 12%

NORMARI

Early shift detected by

NORMARI

% Change

Glucose

PopRI

PerRI

+ 4%

- 1%

Specificity

Time

Sensitivity

Precision

NPV

Fig. 1 | NORMA framework overview. a) NORMA conditions on laboratory history, time, sex, age, and population reference state to predict the distribution of the next healthy laboratory value. b) Study design: training on EHRSHOT and MIMIC-IV; external validation on Clalit Health Services, the eICU Collaborative Research Database, and the INSPIRE cohort. Across all cohorts: 20 years, 37 million sequences, and 30 analytes spanning four clinical panels. c) Comparison of PopRI , PerRI , and NORMARI . d) Illustrative example: a glucose value classified as normal by PopRI but flagged abnormal by PerRI and NORMARI , with corresponding change in specificity, sensitivity, and precision for type 2 diabetes classification. e) Schematic of early abnormality detection by NORMARI relative to PopRI and PerRI over time. Created with BioRender.com.

16

a Cohort Selection

c Mortality Association eICU-CRD

BASELINE

2005

2010

2015

2020

2025

Inclusion: ≥5 outpatient labs, 90 days apart; index in 2015

FIRST 75%

Day 1

Day 3

Day 5

Day 7

Day 10

Day 14

CO2

Inclusion: first 75% of measurements as baseline; next as index

Low

Normal

High

Inpatient

Intra

CHS

CV (%)

40 30

INSPIRE

LDL

CV (%)

10 0

Individuality Index

1.00

eICU-CRD

CHS

INSPIRE

0.75 0.50 0.25 0.00

40 35 30 25 20 15 10 5 0 Q1

Q2

Q3

Q4

Q5

30 27 24 21 18 15 12 9 6 3 Q1

DBIL

32 28 24 20 16 12 8 4 45 40 35 30 25 20 15 10 36 33 30 27 24 21 18 15 12 9

HGB

MCHC

30 27 24 21 18 15 12 9

GLU

MCV

RBC 27 24 21 18 15 12 9 6

TC 30 28 26 24 22

CL

K

PLT

TBIL

ALT

40 36 32 28 24 20 16 12 8

40 35 30 25 20 15 10 5 0

32 28 24 20 16 12 8 4

TP

26 24 22 20 18 16 14 12 10

NA

50 45 40 35 30 25 20 15 10

RDW

30 27 24 21 18 15 12 9 6

20

45 40 35 30 25 20 15 10 5

MCH

28 26 24 22 20 18 16 14 12 10

MPV

40 30

33 30 27 24 21 18 15 12

HDL

32 30 28 26 24 22 20 24 22 20 18 16 14 12 10 8

0 50

CRE

25.8 25.6 25.4 25.2 25.0 24.8 24.6 24.4

54 48 42 36 30 24 18 12 6

20 10

CA 36 32 28 24 20 16 12 8

HCT

Mortality Rate (%)

CV (%)

Inter

60 eICU-CRD 50 40 30 20 10 0

ALP

BUN

40 36 32 28 24 20 16 12 8

60 54 48 42 36 30 24 18 12

Outpatient

b Intra-Patient vs Inter-Patient Variation

50

INSPIRE

28 24 20 16 12 8 4

45 40 35 30 25 20 15 10 5

27 24 21 18 15 12 9 6 3

eICU & INSPIRE

ALB

48 42 36 30 24 18 12 6 0

AST

REMAINING

INDEX

A1C

36 32 28 24 20 16 12 8 4 0

CHS

2000

CHS

MONITORING

INDEX

TGL 27 24 21 18 15 12 9

WBC

Q2

Q3

Q4

Q5

NA K MCHC CL MPV CA LDL TP TC HDL RDW AST ALT CO2 HCT TBIL HGB ALP PLT RBC GLU BUN MCV A1C MCH ALB CRE TGL WBC

Quintile

Fig. 2 | Within-person biological variability and mortality association. a) Cohort Selection. Cohort selection and inclusion criteria for Clalit Health Services, the eICU Collaborative Research Database, and the INSPIRE cohort. b) IntraPatient vs Inter-Patient Variation. Within-person versus between-person coefficient of variation for 30 analytes, with the individuality index shown below. The dashed line indicates an individuality index of 0.6. c) Mortality Association. Observed mortality rate by quintile of z-score from the patient baseline mean in Clalit Health Services, the eICU Collaborative Research Database, and the INSPIRE cohort. Shaded bands denote 95% confidence intervals.

17

a NORMA Architecture

b Forecasting Performance

INPUTS

Context

History

Query

EMBEDDINGS

Context

History

h1

h2

NORMA

NORMA

ARIMA

ARIMA

ARIMA

Mean

Mean

Mean

Query

TOKEN SEQUENCE

CTX

NORMA

hn

Q

Last

Last

Last

DECODER

0

2

4

6

Masked Self-Attention

8

10

MAE

12

14

0

5

10

15

20

MAPE

25

0.0

0.1

0.2

0.3

0.4

0.5

R2

0.6

0.7

c Per-Analyte Performance

Feed-Forward Network × N l ay e r s

1.0

NORMA ARIMA Mean Last

OUTPUT

0.8

Quantile Head

0.6 R²

Query Token Output

OUTPUT QUANTILES

0.4

q97.5

2.0

TP

1.5

WBC

TC

TGL

TBIL

RBC

RDW

NA

PLT

MPV

MCV

MCH

MCHC

K

LDL

HGB

HCT

HDL

GLU

CRE

DBIL

CL

CO2

CA

A1C

0.0

BUN

0.2 ALT

q75

AST

q50

Reference interval: [q2.5, q97.5]

ALP

q25

ALB

q2.5

d Sensitivity Analysis HFP

12

120

17.5

10

100

12.5 10.0 7.5

8

CI Width (% of ref range)

15.0

6 4

5.0 2

2.5 0.0

101

102

Number of Measurements

0

Lipid

2

Deviation from Midpoint (%)

BMP

20.0

CI Width (% of ref range)

CI Width (% of ref range)

CBC

80 60 40

0 2 4

20 101

102

0 0.0

103

Prediction Horizon (days)

6 0.5

1.0

1.5

2.0

Within-Person SD

2.5

3.0

0.0

0.5

1.0

Within-Person SD

2.5

3.0

Fig. 3 | NORMA architecture, forecasting accuracy, and sensitivity analysis. a) NORMA Architecture. Model architecture: demographic context, laboratory history, and query tokens are jointly embedded and processed by a causal transformer decoder to produce quantile predictions for the next laboratory value. b) Forecasting Performance. Held-out test-set forecasting accuracy (mean absolute error, mean absolute percentage error, and coefficient of determination) for NORMA (quantile) versus ARIMA, population mean, and last-value-carried-forward baselines. Error bars denote 95% bootstrap confidence intervals. c) Per-Analyte Performance. Per-analyte coefficient of determination grouped by clinical panel. d) Sensitivity Analysis. Sensitivity of prediction interval width and midpoint deviation (each as a percentage of PopRI width) to history length, forecast horizon, and within-person variability.

18

a Reclassification Rate

b Hazard Ratios NORMARI

PerRI

PopRI

80 Reclassification Rate (%)

PerRI

NORMARI

ALB HGB

60

PLT HCT

40

MCV RDW

20

WBC CRE A1C AL B AL P AL AST BU T N CA C COL CR2 GL E HCU HDT HGL B K LD MC L MC H H MCC MPV V NA PL RB T RD C W TB IL TC TG L T WB P C

0

CL MCH

2

4

Hazard Ratio

6

8

c True Positive Rate vs False Positive Rate NORMARI

PerRI

Anemia

0.8

GLU

True Positive Rate

0.6

HGB TCLDL HDL BUN

A1C

0.2

MCH

TP WBC CL ALP AST K ALB HCT ALT RDW CA RBC PLT MCHC NA HGB LDL TC HDL BUN

CRE

0.6 GLU

0.4

A1C

0.8

CL

0.6

K ALP AST ALB ALT PLT HCT CRE CAMCHC RDW NA RBC HGB

0.2

0.4 0.6 0.8 False Positive Rate

0.2

0.4 0.6 0.8 False Positive Rate

0.2

0.4 0.6 0.8 False Positive Rate

CA

BUN

NA TC

MCHC

RBC

N BU

T AS

NA

ALB

MCV

HDL

HCT

MCH

L HD

V

LDL

CRE

AST

MCHC

HDL

HCT

MC V MC

MCH

MCH

BU

N

TP

TC

LDL

MCHC

TP

A1C MCV

E CR

TP

HDL

TC

TP

RBC

TC

A1C K ALB

K

MC

HC

A1C

W

TP

RBC

TC

A1C K ALB

K A1C

ALB

MC

V

NA

ALB

MCHC

NA

TP

RDW

LDL

MPV

RD ALB

ALB

RB ALT

AST

C RB ALT

E

CA

C

TP H MC

NA

MCHC

CRE

C

NA

C

ALB

MCH

C RB AST

ALT

A1C NA N BU

CR

CL

TBIL

ALT

A1C

AST

MCH

GLU

CA

MCHC

TP

GLU

MCHC

LDL

RBC

TP NA

L HD

HDL

AST

CA

NA

PLT

L HD ALT

MCHC

T AS HDL

WBC

H MC

CRE

U GL

N

TP

RDW

RBC

Anemia

AST

HDL

TC

ALT

BU

CL

BUN

LDL

TC

ALP

WBC

C

MCHC

ALT

N

MCHC

WBC

CL +18.5%

AST

H

TBIL

RBC

PLT

V

TC

NA

MCH

A1C

BU

MC

ALB

W

HGB

MC

GLU

TP

MCV

H

MC

ALB

HDL

ALT

RBC

GLU

CA

H

RBC

T

AS

PLT MPV

NA

CRE MPV

RD

K

TBIL

CA

LDL

CA

MC MCHC

RDW

CRE

HDL

PLT +54.5%

WBC

PLT CL

C

CL

AST HCT

WB

NA

A1C ALT

PLT

MCH

U GL

ALP

LDL

RBC

CRE +15.2%

A1C

MCV

MCH

TP

ALP

HDL

CRE

W

K

V

HCT

HCT

ALP

MPV

MC

HGB

ALB

Chronic Kidney Disease

TP

Mortality

C

P

Type 2 Diabetes

1.0

MPV

TBIL MPV

RD

H

LDL

AL

TBIL

C

ALP

WB

TC

MC MCHC

RDW

AST

CRE

N

H

CL

BU

MC

LDL

C

RDW

AST

ALB

ALT

HDL

0.4 0.6 0.8 False Positive Rate

TC CL

B

PLT

WB

ALT

LDL

HG

GLU

HCT

TC BUN

TBIL

PLT

ALP

0.2

ALP NA

CRE +22.6%

HGB

NA

PLT +50.4%

CA

GLU

W

T

RD

MPV

TBIL +22.8%

CL HGB

MCV

0.0 0.0

1.0

HGB

PLT

C WB

BUN

HGB

H

MC

C

CA

A1C

RDW

W

GLU

V

ALT

L

HC

GLU

TBIL MPV

MC

CL

LD

A1C

T

GLU +10.1%

WBC

WB

RBC

HDL PLT

ALP

K

TC

HC

CRE

PLT MPV

MPV

RD

K

TBIL

C WB

CRE

TBIL

HCT

AST

GLU

HCT

TC TBIL

HGB

ALT

LDL

MCV

HGB

PLT

NA

ALT

MPV

TBIL +23.3%

A1C

A1C

K CA

MCH

LDL

CL

W

TP

WBC

W

AST

LDL

RD

MPV

RD

CA

CL

CRE +12.6%

ALB

N

PLT +51.7%

H

K

BU

HGB

RBC

MCH

V

NA

BUN

HCT

TBIL

HGB

MC

ALP

MC

CRE

MPV

ALT

U

K

N

CRE

CA

HCT

GL

V

CA

BU

CL

CRE +11.4%

ALP

PLT

ALP

MC ALT

K

MCH

ALB

A1C

RBC

C

TBIL

HDL

HGB

PLT

WB

ALB

A1C

W

TC

T

MPV

RD

BUN

HC

B HG

MPV

TBIL +21.2%

WBC

PLT

TBIL

WBC

TC

K

CRE

TBIL

HCT

TP

B

GLU

PLT +48.6%

CA

V

0.6 0.4

TBIL PLT

HG

BUN

CL

MC

CA

L

HD

HGB

U

RDW

NA

CRE +23.2%

CL

ALT

GL

T

TC

TP K

HCT

PLT

ALP

HC HGB

MCV

LDL

BUN

PLT

ALP

TBIL

HDL

HGB

CL

RBC

MPV

TBIL +20.8%

CRE

W

ALB

ALP

A1C

MCH

AST ALP ALT HCT ALB K RBC MCHC CAHGB NA LDL PLT CRE HDL TC BUN GLU

Accuracy GLU

RD

MCH

HCT

LDL

AST

MPV

MC

C

TPWBC RDW

CL

TBIL

0.2

0.0 0.0

1.0

P

A1C

TBIL

WBC

AL

A1C +9.3%

LDL

AST

P

CL

V

CA

AL RBC

WB

0.8

A1C

Specificity CRE

K

K

CRE

ALB

RDW

CL

BUN

TC

U

GL

MPV

MCV

e Lead Time

Sensitivity

TP

WBC

MCH

HDL BUNTCLDL

d Performance Improvement with NORMA Precision

TP

0.2

0.0 0.0

1.0

GLU

0.4

0.2

0.0 0.0

Type 2 Diabetes

1.0

MPV MCV TBIL

MCV TBIL

CL K ALP AST ALT MCHCALB RDW PLTCA RBC HCT CRENA

0.4

Mortality

1.0

MPV

TBIL

True Positive Rate

TPWBC MCH

0.8 True Positive Rate

1.0

MPV MCV

True Positive Rate

1.0

Chronic Kidney Disease

CA GLU CL K CO2

0

20 40 Lead Time (months)

60

Fig. 4 | Reclassification performance and early risk stratification in Clalit Health Services. a) Reclassification Rate. Reclassification rate by analyte among PopRI -normal tests. b) Hazard Ratios. Cox hazard ratios for all-cause mortality (adjusted for age and sex); analytes with p < 0.05 shown. c) True Positive Rate vs False Positive Rate. True positive rate versus false positive rate for clinical outcome prediction across anemia, chronic kidney disease, all-cause mortality, and type 2 diabetes, comparing PerRI and NORMARI . d) Performance Improvement with NORMA. Change in precision, sensitivity, specificity, and balanced accuracy (NORMARI versus PerRI ); inner label shows the largest gain. e) Lead Time. Median lead time from NORMARI flag to first PopRI flag, by analyte (months). Error bars denote the interquartile range.

19

a Reclassification Rate

b Hazard Ratios NORMARI

PerRI

1.5 Hazard Ratio

2

WBC CRE

60

RDW TGL

40

CO2

20

CA TP

AL B AL P AL AST BU T N CA C COL CR2 DB E I GL L HCU HDT HGL B K LD MC L MC H H MCC MPV V NA PL RB T RD C W TB IL TC TG L T WB P C

Reclassification Rate (%)

NORMARI

PLT

80

0

PerRI

PopRI

100

BUN CL

0.5

1

2.5

c True Positive Rate vs False Positive Rate NORMARI

PerRI

Acute Kidney Injury

1.0 MPV

0.4

PLT MCH MCV KALPALT CA TPRDW CO2 MCHC ALB RBC TGL CL AST HGB TBIL WBC HCT BUN DBIL GLU CRE

0.2 0.0 0.0

0.2

0.4 0.6 0.8 False Positive Rate

1.0

NA

0.6 0.4

PLT MCV MCH

0.2

ALP K ALT HGB ALB MCHC HCT TGL RDW CA AST TBIL CO2 TP CL RBC BUN CRE WBC GLU

0.0 0.0

DBIL

0.2

0.8 NA

0.6 PLT

0.4 MCH MCV K CA ALT CO2 ALP RDW TPMCHC ALB GLURBC CL BUN ASTTGL HGB TBIL WBC HCT CRE DBIL

0.2 0.4 0.6 0.8 False Positive Rate

0.0 0.0

1.0

0.2

0.4 0.6 0.8 False Positive Rate

ALB

ALT CL

CL MC

V

MCH

TP

ALP

MC

BUN NA

MCV

HC

PLT

CA

T

MPV

TP

CO2

WB C

HC T

NA

ALB TGL

RBC

DB

P

IL

CA

RDW

AL

TP

HCT

BUN

MC

HCT

MC V

BUN

HC

HCT

BUN

V MC

BUN

V MC

HCT

GLU GLU

CA

CRE

2 CO

2 CO

GLU

MCV

RBC

CA

GLU ALB

DBIL

K

RDW

CA

DBIL

V

GLU

DBIL

CA

HGB

CRE

CO2

PLT

BUN WB C

BUN

MC H

CA

GLU

CA

CRE

GLU

ALT

MPV

HCT

CL

TBIL

RDW

TP

PLT

ALB

MCV

PLT

RBC

HGB

HCT

RBC

MCV

MCHC

N

ALB +4.8% CRE

RBC

CO2

ALP

B

K

ALB

PLT

CA

TP 2 CO ALP

PLT

TGL

HGB

BUN

MPV

Acute Kidney Injury Mortality

NA MC HC

ALT

RDW

RBC

CRE

HG

AST

MPV

IL

CRE CO2

BU

MCV MCHC

MCH

NA RDW

DBIL

T

ALP

CO2

TB IL

H

MC

BUN

DBIL

TGL

HCT

RBC

CRE

WB C

GLU

CO2 +8.9%

ALB

BUN

MCV

AST

BUN

TBIL

AST

ALT

CRE

TBIL

V

CL

CA

AST

ALB

T

PL

PL

LDL TP

MP

Prolonged LOS (>7d)

TBIL

CRE

ALP

ALT

NA DB

WBC

MCHC

ALT

K

K

ALP

WBC

CL

B

L

ALP +41.9%

RBC

NA

AL

MPV

MPV

RBC

AST HGB

K

GLU RDW

TP

K

TG

TBIL

WBC

RBC

CO2

TP

MCH

ALB

RDW

TGL

ALP T

ALB

C

NA

AST

MCH

IL

DB

GLU +5.8%

IL

PLT

MCHC

MCH

ALT

DB

HGB

NA

MCH

CRE

CL

AL

C

WB

CO2

U

C

MCH

K

MPV

HGB

NA IL

ALB

BUN

ALP

ALT TBIL

CO2

L

ALP +43.0% TB

MPV +35.3%

TGL

MCV

AST

L

TG

WBC

HCT

MCV

HGB

TP

AST

CR E

W

RD

TBIL

TGL

RBC

MPV

CRE

PLT

CL

MCH

TG

MPV

HGB

T

AST

ALT

T

TP MCH

GLU

MPV TGL

GL

C

ALB

MCHC

AL

K

AST

CRE

K

RB

CL

GLU +5.0%

HGB

NA

ALP

H MC

CL

M

TP

HCT CO2

MCHC

MPV +35.0%

CHC

TBIL

DBIL

CA

NA

WBC

HGB

CO2

1.0

PLT ALP

HC

T

AS

IL

RD W

U

GL

0.4 0.6 0.8 False Positive Rate

TGL

CO2

DBIL

NA

DB

K

PLT

TP

MP V

W

RD

RDW

HGB

RBC

CO2 TGL

DBIL

BUN

CA

HCT

CL

CO2 +10.6%

WBC

TBIL

TGL

CO2 +4.1%

T

ALP +41.8%

AST

ALT

CA

ALP

CR E PLT

MPV

AST

T

TBIL

CL

TP

AL

C

WB

GLU

MCHC

MCH

TGL

HGB

K

NA

MPV

MPV

AL

HGB

ALP

RDW

ALB

MPV

MPV +32.1%

RDW

MCHC

PLT

TP

WBC

CL

GLU

PLT

NA

BUN

TP

MCH CA

ALT

RD W

C

WB

TGL

BUN

ALB

DBIL

ALB

BU

GLU

IL

CL

IL

MCV

ALP

DB

K

DB

MCH

TGL

NA

TBIL

V

MCHC

RBC

GLU +3.8%

K

TP

ALP +41.9%

MC

K

IL

HCT

AST

CA

TB

V MC

ALP

HC T

C

WB

C

HCT

MCH

2

HGB

MCH

AST

RDW

AST

TBIL

CRE

CL

CO

MCH

Sepsis

MCHC

HG B

T

AL

NA

MPV

ALP

DBIL

CRE

TBIL

MPV +35.7%

0.2

TC N

T

CA

ALT

C

WBC

MCH

PL

T

AS

MCH

K

CL

GLU

RDW

MCV

ALB

CO2

PLT

GLU +4.9%

CL

B

AL

WBC

RBC

RBC

IL

PLT MCH MCV CA KALPALT ALB TPMCHC RDW HGB CO2 TGL AST CL TBIL HCT DBIL WBC BUN GLU CRE RBC

Accuracy

Specificity

TP

CL

NA

CRE

TB

K

H

K

TGL

WBC

ALP

MC

0.4

0.0 0.0

1.0

NA

e Lead Time

Sensitivity

RDW

0.6

0.2

d Performance Improvement with NORMA Precision

MPV

0.8 True Positive Rate

NA

Sepsis

1.0 MPV

0.8 True Positive Rate

True Positive Rate

0.8 0.6

Prolonged LOS (>7d)

1.0 MPV

True Positive Rate

1.0

Mortality

CA HCT RBC HGB GLU

0

25

50 75 100 Lead Time (hours)

125

Fig. 5 | Reclassification performance and early risk stratification in the eICU Collaborative Research Database. a) Reclassification Rate. Reclassification rate by analyte among PopRI -normal tests. b) Hazard Ratios. Cox hazard ratios for in-hospital mortality (adjusted for age and sex); analytes with p < 0.05 shown. c) True Positive Rate vs False Positive Rate. True positive rate versus false positive rate for clinical outcome prediction across in-hospital mortality, acute kidney injury, sepsis, and prolonged ICU stay, comparing PerRI and NORMARI . d) Performance Improvement with NORMA. Change in precision, sensitivity, specificity, and balanced accuracy (NORMARI versus PerRI ); inner label shows the largest gain. e) Lead Time. Median lead time from NORMARI flag to first PopRI flag, by analyte (hours). Error bars denote the interquartile range.

20

a Reclassification Rate

b Hazard Ratios NORMARI

PerRI

60 50

PerRI

NORMARI

4

5 Hazard Ratio

6

TP

40

K

30

ALB CL

20

CA

TP WB C

B AL P AL T AS T BU N

AL

K NA PL T TB IL

WBC CL CO 2 CR E GL U HC T HG B

TBIL

0

CA

10 A1C

Reclassification Rate (%)

PopRI PLT

ALP AST

1

2

3

7

8

9

c True Positive Rate vs False Positive Rate NORMARI

PerRI

1.0

1.0

0.4

CO2

NA K

0.2

PLT GLU

0.0 0.0

0.2

0.6 0.4 GLU CO2 PLT

0.2

CREHCTALP ALT CABUN CLAST TBIL WBC HGB ALB TP

0.4 0.6 0.8 False Positive Rate

K ALT NA ASTALP HCT TBIL TP HGB WBC CRE ALB CA CL BUN

0.0 0.0

1.0

0.2

0.6 0.4 GLU CO2

0.2 KTBIL AST NA WBC CLHCT HGB BUN TP ALB

0.4 0.6 0.8 False Positive Rate

PLT

0.2

0.4 0.6 0.8 False Positive Rate

NA

K ALP

AST HGB

HCT

ALT

CL

TP

HCT

TP CRE

NA

CL

GLU +5.9%

GLU

HCT CO2

CRE TB IL

N

BU

ALB

CA

CA ALB

NA

CA C WB

NA

HCT

CA

K

CL

CO2

HCT

PLT

NA

CA ALB

CA ALB

TBIL

HGB

CA

NA

WBC

HCT

PLT

ALB

CRE

BUN

CL

HCT TBIL

TP

K K

CL

CRE

K CA

K

PLT

ALB ALT

ALP

GLU

BUN

K

K HGB

NA

ALP

PLT

ALP

GLU

AST

GLU

HGB

ALT

Mortality

CRE

ALB

BUN

PLT

TBIL

AST

E

TBIL

TBIL

CO2

ALP

TP

AST

PLT GLU

CO2 +6.3% B

CO2

HGB

BUN

ALB

AST

CRE

ALB

HG

CR

K

ALB

ALP +36.4%

HCT

T

WBC

GLU

T

HC

TP

CA

CL TBIL

C

C WB

AS

GLU

ALP

E

PLT

GLU +12.2%

ALB

NA

CL

ALT

PLT

B

CO2

CR

WB

TBIL

P

TP

CR

ALP

GLU

BUN

CO2

AST

E

ALT

AL

CL

CA

K

CR

HCT

PLT TP

BUN

CL +14.3%

HGB

TBIL

PLT

TBIL

N

AST

HG

E

B

NA

AST

CRE

WBC

U

CO2 +5.3% CA

ALP +31.5%

PLT

CA AL

T

HC

C

GLU

BU

2 CO

T

K

AST

TBIL

BUN

CO2

WB

ALT

TP

AS

B

HG

B

WBC

ALT

ALT

CA

CRE

PLT

GLU

ALP

2

CL

ALP

1.0

NA GL

N

TP

NA

ALT

BU

AST

BUN

TP

CO2

GLU +5.0%

0.4 0.6 0.8 False Positive Rate

TBIL B

K

GLU

ALP HG

T

CO2

HG

CL

CL

0.2

ALP PLT

TP

CO TBIL

HCT

AL

ALB

CRE

WBC

GLU +7.3%

E

CR

C

TBIL

ALB

K

HGB

WB

WBC

ALP +36.5%

T

GLU

BU

PLT

2 CO

K

T

ALP AS BUN

PLT

CL

HCT

N

ALB

HCT

CL

BU

GLU

0.0 0.0

CL

CO2 +11.1%

CRE

ALT

HGB

TBIL

GLU +11.8%

PLT

NA

AL

CA

NA

ALP

N

T

TP

ALP

2

AL

CO2

AST

WBC

CO

K

HCT

C

WB

TBIL

HGB

TP

WBC

BU

ALP +36.1%

AST

TP

TP +4.2%

ALB

ALB

HGB

K

C

CO2

TP

P

WB

NA

CL

CO2

AL

TP

CA

TP

AS T

CL

GLU

HC T

N

T

CO2 +13.9%

HGB

CRE

Perioperative Infection

NA

U

GLU

Prolonged LOS (>7d)

TBIL GL

T

AL

Unplanned ICU Admission

ALT

BUN

NA

AL

K

ALT

AST

CA

HGB

NA +4.6%

ALP

CRE

1.0

GLU CO2 PLT ALT K ALP HCT NA TBIL CL AST HGB BUN WBC ALB CRE TP CA

Accuracy

Specificity NA

PLT

BUN

CA

2 CO

C

WBC

CL

WB

0.4

e Lead Time

Sensitivity

HCT

0.6

ALT ALP

CA 0.0 CRE 0.0 0.2

1.0

Unplanned ICU Admission

0.8

d Performance Improvement with NORMA Precision

1.0

0.8 True Positive Rate

0.6

Prolonged LOS (>7d)

1.0

0.8 True Positive Rate

0.8 True Positive Rate

Perioperative Infection

True Positive Rate

Mortality

0

20

40 60 80 100 Lead Time (hours)

120

Fig. 6 | Reclassification performance and early risk stratification in the INSPIRE cohort. a) Reclassification Rate. Reclassification rate by analyte among PopRI -normal tests. b) Hazard Ratios. Cox hazard ratios for in-hospital mortality (adjusted for age and sex); analytes with p < 0.05 shown. c) True Positive Rate vs False Positive Rate. True positive rate versus false positive rate for clinical outcome prediction across in-hospital mortality, perioperative infection, prolonged hospital stay, and unplanned ICU admission, comparing PerRI and NORMARI . d) Performance Improvement with NORMA. Change in precision, sensitivity, specificity, and balanced accuracy (NORMARI versus PerRI ); inner label shows the largest gain. e) Lead Time. Median lead time from NORMARI flag to first PopRI flag, by analyte (hours). Error bars denote the interquartile range.

21

Supplementary Figures a Forecasting Performance NORMA

NORMA

NORMA

ARIMA

ARIMA

ARIMA

Mean

Mean

Mean

Last

Last

Last

0

2

4

6

8

10

MAE

12

14

0

5

10

15

20

MAPE

25

0.0

0.2

0.4

0.6

R2

0.8

b Per-Analyte Performance

1.0

NORMA ARIMA Mean Last

0.8

0.6 0.4

TP

WBC

TC

1.5

2.0

TGL

TBIL

RBC

RDW

NA

PLT

MPV

MCV

MCH

MCHC

K

LDL

HGB

HCT

HDL

GLU

CRE

DBIL

CO2

CL

CA

BUN

ALT

AST

ALP

A1C

0.0

ALB

0.2

c Sensitivity Analysis CBC

BMP

HFP

80

60 40

60 40 20

20 101

102

Number of Measurements

0

100

0

80

10

Deviation from Midpoint (%)

80

CI Width (% of ref range)

100

0

10

120

100

CI Width (% of ref range)

CI Width (% of ref range)

120

Lipid

60 40 20

101

102

0 0.0

103

Prediction Horizon (days)

20 30 40

0.5

1.0

1.5

2.0

Within-Person SD

2.5

3.0

0.0

0.5

1.0

Within-Person SD

2.5

3.0

Supplementary Fig. 1 | NORMA forecasting performance with Gaussian loss. a) Forecasting Performance. Held-out test-set accuracy (mean absolute error, mean absolute percentage error, and coefficient of determination) for the Gaussian variant versus ARIMA, population mean, and last-value-carried-forward baselines. Error bars denote 95% bootstrap confidence intervals. b) Per-Analyte Performance. Per-analyte coefficient of determination grouped by clinical panel. c) Sensitivity Analysis. Sensitivity of Gaussian-head prediction interval width and midpoint deviation (each as a percentage of PopRI width) to history length, forecast horizon, and within-person variability.

22

MAE

MAPE

A1C

0.4

0.4

0.5

6.1

5.7

7.7

0.64

0.70

ALB

0.3

0.3

0.3

8.5

7.7

10.2

0.74

0.84

0.37 0.75

ALP

16.0

15.6

24.3

15.4

15.2

19.9

0.74

0.80

0.78 0.57

ALT

7.2

6.2

14.5

27.8

24.7

40.4

0.62

0.79

AST

7.0

6.0

20.4

23.5

21.0

44.5

0.58

0.77

0.41

BUN

3.3

3.3

4.4

18.4

17.2

23.2

0.83

0.83

0.82

CA

0.3

0.3

0.4

3.3

3.0

4.4

0.66

0.76

CL

1.9

1.8

2.4

1.8

1.8

2.3

0.69

0.75

CO2

1.5

1.3

2.1

6.3

5.4

8.3

0.70

0.80

CRE

0.1

0.1

0.2

10.6

12.3

12.9

0.88

0.83

DBIL

0.3

0.3

0.4

48.0

36.5

88.0

0.51

0.64

GLU

15.7

14.0

21.1

12.9

11.3

17.7

0.48

0.61

HCT

2.1

2.0

3.0

6.2

6.1

8.7

12.1

11.1

18.8

6.3

6.2

8.5

HDL

6.5

6.0

8.8

HGB

0.7

0.7

1.0

K

0.3

0.3

0.4

LDL

17.7

15.0

18.6

MCH

0.5

0.6

0.6

MCHC

0.6

0.6

0.6

MCV

1.6

1.9

1.7

MPV

0.6

0.7

0.9

NA

1.8

1.8

20.0 17.5 15.0 12.5 10.0

25

20

15

0.79

0.83

0.73

0.80

0.79

0.84

0.40

1.0

0.55 0.48 0.87

0.8

0.57 0.23 0.60

0.6

0.53 0.63

6.7

6.3

8.5

0.45

0.56

21.8

20.0

33.7

0.55

0.72

1.8

2.1

2.2

0.90

0.90

1.8

1.9

1.9

0.72

0.72

1.8

2.2

1.9

0.88

0.85

6.5

6.7

9.9

0.59

0.66

0.21

2.2

1.3

1.3

1.6

0.62

0.64

0.49

7.5 5.0 2.5

10

5

0.15

0.4

0.52 0.83 0.58

0.2

0.82

0.0

PLT

30.3

52.8

36.6

13.5

19.3

18.1

0.79

0.29

RBC

0.2

0.2

0.3

6.3

6.4

8.5

0.81

0.84

0.68

RDW

0.5

0.5

0.7

3.0

3.4

4.6

0.87

0.86

0.81

TBIL

0.2

0.1

0.3

28.0

22.5

32.1

0.71

0.81

0.85

24.8

12.0

9.8

17.5

0.57

0.76

0.44

28.7

22.8

0.48

0.71

0.64

4.7

6.9

0.54

0.74

0.57

24.2

19.6

26.0

0.62

0.79

0.69

IM AR

us sia NGa

NQu an

IM AR

us sia NGa

NQu an

IM AR

us sia NGa

NQu an

A

5.7

2.1

n

0.5

1.3

tile

0.3

1.5

A

0.4

n

TP WBC

tile

28.7

A

17.0

36.2

n

20.3

tile

TC TGL

Supplementary Fig. 2 | Per-analyte forecasting performance. NORMA using quantile and Gaussian loss, compared to ARIMA, the best performing baseline. Color intensity encodes mean absolute error (MAE), mean absolute percentage error (MAPE), and coefficient of determination (R2 ) on the held-out test set.

23

eICU-CRD

A1C

ALB

30

30

30

25

20

20 15

10

10

CA

CL

35 30 25

25

25

20

20

15

15

10

10 5

CO2

20

HCT

15

20

10

50

30

40 30

35

35

26

30

30

24

25

25

20

20

15

15

20

10

MCHC

25.0 22.5 20.0 17.5 15.0 12.5 10.0

10

MCV

25.0 22.5 20.0 17.5 15.0 12.5 10.0

RDW 20 15 10 0

1

2

3

4

5

27 26

26

25

4

5

25

27

20 15 10

RBC

25

16

20

14 12

10

10

TP

0

1

2

3

4

5

24

WBC

30

25

24 3

MCH

28

TGL

25

2

10

10

20

1

15

PLT

20

27

0

15

15

TC

5

20

24

30

TBIL

10

20

25

40

10

15

25

NA

15

30

25

25

LDL

10

20

GLU

26

MPV

25

10

K

40

28

22

20

HGB

40

15

5

10

5

HDL

10

DBIL

20

10

10

25 20

CRE

30

0

1

2

BUN

30

15

25

15

Mortality Rate (%)

25

40

30

20

AST

20

50

40

INSPIRE

ALT

5

0 40

CHS

ALP

3

4

5

30

25

25

20

20

15

15

10

10 0

1

2

3

4

5

0

1

2

3

4

5

Standardized Distance from Baseline (|z|)

Supplementary Fig. 3 | Association of absolute deviation from the personalized baseline with observed mortality. Each panel shows the observed mortality rate by standardized absolute deviation from the personalized baseline, binned by decile, for Clalit Health Services, the eICU Collaborative Research Database, and the INSPIRE cohort.

24

eICU-CRD

A1C

ALB

7.5

25

90

7.0

20 15

6.0

10

5.5

60

5

50

80 70

CA

CL

CO2

115

AST

10

NORMA Midpoint Prediction

HDL

60

5

47.5 45.0 42.5 40.0 37.5 35.0

55

40

50

30

45

20

40

10

MCHC

47.5

MCV

42.5 40.0 37.5 35.0

60

80

NA

145

80

15

PLT

135

TC

30 20 10

TGL

TP

WBC

150

20 15

140

125

20

40

60

80

100

60

25

80

15 10

50 40

20

20

75

20

25

30

100

120

RBC

40

325 300 275 250 225 200 175

140

TBIL 160

0

85

155 150

5 40

MCH 40.0 37.5 35.0 32.5 30.0 27.5

90

MPV

25

10

20

75

100

10

20

25

125 100

95

10

RDW 40 35 30 25 20 15 10

175 150

LDL

20

90

32.5

GLU 200

K 30

30

100

45.0

HGB

50

DBIL

15.0 12.5 10.0 7.5 5.0 2.5 0.0

0

HCT

50.0

15 10

10

15

25

25 20

20

20

100

30

30

CRE

BUN

35

40

25

30

105

10

ALT

30

35

110 15

INSPIRE

45 40 35 30 25 20 15

100

6.5

20

CHS

ALP

10 20

40

60

80

5 20

40

60

80

20

40

60

80

Age (years)

Supplementary Fig. 4 | Age-dependent NORMARI midpoints. Each panel shows individual midpoint predictions with a polynomial smooth overlaid, for Clalit Health Services, the eICU Collaborative Research Database, and the INSPIRE cohort.

25

PopRI

PerRI

NORMARI

100

eICU-CRD

80

60

40

20

0 100

60

CHS

Abnormality Prevalence (%)

80

40

20

0 100

INSPIRE

80

60

40

20

TC TG L TP WB C

K LD L MC H MC HC MC V MP V NA PL T RB C RD W TB IL

CL CO 2 CR E DB IL GL U HC T HD L HG B

CA

A1C AL B AL P AL T AS T BU N

0

Supplementary Fig. 5 | Abnormality prevalence by analyte and reference interval method. Grouped bars show the percentage of measurements classified as abnormal by PopRI , PerRI , and NORMARI , for the eICU Collaborative Research Database, Clalit Health Services, and the INSPIRE cohort.

26

a Hazard Ratios Across All Outcomes PopRI

Chronic Kidney Disease

NORMARI

PerRI

Mortality

Type 2 Diabetes

A1C ALB ALP ALT AST BUN CA CRE GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TP WBC 1

10

Hazard Ratio

0.1

1

10

Hazard Ratio

1

10

Hazard Ratio

b Concordance Index Comparison PopRI

Chronic Kidney Disease MCHC

PerRI

Mortality

NORMARI

Type 2 Diabetes

MPV

ALP

RDW

MCH

TP

MCH

RDW

CA

MCV

MCV

ALT

BUN

MCHC

AST

HCT

CRE

NA

WBC

WBC

WBC

PLT

HCT

CRE

HGB

HGB

GLU

CRE

PLT 0.65

0.70

0.75

C-index

0.80

0.85

CL 0.65

0.70

0.75

C-index

0.80

0.85

0.500 0.525 0.550 0.575 0.600 0.625 0.650

C-index

Supplementary Fig. 6 | Proportional hazards analysis in Clalit Health Services. a) Hazard Ratios Across All Outcomes. Cox hazard ratios for chronic kidney disease, all-cause mortality, type 2 diabetes, and anemia, comparing PopRI , PerRI , and NORMARI abnormality flags; only analytes with p < 0.05 are shown. b) Concordance Index Comparison. Concordance index for the top 10 analytes ranked by NORMARI concordance per outcome.

27

a Hazard Ratios Across All Outcomes Acute Kidney Injury

PopRI

Mortality

PerRI

NORMARI

Prolonged LOS (>7d)

Sepsis

ALB ALP AST BUN CA CL CO2 CRE DBIL HCT HGB K MCH MCHC MCV MPV NA PLT RBC RDW TBIL TP WBC 1

2 × 100 Hazard Ratio

3 × 100

1

4 × 100

Hazard Ratio

2 × 100

3 × 100

4 × 10 1

PerRI

NORMARI

1

6 × 10 1 Hazard Ratio

1

2 × 100 Hazard Ratio

3 × 100

b Concordance Index Comparison Acute Kidney Injury

PopRI

Mortality

Prolonged LOS (>7d)

Sepsis

WBC

CA

RBC

ALP

TBIL

BUN

CL

MCH

ALP

CL

BUN

CO2

CL

AST

ALP

WBC

CA

CO2

NA

CL

K

PLT

WBC

CRE

BUN

RDW

RDW

MCHC

CRE

TGL

K

CA

RDW

CRE

MCHC

RDW

TC

WBC 0.50

0.55

0.60

C-index

0.65

0.70

CRE 0.50

0.55

0.60

C-index

0.65

0.70

0.45

K 0.50

0.55

C-index

0.60

0.65

0.45

0.50

0.55

C-index

0.60

Supplementary Fig. 7 | Proportional hazards analysis in the eICU Collaborative Research Database. a) Hazard Ratios Across All Outcomes. Cox hazard ratios for acute kidney injury, in-hospital mortality, prolonged ICU stay, and sepsis, comparing PopRI , PerRI , and NORMARI abnormality flags; only analytes with p < 0.05 are shown. b) Concordance Index Comparison. Concordance index for the top 10 analytes ranked by NORMARI concordance per outcome.

28

a Hazard Ratios Across All Outcomes Mortality

PopRI

Perioperative Infection

PerRI

NORMARI

Prolonged LOS (>7d)

Unplanned ICU Admission

ALB ALP ALT AST BUN CA CL CO2 CRE GLU HCT HGB K NA PLT TBIL TP WBC 1

1

Hazard Ratio

1.2 × 100

1.4 × 100 Hazard Ratio

1.6 × 100

7 × 10 1

1.8 × 100

8 × 10 1 9 × 10 1 Hazard Ratio

1

8 × 10 1

9 × 10 1

1

Hazard Ratio

b Concordance Index Comparison Mortality

Perioperative Infection

PopRI

PerRI

NORMARI

Prolonged LOS (>7d)

Unplanned ICU Admission

ALP

NA

CL

ALT

NA

HGB

PLT

TBIL

TBIL

CL

WBC

ALP

CL

TP

NA

CA

CA

K

TP

TP

ALB

ALT

BUN

NA

TP

HCT

HCT

ALB

WBC

ALB

HGB

WBC

K

PLT

ALB

CL

PLT

AST

CRE

CRE

0.5

0.6

0.7

C-index

0.8

0.9

0.50

0.55

0.60

C-index

0.65

0.70

0.45

0.50

0.55

C-index

0.60

0.65

0.45

0.50

0.55

C-index

0.60

Supplementary Fig. 8 | Proportional hazards analysis in the INSPIRE cohort. a) Hazard Ratios Across All Outcomes. Cox hazard ratios for in-hospital mortality, perioperative infection, prolonged hospital stay, and unplanned ICU admission, comparing PopRI , PerRI , and NORMARI abnormality flags; only analytes with p < 0.05 are shown. b) Concordance Index Comparison. Concordance index for the top 10 analytes ranked by NORMARI concordance per outcome.

29

Anemia

WBC TP TC TBIL RDW RBC PLT NA MPV MCV MCHC MCH LDL K HGB HDL HCT GLU CRE CL CA BUN AST ALT ALP ALB A1C

Chronic Kidney Disease

WBC TP TC TBIL RDW RBC PLT NA MPV MCV MCHC MCH LDL K HGB HDL HCT GLU CRE CL CA BUN AST ALT ALP ALB A1C

Mortality

WBC TP TC TBIL RDW RBC PLT NA MPV MCV MCHC MCH LDL K HGB HDL HCT GLU CRE CL CA BUN AST ALT ALP ALB A1C

Type 2 Diabetes

Precision

WBC TP TC TBIL RDW RBC PLT NA MPV MCV MCHC MCH LDL K HGB HDL HCT GLU CRE CL CA BUN AST ALT ALP ALB A1C

0.2

0.4

0.6

0.8

1.0

0.4

0.6

NORMARI

PerRI

Sensitivity

0.8

1.0

0.0

Specificity

0.2

0.4

Accuracy

0.6

0.8

0.45 0.50 0.55 0.60 0.65 0.70

Supplementary Fig. 9 | Clinical outcome prediction performance in Clalit Health Services. Each point shows PerRI and NORMARI values for precision, sensitivity, specificity, and balanced accuracy across outcomes. Connected pairs show the direction and magnitude of change between methods for each analyte. Analytes with fewer than 100 measurements are omitted.

30

Acute Kidney Injury

WBC TP TGL TBIL RDW RBC PLT NA MPV MCV MCHC MCH K HGB HCT GLU DBIL CRE CO2 CL CA BUN AST ALT ALP ALB

Mortality

WBC TP TGL TBIL RDW RBC PLT NA MPV MCV MCHC MCH K HGB HCT GLU DBIL CRE CO2 CL CA BUN AST ALT ALP ALB

Prolonged LOS (>7d)

WBC TP TGL TBIL RDW RBC PLT NA MPV MCV MCHC MCH K HGB HCT GLU DBIL CRE CO2 CL CA BUN AST ALT ALP ALB

Sepsis

Precision

WBC TP TGL TBIL RDW RBC PLT NA MPV MCV MCHC MCH K HGB HCT GLU DBIL CRE CO2 CL CA BUN AST ALT ALP ALB

0.1

0.2

0.3

0.4

Sensitivity

0.5

0.6

0.0

0.2

0.4

0.6

PerRI

NORMARI

0.8

0.2

Specificity

0.4

0.6

Accuracy

0.8

0.44 0.46 0.48 0.50 0.52 0.54

Supplementary Fig. 10 | Clinical outcome prediction performance in the eICU Collaborative Research Database. Each point shows PerRI and NORMARI values for precision, sensitivity, specificity, and balanced accuracy across outcomes. Connected pairs show the direction and magnitude of change between methods for each analyte. Analytes with fewer than 100 measurements are omitted.

31

Precision

Sensitivity

PerRI

0.2

0.4

NORMARI

Specificity

Accuracy

WBC TP TBIL PLT NA K

Mortality

HGB HCT GLU CRE CO2 CL CA BUN AST ALT ALP ALB

WBC TP TBIL

Perioperative Infection

PLT NA K HGB HCT GLU CRE CO2 CL CA BUN AST ALT ALP ALB

WBC TP TBIL

Prolonged LOS (>7d)

PLT NA K HGB HCT GLU CRE CO2 CL CA BUN AST ALT ALP ALB

WBC TP

Unplanned ICU Admission

TBIL PLT NA K HGB HCT GLU CRE CO2 CL CA BUN AST ALT ALP ALB

0.0

0.2

0.4

0.6

0.8

0.0

0.1

0.3

0.5

0.5

0.6

0.7

0.8

0.9

0.45

0.50

0.55

0.60

Supplementary Fig. 11 | Clinical outcome prediction performance in the INSPIRE cohort. Each point shows PerRI and NORMARI values for precision, sensitivity, specificity, and balanced accuracy across outcomes. Connected pairs show the direction and magnitude of change between methods for each analyte. Analytes with fewer than 100 measurements are omitted.

32

Supplementary Tables Supplementary Table 1 | Analyte reference information. Full analyte name, abbreviation, unit of measurement, and conventional PopRI for all 30 laboratory analytes included in the study. PopRI were obtained from the American Board of Internal Medicine Laboratory Test Reference Ranges (January 2025). Analyte

Full Name

Unit

Population RI

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

Hemoglobin A1c Albumin Alkaline Phosphatase Alanine Aminotransferase Aspartate Aminotransferase Blood Urea Nitrogen Calcium Chloride Bicarbonate Creatinine Direct Bilirubin Glucose Hematocrit HDL Cholesterol Hemoglobin Potassium LDL Cholesterol Mean Corpuscular Hemoglobin MCH Concentration Mean Corpuscular Volume Mean Platelet Volume Sodium Platelet Count Red Blood Cell Count Red Cell Distribution Width Total Bilirubin Total Cholesterol Triglycerides Total Protein White Blood Cell Count

% g/dL U/L U/L U/L mg/dL mg/dL mEq/L mEq/L mg/dL mg/dL mg/dL % mg/dL g/dL mEq/L mg/dL pg g/dL fL fL mEq/L 103 /µL 106 /µL % mg/dL mg/dL mg/dL g/dL 103 /µL

4.0–5.6 3.5–5.5 30–120 0–35 0–35 8–20 8.6–10.2 98–106 23–28 F: 0.5–1.1, M: 0.7–1.3 0.1–0.3 70–99 F: 37–47, M: 42–50 40–100 F: 12.0–16.0, M: 14.0–18.0 3.5–5.0 <130 28–32 33–36 80–98 7.5–12.5 136–145 150–450 F: 4.0–5.2, M: 4.5–5.9 9–14.5 0.3–1.0 100–200 <150 6.0–8.0 4.5–11.0

33

Supplementary Table 2 | Cohort characteristics. Patient counts, demographics (age, sex, follow-up duration), and per-analyte measurement counts and mean values (± standard deviation) across all five cohorts: EHRSHOT, MIMIC-IV, Clalit Health Services, the eICU Collaborative Research Database, and the INSPIRE cohort.

Patients Age Sex Time

EHRSHOT

MIMIC-IV

eICU-CRD

CHS

INSPIRE

5,676 54.9 ± 17.2 52% F / 48% M 2.49 [0.28, 7.20] yr

179,601 58.2 ± 19.0 54% F / 46% M 1.01 [0.06, 4.08] yr

98,432 63.3 ± 16.1 46% F / 54% M 0.02 [0.01, 0.03] yr

1,450,862 55.1 ± 18.6 63% F / 37% M 20.24 [17.51, 21.47] yr

51,159 57.6 ± 14.9 52% F / 48% M 0.48 [0.17, 0.84] yr

Analyte

N

Value

N

Value

N

Value

N

Value

N

Value

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

1,745 4,282 4,206 4,219 4,220 5,180 5,162 5,172 5,187 5,161 1,429 5,150 5,470 1,549 5,369 5,236 821 5,294 5,317 5,304 371 5,193 5,320 5,312 5,315 4,169 1,635 4 4,213 5,345

6.3 ± 1.2 3.3 ± 0.8 112.5 ± 58.8 37.8 ± 25.2 31.9 ± 19.5 23.0 ± 15.2 8.8 ± 0.7 102.4 ± 5.0 25.9 ± 3.8 1.1 ± 0.5 0.3 ± 0.3 124.9 ± 39.9 32.2 ± 6.8 53.6 ± 17.5 10.7 ± 2.3 4.1 ± 0.5 94.9 ± 34.9 30.4 ± 2.7 33.2 ± 1.4 91.5 ± 6.9 9.6 ± 1.6 137.4 ± 4.2 198.0 ± 111.4 3.5 ± 0.8 16.1 ± 2.9 0.6 ± 0.4 173.5 ± 45.0 165.7 ± 100.8 6.6 ± 1.0 7.8 ± 4.6

39,996 55,220 78,930 92,599 90,639 153,329 121,478 148,304 145,386 156,044 6,635 147,195 164,903 49,011 162,168 150,333 48,688 161,671 161,806 161,661 — 149,303 162,308 161,825 161,640 73,002 49,762 48,445 18,673 164,291

6.6 ± 1.3 3.8 ± 0.8 107.0 ± 58.4 30.2 ± 21.9 32.7 ± 21.4 22.2 ± 14.3 8.8 ± 0.7 101.8 ± 5.2 25.6 ± 4.1 1.1 ± 0.6 0.7 ± 0.9 118.3 ± 40.4 33.6 ± 6.8 55.8 ± 17.7 11.0 ± 2.3 4.2 ± 0.5 104.2 ± 37.8 29.7 ± 2.8 32.6 ± 1.6 91.2 ± 7.3 — 138.7 ± 4.2 229.4 ± 113.3 3.7 ± 0.8 15.1 ± 2.4 0.7 ± 0.6 187.5 ± 44.6 140.4 ± 77.5 7.0 ± 0.7 7.7 ± 4.0

— 19,075 13,582 12,425 12,704 57,540 57,067 59,172 70,208 54,820 1,147 90,958 60,110 11 60,482 66,569 3 48,432 51,710 50,841 37,083 62,419 53,371 53,584 48,715 12,738 35 312 15,134 52,141

— 2.6 ± 0.7 97.0 ± 43.2 33.8 ± 23.4 37.1 ± 25.6 25.3 ± 15.5 8.3 ± 0.7 103.8 ± 6.2 25.0 ± 4.7 1.1 ± 0.6 0.6 ± 0.5 139.6 ± 44.3 30.1 ± 5.9 26.0 ± 14.3 9.9 ± 2.0 4.0 ± 0.6 92.1 ± 56.2 29.7 ± 2.1 32.8 ± 1.3 90.4 ± 5.7 9.7 ± 1.4 138.6 ± 5.0 199.0 ± 98.4 3.4 ± 0.7 15.8 ± 2.0 0.8 ± 0.5 108.3 ± 44.0 144.9 ± 68.0 5.8 ± 1.0 10.6 ± 4.6

336,493 790,253 1,099,044 1,233,902 1,236,781 1,251,453 909,036 31,171 299 1,310,465 — 1,351,168 1,407,912 1,002,204 1,408,118 1,182,544 1,187,844 1,250,759 1,407,568 1,400,648 1,304,857 1,187,537 1,407,716 1,404,794 1,403,339 815,774 1,260,730 1,244,518 791,481 1,409,027

7.1 ± 1.3 4.1 ± 0.5 82.1 ± 34.3 21.9 ± 17.3 23.6 ± 16.3 17.4 ± 8.3 9.3 ± 0.5 103.1 ± 3.9 31.4 ± 12.1 0.9 ± 0.4 — 110.0 ± 36.1 39.1 ± 5.2 48.0 ± 15.9 12.8 ± 1.8 4.5 ± 0.5 105.1 ± 33.6 28.8 ± 2.3 32.8 ± 1.3 87.7 ± 6.1 9.5 ± 1.4 139.9 ± 3.0 243.4 ± 73.6 4.5 ± 0.6 14.1 ± 1.5 0.6 ± 0.4 181.7 ± 40.7 133.5 ± 65.3 7.1 ± 0.6 7.5 ± 3.7

391 30,728 27,797 26,144 26,374 29,912 29,788 30,952 6,247 42,255 — 26,168 39,739 — 36,187 36,363 — — — — — 36,127 34,006 — — 28,571 — — 28,603 34,273

6.7 ± 1.0 3.4 ± 0.6 73.4 ± 29.1 21.0 ± 12.5 24.2 ± 10.0 16.1 ± 8.0 8.5 ± 0.7 102.5 ± 4.8 24.8 ± 3.8 0.8 ± 0.2 — 143.8 ± 48.4 32.9 ± 5.8 — 11.1 ± 2.0 4.0 ± 0.5 — — — — — 137.4 ± 3.5 206.0 ± 92.3 — — 0.8 ± 0.4 — — 6.1 ± 0.9 8.2 ± 3.5

34

Supplementary Table 3 | NORMA architectural and training design choices. Summary of configurable components including temporal encoding, health state encoding, age and laboratory value representations, output parameterization, and context token usage. Component Time embedding

Design Choice

Description

Time2Vec

Encodes the number of days elapsed since the first measurement using learned linear and sinusoidal functions. Captures long-term temporal trends and potential periodic patterns.

Delta-t

Encodes the log-compressed time difference between consecutive measurements. Explicitly represents irregular sampling and emphasizes how much time has passed since the last value.

Health state encoding

Binary

Encodes each lab value as either within or outside the population reference range; does not distinguish high vs. low and can struggle with J-shaped biomarkers.

Ternary

Encodes whether the lab value is low, normal, or high relative to the reference range; handles asymmetric risk patterns; however, can increase overfitting risk.

Age encoding

Raw

Uses age directly as a numeric input; assumes linear effects.

Binned

Discretizes age into groups (e.g., decades) and maps each bin to a learned embedding vector; captures nonlinear and threshold effects.

Lab value encoding

Raw

Uses observed biomarker values directly as model input; preserves absolute clinical scale and keeps preprocessing minimal.

Within-sequence nor-

Applies per-sequence z-scoring with a learnable scale and

malization

shift, then denormalizes at output; stabilizes optimization and improves cross-analyte sharing.

Training objective

Gaussian NLL

Minimizes the negative log-likelihood of a normal distribution by predicting mean and log-variance; aligns with conventional reference interval assumptions but may underperform for skewed or heavy-tailed analytes.

Quantile (pinball) loss

Minimizes pinball loss for fixed quantiles (2.5, 25, 50, 75, 97.5%); distribution-free and robust to skewness or heteroscedasticity.

35

Supplementary Table 4 | Within-person and between-person biological variability. Per-analyte within-person coefficient of variation, between-person coefficient of variation, and individuality index for the eICU Collaborative Research Database and Clalit Health Services. Values are reported as median [95% confidence interval]. eICU-CRD

CHS

INSPIRE

Analyte

CVintra

CVinter

II

CVintra

CVinter

II

CVintra

CVinter

II

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

— 8.9 [8.8, 8.9] 10.6 [10.5, 10.7] 15.5 [15.3, 15.7] 18.3 [18.1, 18.5] 16.5 [16.4, 16.6] 3.8 [3.7, 3.8] 2.0 [2.0, 2.0] 7.1 [7.1, 7.2] 13.5 [13.4, 13.5] — 14.8 [14.8, 14.9] 6.2 [6.2, 6.3] 7.6 [6.0, 9.5] 6.7 [6.7, 6.7] 7.6 [7.6, 7.6] 8.6 [6.0, 11.5] 1.1 [1.1, 1.1] 1.3 [1.3, 1.3] 1.0 [1.0, 1.0] 2.9 [2.9, 2.9] 1.3 [1.3, 1.3] 11.8 [11.7, 11.9] 7.2 [7.2, 7.3] 2.1 [2.0, 2.1] 21.9 [21.8, 22.1] 7.6 [6.6, 8.8] 16.0 [15.3, 16.6] 6.5 [6.4, 6.5] 15.6 [15.5, 15.7]

— 22.6 [22.4, 22.7] 42.1 [41.8, 42.4] 65.3 [64.8, 65.8] 63.4 [62.9, 63.8] 58.7 [58.5, 58.9] 7.2 [7.1, 7.2] 5.1 [5.0, 5.1] 15.9 [15.9, 16.0] 47.7 [47.5, 47.9] 82.8 [81.0, 84.7] 24.8 [24.7, 25.0] 18.4 [18.3, 18.5] 39.7 [33.9, 44.8] 19.1 [19.0, 19.1] 10.5 [10.4, 10.5] 45.6 [38.4, 52.0] 6.9 [6.9, 7.0] 3.6 [3.6, 3.7] 6.2 [6.2, 6.3] 12.8 [12.8, 12.9] 2.9 [2.9, 2.9] 41.6 [41.4, 41.8] 18.6 [18.5, 18.7] 12.6 [12.6, 12.7] 58.2 [57.7, 58.6] 38.4 [34.9, 41.4] 43.2 [41.8, 44.6] 14.3 [14.2, 14.4] 37.5 [37.3, 37.6]

— 0.4 [0.4, 0.4] 0.2 [0.2, 0.3] 0.2 [0.2, 0.2] 0.3 [0.3, 0.3] 0.3 [0.3, 0.3] 0.5 [0.5, 0.5] 0.4 [0.4, 0.4] 0.5 [0.5, 0.5] 0.3 [0.3, 0.3] — 0.6 [0.6, 0.6] 0.3 [0.3, 0.3] 0.2 [0.2, 0.2] 0.3 [0.3, 0.3] 0.7 [0.7, 0.7] 0.2 [0.1, 0.3] 0.1 [0.1, 0.1] 0.3 [0.3, 0.3] 0.2 [0.2, 0.2] 0.2 [0.2, 0.2] 0.4 [0.4, 0.4] 0.3 [0.3, 0.3] 0.4 [0.4, 0.4] 0.2 [0.2, 0.2] 0.4 [0.4, 0.4] 0.2 [0.2, 0.2] 0.4 [0.3, 0.4] 0.5 [0.5, 0.5] 0.4 [0.4, 0.4]

8.6 [8.5, 8.6] 6.7 [6.7, 6.8] 17.8 [17.7, 17.8] 34.2 [34.2, 34.3] 24.7 [24.6, 24.7] 20.5 [20.4, 20.5] 3.5 [3.5, 3.5] 2.4 [2.3, 2.4] 17.7 [17.3, 18.2] 12.2 [12.2, 12.2] — 13.7 [13.7, 13.7] 6.3 [6.3, 6.4] 24.4 [24.4, 24.5] 6.3 [6.3, 6.3] 7.1 [7.1, 7.1] 20.6 [20.5, 20.6] 3.6 [3.6, 3.6] 2.9 [2.9, 2.9] 3.2 [3.2, 3.2] 10.3 [10.3, 10.3] 1.5 [1.5, 1.5] 14.1 [14.1, 14.1] 5.9 [5.9, 5.9] 5.7 [5.7, 5.7] 26.0 [25.9, 26.0] 13.9 [13.9, 14.0] — 4.8 [4.8, 4.8] 19.3 [19.2, 19.3]

16.5 [16.5, 16.6] 18.9 [18.3, 19.4] 30.1 [30.0, 30.2] 50.2 [50.0, 50.5] 35.1 [34.9, 35.4] 36.2 [36.2, 36.3] 3.7 [3.7, 3.7] 2.3 [2.2, 2.3] 27.0 [25.9, 28.1] 34.7 [34.4, 34.9] — 24.0 [23.9, 24.0] 9.6 [9.6, 9.6] 30.1 [30.1, 30.2] 10.5 [10.5, 10.5] 6.5 [6.4, 6.5] 23.7 [23.7, 23.8] 7.1 [7.1, 7.1] 2.7 [2.7, 2.7] 6.1 [6.1, 6.1] 10.8 [10.8, 10.8] 1.2 [1.2, 1.2] 23.8 [23.7, 23.8] 10.1 [10.1, 10.1] 7.1 [7.1, 7.1] 42.8 [42.6, 43.0] 16.6 [16.6, 16.6] — 5.6 [5.5, 5.6] —

0.5 [0.5, 0.5] 0.4 [0.3, 0.4] 0.6 [0.6, 0.6] 0.7 [0.7, 0.7] 0.7 [0.7, 0.7] 0.6 [0.6, 0.6] 0.9 [0.9, 0.9] 1.0 [1.0, 1.1] 0.7 [0.6, 0.7] 0.3 [0.3, 0.4] — 0.6 [0.6, 0.6] 0.7 [0.7, 0.7] 0.8 [0.8, 0.8] 0.6 [0.6, 0.6] 1.1 [1.1, 1.1] 0.9 [0.9, 0.9] 0.5 [0.5, 0.5] 1.1 [1.1, 1.1] 0.5 [0.5, 0.5] 0.9 [0.9, 1.0] 1.2 [1.2, 1.2] 0.6 [0.6, 0.6] 0.6 [0.6, 0.6] 0.8 [0.8, 0.8] 0.6 [0.6, 0.6] 0.8 [0.8, 0.8] — 0.9 [0.9, 0.9] 0.0 [0.0, 0.7]

4.8 [4.4, 5.3] 8.4 [8.3, 8.4] 10.9 [10.7, 11.0] 19.2 [19.0, 19.4] 14.4 [14.3, 14.6] 15.4 [15.2, 15.5] 4.4 [4.3, 4.4] 1.9 [1.8, 1.9] 5.5 [5.4, 5.6] 11.0 [11.0, 11.1] — 12.8 [12.7, 12.9] 6.8 [6.7, 6.8] — 6.9 [6.8, 6.9] 7.5 [7.5, 7.6] — — — — — 1.0 [1.0, 1.0] 12.0 [11.9, 12.2] — — 21.7 [21.6, 21.9] — — 7.5 [7.4, 7.5] 18.3 [18.2, 18.5]

13.0 [11.8, 13.8] 13.4 [13.3, 13.5] 34.0 [33.7, 34.4] 49.8 [49.3, 50.3] 32.6 [32.2, 32.9] 42.0 [41.5, 42.4] 5.6 [5.5, 5.6] 3.3 [3.3, 3.3] 10.2 [10.1, 10.4] 25.3 [25.2, 25.5] — 23.7 [23.5, 23.9] 14.1 [14.0, 14.2] — 14.6 [14.5, 14.7] 8.6 [8.6, 8.7] — — — — — 1.7 [1.7, 1.7] 33.5 [33.2, 33.8] — — 44.6 [44.2, 45.0] — — 10.9 [10.8, 11.0] 31.8 [31.5, 32.0]

0.4 [0.3, 0.4] 0.6 [0.6, 0.6] 0.3 [0.3, 0.3] 0.4 [0.4, 0.4] 0.4 [0.4, 0.5] 0.4 [0.4, 0.4] 0.8 [0.8, 0.8] 0.6 [0.6, 0.6] 0.5 [0.5, 0.6] 0.4 [0.4, 0.4] — 0.5 [0.5, 0.6] 0.5 [0.5, 0.5] — 0.5 [0.5, 0.5] 0.9 [0.9, 0.9] — — — — — 0.6 [0.6, 0.6] 0.4 [0.3, 0.4] — — 0.5 [0.5, 0.5] — — 0.7 [0.7, 0.7] 0.6 [0.6, 0.6]

36

Supplementary Table 5 | NORMA forecasting performance. Mean absolute error (MAE), mean absolute percentage error (MAPE), and coefficient of determination (R2 ) for NORMA Gaussian, NORMA quantile, ARIMA, population mean, and last-value-carried-forward, reported on training, validation, and held-out test sets. Values are reported as median [interquartile range] across analytes. Metric

NORMA Gaussian

NORMA Quantile

ARIMA

Mean

Last

MAE

Train Val Test

6.0 [3.2, 11.1] 6.0 [2.6, 10.7] 6.0 [2.6, 10.2]

5.7 [2.9, 10.1] 5.8 [2.9, 8.7] 5.9 [3.6, 9.1]

6.8 [2.5, 10.2] 6.8 [3.6, 10.0] 6.7 [4.3, 11.9]

8.4 [4.5, 15.0] 8.3 [4.9, 13.5] 8.4 [4.5, 13.6]

7.4 [3.7, 11.0] 7.3 [3.8, 10.3] 7.3 [3.9, 10.7]

MAPE

Train Val Test

10.9 [8.8, 14.0] 11.1 [7.9, 13.2] 11.1 [8.5, 14.3]

11.9 [8.6, 14.7] 12.6 [8.7, 15.6] 12.3 [7.9, 17.3]

15.9 [11.3, 20.1] 15.7 [12.1, 21.5] 16.9 [11.3, 23.3]

19.4 [14.8, 25.4] 19.7 [14.0, 27.4] 19.7 [13.5, 26.7]

14.8 [10.9, 19.6] 15.2 [10.6, 20.1] 15.2 [10.6, 21.1]

R2

Train Val Test

0.76 [0.72, 0.80] 0.75 [0.69, 0.78] 0.75 [0.70, 0.78]

0.70 [0.66, 0.75] 0.69 [0.66, 0.73] 0.68 [0.64, 0.72]

0.62 [0.57, 0.70] 0.62 [0.53, 0.67] 0.58 [0.51, 0.64]

0.46 [0.41, 0.49] 0.46 [0.43, 0.49] 0.45 [0.41, 0.50]

0.56 [0.47, 0.62] 0.55 [0.49, 0.62] 0.54 [0.49, 0.61]

37

Supplementary Table 6 | Sensitivity of NORMA prediction intervals to patient features. Median and interquartile range of prediction interval width change (as a percentage of PopRI width) and midpoint shift (as a percentage of PopRI width) in response to history length, forecast horizon, and within-person variability, for both the quantile and Gaussian parameterizations. CI Width Change (%)

Midpoint Shift (%)

Median

IQR

Median

IQR

Model

Feature

Quantile Quantile Quantile

History Length Prediction Horizon Within-Person Variability

4.8 1.6 115.2

[1.0, 14.4] [0.1, 11.1] [111.4, 120.5]

0.0 0.0 0.7

[0.0, 0.0] [0.0, 0.0] [0.3, 1.7]

Gaussian Gaussian Gaussian

History Length Prediction Horizon Within-Person Variability

28.3 7.2 7.2

[20.6, 41.8] [3.0, 9.9] [3.7, 12.4]

12.8 1.6 1.3

[5.6, 20.6] [0.8, 3.2] [0.3, 3.6]

38

Supplementary Table 7 | Abnormality prevalence and reclassification rates. Overall percentage of measurements classified as abnormal by PopRI , PerRI , and NORMARI , and the reclassification rate (the proportion of PopRI -normal tests flagged abnormal by PerRI or NORMARI ) in the eICU Collaborative Research Database and Clalit Health Services. Dataset

N

PopRI (%)

PerRI (%)

NORMARI (%)

PerRI RR (%)

NORMARI RR (%)

eICU-CRD CHS INSPIRE

29 31 19

50.2 29.6 35.1

68.1 46.8 57.0

55.8 39.1 42.5

37.3 27.2 34.3

9.2 12.0 14.4

39

Supplementary Table 8 | Per-analyte reclassification. Among tests with a PerRI setpoint within PopRI , the number per 1,000 flagged abnormal by PopRI , PerRI , and NORMARI for each analyte, in the eICU Collaborative Research Database and Clalit Health Services. CHS

eICU-CRD

INSPIRE

Analyte

PopRI

PerRI

NORMARI

PopRI

PerRI

NORMARI

PopRI

PerRI

NORMARI

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

240 146 119 136 71 235 140 229 545 141 — 326 322 157 279 122 118 184 400 94 76 111 75 240 234 119 184 — 102 146

315 284 435 414 383 458 290 415 602 256 — 517 497 445 424 250 414 363 498 375 314 361 412 486 378 220 443 — 273 377

282 244 218 210 142 267 199 282 562 191 — 350 377 217 331 186 195 290 439 266 880 150 125 291 304 489 240 — 220 228

— 466 224 252 186 360 245 322 428 231 136 541 535 — 518 148 — 83 289 48 44 246 122 475 248 149 286 212 277 256

— 551 535 531 441 546 378 461 512 333 215 586 677 1,000 643 281 800 308 452 372 320 418 524 570 456 293 743 519 427 435

— 504 294 327 240 400 300 360 467 266 172 556 566 — 555 200 — 177 339 141 616 300 310 509 312 200 400 285 329 298

278 121 123 225 45 254 194 274 475 123 — 619 304 — 282 145 — — — — — 160 88 — — 72 — — 120 230

472 284 534 492 396 479 345 430 577 234 — 696 517 — 478 331 — — — — — 409 540 — — 266 — — 310 395

333 151 225 326 118 285 220 310 562 160 — 719 351 — 315 229 — — — — — 230 312 — — 138 — — 150 266

40

Supplementary Table 9 | In-hospital mortality prediction performance in the eICU Collaborative Research Database. Per-analyte precision, sensitivity, specificity, and balanced accuracy for PerRI and NORMARI abnormality flags, restricted to measurements classified as normal by PopRI . Analytes with fewer than 100 measurements are omitted (—). PerRI

NORMARI

Analyte

N

Events

Precision

Sensitivity

Specificity

Accuracy

Precision

Sensitivity

Specificity

Accuracy

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

— 891 8,843 7,391 8,035 19,881 16,083 26,608 29,891 27,847 289 9,534 2,919 — 3,383 57,801 — 32,416 22,846 43,792 34,479 40,769 33,546 4,595 16,275 8,985 11 176 4,859 25,827

— 49 1,143 913 793 1,058 1,214 2,149 2,393 1,883 49 696 254 — 264 5,680 — 3,272 1,998 4,335 3,564 3,828 2,580 355 889 1,041 2 34 486 1,714

— 0.05 0.13 0.12 0.10 0.06 0.08 0.08 0.06 0.08 0.06 0.04 0.10 — 0.09 0.11 — 0.11 0.09 0.11 0.12 0.10 0.07 0.10 0.06 0.14 0.29 0.20 0.11 0.07

— 0.20 0.60 0.56 0.53 0.48 0.33 0.37 0.23 0.31 0.06 0.15 0.49 — 0.42 0.38 — 0.49 0.41 0.61 0.59 0.44 0.65 0.34 0.49 0.35 1.00 0.56 0.38 0.42

— 0.75 0.42 0.43 0.48 0.54 0.67 0.62 0.69 0.75 0.80 0.74 0.58 — 0.64 0.66 — 0.54 0.61 0.43 0.48 0.57 0.31 0.73 0.57 0.70 0.44 0.46 0.67 0.58

— 0.48 0.51 0.50 0.50 0.51 0.50 0.49 0.46 0.53 0.43 0.45 0.53 — 0.53 0.52 — 0.51 0.51 0.52 0.53 0.50 0.48 0.54 0.53 0.53 0.72 0.51 0.52 0.50

— 0.08 0.16 0.12 0.11 0.06 0.08 0.10 0.08 0.10 0.04 0.08 0.13 — 0.12 0.12 — 0.12 0.10 0.14 0.11 0.10 0.07 0.10 0.06 0.15 0.67 0.22 0.11 0.07

— 0.18 0.21 0.20 0.17 0.15 0.17 0.16 0.16 0.15 0.02 0.10 0.18 — 0.20 0.20 — 0.29 0.18 0.30 0.91 0.65 0.37 0.15 0.17 0.16 1.00 0.18 0.16 0.15

— 0.88 0.83 0.80 0.85 0.87 0.83 0.88 0.83 0.90 0.90 0.91 0.89 — 0.87 0.83 — 0.76 0.85 0.79 0.13 0.37 0.59 0.88 0.84 0.88 0.89 0.85 0.86 0.87

— 0.53 0.52 0.50 0.51 0.51 0.50 0.52 0.50 0.53 0.46 0.51 0.53 — 0.54 0.52 — 0.53 0.51 0.54 0.52 0.51 0.48 0.52 0.51 0.52 0.94 0.51 0.51 0.51

41

Supplementary Table 10 | Acute kidney injury prediction performance in the eICU Collaborative Research Database. Per-analyte precision, sensitivity, specificity, and balanced accuracy for PerRI and NORMARI abnormality flags, restricted to measurements classified as normal by PopRI . Analytes with fewer than 100 measurements are omitted (—). PerRI

NORMARI

Analyte

N

Events

Precision

Sensitivity

Specificity

Accuracy

Precision

Sensitivity

Specificity

Accuracy

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

— 891 8,843 7,391 8,035 19,881 16,083 26,608 29,891 27,847 289 9,534 2,919 — 3,383 57,801 — 32,416 22,846 43,792 34,479 40,769 33,546 4,595 16,275 8,985 11 176 4,859 25,827

— 130 1,712 1,408 1,384 1,164 2,079 3,839 3,601 2,062 78 1,257 266 — 255 8,390 — 4,932 3,339 6,744 5,045 6,154 4,508 429 1,681 1,660 3 42 864 3,491

— 0.14 0.19 0.17 0.17 0.05 0.13 0.13 0.09 0.07 0.36 0.11 0.10 — 0.09 0.14 — 0.16 0.14 0.15 0.15 0.14 0.13 0.11 0.10 0.19 0.29 0.21 0.19 0.12

— 0.23 0.58 0.51 0.50 0.38 0.34 0.34 0.23 0.23 0.23 0.20 0.47 — 0.43 0.34 — 0.48 0.37 0.56 0.53 0.40 0.65 0.31 0.42 0.31 0.67 0.48 0.36 0.39

— 0.75 0.41 0.42 0.47 0.53 0.67 0.62 0.68 0.74 0.85 0.74 0.58 — 0.64 0.66 — 0.54 0.61 0.43 0.47 0.56 0.30 0.73 0.57 0.70 0.38 0.43 0.67 0.57

— 0.49 0.50 0.47 0.49 0.46 0.51 0.48 0.45 0.48 0.54 0.47 0.53 — 0.54 0.50 — 0.51 0.49 0.49 0.50 0.48 0.48 0.52 0.50 0.51 0.52 0.46 0.52 0.48

— 0.17 0.21 0.18 0.16 0.05 0.15 0.16 0.11 0.05 0.33 0.15 0.10 — 0.08 0.17 — 0.17 0.14 0.16 0.15 0.15 0.13 0.11 0.11 0.20 0.33 0.22 0.21 0.14

— 0.15 0.18 0.19 0.14 0.12 0.20 0.14 0.16 0.06 0.10 0.10 0.13 — 0.14 0.20 — 0.28 0.15 0.23 0.88 0.61 0.39 0.14 0.17 0.13 0.33 0.14 0.17 0.13

— 0.88 0.83 0.80 0.85 0.87 0.83 0.88 0.83 0.89 0.92 0.92 0.89 — 0.87 0.83 — 0.76 0.84 0.78 0.13 0.37 0.59 0.88 0.84 0.88 0.75 0.84 0.87 0.87

— 0.51 0.51 0.49 0.49 0.49 0.52 0.51 0.49 0.48 0.51 0.51 0.51 — 0.50 0.52 — 0.52 0.50 0.50 0.51 0.49 0.49 0.51 0.51 0.51 0.54 0.49 0.52 0.50

42

Supplementary Table 11 | Sepsis prediction performance in the eICU Collaborative Research Database. Peranalyte precision, sensitivity, specificity, and balanced accuracy for PerRI and NORMARI abnormality flags, restricted to measurements classified as normal by PopRI . Analytes with fewer than 100 measurements are omitted (—). PerRI

NORMARI

Analyte

N

Events

Precision

Sensitivity

Specificity

Accuracy

Precision

Sensitivity

Specificity

Accuracy

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

— 891 8,843 7,391 8,035 19,881 16,083 26,608 29,891 27,847 289 9,534 2,919 — 3,383 57,801 — 32,416 22,846 43,792 34,479 40,769 33,546 4,595 16,275 8,985 11 176 4,859 25,827

— 87 2,106 1,925 2,036 2,759 2,034 4,164 4,190 4,045 72 1,557 332 — 363 9,268 — 5,623 3,540 7,707 5,542 6,847 5,676 594 2,064 2,286 2 56 1,118 3,531

— 0.08 0.23 0.25 0.24 0.12 0.13 0.14 0.11 0.14 0.32 0.13 0.11 — 0.10 0.16 — 0.17 0.15 0.17 0.16 0.15 0.17 0.13 0.12 0.26 0.29 0.35 0.24 0.12

— 0.20 0.57 0.54 0.49 0.40 0.34 0.34 0.25 0.25 0.22 0.20 0.40 — 0.35 0.34 — 0.47 0.36 0.57 0.52 0.40 0.67 0.27 0.41 0.30 1.00 0.61 0.35 0.36

— 0.75 0.41 0.42 0.47 0.53 0.67 0.62 0.68 0.74 0.84 0.74 0.57 — 0.63 0.65 — 0.53 0.61 0.43 0.47 0.56 0.31 0.73 0.57 0.70 0.44 0.48 0.67 0.57

— 0.47 0.49 0.48 0.48 0.47 0.51 0.48 0.47 0.50 0.53 0.47 0.49 — 0.49 0.49 — 0.50 0.49 0.50 0.50 0.48 0.49 0.50 0.49 0.50 0.72 0.55 0.51 0.47

— 0.13 0.23 0.25 0.23 0.12 0.14 0.17 0.12 0.13 0.33 0.18 0.12 — 0.12 0.17 — 0.19 0.15 0.19 0.16 0.17 0.17 0.10 0.12 0.25 0.00 0.30 0.26 0.12

— 0.16 0.16 0.19 0.14 0.11 0.19 0.13 0.14 0.10 0.11 0.10 0.12 — 0.15 0.18 — 0.27 0.15 0.23 0.88 0.62 0.41 0.09 0.16 0.12 0.00 0.14 0.16 0.11

— 0.88 0.83 0.80 0.85 0.87 0.83 0.87 0.83 0.90 0.93 0.92 0.89 — 0.87 0.83 — 0.76 0.84 0.79 0.13 0.37 0.59 0.88 0.84 0.88 0.67 0.84 0.87 0.87

— 0.52 0.50 0.49 0.49 0.49 0.51 0.50 0.49 0.49 0.52 0.51 0.50 — 0.51 0.51 — 0.51 0.50 0.51 0.50 0.50 0.50 0.48 0.50 0.50 0.33 0.49 0.51 0.49

43

Supplementary Table 12 | Prolonged ICU stay prediction performance in the eICU Collaborative Research Database. Per-analyte precision, sensitivity, specificity, and balanced accuracy for PerRI and NORMARI abnormality flags, restricted to measurements classified as normal by PopRI . Analytes with fewer than 100 measurements are omitted (—). PerRI

NORMARI

Analyte

N

Events

Precision

Sensitivity

Specificity

Accuracy

Precision

Sensitivity

Specificity

Accuracy

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

— 891 8,843 7,391 8,035 19,881 16,083 26,608 29,891 27,847 289 9,534 2,919 — 3,383 57,801 — 32,416 22,846 43,792 34,479 40,769 33,546 4,595 16,275 8,985 11 176 4,859 25,827

— 158 2,882 2,412 2,561 3,352 2,941 5,167 5,495 5,026 137 718 467 — 486 11,154 — 7,342 4,454 9,765 7,956 8,509 7,568 778 2,917 2,983 1 117 1,495 5,040

— 0.17 0.32 0.32 0.30 0.15 0.20 0.15 0.12 0.19 0.56 0.06 0.16 — 0.13 0.20 — 0.25 0.18 0.22 0.24 0.17 0.22 0.18 0.19 0.35 0.00 0.66 0.33 0.17

— 0.24 0.57 0.55 0.48 0.41 0.35 0.28 0.20 0.27 0.20 0.18 0.41 — 0.33 0.35 — 0.51 0.35 0.57 0.55 0.36 0.67 0.29 0.45 0.32 0.00 0.54 0.36 0.38

— 0.75 0.41 0.42 0.46 0.53 0.68 0.60 0.67 0.75 0.85 0.75 0.57 — 0.63 0.66 — 0.55 0.60 0.43 0.48 0.55 0.30 0.73 0.57 0.71 0.30 0.44 0.67 0.57

— 0.50 0.49 0.48 0.47 0.47 0.51 0.44 0.43 0.51 0.53 0.46 0.49 — 0.48 0.50 — 0.53 0.48 0.50 0.51 0.45 0.49 0.51 0.51 0.51 0.15 0.49 0.52 0.47

— 0.23 0.35 0.33 0.31 0.19 0.23 0.23 0.23 0.19 0.50 0.13 0.17 — 0.15 0.26 — 0.28 0.19 0.27 0.24 0.21 0.27 0.19 0.20 0.35 0.33 0.59 0.37 0.20

— 0.16 0.19 0.20 0.14 0.15 0.21 0.15 0.20 0.11 0.09 0.15 0.12 — 0.14 0.23 — 0.31 0.15 0.27 0.90 0.64 0.48 0.14 0.18 0.13 1.00 0.14 0.17 0.13

— 0.88 0.84 0.80 0.85 0.87 0.84 0.88 0.84 0.90 0.92 0.92 0.89 — 0.87 0.84 — 0.77 0.84 0.80 0.14 0.37 0.62 0.89 0.84 0.88 0.80 0.81 0.87 0.87

— 0.52 0.51 0.50 0.50 0.51 0.53 0.52 0.52 0.50 0.50 0.54 0.51 — 0.50 0.54 — 0.54 0.50 0.53 0.52 0.51 0.55 0.51 0.51 0.51 0.90 0.47 0.52 0.50

44

Supplementary Table 13 | In-hospital mortality prediction performance in the INSPIRE cohort. Per-analyte precision, sensitivity, specificity, and balanced accuracy for PerRI and NORMARI abnormality flags, restricted to measurements classified as normal by PopRI . Analytes with fewer than 100 measurements are omitted (—). PerRI

NORMARI

Analyte

N

Events

Precision

Sensitivity

Specificity

Accuracy

Precision

Sensitivity

Specificity

Accuracy

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

26 15,833 23,817 17,229 24,951 18,100 12,367 19,435 3,345 32,702 — 3,293 9,823 — 11,389 31,429 — — — — — 26,515 27,257 — — 23,415 — — 16,931 23,905

0 62 387 383 513 225 197 310 188 355 — 58 45 — 61 580 — — — — — 367 260 — — 397 — — 178 281

0.00 0.00 0.02 0.02 0.03 0.01 0.02 0.01 0.04 0.03 — 0.02 0.01 — 0.01 0.03 — — — — — 0.02 0.01 — — 0.03 — — 0.01 0.01

— 0.19 0.50 0.31 0.45 0.32 0.28 0.20 0.21 0.35 — 0.31 0.40 — 0.43 0.39 — — — — — 0.50 0.43 — — 0.38 — — 0.21 0.23

0.73 0.81 0.53 0.65 0.63 0.69 0.82 0.77 0.68 0.89 — 0.78 0.69 — 0.73 0.76 — — — — — 0.69 0.49 — — 0.79 — — 0.78 0.78

— 0.50 0.51 0.48 0.54 0.51 0.55 0.48 0.44 0.62 — 0.54 0.55 — 0.58 0.57 — — — — — 0.60 0.46 — — 0.59 — — 0.49 0.51

0.00 0.01 0.02 0.02 0.03 0.03 0.07 0.04 0.08 0.07 — 0.02 0.01 — 0.01 0.07 — — — — — 0.07 0.01 — — 0.02 — — 0.01 0.02

— 0.07 0.17 0.14 0.12 0.10 0.12 0.12 0.35 0.18 — 0.28 0.20 — 0.07 0.30 — — — — — 0.33 0.28 — — 0.09 — — 0.03 0.09

0.92 0.96 0.89 0.87 0.92 0.96 0.97 0.95 0.76 0.97 — 0.75 0.94 — 0.96 0.92 — — — — — 0.94 0.75 — — 0.93 — — 0.96 0.95

— 0.52 0.53 0.51 0.52 0.53 0.55 0.54 0.55 0.58 — 0.51 0.57 — 0.51 0.61 — — — — — 0.63 0.52 — — 0.51 — — 0.50 0.52

45

Supplementary Table 14 | Perioperative infection prediction performance in the INSPIRE cohort. Per-analyte precision, sensitivity, specificity, and balanced accuracy for PerRI and NORMARI abnormality flags, restricted to measurements classified as normal by PopRI . Analytes with fewer than 100 measurements are omitted (—). PerRI

NORMARI

Analyte

N

Events

Precision

Sensitivity

Specificity

Accuracy

Precision

Sensitivity

Specificity

Accuracy

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

26 15,833 23,817 17,229 24,951 18,100 12,367 19,435 3,345 32,702 — 3,293 9,823 — 11,389 31,429 — — — — — 26,515 27,257 — — 23,415 — — 16,931 23,905

2 437 1,174 1,005 1,439 736 481 854 421 1,124 — 161 201 — 277 1,374 — — — — — 975 1,137 — — 1,292 — — 763 1,043

0.14 0.03 0.04 0.05 0.05 0.03 0.03 0.03 0.09 0.04 — 0.04 0.02 — 0.02 0.04 — — — — — 0.03 0.03 — — 0.07 — — 0.04 0.03

0.50 0.17 0.37 0.29 0.33 0.23 0.15 0.14 0.22 0.14 — 0.19 0.31 — 0.24 0.22 — — — — — 0.28 0.36 — — 0.25 — — 0.19 0.17

0.75 0.81 0.52 0.64 0.62 0.68 0.81 0.76 0.67 0.89 — 0.77 0.69 — 0.73 0.76 — — — — — 0.69 0.49 — — 0.79 — — 0.78 0.78

0.62 0.49 0.44 0.46 0.48 0.46 0.48 0.45 0.44 0.51 — 0.48 0.50 — 0.49 0.49 — — — — — 0.49 0.42 — — 0.52 — — 0.48 0.48

0.00 0.03 0.04 0.05 0.06 0.02 0.06 0.04 0.12 0.05 — 0.06 0.03 — 0.03 0.07 — — — — — 0.07 0.03 — — 0.05 — — 0.08 0.04

0.00 0.04 0.10 0.12 0.09 0.02 0.04 0.04 0.24 0.04 — 0.30 0.08 — 0.05 0.13 — — — — — 0.12 0.20 — — 0.07 — — 0.06 0.04

0.92 0.96 0.89 0.87 0.92 0.96 0.97 0.95 0.75 0.97 — 0.75 0.94 — 0.96 0.92 — — — — — 0.94 0.75 — — 0.93 — — 0.97 0.95

0.46 0.50 0.49 0.49 0.50 0.49 0.51 0.50 0.50 0.51 — 0.53 0.51 — 0.51 0.53 — — — — — 0.53 0.47 — — 0.50 — — 0.51 0.50

46

Supplementary Table 15 | Prolonged hospital stay prediction performance in the INSPIRE cohort. Per-analyte precision, sensitivity, specificity, and balanced accuracy for PerRI and NORMARI abnormality flags, restricted to measurements classified as normal by PopRI . Analytes with fewer than 100 measurements are omitted (—). PerRI

NORMARI

Analyte

N

Events

Precision

Sensitivity

Specificity

Accuracy

Precision

Sensitivity

Specificity

Accuracy

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

26 15,833 23,817 17,229 24,951 18,100 12,367 19,435 3,345 32,702 — 3,293 9,823 — 11,389 31,429 — — — — — 26,515 27,257 — — 23,415 — — 16,931 23,905

19 10,599 18,548 13,384 19,532 13,840 8,496 15,298 2,885 22,812 — 2,352 5,503 — 6,945 23,052 — — — — — 18,858 19,942 — — 17,682 — — 12,395 17,474

0.71 0.68 0.78 0.76 0.76 0.74 0.67 0.74 0.83 0.70 — 0.67 0.52 — 0.60 0.71 — — — — — 0.69 0.71 — — 0.75 — — 0.73 0.69

0.26 0.19 0.48 0.34 0.36 0.30 0.18 0.22 0.31 0.12 — 0.21 0.29 — 0.27 0.24 — — — — — 0.30 0.50 — — 0.21 — — 0.22 0.21

0.71 0.81 0.53 0.61 0.59 0.65 0.81 0.72 0.62 0.89 — 0.74 0.66 — 0.72 0.73 — — — — — 0.67 0.46 — — 0.78 — — 0.78 0.75

0.49 0.50 0.50 0.48 0.48 0.48 0.49 0.47 0.46 0.50 — 0.48 0.47 — 0.49 0.48 — — — — — 0.48 0.48 — — 0.50 — — 0.50 0.48

1.00 0.66 0.70 0.74 0.72 0.70 0.65 0.73 0.89 0.67 — 0.74 0.49 — 0.53 0.72 — — — — — 0.68 0.68 — — 0.70 — — 0.74 0.76

0.10 0.03 0.10 0.13 0.07 0.04 0.03 0.04 0.25 0.03 — 0.26 0.06 — 0.04 0.08 — — — — — 0.06 0.23 — — 0.07 — — 0.04 0.05

1.00 0.96 0.85 0.84 0.90 0.95 0.97 0.94 0.80 0.97 — 0.78 0.93 — 0.95 0.92 — — — — — 0.93 0.71 — — 0.91 — — 0.97 0.96

0.55 0.50 0.47 0.48 0.48 0.49 0.50 0.49 0.53 0.50 — 0.52 0.49 — 0.49 0.50 — — — — — 0.49 0.47 — — 0.49 — — 0.50 0.50

47

Supplementary Table 16 | Unplanned ICU admission prediction performance in the INSPIRE cohort. Per-analyte precision, sensitivity, specificity, and balanced accuracy for PerRI and NORMARI abnormality flags, restricted to measurements classified as normal by PopRI . Analytes with fewer than 100 measurements are omitted (—). PerRI

NORMARI

Analyte

N

Events

Precision

Sensitivity

Specificity

Accuracy

Precision

Sensitivity

Specificity

Accuracy

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

26 15,833 23,817 17,229 24,951 18,100 12,367 19,435 3,345 32,702 — 3,293 9,823 — 11,389 31,429 — — — — — 26,515 27,257 — — 23,415 — — 16,931 23,905

15 4,067 7,129 5,695 7,856 5,436 3,418 5,038 1,870 6,625 — 622 1,745 — 2,448 8,519 — — — — — 6,473 6,866 — — 6,578 — — 4,750 6,052

0.43 0.22 0.28 0.31 0.27 0.29 0.23 0.19 0.51 0.21 — 0.14 0.15 — 0.21 0.25 — — — — — 0.20 0.20 — — 0.28 — — 0.23 0.19

0.20 0.16 0.45 0.34 0.32 0.30 0.15 0.17 0.29 0.12 — 0.16 0.27 — 0.27 0.22 — — — — — 0.26 0.41 — — 0.21 — — 0.18 0.17

0.64 0.80 0.52 0.64 0.60 0.68 0.81 0.74 0.65 0.89 — 0.76 0.68 — 0.73 0.75 — — — — — 0.68 0.46 — — 0.79 — — 0.76 0.76

0.42 0.48 0.48 0.49 0.46 0.49 0.48 0.46 0.47 0.50 — 0.46 0.47 — 0.50 0.48 — — — — — 0.47 0.44 — — 0.50 — — 0.47 0.47

1.00 0.20 0.26 0.35 0.25 0.33 0.20 0.34 0.57 0.19 — 0.22 0.25 — 0.24 0.32 — — — — — 0.32 0.22 — — 0.27 — — 0.22 0.23

0.13 0.03 0.10 0.14 0.06 0.04 0.02 0.06 0.25 0.03 — 0.29 0.09 — 0.05 0.09 — — — — — 0.09 0.21 — — 0.07 — — 0.03 0.04

1.00 0.96 0.88 0.87 0.92 0.96 0.97 0.96 0.76 0.97 — 0.76 0.94 — 0.96 0.93 — — — — — 0.94 0.74 — — 0.93 — — 0.96 0.95

0.57 0.49 0.49 0.51 0.49 0.50 0.49 0.51 0.51 0.50 — 0.52 0.52 — 0.50 0.51 — — — — — 0.51 0.48 — — 0.50 — — 0.49 0.50

48

Supplementary Table 17 | Positive predictive value across outcomes in Clalit Health Services. Measurements classified as normal by PopRI are excluded. Per 100 patients flagged abnormal by PerRI or NORMARI , the number who experienced each outcome (anemia, chronic kidney disease, all-cause mortality, and type 2 diabetes), by analyte. Analytes with fewer than 100 measurements are omitted (—). Anemia

Chronic Kidney Disease

Mortality

Type 2 Diabetes

Analyte

PerRI

NORMARI

PerRI

NORMARI

PerRI

NORMARI

PerRI

NORMARI

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

21 21 19 19 19 14 21 30 100 16 — 16 10 17 9 21 20 16 14 17 19 19 18 13 14 24 21 — 23 18

30 25 24 22 23 20 27 38 100 24 — 23 12 21 12 25 24 18 17 20 18 26 25 18 15 22 26 — 27 22

39 44 40 40 40 26 45 52 75 34 — 30 28 38 28 41 43 38 30 37 37 41 35 30 32 46 43 — 47 37

48 51 47 43 45 34 53 62 100 57 — 38 34 42 33 48 50 41 34 40 36 51 45 36 33 44 50 — 54 41

23 26 25 24 26 18 28 36 50 21 — 18 15 22 15 27 26 23 18 23 24 26 23 17 18 30 27 — 30 23

34 33 30 27 29 24 35 46 80 32 — 26 19 26 19 34 30 27 22 26 24 34 32 22 19 29 31 — 37 27

70 78 74 74 74 73 77 93 100 75 — 53 70 71 70 78 79 71 73 73 74 76 73 71 71 79 80 — 78 73

78 81 79 78 79 80 82 97 100 84 — 64 75 74 76 83 86 74 77 75 73 84 79 78 72 77 87 — 81 78

Overall

21

25

39

46

24

30

75

80

49

Supplementary Table 18 | Positive predictive value across outcomes in the eICU Collaborative Research Database. Measurements classified as normal by PopRI are excluded. Per 100 patients flagged abnormal by PerRI or NORMARI , the number who experienced each outcome (acute kidney injury, in-hospital mortality, prolonged ICU stay, and sepsis), by analyte. Analytes with fewer than 100 measurements are omitted (—). Acute Kidney Injury

Mortality

Prolonged LOS (>7d)

Sepsis

Analyte

PerRI

NORMARI

PerRI

NORMARI

PerRI

NORMARI

PerRI

NORMARI

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

— 14 19 17 16 5 13 13 9 7 36 11 10 — 9 14 — 16 14 15 15 14 13 11 10 19 29 21 19 12

— 17 21 18 16 5 15 16 11 5 33 16 10 — 8 17 — 17 14 16 15 15 13 11 11 20 33 22 22 14

— 5 13 12 10 6 8 8 6 8 6 4 10 — 9 11 — 11 9 11 12 10 7 10 6 14 29 20 11 7

— 8 16 12 11 6 8 10 8 10 4 8 13 — 12 12 — 12 10 14 11 10 7 10 6 15 67 22 11 8

— 17 32 32 30 15 20 15 12 19 56 6 16 — 13 20 — 25 18 22 24 17 22 18 19 35 0 66 33 17

— 23 35 33 31 19 23 23 23 19 50 13 17 — 15 26 — 28 19 27 24 21 27 19 20 35 33 59 37 20

— 8 23 25 24 12 13 14 11 14 32 13 11 — 10 16 — 18 15 18 16 15 16 13 12 26 29 35 24 12

— 13 23 25 23 12 14 16 12 13 33 18 12 — 12 17 — 19 15 19 16 17 17 10 12 25 0 30 26 12

Overall

15

16

10

13

23

27

18

17

50

Supplementary Table 19 | Positive predictive value across outcomes in the INSPIRE cohort. Measurements classified as normal by PopRI are excluded. Per 100 patients flagged abnormal by PerRI or NORMARI , the number who experienced each outcome (in-hospital mortality, perioperative infection, prolonged hospital stay, and unplanned ICU admission), by analyte. Analytes with fewer than 100 measurements are omitted (—). Mortality

Perioperative Infection

Prolonged LOS (>7d)

Unplanned ICU Admission

Analyte

PerRI

NORMARI

PerRI

NORMARI

PerRI

NORMARI

PerRI

NORMARI

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

0 0 2 2 2 1 2 1 4 3 — 2 1 — 1 3 — — — — — 2 1 — — 3 — — 1 1

0 1 2 2 3 3 7 4 8 7 — 2 1 — 1 7 — — — — — 7 1 — — 2 — — 1 2

14 2 4 5 5 3 3 3 9 4 — 4 2 — 2 4 — — — — — 3 3 — — 7 — — 4 3

0 3 4 5 6 2 6 4 12 5 — 6 2 — 3 7 — — — — — 7 3 — — 5 — — 8 4

71 68 78 76 76 74 67 74 83 70 — 67 52 — 60 71 — — — — — 69 71 — — 75 — — 73 69

100 66 70 74 72 70 65 73 89 67 — 74 49 — 52 72 — — — — — 68 68 — — 70 — — 74 76

43 22 28 31 27 29 23 19 51 21 — 14 15 — 21 25 — — — — — 20 20 — — 28 — — 22 19

100 20 26 35 25 33 20 34 57 19 — 22 25 — 24 32 — — — — — 32 22 — — 27 — — 22 23

Overall

2

3

4

5

71

71

25

31

51

Supplementary Table 20 | Lead time for NORMARI abnormality detection relative to PopRI . Per analyte and cohort: number of tests flagged by NORMARI alone (NORMARI -only) and the subset later confirmed by PopRI (later population), with median lead time and interquartile range (hours for eICU-CRD and INSPIRE, months for CHS). eICU-CRD

CHS

INSPIRE

Analyte

NORMA-only

Later PopRI

Median (h)

IQR (h)

NORMA-only

Later PopRI

Median (mo)

IQR (mo)

NORMA-only

Later PopRI

Median (h)

IQR (h)

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC Overall

— 2,185 4,362 6,749 7,720 25,578 45,543 24,313 52,641 40,464 425 20,333 2,406 2 4,408 62,211 3 40,943 38,761 19,132 101,702 115,529 46,110 6,592 18,706 11,136 28 317 11,311 31,253 740,863

— 865 757 1,235 1,317 7,039 13,222 8,477 26,365 8,744 72 17,921 998 0 1,526 16,970 2 5,361 11,733 1,789 4,805 31,882 5,482 2,129 3,308 1,721 7 38 2,592 8,381 184,738

— 30.0 49.6 47.6 47.4 28.0 25.1 34.7 25.6 26.0 25.2 5.6 24.2 — 23.7 40.9 35.1 45.5 25.1 48.2 37.5 26.6 49.7 23.9 26.3 47.4 70.3 60.0 34.8 35.3 34.8

— 23.1–73.9 24.2–120.2 23.9–99.4 23.9–95.5 23.1–71.3 19.6–69.2 23.0–72.8 19.3–64.5 21.6–67.7 20.8–49.3 3.2–12.1 17.7–48.1 — 14.3–47.1 23.0–78.6 29.2–41.0 23.4–93.1 22.5–49.9 23.7–120.7 23.2–72.5 21.1–69.9 24.0–120.8 16.5–47.6 22.8–70.0 23.8–95.5 42.0–120.9 25.4–145.5 22.9–74.1 23.3–72.7 23.1–72.6

14,227 807,619 1,196,400 1,199,702 1,195,620 568,660 574,875 44,929 96 1,070,918 — 358,479 1,083,585 515,071 1,065,707 1,337,595 764,300 2,300,580 840,255 4,464,421 20,300,446 805,446 1,356,487 1,202,775 1,580,050 2,966,137 675,432 — 962,840 2,183,898 51,436,550

5,064 436,471 543,146 703,509 551,419 402,995 356,694 36,351 80 687,844 — 248,307 762,204 120,646 585,178 881,518 172,816 1,420,584 776,047 1,941,232 8,370,452 506,896 660,368 769,663 933,310 1,268,891 330,110 — 530,403 1,402,700 25,404,898

12.4 7.9 15.5 11.9 10.0 7.6 6.4 4.9 2.9 7.5 — 5.5 8.5 21.9 16.1 4.5 17.6 11.2 7.6 18.8 20.0 8.0 8.5 8.9 16.2 26.4 14.8 — 6.9 7.4 8.7

5.7–27.5 1.2–29.9 3.1–50.6 2.4–41.0 1.6–38.0 1.1–29.2 0.8–29.0 0.5–22.7 0.3–8.7 1.4–27.1 — 0.9–20.6 1.8–30.2 7.3–53.3 3.4–44.3 0.4–25.8 6.4–44.5 2.8–37.1 1.6–29.9 4.1–62.8 4.8–53.3 0.9–34.7 1.1–40.3 1.8–32.9 3.1–52.3 6.0–68.5 5.0–39.6 — 0.9–28.6 1.0–33.0 1.7–33.8

24 1,493 3,442 3,210 2,204 1,983 1,419 3,124 3,787 2,321 — 14,216 2,240 — 1,576 5,722 — — — — — 4,825 8,493 — — 2,742 — — 1,227 2,389 66,437

0 29 201 187 29 409 342 932 2,418 826 — 8,944 484 — 301 2,611 — — — — — 1,308 277 — — 71 — — 18 348 19,735

— 23.5 60.4 47.6 16.6 24.9 39.8 24.0 8.4 43.2 — 4.6 10.4 — 14.8 15.7 — — — — — 29.8 23.3 — — 23.8 — — 13.0 26.0 23.7

— 13.3–25.7 24.7–102.3 19.2–109.0 8.2–24.0 13.0–55.2 16.3–77.0 11.4–54.2 3.5–19.2 16.1–121.8 — 2.5–8.7 3.4–21.5 — 5.7–32.6 5.8–37.2 — — — — — 11.3–75.0 8.9–64.1 — — 9.7–83.0 — — 11.9–24.0 13.8–57.4 11.4–54.7

52

Supplementary Table 21 | Proportional hazards analysis for all-cause mortality in Clalit Health Services. Per-analyte hazard ratios (with 95% confidence intervals) and concordance indices for PopRI , PerRI , and NORMARI abnormality flags, adjusted for age and sex. PopRI

PerRI

NORMARI

Analyte

N train (events)

N test (events)

HR [95% CI]

C

HR [95% CI]

C

HR [95% CI]

C

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

9517 (3118) 13612 (4221) 20262 (5615) 25254 (6667) 25296 (6662) 26980 (7187) 15946 (4688) 846 (307) — 29331 (7624) — 28928 (7343) 31026 (7576) 20800 (5296) 31020 (7578) 25819 (7121) 23706 (6093) 28147 (6984) 31009 (7580) 30848 (7546) 28679 (7094) 25977 (7133) 30976 (7564) 30946 (7556) 30874 (7528) 13506 (3888) 25805 (6534) — 13638 (4217) 31053 (7583)

6345 (2079) 9076 (2815) 13508 (3743) 16836 (4445) 16864 (4442) 17987 (4792) 10631 (3125) 564 (205) — 19555 (5083) — 19286 (4895) 20685 (5051) 13867 (3530) 20680 (5052) 17213 (4748) 15804 (4062) 18765 (4656) 20673 (5053) 20566 (5031) 19120 (4729) 17318 (4756) 20652 (5043) 20631 (5038) 20583 (5018) 9005 (2593) 17204 (4357) — 9092 (2812) 20702 (5055)

0.84 [0.65, 1.09] 4.15 [3.83, 4.49] 2.05 [1.94, 2.17] 2.08 [1.95, 2.22] 1.85 [1.76, 1.94] 2.82 [2.53, 3.15] 1.93 [1.81, 2.07] 3.57 [1.67, 7.61] — 3.00 [2.82, 3.18] — 1.60 [1.39, 1.85] 3.60 [3.15, 4.12] 1.68 [1.56, 1.81] 3.70 [3.32, 4.12] 1.63 [1.53, 1.73] 0.81 [0.77, 0.85] 2.04 [1.89, 2.20] 1.20 [0.89, 1.61] 1.70 [1.62, 1.78] 1.28 [1.22, 1.35] 2.17 [2.04, 2.31] 1.98 [1.89, 2.08] — 3.09 [2.82, 3.39] 1.61 [1.50, 1.72] 0.93 [0.87, 0.99] — 2.23 [2.08, 2.38] 1.86 [1.75, 1.97]

0.689 0.785 0.760 0.761 0.758 0.747 0.754 0.767 — 0.770 — 0.749 0.774 0.733 0.779 0.748 0.735 0.767 0.767 0.774 0.765 0.753 0.779 — 0.774 0.744 0.739 — 0.756 0.773

0.58 [0.54, 0.63] 1.21 [1.10, 1.33] 0.96 [0.87, 1.05] 0.94 [0.86, 1.03] 1.02 [0.94, 1.12] 0.81 [0.75, 0.88] 1.08 [0.99, 1.17] 1.90 [0.97, 3.71] — 1.04 [0.97, 1.12] — 0.77 [0.70, 0.84] 0.79 [0.73, 0.86] 0.73 [0.67, 0.81] 0.83 [0.77, 0.90] 1.07 [0.99, 1.15] 0.74 [0.69, 0.79] 0.98 [0.91, 1.07] 0.96 [0.90, 1.03] 0.90 [0.83, 0.97] 0.93 [0.87, 0.99] 1.03 [0.95, 1.12] 0.92 [0.83, 1.01] 0.78 [0.71, 0.85] 0.90 [0.83, 0.97] 1.13 [1.04, 1.22] 0.80 [0.75, 0.86] — 1.07 [0.98, 1.16] 0.98 [0.89, 1.07]

0.699 0.740 0.741 0.749 0.744 0.739 0.740 0.760 — 0.744 — 0.748 0.765 0.726 0.764 0.742 0.736 0.760 0.767 0.762 0.762 0.738 0.765 0.764 0.762 0.737 0.739 — 0.736 0.763

0.82 [0.65, 1.03] 3.11 [2.80, 3.46] 2.01 [1.86, 2.18] 1.35 [1.27, 1.43] 1.55 [1.47, 1.65] 2.58 [2.31, 2.87] 1.88 [1.73, 2.05] 4.10 [1.82, 9.27] — 3.30 [3.07, 3.54] — 1.61 [1.38, 1.86] 3.36 [2.86, 3.94] 1.72 [1.57, 1.89] 3.72 [3.24, 4.26] 1.69 [1.56, 1.83] 0.87 [0.82, 0.92] 2.18 [1.90, 2.50] 0.99 [0.76, 1.28] 1.95 [1.77, 2.16] 0.10 [0.04, 0.23] 2.02 [1.88, 2.16] 1.88 [1.77, 1.98] — 2.23 [1.96, 2.52] 1.09 [0.98, 1.21] 0.96 [0.90, 1.03] — 2.73 [2.46, 3.04] 1.90 [1.75, 2.07]

0.688 0.757 0.749 0.752 0.750 0.746 0.749 0.755 — 0.768 — 0.749 0.771 0.730 0.772 0.746 0.734 0.763 0.767 0.766 0.762 0.750 0.775 — 0.765 0.738 0.739 — 0.745 0.768

53

Supplementary Table 22 | Proportional hazards analysis for in-hospital mortality in the eICU Collaborative Research Database. Per-analyte hazard ratios (with 95% confidence intervals) and concordance indices for PopRI , PerRI , and NORMARI abnormality flags, adjusted for age and sex. PopRI

PerRI

NORMARI

Analyte

N train (events)

N test (events)

HR [95% CI]

C

HR [95% CI]

C

HR [95% CI]

C

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

— 11151 (1560) 7966 (1118) 7288 (976) 7446 (912) 33379 (3378) 33080 (3474) 34320 (3560) 40679 (4792) 31723 (3152) 684 (122) 52615 (5182) 34809 (3553) — 35018 (3622) 38603 (3943) — 28067 (2889) 29987 (3063) 29488 (3017) 21532 (2191) 36226 (3839) 30955 (3214) 31071 (3221) 28255 (2852) 7465 (930) — 186 (37) 8881 (1307) 30250 (2917)

— 7435 (1040) 5311 (746) 4859 (650) 4964 (608) 22253 (2252) 22054 (2316) 22881 (2374) 27120 (3194) 21149 (2102) 456 (82) 35077 (3455) 23206 (2368) — 23346 (2415) 25736 (2628) — 18712 (1926) 19992 (2042) 19659 (2011) 14355 (1461) 24152 (2559) 20638 (2142) 20714 (2148) 18838 (1901) 4978 (620) — 124 (24) 5921 (871) 20167 (1945)

— 1.04 [0.75, 1.43] 0.97 [0.86, 1.09] 0.98 [0.86, 1.11] 1.68 [1.47, 1.93] 1.74 [1.58, 1.92] 1.81 [1.64, 1.98] 1.52 [1.40, 1.65] 1.94 [1.76, 2.13] 1.89 [1.74, 2.05] 1.23 [0.81, 1.87] 2.14 [1.73, 2.65] 1.11 [0.89, 1.38] — 1.01 [0.82, 1.23] 1.34 [1.26, 1.43] — 0.99 [0.92, 1.07] 0.85 [0.78, 0.92] 1.26 [1.16, 1.37] 0.95 [0.84, 1.07] 1.22 [1.14, 1.31] 2.50 [2.32, 2.68] 1.05 [0.88, 1.25] 1.85 [1.65, 2.07] 1.39 [1.22, 1.59] — 1.09 [0.54, 2.18] 1.67 [1.44, 1.93] 2.39 [2.17, 2.63]

— 0.544 0.534 0.538 0.571 0.573 0.577 0.572 0.586 0.602 0.546 0.565 0.558 — 0.561 0.563 — 0.558 0.563 0.559 0.568 0.551 0.654 0.562 0.601 0.552 — 0.594 0.576 0.619

— 1.01 [0.65, 1.57] 1.09 [0.93, 1.28] 1.05 [0.89, 1.24] 1.82 [1.52, 2.19] 1.84 [1.59, 2.14] 1.68 [1.48, 1.91] 1.69 [1.51, 1.89] 1.82 [1.60, 2.08] 1.97 [1.78, 2.17] 1.07 [0.69, 1.65] 1.65 [1.26, 2.15] 1.45 [0.98, 2.15] — 1.11 [0.83, 1.50] 1.31 [1.22, 1.41] — 0.95 [0.88, 1.03] 0.87 [0.78, 0.97] 1.35 [1.25, 1.47] 1.17 [1.08, 1.28] 1.37 [1.25, 1.50] 1.77 [1.58, 1.99] 1.07 [0.85, 1.34] 2.23 [1.89, 2.63] 1.28 [1.12, 1.47] — 1.03 [0.43, 2.48] 1.68 [1.37, 2.06] 2.48 [2.16, 2.85]

— 0.544 0.534 0.539 0.562 0.559 0.564 0.568 0.566 0.599 0.512 0.563 0.556 — 0.561 0.563 — 0.558 0.558 0.564 0.575 0.557 0.570 0.561 0.589 0.550 — 0.590 0.554 0.598

— 1.14 [0.78, 1.65] 1.12 [0.99, 1.26] 0.99 [0.87, 1.13] 1.61 [1.40, 1.86] 1.77 [1.59, 1.97] 1.75 [1.58, 1.94] 1.58 [1.44, 1.72] 1.90 [1.71, 2.11] 1.98 [1.81, 2.16] 1.14 [0.75, 1.75] 2.19 [1.75, 2.74] 1.31 [1.01, 1.71] — 1.07 [0.85, 1.35] 1.36 [1.27, 1.45] — 0.99 [0.92, 1.06] 0.89 [0.82, 0.98] 1.35 [1.25, 1.45] 1.20 [1.02, 1.40] 1.22 [1.11, 1.35] 1.67 [1.54, 1.82] 1.10 [0.90, 1.33] 1.91 [1.68, 2.16] 1.38 [1.22, 1.58] — 1.22 [0.59, 2.49] 1.66 [1.42, 1.95] 2.41 [2.17, 2.67]

— 0.544 0.535 0.537 0.575 0.571 0.571 0.574 0.576 0.604 0.537 0.565 0.558 — 0.561 0.564 — 0.558 0.561 0.569 0.571 0.552 0.579 0.561 0.595 0.552 — 0.598 0.562 0.617

54

Supplementary Table 23 | Proportional hazards analysis for acute kidney injury in the eICU Collaborative Research Database. Per-analyte hazard ratios (with 95% confidence intervals) and concordance indices for PopRI , PerRI , and NORMARI abnormality flags, adjusted for age and sex. PopRI

PerRI

NORMARI

Analyte

N train (events)

N test (events)

HR [95% CI]

C

HR [95% CI]

C

HR [95% CI]

C

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

— 11121 (2490) 7949 (1702) 7267 (1472) 7429 (1464) 33279 (4964) 32980 (5407) 34213 (5471) 40566 (6164) 31638 (4454) 682 (211) 52476 (6866) 34716 (5015) — 34923 (5009) 38496 (5647) — 27981 (4313) 29898 (4621) 29400 (4528) 21459 (3144) 36123 (5599) 30866 (4735) 30981 (4736) 28174 (4340) 7447 (1482) 20 (4) 186 (54) 8860 (1890) 30162 (4532)

— 7414 (1660) 5300 (1134) 4846 (982) 4954 (976) 22187 (3310) 21987 (3604) 22809 (3647) 27044 (4109) 21092 (2969) 456 (141) 34984 (4577) 23144 (3343) — 23282 (3339) 25664 (3764) — 18654 (2875) 19933 (3080) 19601 (3018) 14307 (2096) 24082 (3733) 20578 (3156) 20654 (3158) 18783 (2894) 4966 (988) 14 (2) 124 (36) 5908 (1260) 20108 (3021)

— 1.28 [1.01, 1.63] 1.25 [1.14, 1.37] 1.08 [0.97, 1.19] 1.19 [1.08, 1.32] 1.89 [1.76, 2.04] 1.43 [1.34, 1.53] 1.43 [1.34, 1.52] 1.40 [1.31, 1.50] 1.77 [1.66, 1.89] 1.32 [0.97, 1.79] 1.06 [0.94, 1.19] 1.91 [1.55, 2.36] — 1.97 [1.62, 2.40] 1.40 [1.33, 1.48] — 1.12 [1.05, 1.19] 1.17 [1.10, 1.25] 1.12 [1.04, 1.21] 1.30 [1.18, 1.43] 1.13 [1.07, 1.19] 1.38 [1.30, 1.46] 1.75 [1.50, 2.04] 1.73 [1.60, 1.87] 1.15 [1.04, 1.28] 0.52 [0.07, 3.77] 1.18 [0.67, 2.09] 1.28 [1.15, 1.43] 1.27 [1.19, 1.35]

— 0.516 0.544 0.504 0.529 0.566 0.541 0.540 0.534 0.576 0.520 0.519 0.524 — 0.527 0.557 — 0.528 0.531 0.523 0.535 0.527 0.543 0.520 0.571 0.526 0.640 0.551 0.517 0.537

— 1.45 [1.03, 2.03] 1.12 [0.99, 1.26] 0.94 [0.82, 1.07] 1.07 [0.95, 1.21] 1.63 [1.47, 1.82] 1.40 [1.28, 1.52] 1.24 [1.14, 1.34] 1.35 [1.23, 1.49] 1.55 [1.44, 1.66] 1.32 [0.94, 1.86] 0.96 [0.82, 1.13] 2.59 [1.78, 3.79] — 2.18 [1.61, 2.96] 1.33 [1.26, 1.41] — 1.08 [1.01, 1.16] 1.14 [1.05, 1.25] 1.02 [0.96, 1.08] 1.07 [1.00, 1.15] 1.06 [0.99, 1.13] 1.03 [0.95, 1.12] 1.71 [1.40, 2.09] 1.65 [1.50, 1.83] 1.11 [1.00, 1.23] — 0.97 [0.47, 2.01] 1.34 [1.15, 1.55] 1.17 [1.08, 1.26]

— 0.515 0.516 0.505 0.504 0.535 0.533 0.525 0.524 0.555 0.520 0.518 0.521 — 0.524 0.544 — 0.527 0.529 0.522 0.523 0.528 0.517 0.521 0.553 0.527 0.600 0.515 0.510 0.524

— 1.36 [1.04, 1.79] 1.24 [1.12, 1.36] 1.06 [0.95, 1.18] 1.14 [1.03, 1.27] 1.73 [1.60, 1.87] 1.39 [1.29, 1.49] 1.43 [1.34, 1.52] 1.33 [1.24, 1.44] 1.59 [1.49, 1.69] 1.40 [1.01, 1.93] 1.04 [0.92, 1.17] 1.97 [1.56, 2.49] — 2.01 [1.61, 2.50] 1.39 [1.32, 1.47] — 1.14 [1.07, 1.21] 1.13 [1.05, 1.21] 1.06 [1.00, 1.13] 1.10 [0.98, 1.24] 1.01 [0.94, 1.08] 1.15 [1.08, 1.22] 1.78 [1.50, 2.11] 1.71 [1.58, 1.86] 1.12 [1.01, 1.24] 0.89 [0.08, 10.23] 1.02 [0.58, 1.82] 1.31 [1.16, 1.48] 1.28 [1.20, 1.36]

— 0.515 0.534 0.504 0.522 0.555 0.539 0.538 0.528 0.560 0.523 0.519 0.524 — 0.526 0.554 — 0.528 0.530 0.523 0.526 0.525 0.524 0.521 0.568 0.532 0.640 0.525 0.518 0.531

55

Supplementary Table 24 | Proportional hazards analysis for sepsis in the eICU Collaborative Research Database. Per-analyte hazard ratios (with 95% confidence intervals) and concordance indices for PopRI , PerRI , and NORMARI abnormality flags, adjusted for age and sex. PopRI

PerRI

NORMARI

Analyte

N train (events)

N test (events)

HR [95% CI]

C

HR [95% CI]

C

HR [95% CI]

C

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

— 11106 (2754) 7938 (2020) 7256 (1877) 7415 (1904) 33240 (5690) 32941 (5802) 34176 (5909) 40521 (6798) 31595 (5314) 682 (186) 52429 (7603) 34671 (5744) — 34876 (5787) 38461 (6288) — 27943 (5010) 29860 (5363) 29362 (5220) 21432 (3423) 36085 (6030) 30826 (5403) 30939 (5515) 28135 (5022) 7437 (1878) 20 (6) 186 (63) 8848 (2278) 30127 (5170)

— 7405 (1836) 5292 (1347) 4838 (1252) 4944 (1270) 22161 (3793) 21962 (3868) 22785 (3939) 27015 (4532) 21064 (3542) 456 (125) 34953 (5068) 23114 (3829) — 23252 (3858) 25642 (4192) — 18630 (3340) 19908 (3575) 19575 (3480) 14289 (2282) 24058 (4020) 20552 (3603) 20626 (3677) 18758 (3349) 4959 (1252) 14 (5) 124 (42) 5899 (1518) 20086 (3447)

— 2.27 [1.72, 3.00] 1.17 [1.07, 1.28] 1.01 [0.92, 1.11] 1.08 [0.98, 1.18] 1.22 [1.15, 1.30] 1.53 [1.44, 1.63] 1.34 [1.26, 1.41] 1.26 [1.18, 1.34] 1.28 [1.21, 1.35] 0.98 [0.72, 1.33] 0.98 [0.88, 1.09] 1.50 [1.26, 1.79] — 1.50 [1.28, 1.76] 1.48 [1.40, 1.55] — 1.20 [1.14, 1.27] 1.32 [1.24, 1.41] 1.24 [1.16, 1.33] 1.27 [1.16, 1.39] 1.05 [1.00, 1.11] 1.16 [1.10, 1.23] 1.38 [1.22, 1.57] 1.83 [1.70, 1.96] 1.02 [0.93, 1.12] 0.71 [0.13, 3.89] 0.95 [0.56, 1.59] 1.07 [0.97, 1.18] 1.48 [1.39, 1.57]

— 0.519 0.544 0.510 0.502 0.533 0.549 0.539 0.540 0.547 0.513 0.527 0.528 — 0.531 0.560 — 0.533 0.547 0.529 0.520 0.519 0.532 0.522 0.557 0.517 0.425 0.409 0.529 0.545

— 2.42 [1.64, 3.56] 0.98 [0.88, 1.10] 1.01 [0.90, 1.13] 0.96 [0.87, 1.07] 1.08 [0.99, 1.17] 1.56 [1.43, 1.71] 1.29 [1.20, 1.39] 1.17 [1.07, 1.27] 1.21 [1.14, 1.29] 1.01 [0.72, 1.41] 0.93 [0.80, 1.08] 1.75 [1.31, 2.34] — 1.53 [1.22, 1.93] 1.32 [1.25, 1.40] — 1.11 [1.04, 1.18] 1.30 [1.20, 1.41] 1.12 [1.05, 1.18] 1.05 [0.98, 1.12] 1.00 [0.93, 1.06] 1.03 [0.95, 1.11] 1.47 [1.24, 1.74] 1.67 [1.52, 1.83] 0.98 [0.89, 1.08] — 1.28 [0.63, 2.63] 1.08 [0.95, 1.22] 1.28 [1.19, 1.38]

— 0.518 0.503 0.510 0.517 0.527 0.538 0.528 0.532 0.536 0.515 0.527 0.527 — 0.527 0.541 — 0.525 0.529 0.510 0.513 0.519 0.526 0.518 0.535 0.511 0.375 0.439 0.527 0.523

— 2.24 [1.66, 3.02] 1.12 [1.02, 1.22] 0.99 [0.90, 1.09] 1.08 [0.98, 1.18] 1.16 [1.09, 1.24] 1.51 [1.40, 1.62] 1.32 [1.24, 1.40] 1.19 [1.12, 1.28] 1.21 [1.15, 1.28] 1.09 [0.79, 1.50] 0.96 [0.86, 1.08] 1.47 [1.22, 1.78] — 1.50 [1.26, 1.79] 1.41 [1.34, 1.48] — 1.19 [1.13, 1.26] 1.31 [1.23, 1.40] 1.22 [1.15, 1.29] 1.04 [0.93, 1.16] 0.98 [0.92, 1.05] 1.07 [1.01, 1.13] 1.36 [1.19, 1.56] 1.77 [1.64, 1.91] 1.01 [0.92, 1.11] 0.76 [0.08, 7.71] 0.89 [0.53, 1.50] 1.11 [1.00, 1.24] 1.39 [1.31, 1.48]

— 0.518 0.532 0.512 0.495 0.527 0.544 0.536 0.534 0.537 0.522 0.527 0.527 — 0.529 0.551 — 0.533 0.541 0.520 0.514 0.519 0.530 0.519 0.545 0.515 0.425 0.405 0.531 0.536

56

Supplementary Table 25 | Proportional hazards analysis for prolonged ICU stay in the eICU Collaborative Research Database. Prolonged ICU stay is defined as ICU stay greater than 7 days. Per-analyte hazard ratios (with 95% confidence intervals) and concordance indices for PopRI , PerRI , and NORMARI abnormality flags, adjusted for age and sex. PopRI

PerRI

NORMARI

Analyte

N train (events)

N test (events)

HR [95% CI]

C

HR [95% CI]

C

HR [95% CI]

C

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

— 11140 (3587) 7956 (2821) 7276 (2596) 7437 (2658) 33336 (7040) 33039 (6996) 34278 (7202) 40636 (7193) 31684 (6787) 684 (328) 52561 (7373) 34765 (7128) — 34972 (7128) 38559 (7280) — 28026 (6451) 29944 (6867) 29445 (6783) 21492 (4995) 36186 (7220) 30913 (7030) 31027 (7105) 28216 (6496) 7456 (2547) 20 (8) 186 (134) 8871 (3098) 30209 (7003)

— 7428 (2392) 5305 (1881) 4851 (1731) 4958 (1772) 22225 (4694) 22026 (4664) 22853 (4801) 27092 (4795) 21123 (4525) 456 (218) 35041 (4915) 23178 (4753) — 23316 (4752) 25707 (4854) — 18684 (4301) 19964 (4578) 19631 (4523) 14328 (3330) 24125 (4814) 20609 (4686) 20686 (4737) 18812 (4331) 4972 (1699) 14 (5) 124 (90) 5914 (2065) 20140 (4668)

— 0.75 [0.61, 0.93] 0.68 [0.63, 0.73] 0.82 [0.76, 0.88] 0.86 [0.80, 0.93] 0.66 [0.62, 0.71] 0.79 [0.75, 0.84] 0.74 [0.70, 0.78] 0.63 [0.58, 0.68] 0.68 [0.65, 0.72] 1.02 [0.81, 1.29] 0.44 [0.36, 0.53] 0.53 [0.46, 0.61] — 0.50 [0.43, 0.58] 0.77 [0.73, 0.80] — 0.92 [0.87, 0.97] 0.67 [0.63, 0.71] 0.88 [0.83, 0.94] 0.85 [0.78, 0.92] 0.74 [0.71, 0.78] 0.90 [0.86, 0.94] 0.55 [0.49, 0.62] 0.67 [0.63, 0.71] 0.86 [0.79, 0.93] 0.67 [0.11, 4.11] 0.82 [0.57, 1.18] 1.03 [0.95, 1.12] 0.75 [0.71, 0.79]

— 0.535 0.551 0.533 0.545 0.550 0.541 0.552 0.547 0.566 0.512 0.537 0.541 — 0.543 0.553 — 0.537 0.559 0.539 0.535 0.560 0.534 0.549 0.557 0.537 0.600 0.527 0.524 0.557

— 0.70 [0.51, 0.96] 0.84 [0.76, 0.92] 0.87 [0.78, 0.96] 0.98 [0.89, 1.07] 0.76 [0.70, 0.83] 0.77 [0.71, 0.83] 0.84 [0.79, 0.90] 0.72 [0.66, 0.80] 0.66 [0.62, 0.70] 0.97 [0.75, 1.25] 0.53 [0.41, 0.70] 0.64 [0.51, 0.81] — 0.51 [0.42, 0.62] 0.81 [0.77, 0.85] — 0.86 [0.81, 0.91] 0.67 [0.62, 0.72] 0.97 [0.92, 1.02] 0.98 [0.92, 1.04] 0.80 [0.75, 0.85] 1.22 [1.14, 1.30] 0.53 [0.45, 0.63] 0.70 [0.65, 0.76] 0.90 [0.83, 0.98] — 0.74 [0.46, 1.19] 0.97 [0.87, 1.09] 0.79 [0.74, 0.85]

— 0.536 0.533 0.520 0.534 0.535 0.537 0.538 0.540 0.562 0.482 0.535 0.537 — 0.537 0.543 — 0.544 0.552 0.540 0.529 0.543 0.542 0.541 0.546 0.542 — 0.533 0.527 0.545

— 0.77 [0.60, 0.97] 0.73 [0.68, 0.79] 0.85 [0.78, 0.92] 0.87 [0.81, 0.94] 0.67 [0.63, 0.72] 0.75 [0.71, 0.80] 0.75 [0.71, 0.80] 0.61 [0.56, 0.66] 0.69 [0.66, 0.73] 0.95 [0.75, 1.22] 0.47 [0.38, 0.57] 0.53 [0.45, 0.62] — 0.52 [0.44, 0.61] 0.76 [0.73, 0.80] — 0.83 [0.79, 0.87] 0.65 [0.61, 0.69] 0.83 [0.80, 0.88] 0.93 [0.84, 1.03] 0.76 [0.71, 0.82] 0.96 [0.91, 1.01] 0.54 [0.47, 0.62] 0.67 [0.63, 0.72] 0.87 [0.81, 0.95] 0.20 [0.00, 12.22] 0.91 [0.63, 1.31] 0.98 [0.90, 1.07] 0.74 [0.69, 0.78]

— 0.535 0.548 0.527 0.545 0.548 0.544 0.548 0.545 0.566 0.461 0.536 0.539 — 0.542 0.554 — 0.541 0.559 0.546 0.532 0.549 0.537 0.547 0.554 0.534 0.500 0.528 0.527 0.552

57

Supplementary Table 26 | Proportional hazards analysis for in-hospital mortality in the INSPIRE cohort. Per-analyte hazard ratios (with 95% confidence intervals) and concordance indices for PopRI , PerRI , and NORMARI abnormality flags, adjusted for age and sex. PopRI

PerRI

NORMARI

Analyte

N train (events)

N test (events)

HR [95% CI]

C

HR [95% CI]

C

HR [95% CI]

C

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

233 (13) 18427 (470) 16668 (437) 15676 (441) 15814 (415) 17938 (447) 17863 (456) 18561 (442) 3742 (331) 24427 (421) — 15690 (419) 23831 (460) — 21702 (466) 21804 (440) — — — — — 21663 (455) 20394 (466) — — 17133 (424) — — 17152 (470) 20554 (458)

156 (8) 12285 (314) 11113 (292) 10452 (294) 10544 (277) 11960 (298) 11909 (304) 12374 (294) 2495 (221) 16285 (281) — 10461 (279) 15888 (307) — 14469 (310) 14537 (294) — — — — — 14442 (303) 13597 (310) — — 11423 (282) — — 11436 (313) 13703 (305)

0.13 [0.01, 2.06] 6.58 [4.90, 8.83] 2.10 [1.72, 2.56] 1.66 [1.37, 2.02] 2.17 [1.76, 2.68] 2.40 [1.90, 3.02] 3.08 [2.48, 3.83] 2.79 [2.21, 3.53] 3.71 [2.58, 5.34] 2.63 [2.12, 3.26] — 1.87 [1.41, 2.49] 4.33 [3.01, 6.23] — 4.38 [3.06, 6.29] 2.98 [2.43, 3.67] — — — — — 2.37 [1.95, 2.88] 4.55 [3.66, 5.65] — — 3.36 [2.76, 4.10] — — 4.44 [3.57, 5.52] 3.80 [3.03, 4.78]

0.446 0.805 0.745 0.632 0.742 0.724 0.764 0.791 0.643 0.731 — 0.681 0.730 — 0.719 0.821 — — — — — 0.741 0.846 — — 0.758 — — 0.822 0.755

0.13 [0.01, 2.06] 4.28 [3.06, 6.00] 2.40 [1.89, 3.04] 1.48 [1.20, 1.82] 1.99 [1.62, 2.45] 2.21 [1.71, 2.86] 2.89 [2.25, 3.71] 2.61 [2.02, 3.39] 3.25 [2.10, 5.03] 2.47 [1.95, 3.13] — 1.91 [1.37, 2.65] 3.96 [2.49, 6.29] — 3.82 [2.48, 5.89] 2.61 [2.10, 3.24] — — — — — 2.90 [2.31, 3.65] 3.75 [2.89, 4.86] — — 2.78 [2.25, 3.43] — — 3.20 [2.54, 4.04] 3.42 [2.68, 4.37]

0.446 0.719 0.727 0.636 0.713 0.670 0.713 0.718 0.616 0.745 — 0.658 0.676 — 0.683 0.747 — — — — — 0.738 0.752 — — 0.761 — — 0.733 0.751

— 5.87 [4.30, 8.02] 2.43 [1.98, 2.98] 1.51 [1.24, 1.83] 2.21 [1.81, 2.70] 2.41 [1.91, 3.06] 3.24 [2.58, 4.07] 2.81 [2.20, 3.60] 3.20 [2.15, 4.76] 2.70 [2.16, 3.38] — 2.07 [1.39, 3.09] 4.07 [2.76, 6.00] — 4.31 [2.94, 6.31] 2.99 [2.42, 3.69] — — — — — 2.64 [2.15, 3.24] 3.96 [3.15, 4.96] — — 2.87 [2.36, 3.51] — — 3.93 [3.15, 4.90] 3.83 [3.03, 4.84]

0.482 0.785 0.754 0.637 0.732 0.707 0.776 0.773 0.627 0.750 — 0.655 0.724 — 0.717 0.806 — — — — — 0.765 0.816 — — 0.765 — — 0.799 0.799

58

Supplementary Table 27 | Proportional hazards analysis for perioperative infection in the INSPIRE cohort. Peranalyte hazard ratios (with 95% confidence intervals) and concordance indices for PopRI , PerRI , and NORMARI abnormality flags, adjusted for age and sex. PopRI

PerRI

NORMARI

Analyte

N train (events)

N test (events)

HR [95% CI]

C

HR [95% CI]

C

HR [95% CI]

C

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

226 (32) 18112 (691) 16373 (656) 15382 (665) 15519 (665) 17637 (682) 17550 (687) 18256 (689) 3583 (413) 24107 (694) — 15420 (629) 23504 (722) — 21375 (716) 21492 (704) — — — — — 21349 (703) 20078 (706) — — 16829 (676) — — 16841 (680) 20235 (707)

152 (21) 12076 (461) 10916 (438) 10255 (443) 10347 (444) 11758 (455) 11700 (458) 12172 (460) 2390 (275) 16072 (462) — 10281 (419) 15670 (482) — 14251 (478) 14329 (469) — — — — — 14234 (469) 13386 (471) — — 11220 (450) — — 11228 (454) 13491 (471)

0.82 [0.35, 1.97] 1.22 [1.05, 1.43] 1.16 [0.98, 1.38] 1.03 [0.88, 1.21] 1.45 [1.15, 1.84] 1.09 [0.94, 1.27] 1.17 [1.00, 1.36] 1.09 [0.94, 1.27] 1.00 [0.81, 1.22] 1.08 [0.92, 1.26] — 0.97 [0.81, 1.16] 1.09 [0.92, 1.28] — 1.06 [0.90, 1.25] 1.04 [0.88, 1.24] — — — — — 1.12 [0.95, 1.33] 1.06 [0.90, 1.26] — — 1.07 [0.87, 1.31] — — 1.27 [1.08, 1.49] 1.18 [1.01, 1.37]

0.559 0.624 0.593 0.610 0.642 0.557 0.582 0.603 0.575 0.573 — 0.547 0.616 — 0.602 0.610 — — — — — 0.604 0.611 — — 0.576 — — 0.601 0.558

1.30 [0.43, 4.00] 1.14 [0.97, 1.34] 1.12 [0.96, 1.30] 1.02 [0.87, 1.19] 1.20 [1.02, 1.40] 1.06 [0.90, 1.24] 1.21 [1.04, 1.42] 1.03 [0.88, 1.20] 0.98 [0.78, 1.23] 0.97 [0.84, 1.13] — 1.00 [0.81, 1.22] 1.10 [0.90, 1.35] — 1.00 [0.82, 1.20] 1.05 [0.90, 1.22] — — — — — 1.03 [0.88, 1.19] 1.15 [0.99, 1.33] — — 1.10 [0.94, 1.29] — — 1.41 [1.21, 1.64] 1.21 [1.05, 1.41]

0.572 0.626 0.598 0.610 0.625 0.548 0.593 0.595 0.577 0.571 — 0.547 0.612 — 0.592 0.609 — — — — — 0.600 0.603 — — 0.584 — — 0.582 0.564

1.80 [0.58, 5.56] 1.15 [0.98, 1.34] 1.16 [0.99, 1.36] 1.08 [0.93, 1.26] 1.26 [1.03, 1.54] 1.12 [0.96, 1.31] 1.15 [0.99, 1.34] 1.14 [0.98, 1.32] 1.02 [0.82, 1.27] 1.07 [0.91, 1.25] — 1.00 [0.79, 1.27] 1.05 [0.88, 1.26] — 1.06 [0.89, 1.26] 1.02 [0.87, 1.19] — — — — — 0.97 [0.83, 1.13] 1.22 [1.05, 1.41] — — 1.02 [0.85, 1.22] — — 1.25 [1.08, 1.46] 1.26 [1.09, 1.47]

0.572 0.631 0.590 0.614 0.644 0.556 0.585 0.605 0.577 0.571 — 0.547 0.615 — 0.601 0.611 — — — — — 0.593 0.638 — — 0.578 — — 0.607 0.559

59

Supplementary Table 28 | Proportional hazards analysis for prolonged hospital stay in the INSPIRE cohort. Prolonged hospital stay is defined as hospital stay greater than 7 days. Per-analyte hazard ratios (with 95% confidence intervals) and concordance indices for PopRI , PerRI , and NORMARI abnormality flags, adjusted for age and sex. PopRI

PerRI

NORMARI

Analyte

N train (events)

N test (events)

HR [95% CI]

C

HR [95% CI]

C

HR [95% CI]

C

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

233 (164) 18427 (14140) 16668 (13163) 15676 (12343) 15814 (12427) 17938 (14031) 17863 (13403) 18561 (14572) 3742 (3239) 24427 (17097) — 15690 (11167) 23831 (16876) — 21702 (15789) 21804 (15982) — — — — — 21663 (15923) 20394 (15193) — — 17133 (13177) — — 17152 (13544) 20554 (15316)

156 (109) 12285 (9427) 11113 (8776) 10452 (8230) 10544 (8286) 11960 (9355) 11909 (8936) 12374 (9714) 2495 (2160) 16285 (11399) — 10461 (7446) 15888 (11251) — 14469 (10527) 14537 (10655) — — — — — 14442 (10615) 13597 (10129) — — 11423 (8786) — — 11436 (9031) 13703 (10211)

0.95 [0.62, 1.45] 0.70 [0.67, 0.73] 0.67 [0.64, 0.71] 0.85 [0.82, 0.88] 0.71 [0.65, 0.77] 0.75 [0.72, 0.77] 0.89 [0.86, 0.92] 0.77 [0.74, 0.79] 0.83 [0.77, 0.89] 0.65 [0.63, 0.68] — 1.03 [0.99, 1.07] 0.81 [0.78, 0.83] — 0.74 [0.71, 0.76] 0.73 [0.69, 0.76] — — — — — 0.71 [0.68, 0.74] 0.69 [0.66, 0.73] — — 0.87 [0.82, 0.92] — — 0.79 [0.76, 0.83] 0.80 [0.78, 0.83]

0.511 0.566 0.558 0.545 0.539 0.553 0.539 0.552 0.516 0.580 — 0.536 0.560 — 0.566 0.538 — — — — — 0.546 0.553 — — 0.534 — — 0.542 0.545

0.82 [0.46, 1.45] 0.73 [0.71, 0.76] 1.10 [1.06, 1.13] 1.02 [0.99, 1.06] 1.12 [1.08, 1.16] 0.96 [0.92, 0.99] 0.95 [0.92, 0.98] 0.97 [0.93, 1.00] 0.98 [0.91, 1.06] 0.71 [0.68, 0.73] — 1.07 [1.02, 1.11] 0.85 [0.82, 0.88] — 0.75 [0.72, 0.78] 1.02 [0.99, 1.05] — — — — — 0.95 [0.92, 0.98] 1.22 [1.18, 1.26] — — 0.95 [0.92, 0.99] — — 0.96 [0.93, 0.99] 0.97 [0.94, 1.00]

0.500 0.568 0.532 0.533 0.539 0.538 0.538 0.525 0.507 0.574 — 0.538 0.555 — 0.557 0.533 — — — — — 0.528 0.553 — — 0.532 — — 0.535 0.537

0.88 [0.52, 1.48] 0.66 [0.64, 0.69] 0.82 [0.79, 0.85] 0.92 [0.89, 0.95] 0.95 [0.90, 1.00] 0.76 [0.74, 0.79] 0.88 [0.85, 0.91] 0.83 [0.81, 0.86] 0.85 [0.79, 0.92] 0.68 [0.66, 0.71] — 1.02 [0.97, 1.07] 0.80 [0.78, 0.83] — 0.71 [0.69, 0.74] 0.85 [0.82, 0.89] — — — — — 0.76 [0.73, 0.79] 1.04 [1.00, 1.07] — — 0.95 [0.91, 0.99] — — 0.78 [0.75, 0.81] 0.84 [0.81, 0.87]

0.523 0.572 0.538 0.536 0.530 0.551 0.540 0.540 0.515 0.578 — 0.536 0.560 — 0.568 0.532 — — — — — 0.545 0.542 — — 0.532 — — 0.545 0.543

60

Supplementary Table 29 | Proportional hazards analysis for unplanned ICU admission in the INSPIRE cohort. Per-analyte hazard ratios (with 95% confidence intervals) and concordance indices for PopRI , PerRI , and NORMARI abnormality flags, adjusted for age and sex. PopRI

PerRI

NORMARI

Analyte

N train (events)

N test (events)

HR [95% CI]

C

HR [95% CI]

C

HR [95% CI]

C

A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC

233 (105) 18427 (5221) 16668 (5096) 15676 (5066) 15814 (5026) 17938 (5376) 17863 (5265) 18561 (5536) 3742 (2059) 24427 (5246) — 15690 (4539) 23831 (6037) — 21702 (5792) 21804 (5871) — — — — — 21663 (5875) 20394 (5724) — — 17133 (5043) — — 17152 (5154) 20554 (5663)

156 (70) 12285 (3481) 11113 (3397) 10452 (3378) 10544 (3351) 11960 (3585) 11909 (3510) 12374 (3691) 2495 (1373) 16285 (3498) — 10461 (3026) 15888 (4025) — 14469 (3862) 14537 (3915) — — — — — 14442 (3916) 13597 (3817) — — 11423 (3363) — — 11436 (3437) 13703 (3775)

1.09 [0.68, 1.74] 0.85 [0.80, 0.91] 1.07 [0.99, 1.15] 0.95 [0.90, 1.02] 1.14 [1.01, 1.28] 1.00 [0.94, 1.06] 0.85 [0.80, 0.90] 1.29 [1.22, 1.37] 0.93 [0.86, 1.02] 1.26 [1.18, 1.35] — 1.03 [0.97, 1.10] 0.94 [0.89, 0.99] — 1.01 [0.96, 1.07] 1.11 [1.02, 1.20] — — — — — 1.31 [1.22, 1.40] 1.35 [1.25, 1.45] — — 1.11 [1.02, 1.21] — — 0.87 [0.81, 0.94] 1.16 [1.10, 1.23]

0.495 0.541 0.535 0.535 0.530 0.523 0.534 0.530 0.521 0.563 — 0.529 0.529 — 0.532 0.520 — — — — — 0.532 0.537 — — 0.534 — — 0.539 0.540

1.00 [0.59, 1.69] 0.90 [0.86, 0.96] 0.96 [0.91, 1.01] 0.94 [0.88, 0.99] 0.84 [0.79, 0.89] 0.94 [0.89, 0.99] 0.87 [0.82, 0.91] 1.10 [1.04, 1.15] 0.93 [0.84, 1.02] 1.19 [1.12, 1.26] — 0.93 [0.87, 1.00] 0.98 [0.92, 1.04] — 1.10 [1.04, 1.17] 0.93 [0.88, 0.98] — — — — — 0.98 [0.93, 1.03] 0.83 [0.79, 0.87] — — 1.07 [1.01, 1.14] — — 0.87 [0.82, 0.92] 1.00 [0.95, 1.05]

0.499 0.542 0.539 0.537 0.540 0.524 0.536 0.521 0.520 0.562 — 0.531 0.529 — 0.532 0.522 — — — — — 0.511 0.540 — — 0.535 — — 0.542 0.537

1.17 [0.66, 2.07] 0.94 [0.89, 1.00] 1.01 [0.95, 1.08] 1.01 [0.96, 1.07] 0.90 [0.82, 0.98] 1.01 [0.96, 1.07] 0.84 [0.79, 0.89] 1.42 [1.34, 1.49] 0.90 [0.82, 0.99] 1.22 [1.15, 1.30] — 1.26 [1.15, 1.37] 1.05 [0.99, 1.10] — 1.05 [0.99, 1.11] 1.13 [1.05, 1.20] — — — — — 1.37 [1.29, 1.45] 0.99 [0.93, 1.04] — — 1.09 [1.02, 1.17] — — 0.89 [0.83, 0.95] 1.15 [1.09, 1.22]

0.528 0.539 0.536 0.534 0.532 0.523 0.536 0.544 0.518 0.562 — 0.529 0.528 — 0.533 0.518 — — — — — 0.539 0.531 — — 0.535 — — 0.538 0.541

61

Record · ID 200480 · SHA-256 3dcd968156aae199
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.