Learning Normal Representations for Blood Biomarkers Aashna P. Shah1,2,6 , Michelle M. Li1,6 , Yash Lal4 , Seffi Cohen1,6,7 , Liat F. Antwarg1,6,7 , Morgan Sanchez1,6 , James A. Diao1,3,6 , Chirag J. Patel1 , Ben Y. Reis1,5,6,7 , Ran D. Balicer6,7 *, Noa Dagan6,7,8 *, and Arjun K. Manrai1,6 * 1. Department of Biomedical Informatics, Harvard Medical School, Boston, MA, USA 2. Department of Systems Biology, Harvard Medical School, Boston, MA, USA 3. Department of Medicine, Brigham and Women’s Hospital, Boston, MA, USA 4. Department of Mathematics, Johns Hopkins University, Baltimore, MD, USA 5. Computational Health Informatics Program (CHIP), Boston Children’s Hospital, Boston, MA, USA 6. The Ivan and Francesca Berkowitz Family Living Laboratory Collaboration at Harvard Medical School and Clalit Research Institute,
arXiv:2605.18701v1 [cs.LG] 18 May 2026
USA and Israel 7. Clalit Research Institute, Innovation Division, Clalit Health Services, Ramat-Gan, Israel 8. Faculty of Computer and Information Science, Ben Gurion University, Be’er Sheva, Israel
* Co-senior authors Correspondence: [email protected]
Abstract Blood-based biomarkers underpin clinical diagnosis and management, yet their interpretation relies largely on fixed population reference intervals that ignore stable, intra-patient variability. As such, population-based interpretation can mask meaningful deviation from an individual’s baseline, risking delayed disease detection. To remedy this, there have been increasing efforts to personalize blood biomarker interpretation using individual testing histories. However, these methods may overfit to sparse data, inflating false-positive rates and unnecessary follow-up, and can also unwittingly include unrecognized or subclinical disease. Here, we leverage nearly 2 billion longitudinal laboratory measurements from over 1.6 million individuals across North America, the Middle East, and East Asia, to show that while laboratory values are highly individual, purely personalized intervals routinely overfit, classifying up to 68% of measurements as abnormal, without corresponding associations with adverse clinical outcomes. We then introduce NORMA, a conditional transformer-based framework that generates reference intervals by conditioning on both a patient’s history and population-level data about “normal” variation. NORMA-derived intervals achieve higher precision for predicting outcomes, including mortality, acute kidney injury, and chronic disease. These findings caution against over-personalization in laboratory medicine and demonstrate that anchoring individual trajectories to population-level priors outperforms either approach alone. To promote transparency, we publicly release the model, code, and an interactive user interface for accessible, individualized laboratory interpretation.
Introduction Laboratory testing is among the most frequently performed medical tests, with more than 14 billion tests ordered annually in the United States alone 1 . Yet interpretation has changed remarkably little since the 1960s; values are classified as “low,” “normal,” or “high” against fixed population reference intervals (PopRI ) that represent the central 95% of measurements from ostensibly healthy individuals 2–8 . Prior studies have recognized that many routine biomarkers fluctuate around narrow individual “setpoints” and that within-person change carries prognostic information that population intervals miss 9–20 . Despite this, clinical practice often still relies on universal thresholds, partly because of the ease of a one-size-fits-all approach, and because operationalizing individualized interpretation requires modeling patient-specific trajectories at scale.
1
Efforts to move beyond universal thresholds have taken two directions. The first refines reference intervals for predefined subgroups, for example, sex-specific hematologic ranges or adjusted HbA1c thresholds for patients with hemoglobin variants 12,14,21–24 . Defining population subgroups introduces several challenges— such groups are often defined coarsely or arbitrarily, may encode existing biases, and cannot capture the full spectrum of individual variation embedded in longitudinal trajectories. The practice of race adjustment, in particular, has been increasingly challenged with major clinical and non-clinical implications for patients 25–27 . The second approach derives personalized reference intervals (PerRI ) directly from each patient’s laboratory testing history. Foy et al., for example, showed that deviations from individually derived hematologic setpoints associate more strongly with mortality than deviations from population thresholds, underscoring the prognostic value of within-person change 16 . However, purely individualized approaches that derive baselines entirely from a patient’s own data risk incorporating unrecognized chronic disease into the estimated setpoint and overfitting to sparse histories, producing overly narrow intervals that label benign physiological variation as pathological or vice versa 16,18,28,29 . Here, we ask whether these competing approaches can be combined using nearly 2 billion measurements across 30 analytes, spanning two decades of national longitudinal care, dense ICU monitoring, and perioperative data across three countries (Fig. 1). We first quantify the individuality of routine blood tests and then systematically compare population and personalized reference intervals. We then introduce NORMA (Normal Outcome Range Modeling with Attention), a new conditional transformer framework that combines both approaches to generate individualized reference intervals by anchoring each patient’s trajectory to populationlevel expectations for a healthy state. Rather than choosing between the false dichotomy of purely population and personalized approaches, NORMA learns where each patient’s interval is expected to fall between the two extremes. We show that purely personalized intervals overcall abnormalities and lose prognostic signal, that NORMA resolves this tradeoff, and test whether the resulting intervals detect clinically meaningful change months to years before population intervals.
Results Study populations We assessed population-based (PopRI ), personalized (PerRI ), and NORMA-derived (NORMARI ) reference intervals across 30 routine laboratory analytes spanning four clinical panels: complete blood count, comprehensive metabolic panel, lipid panel, and hemoglobin A1c (HbA1c). These panels are used in routine health evaluation, cardiometabolic risk assessment, and diabetes screening or monitoring, as detailed in the Methods. Within this analyte set, NORMA was trained on 3.4 million longitudinal laboratory sequences from two development cohorts: MIMIC-IV (179,601 patients) and EHRSHOT (5,676 patients); no development data were used in external validation (Supplementary Tables 1–3). We validated NORMA in three independent cohorts that differed in temporal scale and clinical setting. The Clalit Health Services (CHS) cohort included 1,450,862 adults (63% female; mean age 55.1 ± 18.6 years) contributing approximately 1.9 billion laboratory measurements over more than two decades of outpatient follow-up (Fig. 2a; Supplementary Table 3). To ensure stable baseline estimation, we required at least five outpatient measurements per analyte spaced at least 90 days apart, yielding trajectories spanning a median of more than four years. The eICU Collaborative Research Database (eICU-CRD) cohort included 98,432 ICU patients (46% female; mean age 63.3 ± 16.1 years) across 208 U.S. hospitals, contributing over 20 million measurements with dense sampling (median number of measurements per patient ranged between 8-35 for high-frequency analytes) over a median stay of approximately 8 days.
2
The INSPIRE cohort included 51,159 surgical patients (52% female; mean age 57.6 ± 14.9 years) undergoing perioperative care in South Korea, contributing over 10 million laboratory measurements collected across preoperative, intraoperative, and postoperative timepoints, with a median follow-up of approximately 175 days. Laboratory distributions in CHS, EHRSHOT, and INSPIRE largely overlapped population reference intervals, whereas eICU-CRD and MIMIC-IV exhibited shifts consistent with greater rates of critical illness, including lower albumin (ALB) and higher glucose (GLU) variability (Supplementary Table 3).
Lab values are individualized Across all three cohorts, within-person variability was substantially smaller than between-person variability for most analytes, confirming that laboratory values fluctuate around narrow individual setpoints (Fig. 2b; Supplementary Table 4). We quantified this finding using the individuality index (II), the ratio of within-person to between-person coefficient of variation, where values below 0.6 indicate high biological individuality. Hematologic indices showed the lowest intra-individual variability relative to inter-individual variability. Mean corpuscular hemoglobin (MCH), mean corpuscular volume (MCV), and red blood cell count (RBC) had individuality indices ranging from approximately 0.2 to 0.6 across CHS, eICU-CRD, and INSPIRE, with values in eICU-CRD generally equal to or lower than those observed in CHS and INSPIRE (Supplementary Table 4). In contrast, electrolytes and metabolic analytes exhibited lower biological individuality; potassium (K), glucose (GLU), sodium (NA), and calcium (CA) had individuality indices ranging from approximately 0.7 to 1.2 across cohorts, with within-person variability approaching the population spread (Fig. 2b; Supplementary Table 4). INSPIRE showed patterns consistent with an intermediate regime between outpatient and ICU settings. Within-person stability carried prognostic signal. Greater deviation from a patient’s baseline, quantified as the absolute z-score of each index measurement relative to that patient’s baseline average, was monotonically associated with mortality across all cohorts (Fig. 2c; Supplementary Fig. 3). When stratified by raw analyte values, several biomarkers, including white blood cell count (WBC), glucose (GLU), aspartate aminotransferase (AST), and hemoglobin A1c (HbA1c), exhibited U-shaped mortality curves, in which both abnormally high and low values were associated with elevated risk.
NORMA generates calibrated individualized intervals NORMA is a conditional, autoregressive transformer framework with multiple configurations that models the distribution of a patient’s next laboratory value given their longitudinal measurement history (Fig. 3a; Methods) and population data on “healthy” variation. The model is trained to forecast the next observed value. To derive a personalized reference interval, we condition the query token on values within the population reference range, so that the resulting 95% prediction interval represents the expected range for that patient given their prior trajectory and population context, rather than a strictly disease-free state. We trained NORMA on 3.4 million longitudinal sequences from MIMIC-IV and EHRSHOT under two output parameterizations (Gaussian and quantile), which share the same architecture but differ in distributional assumptions (Methods and Supplementary Table 2). On a held-out test set, NORMA outperformed all baselines in next-step forecasting accuracy across both output parameterizations. Baseline models included last value carried forward, autoregressive integrated moving average (ARIMA), and patient-specific mean. Mean absolute error was lowest for NORMA quantile (5.9 [IQR 3.6–9.1]) and NORMA Gaussian (6.0 [2.6–10.2]), compared with ARIMA (6.7 [4.3–11.9]), last value carried forward (7.3 [3.9–10.7]), and patient-specific mean (8.4 [4.5–13.6]; Supplementary Table 5; Fig. 3b). Forecasting accuracy varied by analyte, with mean platelet volume (MPV) showing the weakest performance, consistent with its having the smallest training sample size (Fig. 3c; Supplementary Fig. 2).
3
NORMA’s prediction intervals responded appropriately to the factors that govern uncertainty (Fig. 3d; Supplementary Table 6). More prior measurements narrowed intervals, longer prediction horizons widened them, and greater within-person variability produced substantially wider intervals. The quantile parameterization was more responsive than the Gaussian to changes in input features; for example, doubling within-person variability roughly doubled the predicted interval width under the quantile model, compared with minimal change under the Gaussian. Importantly, interval width stabilized after approximately 30 prior measurements and did not continue to narrow with additional data, indicating that the model does not overweight longer histories.
NORMARI detects abnormalities earlier than population intervals In all three external validation cohorts, PerRI flagged the most tests as abnormal, PopRI the fewest, and NORMARI was intermediate. In CHS, abnormality rates were 29.6% (PopRI ), 39.1% (NORMARI ), and 46.8% (PerRI ). In eICU-CRD, where baseline abnormality rates were higher due to acute illness, the same ordering held: 50.2% (PopRI ), 55.8% (NORMARI ), and 68.1% (PerRI ). In INSPIRE, abnormality rates followed the same pattern, with values of 35.1% (PopRI ), 42.5% (NORMARI ), and 57.0% (PerRI ), consistent with its intermediate clinical setting (Fig. 4a; Fig. 5a; Fig. 6a; Supplementary Table 7). Among tests classified as normal by population intervals, NORMARI reclassified 12, 9, and 14 per 100 as abnormal in CHS, eICU-CRD, and INSPIRE, respectively, compared with 27, 37, and 34 per 100 for PerRI (Supplementary Tables 7–8). NORMARI flagged abnormalities earlier than population intervals across all cohorts. In CHS, 49% of tests NORMARI flagged as abnormal were later also flagged by PopRI , with a median lead time of 8.7 months (IQR 1.7–33.8; Supplementary Table 20; Fig. 4e). In eICU-CRD, 25% were later confirmed by PopRI , with a median lead time of 34.8 hours (IQR 23.1–72.6; Supplementary Table 20; Fig. 5e). In INSPIRE, 30% were later confirmed by PopRI , with a median lead time of 23.7 hours (IQR 11.4–54.7), consistent with its perioperative timescale (Fig. 6e; Supplementary Table 20).
NORMARI detects clinically meaningful abnormalities missed by population intervals We next asked whether the additional abnormalities detected by NORMARI carried clinical meaning. Restricting to measurements classified as normal by population intervals, where personalization could add signal beyond standard practice, NORMARI reclassifications carried substantially higher positive predictive value than PerRI . For every 100 patients PopRI classified as normal but NORMARI flagged as abnormal in eICU-CRD, 13 died in hospital versus 10 under PerRI (∆ = +3), 16 developed AKI versus 15 (∆ = +1), and 27 had prolonged ICU stays versus 23 (∆ = +4). In CHS, 30 died versus 24 under PerRI (∆ = +6), 46 developed CKD versus 39 (∆ = +7), and 80 had type 2 diabetes versus 75 (∆ = +5). In INSPIRE, NORMARI identified 6 more unplanned ICU admissions, 1 more in-hospital death, and 1 more perioperative infection per 100 reclassified patients than PerRI (Supplementary Tables 17–19). Across all assessed outcomes – mortality, type 2 diabetes, and CKD over 10 years in CHS; in-hospital mortality, AKI, sepsis, and prolonged LOS in eICU-CRD; and perioperative mortality, prolonged LOS, unplanned ICU admission, and postoperative infection in INSPIRE – PerRI achieved higher sensitivity but NORMARI achieved substantially higher specificity and positive predictive value (Figs. 4c,d, 5c,d, 6c,d and Supplementary Figs. 9–11).
NORMARI improves clinical prediction In time-to-event analyses across all cohorts, NORMARI and PopRI abnormality flags showed comparable prognostic associations with clinical outcomes, while PerRI flags were substantially weaker (Figs. 4b, 5b, 6b and Supplementary Tables 21–29). This was most pronounced in CHS, where NORMARI abnormality flags
4
were strongly associated with all-cause mortality for hematologic and metabolic analytes, with several showing 2- to 4-fold elevated risk after adjusting for age and sex – including creatinine (HR 3.30), hemoglobin (HR 3.72), and hematocrit (HR 3.36). Effect sizes were comparable to or slightly exceeded those of PopRI for most analytes, while PerRI flags carried essentially no prognostic signal. Mean concordance indices for mortality were similar between NORMARI and PopRI across all three cohorts (CHS: 0.752 vs. 0.758; eICU-CRD: 0.567 vs. 0.571; INSPIRE: 0.730 vs. 0.731), while PerRI was consistently lowest (CHS: 0.748; eICU-CRD: 0.562; INSPIRE: 0.695). The gap was most pronounced in INSPIRE, where PerRI concordance (median 0.718, IQR [0.673, 0.742]) fell well below both NORMARI (0.754 [0.712, 0.780]) and PopRI (0.742 [0.722, 0.778]; Supplementary Tables 21, 22, 26).
Discussion Routine laboratory testing is becoming more frequent, more accessible, and more central to medical decisionmaking 30,31 . Beyond tests that clinicians order within the healthcare system, patients are increasingly seeking out repeated panel-based screening through direct-to-consumer companies. In parallel, interpretation is shifting away from static cutoffs and monolithic “normal ranges” towards “personalized” ranges that leverage trajectories of variation and individual baselines, yet the implications and risks of this shift remain largely unclear 17–19,28 . Across nearly two billion laboratory measurements spanning outpatient and intensive care settings, we find that purely personalized reference intervals consistently overcall abnormalities and are poorly associated with adverse outcomes. NORMA balances purely personalized and population intervals by training a transformer to predict what a patient’s next laboratory value is expected to fall if they remain physiologically stable, and defining the reference interval from this conditional distribution. In doing so, NORMARI inherits the sensitivity of personalized interpretation while maintaining the specificity of population-level definitions of normal variation. In eICU, PerRI flagged the majority of laboratory measurements as abnormal, and similarly elevated abnormality rates were observed in CHS and INSPIRE compared to population-based intervals, indicating that a purely personalized framework would likely generate alerts for a substantial fraction of tests across critical care, outpatient, and perioperative settings. At the scale of a national health system processing millions of tests annually, this could translate to an enormous burden of false alarms, unnecessary follow-up testing, and potential patient anxiety 29,32,33 . By contrast, NORMARI moderated this alert burden, flagging fewer measurements than PerRI while yielding abnormalities that were more strongly enriched for adverse clinical outcomes. This benefit was not apparent from aggregate discrimination alone. Across individual analytes and outcomes, NORMARI and PopRI achieved broadly similar concordance indices, whereas PerRI consistently performed worst. This similarity between NORMARI and PopRI might suggest that personalization adds little value; however, the value of NORMARI is seen in individuals that PopRI labels normal but NORMARI flags as abnormal. In this reclassified subset, NORMARI achieved substantially higher precision and specificity than PerRI . By anchoring individualized intervals to population-level expectations, NORMA reduces the risk that chronic or subclinical disease is absorbed into a patient’s estimated baseline, a limitation of purely personalized approaches such as PerRI . This limitation is not unique to the Gaussian mixture approach used for PerRI . Even with Bayesian updating or hierarchical extensions, a population prior converges toward the patient’s observed distribution as measurements accumulate, eventually washing out the prior and reproducing the same overcalling behavior 28,34 . However, such methods are also nondeterministic, computationally expensive, introduce prior-strength and group-structure hyperparameters, and lack native handling of time or irregular
5
sampling. NORMA avoids this by conditioning on a healthy state at every prediction, regardless of how many measurements are available, which is why its interval width stabilizes rather than continuing to narrow. Furthermore, it does not require explicit pre-filtering of trajectories for stability; it handles irregular measurement spacing natively and its inputs are limited to data already present in most electronic health records. In a clinical deployment, NORMA could provide more "personalized" assessment for patients even without sufficient measurement history. NORMA builds on a parallel line of work that scales sequence modelling over longitudinal health records into clinical foundation models. Early efforts demonstrated that deep learning over raw EHR streams could match or exceed task-specific models for mortality, readmission, and laboratory forecasting 35–38 . More recent transformer-based foundation models, trained on event streams from millions of patient timelines, have demonstrated that longitudinal records can support zero- or few- shot forecasting of diagnoses, procedures, and disease progression at increasing scale, with prediction horizons extending years into the future and architectures expanding to incorporate multimodal clinical data 39–42 . The same paradigm is now extending beyond structured EHRs to continuous physiological streams such as continuous glucose monitoring and other wearables 43,44 . These models demonstrate that longitudinal patient representations can support broad outcome prediction across clinical domains. While NORMA learns from patient trajectories, it is distinct in the prediction task. Rather than learning a general representation of trajectories, such as APOLLO, to predict diagnoses, procedures, or downstream clinical outcomes, it models continuous biomarker distributions and estimates the range of values expected for a given patient under a specified future laboratory state 42 . NORMA therefore draws on counterfactual prediction and sequence-editing frameworks, such as CLEF, which ask how a trajectory would change under an imposed condition or intervention 45–47 . Here, the imposed condition is a future normal laboratory state, allowing NORMA to estimate an individualized healthy-state reference interval rather than simply forecast the next observed measurement. The clinical outcomes evaluated here, including all-cause mortality, acute kidney injury, type 2 diabetes, and chronic kidney disease, were used only for downstream validation rather than model training. This disease-agnostic design suggests that deviations from an individual’s expected healthy trajectory may carry prognostic relevance across diverse disease states without requiring outcome-specific retraining. As longitudinal biomarker monitoring expands further into precision medicine, the challenge of reconciling population-derived and individualized reference standards will broaden, and conditional prediction anchored to a healthy reference state may generalize across these settings 43,44,47,48 . Our study had several limitations. First, NORMA’s current scope is limited by the analytes and clinical inputs used for training. NORMA was trained and evaluated on 30 common blood tests and conditioned only on age, sex, and laboratory trajectories; it does not currently incorporate comorbidities, medications, or other clinical context, which could further refine expected trajectories and abnormality thresholds 49 . Performance was also not uniform across analytes. MPV showed the lowest forecasting accuracy, consistent with its having the smallest training sample size and weakest forecasting accuracy among the 30 analytes. Second, some modeling assumptions may affect calibration. In CHS, we modeled predictive distributions as Gaussian, which may not fully capture skewed or heavy-tailed analytes. However, we implemented quantile regression in eICU, and results were consistent across both parameterizations, suggesting that NORMA’s clinical value does not depend on a strict Gaussian assumption. Broader exploration of output distributions may further improve calibration for skewed or heavy-tailed analytes. Third, generalizability and clinical impact remain to be established prospectively. The CHS analysis was conducted within a single national health system; eICU, although geographically diverse, represents only the ICU segment of hospital care; and INSPIRE reflects perioperative care at a single academic medical center. Performance may differ in emergency departments, primary care clinics, non-surgical inpatient settings, or populations with different demographics and laboratory utilization patterns. Finally, this study evaluated retrospective associations between abnormality flags and
6
clinical outcomes. Prospective trials are needed to determine whether NORMARI -augmented interpretation accelerates time to diagnosis, alters downstream testing, and influences clinician behavior. Such evaluation may further benefit from cohort-specific fine-tuning, for example on Clalit data before deployment, which could improve local performance at the cost of generalizability. Individualized laboratory interpretation can be made more precise and more clinically useful by combining patient-specific trajectories with population-level expectations and uncertainty-aware prediction. By validating across longitudinal outpatient and acute settings, outcome horizons, and patient populations, we show that this framework generalizes beyond a single clinical context. NORMA provides a framework for integrating personalized laboratory interpretation into routine practice, balancing the sensitivity of personalized intervals with the specificity of population-anchored prediction.
Methods Study Populations We used five cohorts spanning model development and external validation. Model development used MIMIC-IV and EHRSHOT. External validation used three independent cohorts representing distinct clinical settings and time horizons: Clalit Health Services (CHS), a national longitudinal health system cohort from Israel; the eICU Collaborative Research Database (eICU-CRD), a multicenter critical care cohort from the United States; and INSPIRE, a perioperative cohort from Seoul National University Hospital in South Korea. Across these cohorts, we evaluated three reference interval frameworks: population-based reference intervals (PopRI), personalized reference intervals (PerRI ), and NORMA-derived reference intervals (NORMARI ). No development data were used in external validation. MIMIC-IV and EHRSHOT. MIMIC-IV contributed 179,601 patients from a tertiary-care hospital system and was used for acute-care, short-horizon model development. EHRSHOT contributed 5,676 patients from longitudinal electronic health records and was used for intermediate-horizon model development. Together, these cohorts provided the development data used for NORMA training, validation, and held-out testing. Clalit Health Services. The CHS cohort comprises nationwide longitudinal electronic healthcare data from Israel’s largest health maintenance organization 50 . We included adults aged 18–99 years with repeated outpatient laboratory testing between 2000 and 2024. For each analyte, eligibility required at least five outpatient measurements spaced at least 90 days apart prior to January 1, 2015 (baseline period) and at least one measurement between January 1, 2015 and January 1, 2016 (classification period). We excluded hospitalizations, emergency department visits, and urgent encounters during baseline ascertainment. The first measurement during the classification period served as the index value and defined the index date. We assessed eligibility independently per analyte, allowing individuals to contribute to multiple analyses. We followed individuals for up to ten years from each analyte-specific index date. For those without events, we censored follow-up at the earliest of health plan disenrollment or death. Patients remaining enrolled without a recorded death were censored at the last available date in the dataset (January 1, 2024). Primary outcomes included all-cause mortality and incident diagnoses of type 2 diabetes (T2D) and chronic kidney disease (CKD). We ascertained outcomes using ICD-9/10 codes supplemented by laboratory criteria. We defined T2D as HbA1c ≥6.5% or fasting glucose ≥126 mg/dL on two occasions; CKD as an estimated glomerular filtration rate <60 mL/min/1.73 m² on two measurements at least 90 days apart. eICU Collaborative Research Database. The eICU Collaborative Research Database (eICU-CRD v2.0) is a multicenter critical care cohort comprising more than 200,859 ICU admissions from 208 hospitals across
7
the United States 51 . We included adults aged 18–99 years with repeated laboratory measurements during their ICU stay. We applied a baseline filter requiring at least five measurements. After filtering, 98,432 unique patients contributed over 20 million measurements across 29 analytes. For each patient–analyte sequence, we used the first 75% of measurements as the baseline to estimate PerRI and NORMARI , and we used the remaining measurements as index values for classification. Primary outcomes included in-hospital mortality, acute kidney injury (AKI), sepsis, and prolonged ICU length of stay (>7 days). Outcomes were ascertained using ICD-9 diagnosis codes and structured clinical documentation. INSPIRE. The INSPIRE dataset is a publicly available perioperative research cohort comprising approximately 130,000 surgical cases from a single academic medical center (Seoul National University Hospital) in South Korea between 2011 and 2020 52 . We included adults aged 20–90 years with repeated laboratory measurements during the perioperative period. We applied a baseline filter requiring at least five measurements. After filtering, 51,159 unique patients contributed over 10 million measurements across 19 analytes. For each patient–analyte sequence, we used the first 75% of measurements as the baseline to estimate PerRI and NORMARI , and we used the remaining measurements as index values for classification. Primary outcomes included in-hospital mortality, prolonged length of stay, unplanned ICU admission, and perioperative infection. Outcomes were ascertained using ICD-10-CM diagnosis codes and structured clinical documentation.
Laboratory Measurements We selected 30 routinely measured blood analytes from common clinical panels based on clinical ubiquity and sufficient repeated measurement across development and validation datasets. We grouped analytes into four standard clinical panels: complete blood count (hematocrit, hemoglobin, mean corpuscular hemoglobin, mean corpuscular hemoglobin concentration, mean corpuscular volume, mean platelet volume, platelet count, red blood cell count, red cell distribution width, white blood cell count); comprehensive metabolic panel (sodium, potassium, chloride, calcium, bicarbonate, blood urea nitrogen, creatinine, glucose, aspartate aminotransferase, alanine aminotransferase, alkaline phosphatase, total bilirubin, direct bilirubin, total protein, albumin); lipid panel (total cholesterol, low-density lipoprotein cholesterol, high-density lipoprotein cholesterol, triglycerides); glycemic control (glycated hemoglobin). We mapped all laboratory results to Logical Observation Identifiers Names and Codes (LOINC) 53 . We retained only positive values, standardized units to canonical formats, and excluded analyte-specific outliers exceeding ±3 standard deviations from the analyte median. When duplicate results shared identical timestamps, we averaged values within each analyte.
Reference Interval Frameworks Population Reference Intervals. We defined PopRI using externally specified laboratory reference ranges and clinical decision thresholds rather than estimating cohort-specific intervals. For each analyte, we used the adult reference range reported by the American Board of Internal Medicine Laboratory Test Reference Ranges 54 . Where sex-specific reference intervals were provided, classifications were made using the patient’s sex-specific range. Supplementary Table 1 provides the reference interval, sex stratification, and canonical unit used for each analyte. Personalized Reference Intervals. We estimated PerRI following Foy et al. For each patient-analyte pair, we fit Gaussian mixture models with one to three components to baseline measurements and selected model order using the Akaike information criterion 16 . From the selected model, we chose the mixture component that explained the largest number of that patient’s baseline measurements and treated it as the individual’s 8
physiological setpoint. We defined the PerRI as the mean of this component ± 2 standard deviations.
NORMA (Normal Outcome Range Modeling with Attention) NORMA is a conditional, decoder-only transformer that models the distribution of a patient’s next laboratory value given their longitudinal measurement history and a query specifying a future health state and prediction horizon. We trained the model on 3.4 million longitudinal sequences from MIMIC-IV and EHRSHOT, requiring at least three repeated measurements per patient without constraints on sampling intervals or clinical setting 55,56 . Data were split at the patient level into training (70%), validation (10%), and test (20%) sets to prevent leakage across partitions. Each sequence was tokenized, padded, and masked to handle variable-length histories. A weighted sampling scheme was used during training to balance representation across 30 laboratory analytes. Alternative input encodings with justifications are summarized (Supplementary Table 3). Input representation. For a given patient-biomarker pair, we denote the ordered sequence of observed laboratory values as x = (x1 , . . . , xT ), with corresponding measurement times t = (t1 , . . . , tT ) and clinical states s = (s1 , . . . , sT ), where each si indicates the laboratory state relative to the population reference interval (low, normal, or high). The model processes these inputs as a sequence of three token types: a context token, history tokens, and a query token. The context token zc encodes static patient covariates: sex g, age a, and laboratory test code c, as a sum of learned embeddings:
zc = eg (g) + ea (a) + ec (c) Each history token zihist represents one prior measurement and combines three components: zihist = fv (x̃i ) + es (si ) + fτ (∆ti ) where fv is a linear projection of the input value x̃i , es is a learned state embedding, and fτ encodes the inter-measurement interval ∆ti = ti − ti −1 . The query token, zquery , specifies the context at the time of prediction: zquery = es (sT +1 ) + fh (tT +1 − tT ) where sT +1 denotes the future health state and tT +1 − tT the prediction horizon. The full sequence is processed using a causally masked Transformer decoder to obtain the predictive distribution of the next observed laboratory value by P(xT +1 |x, t, s, g, a, c, s∗, t ∗). Training Objective. Given the input sequence, the model is optimized to predict the conditional distribution of the next observed value xT +1 . We trained NORMA under two output parameterizations. Under the Gaussian parameterization, the model jointly estimates a predicted mean µ and log-variance log
σ ²: L = 12 [log σ 2 + (xT +1 − µ)2 /σ 2 ] Under the quantile parameterization, the model directly predicts fixed quantiles τ ∈ {0.025, 0.25, 0.50, 0.75, 0.975} using a pinball loss without distributional assumptions:
9
Lτ = τ × max(0, xT +1 − qτ ) + (1 − τ ) × max(0, qτ − xT +1 ) where qτ is the predicted τ -th quantile. Both parameterizations share the same core architecture and were trained on the same development data. NORMA reference intervals. We define the NORMA reference interval (NORMARI ) as the 95% prediction interval obtained by conditioning on a future normal laboratory state. Under the Gaussian parameterization, this corresponds to [µ− 1.96σ , µ + 1.96σ ]; under the quantile parameterization, it corresponds to [q̂0.025 , q̂0.975 ]. Evaluation. We first evaluated forecasting performance for next-step prediction. For point accuracy, we used the 50th percentile (median) prediction and compared it against three baselines: a person-specific historical mean model, last observation carried forward ("Last"), and an autoregressive integrated moving average model (ARIMA). Performance was quantified using mean absolute error (MAE), mean absolute percentage error (MAPE), and coefficient of determination (R²). Sensitivity Analysis. To assess how NORMA captures uncertainty across clinical contexts, we conducted one-at-a-time sensitivity sweeps over five input features: patient age, sex, number of prior measurements, prediction horizon, and within-sequence variability. For each of 30 laboratory tests, we constructed a synthetic baseline sequence: 10 measurements spaced 90 days apart, each set to the midpoint of the sex-specific population reference interval, for a 50-year-old male with a 30-day prediction horizon. We then varied one feature at a time while holding all others at baseline values. Age was swept from 20 to 80 years in 5-year increments; history length from 2 to 300 measurements; and prediction horizon from 7 days to 10 years. For within-sequence variability, we parameterized the standard deviation as a multiplier of one-tenth of the reference range width (0.0 to 3.0) and, at each noise level, drew 30 independent histories from N(midpoint, σ ) to capture sampling variability. For each configuration, we recorded the predicted median (50th quantile), the 90% prediction interval width, and the percent change in interval width relative to baseline, enabling direct comparison of sensitivity magnitude across biomarkers. NORMA configuration. For CHS validation, we applied NORMA with the default configuration: Time2Vec temporal encoding, ternary health-state indicators, raw-value laboratory inputs, raw age projected linearly, and the Gaussian output parameterization. For eICU validation, we used log-delta-t temporal encoding with periodic components, ternary health-state indicators, within-sequence normalization of laboratory values, binned age embeddings in decade-wide bins, a dedicated context token for patient covariates, and the quantile output parameterization. INSPIRE used the same configuration as eICU. All cohort-specific NORMA configurations are summarized in Supplementary Table 3.
Statistical Analysis Mortality association analysis. We assessed the association between laboratory values and mortality using two complementary approaches. First, we grouped index-period measurements into quintiles of raw analyte values within each analyte and estimated mortality rates per quintile. Second, to evaluate whether deviation from a patient’s personal baseline predicts mortality, we computed a deviation score for each index measurement as the absolute z-score from the patient’s baseline mean (|value −µbaseline | / σbaseline ), where
µbaseline and σbaseline are the mean and standard deviation of the patient’s baseline measurements. We grouped deviation scores into deciles and estimated mortality rates within each decile. For both analyses, we computed 95% confidence intervals using Wilson score intervals. Analytes with fewer than 20 observations were excluded.
10
Abnormality classification. We classified each index measurement as normal or abnormal under PopRI , PerRI , and NORMARI . To avoid incorporating pre-existing pathology into PerRI thresholds, we excluded patients whose PerRI mean fell outside the PopRI . Measurements classified as PopRI -abnormal were also classified as abnormal using PerRI and NORMARI . Clinical Risk Stratification and Prediction. We evaluated whether abnormal classifications under each framework identified individuals at elevated risk of future events. For each analyte and framework (PopRI , PerRI , NORMARI ), we defined a binary indicator of abnormality for values outside the corresponding interval. For each analyte–outcome pair, we calculated positive predictive value (probability of the event given abnormal classification), sensitivity (probability of abnormal classification among those with the event), and specificity (probability of normal classification among those without the event). To isolate differences between interval paradigms, we restricted analyses to measurements within the PopRI , because values outside the PopRI necessarily fall outside PerRI and NORMARI . We estimated associations using Cox proportional hazards models with a single binary abnormality indicator adjusted for age and sex. We split data into training (60 percent) and test (40 percent) sets stratified by event status and approximate event time in one-year bins. We evaluated performance at 1, 3, 5, and 10 years using time-dependent receiver operating characteristic curves and quantified uncertainty with 95% confidence intervals from 1,000 bootstrap resamples. We applied Benjamini-Hochberg false discovery rate correction to account for multiple comparisons across outcomes and analytes.
Ethics Approval The use of Clalit Health Services data in this study was approved by the Clalit Health Services Institutional Review Board (Helsinki Committee). MIMIC-IV, eICU-CRD, and INSPIRE were accessed under the PhysioNet Credentialed Health Data Use Agreement; EHRSHOT was accessed under the Stanford Research Use Agreement. All datasets were de-identified prior to release.
Data Availability This study utilizes two publicly available datasets for the development of NORMA: EHRSHOT and MIMIC-IV, both accessible to qualified researchers under their respective data-use agreements. The eICU Collaborative Research Database is also publicly available via PhysioNet. Clalit Health Services data are not publicly available.
Code Availability The full source code for the NORMA model, including data processing, training, and evaluation pipelines, is publicly available at https://github.com/aashnapshah/NORMA. The interactive web-based user interface for individualized laboratory interpretation is available at https://norma-tpy0.onrender.com/.
Acknowledgments We gratefully acknowledge support from the Ivan and Francesca Berkowitz Family Living Laboratory Collaboration at Harvard Medical School and Clalit Research Institute, NIEHS R01ES032470, and NIDDK R01DK137993.
11
References 1. Danielle B Freedman. Towards better test utilization - strategies to improve physician ordering and their impact on patient outcomes. EJIFCC, 26(1):15–30, January 2015. 2. G H Guyatt, A D Oxman, M Ali, A Willan, W McIlroy, and C Patterson. Laboratory diagnosis of irondeficiency anemia: an overview. J. Gen. Intern. Med., 7(2):145–153, March 1992. 3. P Newsome, R Cramb, S Davison, J Dillon, M Foulerton, E Godfrey, Richard Hall, Ulrike Harrower, M Hudson, A Langford, A Mackie, R Mitchell-Thain, K Sennett, N Sheron, J Verne, Martine Walmsley, and A Yeoman. Guidelines on the management of abnormal liver blood tests. Gut, 67:6–19, November 2017. 4. Silvio E Inzucchi. Clinical practice. diagnosis of diabetes. N. Engl. J. Med., 367(6):542–550, August 2012. 5. Kenneth A Sikaris. Enhancing the clinical value of medical laboratory testing. Clin. Biochem. Rev., 38(3): 107–114, November 2017. 6. Ferruccio Ceriotti, Rolf Hinzmann, and Mauro Panteghini. Reference intervals: the way forward. Ann. Clin. Biochem., 46(Pt 1):8–17, January 2009. 7. Richard C Friedberg, Rhona Souers, Elizabeth A Wagar, Ana K Stankovic, Paul N Valenstein, and College of American Pathologists. The origin of reference intervals: A college of american pathologists Q-probes study of “normal ranges” used in 163 clinical laboratories. Arch. Pathol. Lab. Med., 131(3):348–357, March 2007. 8. Joris L J M Müskens, Rudolf Bertijn Kool, Simone A van Dulmen, and Gert P Westert. Overuse of diagnostic testing in healthcare: a systematic review. BMJ Qual. Saf., 31(1):54–63, January 2022. 9. Brian A Ference, Henry N Ginsberg, Ian Graham, Kausik K Ray, Chris J Packard, Eric Bruckert, Robert A Hegele, Ronald M Krauss, Frederick J Raal, Heribert Schunkert, Gerald F Watts, Jan Borén, Sergio Fazio, Jay D Horton, Luis Masana, Stephen J Nicholls, Børge G Nordestgaard, Bart van de Sluis, Marja-Riitta Taskinen, Lale Tokgözoglu, Ulf Landmesser, Ulrich Laufs, Olov Wiklund, Jane K Stock, M John Chapman, and Alberico L Catapano. Low-density lipoproteins cause atherosclerotic cardiovascular disease. 1. evidence from genetic, epidemiologic, and clinical studies. a consensus statement from the european atherosclerosis society consensus panel. Eur. Heart J., 38(32):2459–2472, August 2017. 10. Thore Buergel, Jakob Steinfeldt, Greg Ruyoga, Maik Pietzner, Daniele Bizzarri, Dina Vojinovic, Julius Upmeier Zu Belzen, Lukas Loock, Paul Kittner, Lara Christmann, Noah Hollmann, Henrik Strangalies, Jana M Braunger, Benjamin Wild, Scott T Chiesa, Joachim Spranger, Fabian Klostermann, Erik B van den Akker, Stella Trompet, Simon P Mooijaart, Naveed Sattar, J Wouter Jukema, Birgit Lavrijssen, Maryam Kavousi, Mohsen Ghanbari, Mohammad A Ikram, Eline Slagboom, Mika Kivimaki, Claudia Langenberg, John Deanfield, Roland Eils, and Ulf Landmesser. Metabolomic profiles predict individual multidisease outcomes. Nat. Med., 28(11):2309–2320, November 2022. 11. Edoardo G Giannini, Roberto Testa, and Vincenzo Savarino. Liver enzyme alteration: a guide for clinicians. CMAJ, 172(3):367–379, February 2005. 12. M C Walters and H T Abelson. Interpretation of the complete blood count. Pediatr. Clin. North Am., 43(3): 599–622, June 1996. 13. James C Boyd. Defining laboratory reference values and decision limits: populations, intervals, and interpretations. Asian J. Androl., 12(1):83–90, January 2010. 14. Arjun K Manrai, Chirag J Patel, and John P A Ioannidis. In the era of precision medicine and big data, who is normal? JAMA, 319(19):1981–1982, May 2018. 15. D J Nazir, R S Roberts, S A Hill, and M J McQueen. Monthly intra-individual variation in lipids over a
12
1-year period in 22 normal subjects. Clin. Biochem., 32(5):381–389, July 1999. 16. Brody H Foy, Rachel Petherbridge, Maxwell T Roth, Cindy Zhang, Daniel C De Souza, Christopher Mow, Hasmukh R Patel, Chhaya H Patel, Samantha N Ho, Evie Lam, Camille E Powe, Robert P Hasserjian, Konrad J Karczewski, Veronica Tozzo, and John M Higgins. Haematological setpoints are a stable and patient-specific deep phenotype. Nature, 637(8045):430–438, January 2025. 17. Shuo Wang, Min Zhao, Zihan Su, and Runqing Mu. Annual biological variation and personalized reference intervals of clinical chemistry and hematology analytes. Clin. Chem. Lab. Med., 60(4):606–617, March 2022. 18. Abdurrahman Coskun, Sverre Sandberg, Ibrahim Unsal, Fulya G Yavuz, Coskun Cavusoglu, Mustafa Serteser, Meltem Kilercik, and Aasne K Aarsand. Personalized reference intervals - statistical approaches and considerations. Clin. Chem. Lab. Med., 60(4):629–635, March 2022. 19. Abdurrahman Coşkun, Sverre Sandberg, Ibrahim Unsal, Coskun Cavusoglu, Mustafa Serteser, Meltem Kilercik, and Aasne K Aarsand. Personalized reference intervals in laboratory medicine: A new model based on within-subject biological variation. Clin. Chem., 67(2):374–384, January 2021. 20. A E Obstfeld, K Patel, J C Boyd, J Drees, D T Holmes, J P Ioannidis, and A K Manrai. Data mining approaches to reference interval studies. Clinical Chemistry, 67(9):1175–1181, 2021. 21. Mary E Lacy, Gregory A Wellenius, Anne E Sumner, Adolfo Correa, Mercedes R Carnethon, Robert I Liem, James G Wilson, David B Sacks, David R Jacobs, Jr, April P Carson, Xi Luo, Annie Gjelsvik, Alexander P Reiner, Rakhi P Naik, Simin Liu, Solomon K Musani, Charles B Eaton, and Wen-Chih Wu. Association of sickle cell trait with hemoglobin A1c in african americans. JAMA, 317(5):507–515, February 2017. 22. Kenneth R Feingold. Guidelines for the management of high blood cholesterol. In Endotext [Internet]. MDText. com, Inc., 2025. 23. O Yaw Addo, Emma X Yu, Anne M Williams, Melissa Fox Young, Andrea J Sharma, Zuguo Mei, Nicholas J Kassebaum, Maria Elena D Jefferds, and Parminder S Suchdev. Evaluation of hemoglobin cutoff levels to define anemia among healthy individuals. JAMA Netw. Open, 4(8):e2119123, August 2021. 24. D H Rushton, R Dover, A W Sainsbury, M J Norris, J J Gilkes, and I D Ramsay. Why should women have lower reference limits for haemoglobin and ferritin concentrations than men? BMJ, 322(7298):1355–1357, June 2001. 25. James A Diao, Yixuan He, Rohan Khazanchi, Max Jordan Nguemeni Tiako, Jonathan I Witonsky, Emma Pierson, Pranav Rajpurkar, Jennifer R Elhawary, Luke Melas-Kyriazi, Albert Yen, Alicia R Martin, Sean Levy, Chirag J Patel, Maha Farhat, Luisa N Borrell, Michael H Cho, Edwin K Silverman, Esteban G Burchard, and Arjun K Manrai. Implications of race adjustment in lung-function equations. N. Engl. J. Med., 390(22):2083–2097, June 2024. 26. Darshali A Vyas, Leo G Eisenstein, and David S Jones. Hidden in plain sight—reconsidering the use of race correction in clinical algorithms. New England Journal of Medicine, 383(9):874–882, 2020. 27. Aashna P Shah, James A Diao, Emma Pierson, Chirag J Patel, and Arjun K Manrai. Disentangling proxies of demographic adjustments in clinical equations. arXiv [q-bio.QM], November 2025. 28. Nabihah Tayob and Ziding Feng. Personalized statistical learning algorithms to improve the early detection of cancer using longitudinal biomarkers. Cancer Biomark., 33(2):199–210, 2022. 29. Isaac S Kohane, Daniel R Masys, and Russ B Altman. The incidentalome: a threat to genomic medicine. JAMA, 296(2):212–215, July 2006. 30. Christina Koch, Katherine Roberts, Christopher Petruccelli, and Daniel J Morgan. The frequency of unnecessary testing in hospitalized patients. Am. J. Med., 131(5):500–503, May 2018. 13
31. Henrik L Jørgensen and Bent S Lind. Blood tests - too much of a good thing. Scand. J. Prim. Health Care, 40(2):165–166, June 2022. 32. Christopher Naugler and Irene Ma. More than half of abnormal results from laboratory tests ordered by family physicians could be false-positive. Can. Fam. Physician, 64(3):202–203, March 2018. 33. Tony Badrick, Joe M El-Khoury, and Elvar Theodorsson. Laboratory reference intervals - history and modern approaches for improved utility. Scand. J. Clin. Lab. Invest., 85(4):229–241, June 2025. 34. Davood Roshan, John Ferguson, Charles R Pedlar, Andrew Simpkin, William Wyns, Frank Sullivan, and John Newell. A comparison of methods to generate adaptive reference ranges in longitudinal monitoring. PLoS One, 16(2):e0247338, February 2021. 35. Alvin Rajkomar, Eyal Oren, Kai Chen, Andrew M Dai, Nissan Hajaj, Michaela Hardt, Peter J Liu, Xiaobing Liu, Jake Marcus, Mimi Sun, Patrik Sundberg, Hector Yee, Kun Zhang, Yi Zhang, Gerardo Flores, Gavin E Duggan, Jamie Irvine, Quoc Le, Kurt Litsch, Alexander Mossin, Justin Tansuwan, De Wang, James Wexler, Jimbo Wilson, Dana Ludwig, Samuel L Volchenboum, Katherine Chou, Michael Pearson, Srinivasan Madabushi, Nigam H Shah, Atul J Butte, Michael D Howell, Claire Cui, Greg S Corrado, and Jeffrey Dean. Scalable and accurate deep learning with electronic health records. NPJ Digit. Med., 1(1):18, May 2018. 36. Matthew B A McDermott, Bret Nestor, Peniel Argaw, and Isaac Kohane. Event stream GPT: A data preprocessing and modeling library for generative, pre-trained transformers over continuous-time sequences of complex events. arXiv [cs.LG], June 2023. 37. Zhichao Yang, Avijit Mitra, Weisong Liu, Dan Berlowitz, and Hong Yu. TransformEHR: transformer-based encoder-decoder generative model to enhance prediction of disease outcomes using electronic health records. Nat. Commun., 14(1):7857, November 2023. 38. Lavender Yao Jiang, Xujin Chris Liu, Nima Pour Nejatian, Mustafa Nasir-Moin, Duo Wang, Anas Abidin, Kevin Eaton, Howard Antony Riina, Ilya Laufer, Paawan Punjabi, Madeline Miceli, Nora C Kim, Cordelia Orillac, Zane Schnurman, Christopher Livia, Hannah Weiss, David Kurland, Sean Neifert, Yosef Dastagirzada, Douglas Kondziolka, Alexander T M Cheung, Grace Yang, Ming Cao, Mona Flores, Anthony B Costa, Yindalon Aphinyanaphongs, Kyunghyun Cho, and Eric Karl Oermann. Health system-scale language models are all-purpose prediction engines. Nature, 619(7969):357–362, July 2023. 39. Artem Shmatko, Alexander Wolfgang Jung, Kumar Gaurav, Søren Brunak, Laust Hvas Mortensen, Ewan Birney, Tom Fitzgerald, and Moritz Gerstung. Learning the natural history of human disease with generative transformers. Nature, 647(8088):248–256, November 2025. 40. Shane Waxler, Paul Blazek, Davis White, Daniel Sneider, Kevin Chung, Mani Nagarathnam, Patrick Williams, Hank Voeller, Karen Wong, Matthew Swanhorst, Sheng Zhang, Naoto Usuyama, Cliff Wong, Tristan Naumann, Hoifung Poon, Andrew Loza, Daniella Meeker, Seth Hain, and Rahul Shah. Generative medical event models improve with scale. arXiv [cs.LG], November 2025. 41. Pawel Renc, Yugang Jia, Anthony E Samir, Jaroslaw Was, Quanzheng Li, David W Bates, and Arkadiusz Sitek. Zero shot health trajectory prediction using transformer. NPJ Digit. Med., 7(1):256, September 2024. 42. Andrew Zhang, Tong Ding, Sophia J Wagner, Caiwei Tian, Ming Y Lu, Rowland Pettit, Joshua E Lewis, Alexandre Misrahi, Dandan Mo, Long Phi Le, and Faisal Mahmood. A multimodal and temporal foundation model for virtual patient representations at healthcare system scale. arXiv [cs.LG], April 2026. 43. Guy Lutsker, Gal Sapir, Smadar Shilo, Jordi Merino, Anastasia Godneva, Jerry R Greenfield, Dorit Samocha-Bonet, Raja Dhir, Francisco Gude, Shie Mannor, Eli Meirom, Eric P Xing, Gal Chechik, Hagai Rossman, and Eran Segal. A foundation model for continuous glucose monitoring data. Nature, 650
14
(8103):978–986, February 2026. 44. Ahmed A Metwally, A Ali Heydari, Daniel McDuff, Alexandru Solot, Zeinab Esmaeilpour, Anthony Z Faranesh, Menglian Zhou, Girish Narayanswamy, Maxwell A Xu, Xin Liu, Yuzhe Yang, David B Savage, Mark Malhotra, Conor Heneghan, Shwetak Patel, Cathy Speed, and Javier L Prieto. Insulin resistance prediction from wearables and routine blood biomarkers. Nature, March 2026. 45. Valentyn Melnychuk, Dennis Frauen, and Stefan Feuerriegel. Causal transformer for estimating counterfactual outcomes. arXiv [cs.LG], April 2022. 46. Michelle M Li, Kevin Li, Yasha Ektefaie, Ying Jin, Yepeng Huang, Shvat Messica, Tianxi Cai, and Marinka Zitnik. Controllable sequence editing for biological and clinical trajectories. arXiv [cs.LG], February 2025. 47. Shakson Isaac, Yentl Collin, and Chirag Patel. SSM-CGM: Interpretable state-space forecasting model of continuous glucose monitoring for personalized diabetes management. arXiv [cs.LG], October 2025. 48. Martin W McIntosh, Nicole Urban, and Beth Karlan. Generating longitudinal screening algorithms using novel biomarkers for disease. Cancer Epidemiol. Biomarkers Prev., 11(2):159–166, February 2002. 49. Ruth Johnson, Uri Gottlieb, Galit Shaham, Lihi Eisen, Jacob Waxman, Stav Devons-Sberro, Curtis R Ginder, Peter Hong, Raheel Sayeed, Xiaorui Su, Ben Y Reis, Ran D Balicer, Noa Dagan, and Marinka Zitnik. ClinVec: Unified embeddings of clinical codes enable knowledge-grounded AI in medicine. medRxiv, May 2025. 50. Ran D Balicer, Efrat Shadmi, Nicky Lieberman, Sari Greenberg-Dotan, Margalit Goldfracht, Liora Jana, Arnon D Cohen, Sigal Regev-Rosenberg, and Orit Jacobson. Reducing health disparities: strategy planning and implementation in israel’s largest health care organization. Health Serv. Res., 46(4): 1281–1299, August 2011. 51. Tom J Pollard, Alistair E W Johnson, Jesse D Raffa, Leo A Celi, Roger G Mark, and Omar Badawi. The eICU collaborative research database, a freely available multi-center database for critical care research. Sci. Data, 5(1):180178, September 2018. 52. Leerang Lim, Hyeonhoon Lee, Chul-Woo Jung, Dayeon Sim, Xavier Borrat, Tom J Pollard, Leo A Celi, Roger G Mark, Simon T Vistisen, and Hyung-Chul Lee. INSPIRE, a publicly available research dataset for perioperative medicine. Sci. Data, 11(1):655, June 2024. 53. Clement J McDonald, Stanley M Huff, Jeffrey G Suico, Gilbert Hill, Dennis Leavelle, Raymond Aller, Arden Forrey, Kathy Mercer, Georges DeMoor, John Hook, Warren Williams, James Case, and Pat Maloney. LOINC, a universal standard for identifying laboratory observations: a 5-year update. Clin. Chem., 49(4): 624–633, April 2003. 54. American Board of Internal Medicine. ABIM laboratory test reference ranges. Technical report, January 2025. 55. Alistair Johnson, Lucas Bulgarelli, Tom Pollard, Brian Gow, Benjamin Moody, Steven Horng, Leo Anthony Celi, and Roger Mark. MIMIC-IV, October 2024. 56. Michael Wornow, Rahul Thapa, Ethan Steinberg, Jason A Fries, and Nigam H Shah. EHRSHOT: An EHR benchmark for few-shot evaluation of foundation models. arXiv [cs.LG], July 2023.
15
Figures a Contextual Sequence Modeling
b Training and Validation Cohorts
Labs Time
Sex
NORMA
TRAINING
Autoregressive Transformer
PopRI State
Age
VALIDATION
EHRSHOT
MIMIC-IV
CLALIT
eICU
INSPIRE
5.7k patients 122k sequences
180k patients 3.3M sequences
1.4M patients 32M sequences
98k patients 1M sequences
51K patients 550k sequences
Abn
Predicts distribution for next “healthy” lab
>20 years | 37M sequences | 30 labs
c Reference Interval Paradigms
PopRI
PerRI
NORMARI
Norm
Complete Blood Count HGB (g/dL)
HGB (g/dL)
Population
Hepatic Function
Metabolic Function
Lipid Panel
HGB (g/dL)
Individualized
Contextualized
✗ Personalized
✗ Robust to noise
✓ Personalized
✗ Repeat Labs
✗ Time-Sensitive
✓ Population-Aware
✗ Handles sparsity
✓ Time-Sensitive
d Clinical Utility of NORMARI Glucose → Type 2 Diabetes
✓ Models Uncertainty
PopRI
PerRI
NORMARI
Normal
High
High
e Early Risk Stratification + 12%
+ 12%
NORMARI
Early shift detected by
NORMARI
% Change
Glucose
PopRI
PerRI
+ 4%
- 1%
Specificity
Time
Sensitivity
Precision
NPV
Fig. 1 | NORMA framework overview. a) NORMA conditions on laboratory history, time, sex, age, and population reference state to predict the distribution of the next healthy laboratory value. b) Study design: training on EHRSHOT and MIMIC-IV; external validation on Clalit Health Services, the eICU Collaborative Research Database, and the INSPIRE cohort. Across all cohorts: 20 years, 37 million sequences, and 30 analytes spanning four clinical panels. c) Comparison of PopRI , PerRI , and NORMARI . d) Illustrative example: a glucose value classified as normal by PopRI but flagged abnormal by PerRI and NORMARI , with corresponding change in specificity, sensitivity, and precision for type 2 diabetes classification. e) Schematic of early abnormality detection by NORMARI relative to PopRI and PerRI over time. Created with BioRender.com.
16
a Cohort Selection
c Mortality Association eICU-CRD
BASELINE
2005
2010
2015
2020
2025
Inclusion: ≥5 outpatient labs, 90 days apart; index in 2015
FIRST 75%
Day 1
Day 3
Day 5
Day 7
Day 10
Day 14
CO2
Inclusion: first 75% of measurements as baseline; next as index
Low
Normal
High
Inpatient
Intra
CHS
CV (%)
40 30
INSPIRE
LDL
CV (%)
10 0
Individuality Index
1.00
eICU-CRD
CHS
INSPIRE
0.75 0.50 0.25 0.00
40 35 30 25 20 15 10 5 0 Q1
Q2
Q3
Q4
Q5
30 27 24 21 18 15 12 9 6 3 Q1
DBIL
32 28 24 20 16 12 8 4 45 40 35 30 25 20 15 10 36 33 30 27 24 21 18 15 12 9
HGB
MCHC
30 27 24 21 18 15 12 9
GLU
MCV
RBC 27 24 21 18 15 12 9 6
TC 30 28 26 24 22
CL
K
PLT
TBIL
ALT
40 36 32 28 24 20 16 12 8
40 35 30 25 20 15 10 5 0
32 28 24 20 16 12 8 4
TP
26 24 22 20 18 16 14 12 10
NA
50 45 40 35 30 25 20 15 10
RDW
30 27 24 21 18 15 12 9 6
20
45 40 35 30 25 20 15 10 5
MCH
28 26 24 22 20 18 16 14 12 10
MPV
40 30
33 30 27 24 21 18 15 12
HDL
32 30 28 26 24 22 20 24 22 20 18 16 14 12 10 8
0 50
CRE
25.8 25.6 25.4 25.2 25.0 24.8 24.6 24.4
54 48 42 36 30 24 18 12 6
20 10
CA 36 32 28 24 20 16 12 8
HCT
Mortality Rate (%)
CV (%)
Inter
60 eICU-CRD 50 40 30 20 10 0
ALP
BUN
40 36 32 28 24 20 16 12 8
60 54 48 42 36 30 24 18 12
Outpatient
b Intra-Patient vs Inter-Patient Variation
50
INSPIRE
28 24 20 16 12 8 4
45 40 35 30 25 20 15 10 5
27 24 21 18 15 12 9 6 3
eICU & INSPIRE
ALB
48 42 36 30 24 18 12 6 0
AST
REMAINING
INDEX
A1C
36 32 28 24 20 16 12 8 4 0
CHS
2000
CHS
MONITORING
INDEX
TGL 27 24 21 18 15 12 9
WBC
Q2
Q3
Q4
Q5
NA K MCHC CL MPV CA LDL TP TC HDL RDW AST ALT CO2 HCT TBIL HGB ALP PLT RBC GLU BUN MCV A1C MCH ALB CRE TGL WBC
Quintile
Fig. 2 | Within-person biological variability and mortality association. a) Cohort Selection. Cohort selection and inclusion criteria for Clalit Health Services, the eICU Collaborative Research Database, and the INSPIRE cohort. b) IntraPatient vs Inter-Patient Variation. Within-person versus between-person coefficient of variation for 30 analytes, with the individuality index shown below. The dashed line indicates an individuality index of 0.6. c) Mortality Association. Observed mortality rate by quintile of z-score from the patient baseline mean in Clalit Health Services, the eICU Collaborative Research Database, and the INSPIRE cohort. Shaded bands denote 95% confidence intervals.
17
a NORMA Architecture
b Forecasting Performance
INPUTS
Context
History
Query
EMBEDDINGS
Context
History
h1
h2
NORMA
NORMA
ARIMA
ARIMA
ARIMA
Mean
Mean
Mean
Query
TOKEN SEQUENCE
CTX
NORMA
…
hn
Q
Last
Last
Last
DECODER
0
2
4
6
Masked Self-Attention
8
10
MAE
12
14
0
5
10
15
20
MAPE
25
0.0
0.1
0.2
0.3
0.4
0.5
R2
0.6
0.7
c Per-Analyte Performance
Feed-Forward Network × N l ay e r s
1.0
NORMA ARIMA Mean Last
OUTPUT
0.8
Quantile Head
0.6 R²
Query Token Output
OUTPUT QUANTILES
0.4
q97.5
2.0
TP
1.5
WBC
TC
TGL
TBIL
RBC
RDW
NA
PLT
MPV
MCV
MCH
MCHC
K
LDL
HGB
HCT
HDL
GLU
CRE
DBIL
CL
CO2
CA
A1C
0.0
BUN
0.2 ALT
q75
AST
q50
Reference interval: [q2.5, q97.5]
ALP
q25
ALB
q2.5
d Sensitivity Analysis HFP
12
120
17.5
10
100
12.5 10.0 7.5
8
CI Width (% of ref range)
15.0
6 4
5.0 2
2.5 0.0
101
102
Number of Measurements
0
Lipid
2
Deviation from Midpoint (%)
BMP
20.0
CI Width (% of ref range)
CI Width (% of ref range)
CBC
80 60 40
0 2 4
20 101
102
0 0.0
103
Prediction Horizon (days)
6 0.5
1.0
1.5
2.0
Within-Person SD
2.5
3.0
0.0
0.5
1.0
Within-Person SD
2.5
3.0
Fig. 3 | NORMA architecture, forecasting accuracy, and sensitivity analysis. a) NORMA Architecture. Model architecture: demographic context, laboratory history, and query tokens are jointly embedded and processed by a causal transformer decoder to produce quantile predictions for the next laboratory value. b) Forecasting Performance. Held-out test-set forecasting accuracy (mean absolute error, mean absolute percentage error, and coefficient of determination) for NORMA (quantile) versus ARIMA, population mean, and last-value-carried-forward baselines. Error bars denote 95% bootstrap confidence intervals. c) Per-Analyte Performance. Per-analyte coefficient of determination grouped by clinical panel. d) Sensitivity Analysis. Sensitivity of prediction interval width and midpoint deviation (each as a percentage of PopRI width) to history length, forecast horizon, and within-person variability.
18
a Reclassification Rate
b Hazard Ratios NORMARI
PerRI
PopRI
80 Reclassification Rate (%)
PerRI
NORMARI
ALB HGB
60
PLT HCT
40
MCV RDW
20
WBC CRE A1C AL B AL P AL AST BU T N CA C COL CR2 GL E HCU HDT HGL B K LD MC L MC H H MCC MPV V NA PL RB T RD C W TB IL TC TG L T WB P C
0
CL MCH
2
4
Hazard Ratio
6
8
c True Positive Rate vs False Positive Rate NORMARI
PerRI
Anemia
0.8
GLU
True Positive Rate
0.6
HGB TCLDL HDL BUN
A1C
0.2
MCH
TP WBC CL ALP AST K ALB HCT ALT RDW CA RBC PLT MCHC NA HGB LDL TC HDL BUN
CRE
0.6 GLU
0.4
A1C
0.8
CL
0.6
K ALP AST ALB ALT PLT HCT CRE CAMCHC RDW NA RBC HGB
0.2
0.4 0.6 0.8 False Positive Rate
0.2
0.4 0.6 0.8 False Positive Rate
0.2
0.4 0.6 0.8 False Positive Rate
CA
BUN
NA TC
MCHC
RBC
N BU
T AS
NA
ALB
MCV
HDL
HCT
MCH
L HD
V
LDL
CRE
AST
MCHC
HDL
HCT
MC V MC
MCH
MCH
BU
N
TP
TC
LDL
MCHC
TP
A1C MCV
E CR
TP
HDL
TC
TP
RBC
TC
A1C K ALB
K
MC
HC
A1C
W
TP
RBC
TC
A1C K ALB
K A1C
ALB
MC
V
NA
ALB
MCHC
NA
TP
RDW
LDL
MPV
RD ALB
ALB
RB ALT
AST
C RB ALT
E
CA
C
TP H MC
NA
MCHC
CRE
C
NA
C
ALB
MCH
C RB AST
ALT
A1C NA N BU
CR
CL
TBIL
ALT
A1C
AST
MCH
GLU
CA
MCHC
TP
GLU
MCHC
LDL
RBC
TP NA
L HD
HDL
AST
CA
NA
PLT
L HD ALT
MCHC
T AS HDL
WBC
H MC
CRE
U GL
N
TP
RDW
RBC
Anemia
AST
HDL
TC
ALT
BU
CL
BUN
LDL
TC
ALP
WBC
C
MCHC
ALT
N
MCHC
WBC
CL +18.5%
AST
H
TBIL
RBC
PLT
V
TC
NA
MCH
A1C
BU
MC
ALB
W
HGB
MC
GLU
TP
MCV
H
MC
ALB
HDL
ALT
RBC
GLU
CA
H
RBC
T
AS
PLT MPV
NA
CRE MPV
RD
K
TBIL
CA
LDL
CA
MC MCHC
RDW
CRE
HDL
PLT +54.5%
WBC
PLT CL
C
CL
AST HCT
WB
NA
A1C ALT
PLT
MCH
U GL
ALP
LDL
RBC
CRE +15.2%
A1C
MCV
MCH
TP
ALP
HDL
CRE
W
K
V
HCT
HCT
ALP
MPV
MC
HGB
ALB
Chronic Kidney Disease
TP
Mortality
C
P
Type 2 Diabetes
1.0
MPV
TBIL MPV
RD
H
LDL
AL
TBIL
C
ALP
WB
TC
MC MCHC
RDW
AST
CRE
N
H
CL
BU
MC
LDL
C
RDW
AST
ALB
ALT
HDL
0.4 0.6 0.8 False Positive Rate
TC CL
B
PLT
WB
ALT
LDL
HG
GLU
HCT
TC BUN
TBIL
PLT
ALP
0.2
ALP NA
CRE +22.6%
HGB
NA
PLT +50.4%
CA
GLU
W
T
RD
MPV
TBIL +22.8%
CL HGB
MCV
0.0 0.0
1.0
HGB
PLT
C WB
BUN
HGB
H
MC
C
CA
A1C
RDW
W
GLU
V
ALT
L
HC
GLU
TBIL MPV
MC
CL
LD
A1C
T
GLU +10.1%
WBC
WB
RBC
HDL PLT
ALP
K
TC
HC
CRE
PLT MPV
MPV
RD
K
TBIL
C WB
CRE
TBIL
HCT
AST
GLU
HCT
TC TBIL
HGB
ALT
LDL
MCV
HGB
PLT
NA
ALT
MPV
TBIL +23.3%
A1C
A1C
K CA
MCH
LDL
CL
W
TP
WBC
W
AST
LDL
RD
MPV
RD
CA
CL
CRE +12.6%
ALB
N
PLT +51.7%
H
K
BU
HGB
RBC
MCH
V
NA
BUN
HCT
TBIL
HGB
MC
ALP
MC
CRE
MPV
ALT
U
K
N
CRE
CA
HCT
GL
V
CA
BU
CL
CRE +11.4%
ALP
PLT
ALP
MC ALT
K
MCH
ALB
A1C
RBC
C
TBIL
HDL
HGB
PLT
WB
ALB
A1C
W
TC
T
MPV
RD
BUN
HC
B HG
MPV
TBIL +21.2%
WBC
PLT
TBIL
WBC
TC
K
CRE
TBIL
HCT
TP
B
GLU
PLT +48.6%
CA
V
0.6 0.4
TBIL PLT
HG
BUN
CL
MC
CA
L
HD
HGB
U
RDW
NA
CRE +23.2%
CL
ALT
GL
T
TC
TP K
HCT
PLT
ALP
HC HGB
MCV
LDL
BUN
PLT
ALP
TBIL
HDL
HGB
CL
RBC
MPV
TBIL +20.8%
CRE
W
ALB
ALP
A1C
MCH
AST ALP ALT HCT ALB K RBC MCHC CAHGB NA LDL PLT CRE HDL TC BUN GLU
Accuracy GLU
RD
MCH
HCT
LDL
AST
MPV
MC
C
TPWBC RDW
CL
TBIL
0.2
0.0 0.0
1.0
P
A1C
TBIL
WBC
AL
A1C +9.3%
LDL
AST
P
CL
V
CA
AL RBC
WB
0.8
A1C
Specificity CRE
K
K
CRE
ALB
RDW
CL
BUN
TC
U
GL
MPV
MCV
e Lead Time
Sensitivity
TP
WBC
MCH
HDL BUNTCLDL
d Performance Improvement with NORMA Precision
TP
0.2
0.0 0.0
1.0
GLU
0.4
0.2
0.0 0.0
Type 2 Diabetes
1.0
MPV MCV TBIL
MCV TBIL
CL K ALP AST ALT MCHCALB RDW PLTCA RBC HCT CRENA
0.4
Mortality
1.0
MPV
TBIL
True Positive Rate
TPWBC MCH
0.8 True Positive Rate
1.0
MPV MCV
True Positive Rate
1.0
Chronic Kidney Disease
CA GLU CL K CO2
0
20 40 Lead Time (months)
60
Fig. 4 | Reclassification performance and early risk stratification in Clalit Health Services. a) Reclassification Rate. Reclassification rate by analyte among PopRI -normal tests. b) Hazard Ratios. Cox hazard ratios for all-cause mortality (adjusted for age and sex); analytes with p < 0.05 shown. c) True Positive Rate vs False Positive Rate. True positive rate versus false positive rate for clinical outcome prediction across anemia, chronic kidney disease, all-cause mortality, and type 2 diabetes, comparing PerRI and NORMARI . d) Performance Improvement with NORMA. Change in precision, sensitivity, specificity, and balanced accuracy (NORMARI versus PerRI ); inner label shows the largest gain. e) Lead Time. Median lead time from NORMARI flag to first PopRI flag, by analyte (months). Error bars denote the interquartile range.
19
a Reclassification Rate
b Hazard Ratios NORMARI
PerRI
1.5 Hazard Ratio
2
WBC CRE
60
RDW TGL
40
CO2
20
CA TP
AL B AL P AL AST BU T N CA C COL CR2 DB E I GL L HCU HDT HGL B K LD MC L MC H H MCC MPV V NA PL RB T RD C W TB IL TC TG L T WB P C
Reclassification Rate (%)
NORMARI
PLT
80
0
PerRI
PopRI
100
BUN CL
0.5
1
2.5
c True Positive Rate vs False Positive Rate NORMARI
PerRI
Acute Kidney Injury
1.0 MPV
0.4
PLT MCH MCV KALPALT CA TPRDW CO2 MCHC ALB RBC TGL CL AST HGB TBIL WBC HCT BUN DBIL GLU CRE
0.2 0.0 0.0
0.2
0.4 0.6 0.8 False Positive Rate
1.0
NA
0.6 0.4
PLT MCV MCH
0.2
ALP K ALT HGB ALB MCHC HCT TGL RDW CA AST TBIL CO2 TP CL RBC BUN CRE WBC GLU
0.0 0.0
DBIL
0.2
0.8 NA
0.6 PLT
0.4 MCH MCV K CA ALT CO2 ALP RDW TPMCHC ALB GLURBC CL BUN ASTTGL HGB TBIL WBC HCT CRE DBIL
0.2 0.4 0.6 0.8 False Positive Rate
0.0 0.0
1.0
0.2
0.4 0.6 0.8 False Positive Rate
ALB
ALT CL
CL MC
V
MCH
TP
ALP
MC
BUN NA
MCV
HC
PLT
CA
T
MPV
TP
CO2
WB C
HC T
NA
ALB TGL
RBC
DB
P
IL
CA
RDW
AL
TP
HCT
BUN
MC
HCT
MC V
BUN
HC
HCT
BUN
V MC
BUN
V MC
HCT
GLU GLU
CA
CRE
2 CO
2 CO
GLU
MCV
RBC
CA
GLU ALB
DBIL
K
RDW
CA
DBIL
V
GLU
DBIL
CA
HGB
CRE
CO2
PLT
BUN WB C
BUN
MC H
CA
GLU
CA
CRE
GLU
ALT
MPV
HCT
CL
TBIL
RDW
TP
PLT
ALB
MCV
PLT
RBC
HGB
HCT
RBC
MCV
MCHC
N
ALB +4.8% CRE
RBC
CO2
ALP
B
K
ALB
PLT
CA
TP 2 CO ALP
PLT
TGL
HGB
BUN
MPV
Acute Kidney Injury Mortality
NA MC HC
ALT
RDW
RBC
CRE
HG
AST
MPV
IL
CRE CO2
BU
MCV MCHC
MCH
NA RDW
DBIL
T
ALP
CO2
TB IL
H
MC
BUN
DBIL
TGL
HCT
RBC
CRE
WB C
GLU
CO2 +8.9%
ALB
BUN
MCV
AST
BUN
TBIL
AST
ALT
CRE
TBIL
V
CL
CA
AST
ALB
T
PL
PL
LDL TP
MP
Prolonged LOS (>7d)
TBIL
CRE
ALP
ALT
NA DB
WBC
MCHC
ALT
K
K
ALP
WBC
CL
B
L
ALP +41.9%
RBC
NA
AL
MPV
MPV
RBC
AST HGB
K
GLU RDW
TP
K
TG
TBIL
WBC
RBC
CO2
TP
MCH
ALB
RDW
TGL
ALP T
ALB
C
NA
AST
MCH
IL
DB
GLU +5.8%
IL
PLT
MCHC
MCH
ALT
DB
HGB
NA
MCH
CRE
CL
AL
C
WB
CO2
U
C
MCH
K
MPV
HGB
NA IL
ALB
BUN
ALP
ALT TBIL
CO2
L
ALP +43.0% TB
MPV +35.3%
TGL
MCV
AST
L
TG
WBC
HCT
MCV
HGB
TP
AST
CR E
W
RD
TBIL
TGL
RBC
MPV
CRE
PLT
CL
MCH
TG
MPV
HGB
T
AST
ALT
T
TP MCH
GLU
MPV TGL
GL
C
ALB
MCHC
AL
K
AST
CRE
K
RB
CL
GLU +5.0%
HGB
NA
ALP
H MC
CL
M
TP
HCT CO2
MCHC
MPV +35.0%
CHC
TBIL
DBIL
CA
NA
WBC
HGB
CO2
1.0
PLT ALP
HC
T
AS
IL
RD W
U
GL
0.4 0.6 0.8 False Positive Rate
TGL
CO2
DBIL
NA
DB
K
PLT
TP
MP V
W
RD
RDW
HGB
RBC
CO2 TGL
DBIL
BUN
CA
HCT
CL
CO2 +10.6%
WBC
TBIL
TGL
CO2 +4.1%
T
ALP +41.8%
AST
ALT
CA
ALP
CR E PLT
MPV
AST
T
TBIL
CL
TP
AL
C
WB
GLU
MCHC
MCH
TGL
HGB
K
NA
MPV
MPV
AL
HGB
ALP
RDW
ALB
MPV
MPV +32.1%
RDW
MCHC
PLT
TP
WBC
CL
GLU
PLT
NA
BUN
TP
MCH CA
ALT
RD W
C
WB
TGL
BUN
ALB
DBIL
ALB
BU
GLU
IL
CL
IL
MCV
ALP
DB
K
DB
MCH
TGL
NA
TBIL
V
MCHC
RBC
GLU +3.8%
K
TP
ALP +41.9%
MC
K
IL
HCT
AST
CA
TB
V MC
ALP
HC T
C
WB
C
HCT
MCH
2
HGB
MCH
AST
RDW
AST
TBIL
CRE
CL
CO
MCH
Sepsis
MCHC
HG B
T
AL
NA
MPV
ALP
DBIL
CRE
TBIL
MPV +35.7%
0.2
TC N
T
CA
ALT
C
WBC
MCH
PL
T
AS
MCH
K
CL
GLU
RDW
MCV
ALB
CO2
PLT
GLU +4.9%
CL
B
AL
WBC
RBC
RBC
IL
PLT MCH MCV CA KALPALT ALB TPMCHC RDW HGB CO2 TGL AST CL TBIL HCT DBIL WBC BUN GLU CRE RBC
Accuracy
Specificity
TP
CL
NA
CRE
TB
K
H
K
TGL
WBC
ALP
MC
0.4
0.0 0.0
1.0
NA
e Lead Time
Sensitivity
RDW
0.6
0.2
d Performance Improvement with NORMA Precision
MPV
0.8 True Positive Rate
NA
Sepsis
1.0 MPV
0.8 True Positive Rate
True Positive Rate
0.8 0.6
Prolonged LOS (>7d)
1.0 MPV
True Positive Rate
1.0
Mortality
CA HCT RBC HGB GLU
0
25
50 75 100 Lead Time (hours)
125
Fig. 5 | Reclassification performance and early risk stratification in the eICU Collaborative Research Database. a) Reclassification Rate. Reclassification rate by analyte among PopRI -normal tests. b) Hazard Ratios. Cox hazard ratios for in-hospital mortality (adjusted for age and sex); analytes with p < 0.05 shown. c) True Positive Rate vs False Positive Rate. True positive rate versus false positive rate for clinical outcome prediction across in-hospital mortality, acute kidney injury, sepsis, and prolonged ICU stay, comparing PerRI and NORMARI . d) Performance Improvement with NORMA. Change in precision, sensitivity, specificity, and balanced accuracy (NORMARI versus PerRI ); inner label shows the largest gain. e) Lead Time. Median lead time from NORMARI flag to first PopRI flag, by analyte (hours). Error bars denote the interquartile range.
20
a Reclassification Rate
b Hazard Ratios NORMARI
PerRI
60 50
PerRI
NORMARI
4
5 Hazard Ratio
6
TP
40
K
30
ALB CL
20
CA
TP WB C
B AL P AL T AS T BU N
AL
K NA PL T TB IL
WBC CL CO 2 CR E GL U HC T HG B
TBIL
0
CA
10 A1C
Reclassification Rate (%)
PopRI PLT
ALP AST
1
2
3
7
8
9
c True Positive Rate vs False Positive Rate NORMARI
PerRI
1.0
1.0
0.4
CO2
NA K
0.2
PLT GLU
0.0 0.0
0.2
0.6 0.4 GLU CO2 PLT
0.2
CREHCTALP ALT CABUN CLAST TBIL WBC HGB ALB TP
0.4 0.6 0.8 False Positive Rate
K ALT NA ASTALP HCT TBIL TP HGB WBC CRE ALB CA CL BUN
0.0 0.0
1.0
0.2
0.6 0.4 GLU CO2
0.2 KTBIL AST NA WBC CLHCT HGB BUN TP ALB
0.4 0.6 0.8 False Positive Rate
PLT
0.2
0.4 0.6 0.8 False Positive Rate
NA
K ALP
AST HGB
HCT
ALT
CL
TP
HCT
TP CRE
NA
CL
GLU +5.9%
GLU
HCT CO2
CRE TB IL
N
BU
ALB
CA
CA ALB
NA
CA C WB
NA
HCT
CA
K
CL
CO2
HCT
PLT
NA
CA ALB
CA ALB
TBIL
HGB
CA
NA
WBC
HCT
PLT
ALB
CRE
BUN
CL
HCT TBIL
TP
K K
CL
CRE
K CA
K
PLT
ALB ALT
ALP
GLU
BUN
K
K HGB
NA
ALP
PLT
ALP
GLU
AST
GLU
HGB
ALT
Mortality
CRE
ALB
BUN
PLT
TBIL
AST
E
TBIL
TBIL
CO2
ALP
TP
AST
PLT GLU
CO2 +6.3% B
CO2
HGB
BUN
ALB
AST
CRE
ALB
HG
CR
K
ALB
ALP +36.4%
HCT
T
WBC
GLU
T
HC
TP
CA
CL TBIL
C
C WB
AS
GLU
ALP
E
PLT
GLU +12.2%
ALB
NA
CL
ALT
PLT
B
CO2
CR
WB
TBIL
P
TP
CR
ALP
GLU
BUN
CO2
AST
E
ALT
AL
CL
CA
K
CR
HCT
PLT TP
BUN
CL +14.3%
HGB
TBIL
PLT
TBIL
N
AST
HG
E
B
NA
AST
CRE
WBC
U
CO2 +5.3% CA
ALP +31.5%
PLT
CA AL
T
HC
C
GLU
BU
2 CO
T
K
AST
TBIL
BUN
CO2
WB
ALT
TP
AS
B
HG
B
WBC
ALT
ALT
CA
CRE
PLT
GLU
ALP
2
CL
ALP
1.0
NA GL
N
TP
NA
ALT
BU
AST
BUN
TP
CO2
GLU +5.0%
0.4 0.6 0.8 False Positive Rate
TBIL B
K
GLU
ALP HG
T
CO2
HG
CL
CL
0.2
ALP PLT
TP
CO TBIL
HCT
AL
ALB
CRE
WBC
GLU +7.3%
E
CR
C
TBIL
ALB
K
HGB
WB
WBC
ALP +36.5%
T
GLU
BU
PLT
2 CO
K
T
ALP AS BUN
PLT
CL
HCT
N
ALB
HCT
CL
BU
GLU
0.0 0.0
CL
CO2 +11.1%
CRE
ALT
HGB
TBIL
GLU +11.8%
PLT
NA
AL
CA
NA
ALP
N
T
TP
ALP
2
AL
CO2
AST
WBC
CO
K
HCT
C
WB
TBIL
HGB
TP
WBC
BU
ALP +36.1%
AST
TP
TP +4.2%
ALB
ALB
HGB
K
C
CO2
TP
P
WB
NA
CL
CO2
AL
TP
CA
TP
AS T
CL
GLU
HC T
N
T
CO2 +13.9%
HGB
CRE
Perioperative Infection
NA
U
GLU
Prolonged LOS (>7d)
TBIL GL
T
AL
Unplanned ICU Admission
ALT
BUN
NA
AL
K
ALT
AST
CA
HGB
NA +4.6%
ALP
CRE
1.0
GLU CO2 PLT ALT K ALP HCT NA TBIL CL AST HGB BUN WBC ALB CRE TP CA
Accuracy
Specificity NA
PLT
BUN
CA
2 CO
C
WBC
CL
WB
0.4
e Lead Time
Sensitivity
HCT
0.6
ALT ALP
CA 0.0 CRE 0.0 0.2
1.0
Unplanned ICU Admission
0.8
d Performance Improvement with NORMA Precision
1.0
0.8 True Positive Rate
0.6
Prolonged LOS (>7d)
1.0
0.8 True Positive Rate
0.8 True Positive Rate
Perioperative Infection
True Positive Rate
Mortality
0
20
40 60 80 100 Lead Time (hours)
120
Fig. 6 | Reclassification performance and early risk stratification in the INSPIRE cohort. a) Reclassification Rate. Reclassification rate by analyte among PopRI -normal tests. b) Hazard Ratios. Cox hazard ratios for in-hospital mortality (adjusted for age and sex); analytes with p < 0.05 shown. c) True Positive Rate vs False Positive Rate. True positive rate versus false positive rate for clinical outcome prediction across in-hospital mortality, perioperative infection, prolonged hospital stay, and unplanned ICU admission, comparing PerRI and NORMARI . d) Performance Improvement with NORMA. Change in precision, sensitivity, specificity, and balanced accuracy (NORMARI versus PerRI ); inner label shows the largest gain. e) Lead Time. Median lead time from NORMARI flag to first PopRI flag, by analyte (hours). Error bars denote the interquartile range.
21
Supplementary Figures a Forecasting Performance NORMA
NORMA
NORMA
ARIMA
ARIMA
ARIMA
Mean
Mean
Mean
Last
Last
Last
0
2
4
6
8
10
MAE
12
14
0
5
10
15
20
MAPE
25
0.0
0.2
0.4
0.6
R2
0.8
b Per-Analyte Performance
1.0
NORMA ARIMA Mean Last
0.8
R²
0.6 0.4
TP
WBC
TC
1.5
2.0
TGL
TBIL
RBC
RDW
NA
PLT
MPV
MCV
MCH
MCHC
K
LDL
HGB
HCT
HDL
GLU
CRE
DBIL
CO2
CL
CA
BUN
ALT
AST
ALP
A1C
0.0
ALB
0.2
c Sensitivity Analysis CBC
BMP
HFP
80
60 40
60 40 20
20 101
102
Number of Measurements
0
100
0
80
10
Deviation from Midpoint (%)
80
CI Width (% of ref range)
100
0
10
120
100
CI Width (% of ref range)
CI Width (% of ref range)
120
Lipid
60 40 20
101
102
0 0.0
103
Prediction Horizon (days)
20 30 40
0.5
1.0
1.5
2.0
Within-Person SD
2.5
3.0
0.0
0.5
1.0
Within-Person SD
2.5
3.0
Supplementary Fig. 1 | NORMA forecasting performance with Gaussian loss. a) Forecasting Performance. Held-out test-set accuracy (mean absolute error, mean absolute percentage error, and coefficient of determination) for the Gaussian variant versus ARIMA, population mean, and last-value-carried-forward baselines. Error bars denote 95% bootstrap confidence intervals. b) Per-Analyte Performance. Per-analyte coefficient of determination grouped by clinical panel. c) Sensitivity Analysis. Sensitivity of Gaussian-head prediction interval width and midpoint deviation (each as a percentage of PopRI width) to history length, forecast horizon, and within-person variability.
22
MAE
MAPE
R²
A1C
0.4
0.4
0.5
6.1
5.7
7.7
0.64
0.70
ALB
0.3
0.3
0.3
8.5
7.7
10.2
0.74
0.84
0.37 0.75
ALP
16.0
15.6
24.3
15.4
15.2
19.9
0.74
0.80
0.78 0.57
ALT
7.2
6.2
14.5
27.8
24.7
40.4
0.62
0.79
AST
7.0
6.0
20.4
23.5
21.0
44.5
0.58
0.77
0.41
BUN
3.3
3.3
4.4
18.4
17.2
23.2
0.83
0.83
0.82
CA
0.3
0.3
0.4
3.3
3.0
4.4
0.66
0.76
CL
1.9
1.8
2.4
1.8
1.8
2.3
0.69
0.75
CO2
1.5
1.3
2.1
6.3
5.4
8.3
0.70
0.80
CRE
0.1
0.1
0.2
10.6
12.3
12.9
0.88
0.83
DBIL
0.3
0.3
0.4
48.0
36.5
88.0
0.51
0.64
GLU
15.7
14.0
21.1
12.9
11.3
17.7
0.48
0.61
HCT
2.1
2.0
3.0
6.2
6.1
8.7
12.1
11.1
18.8
6.3
6.2
8.5
HDL
6.5
6.0
8.8
HGB
0.7
0.7
1.0
K
0.3
0.3
0.4
LDL
17.7
15.0
18.6
MCH
0.5
0.6
0.6
MCHC
0.6
0.6
0.6
MCV
1.6
1.9
1.7
MPV
0.6
0.7
0.9
NA
1.8
1.8
20.0 17.5 15.0 12.5 10.0
25
20
15
0.79
0.83
0.73
0.80
0.79
0.84
0.40
1.0
0.55 0.48 0.87
0.8
0.57 0.23 0.60
0.6
0.53 0.63
6.7
6.3
8.5
0.45
0.56
21.8
20.0
33.7
0.55
0.72
1.8
2.1
2.2
0.90
0.90
1.8
1.9
1.9
0.72
0.72
1.8
2.2
1.9
0.88
0.85
6.5
6.7
9.9
0.59
0.66
0.21
2.2
1.3
1.3
1.6
0.62
0.64
0.49
7.5 5.0 2.5
10
5
0.15
0.4
0.52 0.83 0.58
0.2
0.82
0.0
PLT
30.3
52.8
36.6
13.5
19.3
18.1
0.79
0.29
RBC
0.2
0.2
0.3
6.3
6.4
8.5
0.81
0.84
0.68
RDW
0.5
0.5
0.7
3.0
3.4
4.6
0.87
0.86
0.81
TBIL
0.2
0.1
0.3
28.0
22.5
32.1
0.71
0.81
0.85
24.8
12.0
9.8
17.5
0.57
0.76
0.44
28.7
22.8
0.48
0.71
0.64
4.7
6.9
0.54
0.74
0.57
24.2
19.6
26.0
0.62
0.79
0.69
IM AR
us sia NGa
NQu an
IM AR
us sia NGa
NQu an
IM AR
us sia NGa
NQu an
A
5.7
2.1
n
0.5
1.3
tile
0.3
1.5
A
0.4
n
TP WBC
tile
28.7
A
17.0
36.2
n
20.3
tile
TC TGL
Supplementary Fig. 2 | Per-analyte forecasting performance. NORMA using quantile and Gaussian loss, compared to ARIMA, the best performing baseline. Color intensity encodes mean absolute error (MAE), mean absolute percentage error (MAPE), and coefficient of determination (R2 ) on the held-out test set.
23
eICU-CRD
A1C
ALB
30
30
30
25
20
20 15
10
10
CA
CL
35 30 25
25
25
20
20
15
15
10
10 5
CO2
20
HCT
15
20
10
50
30
40 30
35
35
26
30
30
24
25
25
20
20
15
15
20
10
MCHC
25.0 22.5 20.0 17.5 15.0 12.5 10.0
10
MCV
25.0 22.5 20.0 17.5 15.0 12.5 10.0
RDW 20 15 10 0
1
2
3
4
5
27 26
26
25
4
5
25
27
20 15 10
RBC
25
16
20
14 12
10
10
TP
0
1
2
3
4
5
24
WBC
30
25
24 3
MCH
28
TGL
25
2
10
10
20
1
15
PLT
20
27
0
15
15
TC
5
20
24
30
TBIL
10
20
25
40
10
15
25
NA
15
30
25
25
LDL
10
20
GLU
26
MPV
25
10
K
40
28
22
20
HGB
40
15
5
10
5
HDL
10
DBIL
20
10
10
25 20
CRE
30
0
1
2
BUN
30
15
25
15
Mortality Rate (%)
25
40
30
20
AST
20
50
40
INSPIRE
ALT
5
0 40
CHS
ALP
3
4
5
30
25
25
20
20
15
15
10
10 0
1
2
3
4
5
0
1
2
3
4
5
Standardized Distance from Baseline (|z|)
Supplementary Fig. 3 | Association of absolute deviation from the personalized baseline with observed mortality. Each panel shows the observed mortality rate by standardized absolute deviation from the personalized baseline, binned by decile, for Clalit Health Services, the eICU Collaborative Research Database, and the INSPIRE cohort.
24
eICU-CRD
A1C
ALB
7.5
25
90
7.0
20 15
6.0
10
5.5
60
5
50
80 70
CA
CL
CO2
115
AST
10
NORMA Midpoint Prediction
HDL
60
5
47.5 45.0 42.5 40.0 37.5 35.0
55
40
50
30
45
20
40
10
MCHC
47.5
MCV
42.5 40.0 37.5 35.0
60
80
NA
145
80
15
PLT
135
TC
30 20 10
TGL
TP
WBC
150
20 15
140
125
20
40
60
80
100
60
25
80
15 10
50 40
20
20
75
20
25
30
100
120
RBC
40
325 300 275 250 225 200 175
140
TBIL 160
0
85
155 150
5 40
MCH 40.0 37.5 35.0 32.5 30.0 27.5
90
MPV
25
10
20
75
100
10
20
25
125 100
95
10
RDW 40 35 30 25 20 15 10
175 150
LDL
20
90
32.5
GLU 200
K 30
30
100
45.0
HGB
50
DBIL
15.0 12.5 10.0 7.5 5.0 2.5 0.0
0
HCT
50.0
15 10
10
15
25
25 20
20
20
100
30
30
CRE
BUN
35
40
25
30
105
10
ALT
30
35
110 15
INSPIRE
45 40 35 30 25 20 15
100
6.5
20
CHS
ALP
10 20
40
60
80
5 20
40
60
80
20
40
60
80
Age (years)
Supplementary Fig. 4 | Age-dependent NORMARI midpoints. Each panel shows individual midpoint predictions with a polynomial smooth overlaid, for Clalit Health Services, the eICU Collaborative Research Database, and the INSPIRE cohort.
25
PopRI
PerRI
NORMARI
100
eICU-CRD
80
60
40
20
0 100
60
CHS
Abnormality Prevalence (%)
80
40
20
0 100
INSPIRE
80
60
40
20
TC TG L TP WB C
K LD L MC H MC HC MC V MP V NA PL T RB C RD W TB IL
CL CO 2 CR E DB IL GL U HC T HD L HG B
CA
A1C AL B AL P AL T AS T BU N
0
Supplementary Fig. 5 | Abnormality prevalence by analyte and reference interval method. Grouped bars show the percentage of measurements classified as abnormal by PopRI , PerRI , and NORMARI , for the eICU Collaborative Research Database, Clalit Health Services, and the INSPIRE cohort.
26
a Hazard Ratios Across All Outcomes PopRI
Chronic Kidney Disease
NORMARI
PerRI
Mortality
Type 2 Diabetes
A1C ALB ALP ALT AST BUN CA CRE GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TP WBC 1
10
Hazard Ratio
0.1
1
10
Hazard Ratio
1
10
Hazard Ratio
b Concordance Index Comparison PopRI
Chronic Kidney Disease MCHC
PerRI
Mortality
NORMARI
Type 2 Diabetes
MPV
ALP
RDW
MCH
TP
MCH
RDW
CA
MCV
MCV
ALT
BUN
MCHC
AST
HCT
CRE
NA
WBC
WBC
WBC
PLT
HCT
CRE
HGB
HGB
GLU
CRE
PLT 0.65
0.70
0.75
C-index
0.80
0.85
CL 0.65
0.70
0.75
C-index
0.80
0.85
0.500 0.525 0.550 0.575 0.600 0.625 0.650
C-index
Supplementary Fig. 6 | Proportional hazards analysis in Clalit Health Services. a) Hazard Ratios Across All Outcomes. Cox hazard ratios for chronic kidney disease, all-cause mortality, type 2 diabetes, and anemia, comparing PopRI , PerRI , and NORMARI abnormality flags; only analytes with p < 0.05 are shown. b) Concordance Index Comparison. Concordance index for the top 10 analytes ranked by NORMARI concordance per outcome.
27
a Hazard Ratios Across All Outcomes Acute Kidney Injury
PopRI
Mortality
PerRI
NORMARI
Prolonged LOS (>7d)
Sepsis
ALB ALP AST BUN CA CL CO2 CRE DBIL HCT HGB K MCH MCHC MCV MPV NA PLT RBC RDW TBIL TP WBC 1
2 × 100 Hazard Ratio
3 × 100
1
4 × 100
Hazard Ratio
2 × 100
3 × 100
4 × 10 1
PerRI
NORMARI
1
6 × 10 1 Hazard Ratio
1
2 × 100 Hazard Ratio
3 × 100
b Concordance Index Comparison Acute Kidney Injury
PopRI
Mortality
Prolonged LOS (>7d)
Sepsis
WBC
CA
RBC
ALP
TBIL
BUN
CL
MCH
ALP
CL
BUN
CO2
CL
AST
ALP
WBC
CA
CO2
NA
CL
K
PLT
WBC
CRE
BUN
RDW
RDW
MCHC
CRE
TGL
K
CA
RDW
CRE
MCHC
RDW
TC
WBC 0.50
0.55
0.60
C-index
0.65
0.70
CRE 0.50
0.55
0.60
C-index
0.65
0.70
0.45
K 0.50
0.55
C-index
0.60
0.65
0.45
0.50
0.55
C-index
0.60
Supplementary Fig. 7 | Proportional hazards analysis in the eICU Collaborative Research Database. a) Hazard Ratios Across All Outcomes. Cox hazard ratios for acute kidney injury, in-hospital mortality, prolonged ICU stay, and sepsis, comparing PopRI , PerRI , and NORMARI abnormality flags; only analytes with p < 0.05 are shown. b) Concordance Index Comparison. Concordance index for the top 10 analytes ranked by NORMARI concordance per outcome.
28
a Hazard Ratios Across All Outcomes Mortality
PopRI
Perioperative Infection
PerRI
NORMARI
Prolonged LOS (>7d)
Unplanned ICU Admission
ALB ALP ALT AST BUN CA CL CO2 CRE GLU HCT HGB K NA PLT TBIL TP WBC 1
1
Hazard Ratio
1.2 × 100
1.4 × 100 Hazard Ratio
1.6 × 100
7 × 10 1
1.8 × 100
8 × 10 1 9 × 10 1 Hazard Ratio
1
8 × 10 1
9 × 10 1
1
Hazard Ratio
b Concordance Index Comparison Mortality
Perioperative Infection
PopRI
PerRI
NORMARI
Prolonged LOS (>7d)
Unplanned ICU Admission
ALP
NA
CL
ALT
NA
HGB
PLT
TBIL
TBIL
CL
WBC
ALP
CL
TP
NA
CA
CA
K
TP
TP
ALB
ALT
BUN
NA
TP
HCT
HCT
ALB
WBC
ALB
HGB
WBC
K
PLT
ALB
CL
PLT
AST
CRE
CRE
0.5
0.6
0.7
C-index
0.8
0.9
0.50
0.55
0.60
C-index
0.65
0.70
0.45
0.50
0.55
C-index
0.60
0.65
0.45
0.50
0.55
C-index
0.60
Supplementary Fig. 8 | Proportional hazards analysis in the INSPIRE cohort. a) Hazard Ratios Across All Outcomes. Cox hazard ratios for in-hospital mortality, perioperative infection, prolonged hospital stay, and unplanned ICU admission, comparing PopRI , PerRI , and NORMARI abnormality flags; only analytes with p < 0.05 are shown. b) Concordance Index Comparison. Concordance index for the top 10 analytes ranked by NORMARI concordance per outcome.
29
Anemia
WBC TP TC TBIL RDW RBC PLT NA MPV MCV MCHC MCH LDL K HGB HDL HCT GLU CRE CL CA BUN AST ALT ALP ALB A1C
Chronic Kidney Disease
WBC TP TC TBIL RDW RBC PLT NA MPV MCV MCHC MCH LDL K HGB HDL HCT GLU CRE CL CA BUN AST ALT ALP ALB A1C
Mortality
WBC TP TC TBIL RDW RBC PLT NA MPV MCV MCHC MCH LDL K HGB HDL HCT GLU CRE CL CA BUN AST ALT ALP ALB A1C
Type 2 Diabetes
Precision
WBC TP TC TBIL RDW RBC PLT NA MPV MCV MCHC MCH LDL K HGB HDL HCT GLU CRE CL CA BUN AST ALT ALP ALB A1C
0.2
0.4
0.6
0.8
1.0
0.4
0.6
NORMARI
PerRI
Sensitivity
0.8
1.0
0.0
Specificity
0.2
0.4
Accuracy
0.6
0.8
0.45 0.50 0.55 0.60 0.65 0.70
Supplementary Fig. 9 | Clinical outcome prediction performance in Clalit Health Services. Each point shows PerRI and NORMARI values for precision, sensitivity, specificity, and balanced accuracy across outcomes. Connected pairs show the direction and magnitude of change between methods for each analyte. Analytes with fewer than 100 measurements are omitted.
30
Acute Kidney Injury
WBC TP TGL TBIL RDW RBC PLT NA MPV MCV MCHC MCH K HGB HCT GLU DBIL CRE CO2 CL CA BUN AST ALT ALP ALB
Mortality
WBC TP TGL TBIL RDW RBC PLT NA MPV MCV MCHC MCH K HGB HCT GLU DBIL CRE CO2 CL CA BUN AST ALT ALP ALB
Prolonged LOS (>7d)
WBC TP TGL TBIL RDW RBC PLT NA MPV MCV MCHC MCH K HGB HCT GLU DBIL CRE CO2 CL CA BUN AST ALT ALP ALB
Sepsis
Precision
WBC TP TGL TBIL RDW RBC PLT NA MPV MCV MCHC MCH K HGB HCT GLU DBIL CRE CO2 CL CA BUN AST ALT ALP ALB
0.1
0.2
0.3
0.4
Sensitivity
0.5
0.6
0.0
0.2
0.4
0.6
PerRI
NORMARI
0.8
0.2
Specificity
0.4
0.6
Accuracy
0.8
0.44 0.46 0.48 0.50 0.52 0.54
Supplementary Fig. 10 | Clinical outcome prediction performance in the eICU Collaborative Research Database. Each point shows PerRI and NORMARI values for precision, sensitivity, specificity, and balanced accuracy across outcomes. Connected pairs show the direction and magnitude of change between methods for each analyte. Analytes with fewer than 100 measurements are omitted.
31
Precision
Sensitivity
PerRI
0.2
0.4
NORMARI
Specificity
Accuracy
WBC TP TBIL PLT NA K
Mortality
HGB HCT GLU CRE CO2 CL CA BUN AST ALT ALP ALB
WBC TP TBIL
Perioperative Infection
PLT NA K HGB HCT GLU CRE CO2 CL CA BUN AST ALT ALP ALB
WBC TP TBIL
Prolonged LOS (>7d)
PLT NA K HGB HCT GLU CRE CO2 CL CA BUN AST ALT ALP ALB
WBC TP
Unplanned ICU Admission
TBIL PLT NA K HGB HCT GLU CRE CO2 CL CA BUN AST ALT ALP ALB
0.0
0.2
0.4
0.6
0.8
0.0
0.1
0.3
0.5
0.5
0.6
0.7
0.8
0.9
0.45
0.50
0.55
0.60
Supplementary Fig. 11 | Clinical outcome prediction performance in the INSPIRE cohort. Each point shows PerRI and NORMARI values for precision, sensitivity, specificity, and balanced accuracy across outcomes. Connected pairs show the direction and magnitude of change between methods for each analyte. Analytes with fewer than 100 measurements are omitted.
32
Supplementary Tables Supplementary Table 1 | Analyte reference information. Full analyte name, abbreviation, unit of measurement, and conventional PopRI for all 30 laboratory analytes included in the study. PopRI were obtained from the American Board of Internal Medicine Laboratory Test Reference Ranges (January 2025). Analyte
Full Name
Unit
Population RI
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
Hemoglobin A1c Albumin Alkaline Phosphatase Alanine Aminotransferase Aspartate Aminotransferase Blood Urea Nitrogen Calcium Chloride Bicarbonate Creatinine Direct Bilirubin Glucose Hematocrit HDL Cholesterol Hemoglobin Potassium LDL Cholesterol Mean Corpuscular Hemoglobin MCH Concentration Mean Corpuscular Volume Mean Platelet Volume Sodium Platelet Count Red Blood Cell Count Red Cell Distribution Width Total Bilirubin Total Cholesterol Triglycerides Total Protein White Blood Cell Count
% g/dL U/L U/L U/L mg/dL mg/dL mEq/L mEq/L mg/dL mg/dL mg/dL % mg/dL g/dL mEq/L mg/dL pg g/dL fL fL mEq/L 103 /µL 106 /µL % mg/dL mg/dL mg/dL g/dL 103 /µL
4.0–5.6 3.5–5.5 30–120 0–35 0–35 8–20 8.6–10.2 98–106 23–28 F: 0.5–1.1, M: 0.7–1.3 0.1–0.3 70–99 F: 37–47, M: 42–50 40–100 F: 12.0–16.0, M: 14.0–18.0 3.5–5.0 <130 28–32 33–36 80–98 7.5–12.5 136–145 150–450 F: 4.0–5.2, M: 4.5–5.9 9–14.5 0.3–1.0 100–200 <150 6.0–8.0 4.5–11.0
33
Supplementary Table 2 | Cohort characteristics. Patient counts, demographics (age, sex, follow-up duration), and per-analyte measurement counts and mean values (± standard deviation) across all five cohorts: EHRSHOT, MIMIC-IV, Clalit Health Services, the eICU Collaborative Research Database, and the INSPIRE cohort.
Patients Age Sex Time
EHRSHOT
MIMIC-IV
eICU-CRD
CHS
INSPIRE
5,676 54.9 ± 17.2 52% F / 48% M 2.49 [0.28, 7.20] yr
179,601 58.2 ± 19.0 54% F / 46% M 1.01 [0.06, 4.08] yr
98,432 63.3 ± 16.1 46% F / 54% M 0.02 [0.01, 0.03] yr
1,450,862 55.1 ± 18.6 63% F / 37% M 20.24 [17.51, 21.47] yr
51,159 57.6 ± 14.9 52% F / 48% M 0.48 [0.17, 0.84] yr
Analyte
N
Value
N
Value
N
Value
N
Value
N
Value
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
1,745 4,282 4,206 4,219 4,220 5,180 5,162 5,172 5,187 5,161 1,429 5,150 5,470 1,549 5,369 5,236 821 5,294 5,317 5,304 371 5,193 5,320 5,312 5,315 4,169 1,635 4 4,213 5,345
6.3 ± 1.2 3.3 ± 0.8 112.5 ± 58.8 37.8 ± 25.2 31.9 ± 19.5 23.0 ± 15.2 8.8 ± 0.7 102.4 ± 5.0 25.9 ± 3.8 1.1 ± 0.5 0.3 ± 0.3 124.9 ± 39.9 32.2 ± 6.8 53.6 ± 17.5 10.7 ± 2.3 4.1 ± 0.5 94.9 ± 34.9 30.4 ± 2.7 33.2 ± 1.4 91.5 ± 6.9 9.6 ± 1.6 137.4 ± 4.2 198.0 ± 111.4 3.5 ± 0.8 16.1 ± 2.9 0.6 ± 0.4 173.5 ± 45.0 165.7 ± 100.8 6.6 ± 1.0 7.8 ± 4.6
39,996 55,220 78,930 92,599 90,639 153,329 121,478 148,304 145,386 156,044 6,635 147,195 164,903 49,011 162,168 150,333 48,688 161,671 161,806 161,661 — 149,303 162,308 161,825 161,640 73,002 49,762 48,445 18,673 164,291
6.6 ± 1.3 3.8 ± 0.8 107.0 ± 58.4 30.2 ± 21.9 32.7 ± 21.4 22.2 ± 14.3 8.8 ± 0.7 101.8 ± 5.2 25.6 ± 4.1 1.1 ± 0.6 0.7 ± 0.9 118.3 ± 40.4 33.6 ± 6.8 55.8 ± 17.7 11.0 ± 2.3 4.2 ± 0.5 104.2 ± 37.8 29.7 ± 2.8 32.6 ± 1.6 91.2 ± 7.3 — 138.7 ± 4.2 229.4 ± 113.3 3.7 ± 0.8 15.1 ± 2.4 0.7 ± 0.6 187.5 ± 44.6 140.4 ± 77.5 7.0 ± 0.7 7.7 ± 4.0
— 19,075 13,582 12,425 12,704 57,540 57,067 59,172 70,208 54,820 1,147 90,958 60,110 11 60,482 66,569 3 48,432 51,710 50,841 37,083 62,419 53,371 53,584 48,715 12,738 35 312 15,134 52,141
— 2.6 ± 0.7 97.0 ± 43.2 33.8 ± 23.4 37.1 ± 25.6 25.3 ± 15.5 8.3 ± 0.7 103.8 ± 6.2 25.0 ± 4.7 1.1 ± 0.6 0.6 ± 0.5 139.6 ± 44.3 30.1 ± 5.9 26.0 ± 14.3 9.9 ± 2.0 4.0 ± 0.6 92.1 ± 56.2 29.7 ± 2.1 32.8 ± 1.3 90.4 ± 5.7 9.7 ± 1.4 138.6 ± 5.0 199.0 ± 98.4 3.4 ± 0.7 15.8 ± 2.0 0.8 ± 0.5 108.3 ± 44.0 144.9 ± 68.0 5.8 ± 1.0 10.6 ± 4.6
336,493 790,253 1,099,044 1,233,902 1,236,781 1,251,453 909,036 31,171 299 1,310,465 — 1,351,168 1,407,912 1,002,204 1,408,118 1,182,544 1,187,844 1,250,759 1,407,568 1,400,648 1,304,857 1,187,537 1,407,716 1,404,794 1,403,339 815,774 1,260,730 1,244,518 791,481 1,409,027
7.1 ± 1.3 4.1 ± 0.5 82.1 ± 34.3 21.9 ± 17.3 23.6 ± 16.3 17.4 ± 8.3 9.3 ± 0.5 103.1 ± 3.9 31.4 ± 12.1 0.9 ± 0.4 — 110.0 ± 36.1 39.1 ± 5.2 48.0 ± 15.9 12.8 ± 1.8 4.5 ± 0.5 105.1 ± 33.6 28.8 ± 2.3 32.8 ± 1.3 87.7 ± 6.1 9.5 ± 1.4 139.9 ± 3.0 243.4 ± 73.6 4.5 ± 0.6 14.1 ± 1.5 0.6 ± 0.4 181.7 ± 40.7 133.5 ± 65.3 7.1 ± 0.6 7.5 ± 3.7
391 30,728 27,797 26,144 26,374 29,912 29,788 30,952 6,247 42,255 — 26,168 39,739 — 36,187 36,363 — — — — — 36,127 34,006 — — 28,571 — — 28,603 34,273
6.7 ± 1.0 3.4 ± 0.6 73.4 ± 29.1 21.0 ± 12.5 24.2 ± 10.0 16.1 ± 8.0 8.5 ± 0.7 102.5 ± 4.8 24.8 ± 3.8 0.8 ± 0.2 — 143.8 ± 48.4 32.9 ± 5.8 — 11.1 ± 2.0 4.0 ± 0.5 — — — — — 137.4 ± 3.5 206.0 ± 92.3 — — 0.8 ± 0.4 — — 6.1 ± 0.9 8.2 ± 3.5
34
Supplementary Table 3 | NORMA architectural and training design choices. Summary of configurable components including temporal encoding, health state encoding, age and laboratory value representations, output parameterization, and context token usage. Component Time embedding
Design Choice
Description
Time2Vec
Encodes the number of days elapsed since the first measurement using learned linear and sinusoidal functions. Captures long-term temporal trends and potential periodic patterns.
Delta-t
Encodes the log-compressed time difference between consecutive measurements. Explicitly represents irregular sampling and emphasizes how much time has passed since the last value.
Health state encoding
Binary
Encodes each lab value as either within or outside the population reference range; does not distinguish high vs. low and can struggle with J-shaped biomarkers.
Ternary
Encodes whether the lab value is low, normal, or high relative to the reference range; handles asymmetric risk patterns; however, can increase overfitting risk.
Age encoding
Raw
Uses age directly as a numeric input; assumes linear effects.
Binned
Discretizes age into groups (e.g., decades) and maps each bin to a learned embedding vector; captures nonlinear and threshold effects.
Lab value encoding
Raw
Uses observed biomarker values directly as model input; preserves absolute clinical scale and keeps preprocessing minimal.
Within-sequence nor-
Applies per-sequence z-scoring with a learnable scale and
malization
shift, then denormalizes at output; stabilizes optimization and improves cross-analyte sharing.
Training objective
Gaussian NLL
Minimizes the negative log-likelihood of a normal distribution by predicting mean and log-variance; aligns with conventional reference interval assumptions but may underperform for skewed or heavy-tailed analytes.
Quantile (pinball) loss
Minimizes pinball loss for fixed quantiles (2.5, 25, 50, 75, 97.5%); distribution-free and robust to skewness or heteroscedasticity.
35
Supplementary Table 4 | Within-person and between-person biological variability. Per-analyte within-person coefficient of variation, between-person coefficient of variation, and individuality index for the eICU Collaborative Research Database and Clalit Health Services. Values are reported as median [95% confidence interval]. eICU-CRD
CHS
INSPIRE
Analyte
CVintra
CVinter
II
CVintra
CVinter
II
CVintra
CVinter
II
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
— 8.9 [8.8, 8.9] 10.6 [10.5, 10.7] 15.5 [15.3, 15.7] 18.3 [18.1, 18.5] 16.5 [16.4, 16.6] 3.8 [3.7, 3.8] 2.0 [2.0, 2.0] 7.1 [7.1, 7.2] 13.5 [13.4, 13.5] — 14.8 [14.8, 14.9] 6.2 [6.2, 6.3] 7.6 [6.0, 9.5] 6.7 [6.7, 6.7] 7.6 [7.6, 7.6] 8.6 [6.0, 11.5] 1.1 [1.1, 1.1] 1.3 [1.3, 1.3] 1.0 [1.0, 1.0] 2.9 [2.9, 2.9] 1.3 [1.3, 1.3] 11.8 [11.7, 11.9] 7.2 [7.2, 7.3] 2.1 [2.0, 2.1] 21.9 [21.8, 22.1] 7.6 [6.6, 8.8] 16.0 [15.3, 16.6] 6.5 [6.4, 6.5] 15.6 [15.5, 15.7]
— 22.6 [22.4, 22.7] 42.1 [41.8, 42.4] 65.3 [64.8, 65.8] 63.4 [62.9, 63.8] 58.7 [58.5, 58.9] 7.2 [7.1, 7.2] 5.1 [5.0, 5.1] 15.9 [15.9, 16.0] 47.7 [47.5, 47.9] 82.8 [81.0, 84.7] 24.8 [24.7, 25.0] 18.4 [18.3, 18.5] 39.7 [33.9, 44.8] 19.1 [19.0, 19.1] 10.5 [10.4, 10.5] 45.6 [38.4, 52.0] 6.9 [6.9, 7.0] 3.6 [3.6, 3.7] 6.2 [6.2, 6.3] 12.8 [12.8, 12.9] 2.9 [2.9, 2.9] 41.6 [41.4, 41.8] 18.6 [18.5, 18.7] 12.6 [12.6, 12.7] 58.2 [57.7, 58.6] 38.4 [34.9, 41.4] 43.2 [41.8, 44.6] 14.3 [14.2, 14.4] 37.5 [37.3, 37.6]
— 0.4 [0.4, 0.4] 0.2 [0.2, 0.3] 0.2 [0.2, 0.2] 0.3 [0.3, 0.3] 0.3 [0.3, 0.3] 0.5 [0.5, 0.5] 0.4 [0.4, 0.4] 0.5 [0.5, 0.5] 0.3 [0.3, 0.3] — 0.6 [0.6, 0.6] 0.3 [0.3, 0.3] 0.2 [0.2, 0.2] 0.3 [0.3, 0.3] 0.7 [0.7, 0.7] 0.2 [0.1, 0.3] 0.1 [0.1, 0.1] 0.3 [0.3, 0.3] 0.2 [0.2, 0.2] 0.2 [0.2, 0.2] 0.4 [0.4, 0.4] 0.3 [0.3, 0.3] 0.4 [0.4, 0.4] 0.2 [0.2, 0.2] 0.4 [0.4, 0.4] 0.2 [0.2, 0.2] 0.4 [0.3, 0.4] 0.5 [0.5, 0.5] 0.4 [0.4, 0.4]
8.6 [8.5, 8.6] 6.7 [6.7, 6.8] 17.8 [17.7, 17.8] 34.2 [34.2, 34.3] 24.7 [24.6, 24.7] 20.5 [20.4, 20.5] 3.5 [3.5, 3.5] 2.4 [2.3, 2.4] 17.7 [17.3, 18.2] 12.2 [12.2, 12.2] — 13.7 [13.7, 13.7] 6.3 [6.3, 6.4] 24.4 [24.4, 24.5] 6.3 [6.3, 6.3] 7.1 [7.1, 7.1] 20.6 [20.5, 20.6] 3.6 [3.6, 3.6] 2.9 [2.9, 2.9] 3.2 [3.2, 3.2] 10.3 [10.3, 10.3] 1.5 [1.5, 1.5] 14.1 [14.1, 14.1] 5.9 [5.9, 5.9] 5.7 [5.7, 5.7] 26.0 [25.9, 26.0] 13.9 [13.9, 14.0] — 4.8 [4.8, 4.8] 19.3 [19.2, 19.3]
16.5 [16.5, 16.6] 18.9 [18.3, 19.4] 30.1 [30.0, 30.2] 50.2 [50.0, 50.5] 35.1 [34.9, 35.4] 36.2 [36.2, 36.3] 3.7 [3.7, 3.7] 2.3 [2.2, 2.3] 27.0 [25.9, 28.1] 34.7 [34.4, 34.9] — 24.0 [23.9, 24.0] 9.6 [9.6, 9.6] 30.1 [30.1, 30.2] 10.5 [10.5, 10.5] 6.5 [6.4, 6.5] 23.7 [23.7, 23.8] 7.1 [7.1, 7.1] 2.7 [2.7, 2.7] 6.1 [6.1, 6.1] 10.8 [10.8, 10.8] 1.2 [1.2, 1.2] 23.8 [23.7, 23.8] 10.1 [10.1, 10.1] 7.1 [7.1, 7.1] 42.8 [42.6, 43.0] 16.6 [16.6, 16.6] — 5.6 [5.5, 5.6] —
0.5 [0.5, 0.5] 0.4 [0.3, 0.4] 0.6 [0.6, 0.6] 0.7 [0.7, 0.7] 0.7 [0.7, 0.7] 0.6 [0.6, 0.6] 0.9 [0.9, 0.9] 1.0 [1.0, 1.1] 0.7 [0.6, 0.7] 0.3 [0.3, 0.4] — 0.6 [0.6, 0.6] 0.7 [0.7, 0.7] 0.8 [0.8, 0.8] 0.6 [0.6, 0.6] 1.1 [1.1, 1.1] 0.9 [0.9, 0.9] 0.5 [0.5, 0.5] 1.1 [1.1, 1.1] 0.5 [0.5, 0.5] 0.9 [0.9, 1.0] 1.2 [1.2, 1.2] 0.6 [0.6, 0.6] 0.6 [0.6, 0.6] 0.8 [0.8, 0.8] 0.6 [0.6, 0.6] 0.8 [0.8, 0.8] — 0.9 [0.9, 0.9] 0.0 [0.0, 0.7]
4.8 [4.4, 5.3] 8.4 [8.3, 8.4] 10.9 [10.7, 11.0] 19.2 [19.0, 19.4] 14.4 [14.3, 14.6] 15.4 [15.2, 15.5] 4.4 [4.3, 4.4] 1.9 [1.8, 1.9] 5.5 [5.4, 5.6] 11.0 [11.0, 11.1] — 12.8 [12.7, 12.9] 6.8 [6.7, 6.8] — 6.9 [6.8, 6.9] 7.5 [7.5, 7.6] — — — — — 1.0 [1.0, 1.0] 12.0 [11.9, 12.2] — — 21.7 [21.6, 21.9] — — 7.5 [7.4, 7.5] 18.3 [18.2, 18.5]
13.0 [11.8, 13.8] 13.4 [13.3, 13.5] 34.0 [33.7, 34.4] 49.8 [49.3, 50.3] 32.6 [32.2, 32.9] 42.0 [41.5, 42.4] 5.6 [5.5, 5.6] 3.3 [3.3, 3.3] 10.2 [10.1, 10.4] 25.3 [25.2, 25.5] — 23.7 [23.5, 23.9] 14.1 [14.0, 14.2] — 14.6 [14.5, 14.7] 8.6 [8.6, 8.7] — — — — — 1.7 [1.7, 1.7] 33.5 [33.2, 33.8] — — 44.6 [44.2, 45.0] — — 10.9 [10.8, 11.0] 31.8 [31.5, 32.0]
0.4 [0.3, 0.4] 0.6 [0.6, 0.6] 0.3 [0.3, 0.3] 0.4 [0.4, 0.4] 0.4 [0.4, 0.5] 0.4 [0.4, 0.4] 0.8 [0.8, 0.8] 0.6 [0.6, 0.6] 0.5 [0.5, 0.6] 0.4 [0.4, 0.4] — 0.5 [0.5, 0.6] 0.5 [0.5, 0.5] — 0.5 [0.5, 0.5] 0.9 [0.9, 0.9] — — — — — 0.6 [0.6, 0.6] 0.4 [0.3, 0.4] — — 0.5 [0.5, 0.5] — — 0.7 [0.7, 0.7] 0.6 [0.6, 0.6]
36
Supplementary Table 5 | NORMA forecasting performance. Mean absolute error (MAE), mean absolute percentage error (MAPE), and coefficient of determination (R2 ) for NORMA Gaussian, NORMA quantile, ARIMA, population mean, and last-value-carried-forward, reported on training, validation, and held-out test sets. Values are reported as median [interquartile range] across analytes. Metric
NORMA Gaussian
NORMA Quantile
ARIMA
Mean
Last
MAE
Train Val Test
6.0 [3.2, 11.1] 6.0 [2.6, 10.7] 6.0 [2.6, 10.2]
5.7 [2.9, 10.1] 5.8 [2.9, 8.7] 5.9 [3.6, 9.1]
6.8 [2.5, 10.2] 6.8 [3.6, 10.0] 6.7 [4.3, 11.9]
8.4 [4.5, 15.0] 8.3 [4.9, 13.5] 8.4 [4.5, 13.6]
7.4 [3.7, 11.0] 7.3 [3.8, 10.3] 7.3 [3.9, 10.7]
MAPE
Train Val Test
10.9 [8.8, 14.0] 11.1 [7.9, 13.2] 11.1 [8.5, 14.3]
11.9 [8.6, 14.7] 12.6 [8.7, 15.6] 12.3 [7.9, 17.3]
15.9 [11.3, 20.1] 15.7 [12.1, 21.5] 16.9 [11.3, 23.3]
19.4 [14.8, 25.4] 19.7 [14.0, 27.4] 19.7 [13.5, 26.7]
14.8 [10.9, 19.6] 15.2 [10.6, 20.1] 15.2 [10.6, 21.1]
R2
Train Val Test
0.76 [0.72, 0.80] 0.75 [0.69, 0.78] 0.75 [0.70, 0.78]
0.70 [0.66, 0.75] 0.69 [0.66, 0.73] 0.68 [0.64, 0.72]
0.62 [0.57, 0.70] 0.62 [0.53, 0.67] 0.58 [0.51, 0.64]
0.46 [0.41, 0.49] 0.46 [0.43, 0.49] 0.45 [0.41, 0.50]
0.56 [0.47, 0.62] 0.55 [0.49, 0.62] 0.54 [0.49, 0.61]
37
Supplementary Table 6 | Sensitivity of NORMA prediction intervals to patient features. Median and interquartile range of prediction interval width change (as a percentage of PopRI width) and midpoint shift (as a percentage of PopRI width) in response to history length, forecast horizon, and within-person variability, for both the quantile and Gaussian parameterizations. CI Width Change (%)
Midpoint Shift (%)
Median
IQR
Median
IQR
Model
Feature
Quantile Quantile Quantile
History Length Prediction Horizon Within-Person Variability
4.8 1.6 115.2
[1.0, 14.4] [0.1, 11.1] [111.4, 120.5]
0.0 0.0 0.7
[0.0, 0.0] [0.0, 0.0] [0.3, 1.7]
Gaussian Gaussian Gaussian
History Length Prediction Horizon Within-Person Variability
28.3 7.2 7.2
[20.6, 41.8] [3.0, 9.9] [3.7, 12.4]
12.8 1.6 1.3
[5.6, 20.6] [0.8, 3.2] [0.3, 3.6]
38
Supplementary Table 7 | Abnormality prevalence and reclassification rates. Overall percentage of measurements classified as abnormal by PopRI , PerRI , and NORMARI , and the reclassification rate (the proportion of PopRI -normal tests flagged abnormal by PerRI or NORMARI ) in the eICU Collaborative Research Database and Clalit Health Services. Dataset
N
PopRI (%)
PerRI (%)
NORMARI (%)
PerRI RR (%)
NORMARI RR (%)
eICU-CRD CHS INSPIRE
29 31 19
50.2 29.6 35.1
68.1 46.8 57.0
55.8 39.1 42.5
37.3 27.2 34.3
9.2 12.0 14.4
39
Supplementary Table 8 | Per-analyte reclassification. Among tests with a PerRI setpoint within PopRI , the number per 1,000 flagged abnormal by PopRI , PerRI , and NORMARI for each analyte, in the eICU Collaborative Research Database and Clalit Health Services. CHS
eICU-CRD
INSPIRE
Analyte
PopRI
PerRI
NORMARI
PopRI
PerRI
NORMARI
PopRI
PerRI
NORMARI
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
240 146 119 136 71 235 140 229 545 141 — 326 322 157 279 122 118 184 400 94 76 111 75 240 234 119 184 — 102 146
315 284 435 414 383 458 290 415 602 256 — 517 497 445 424 250 414 363 498 375 314 361 412 486 378 220 443 — 273 377
282 244 218 210 142 267 199 282 562 191 — 350 377 217 331 186 195 290 439 266 880 150 125 291 304 489 240 — 220 228
— 466 224 252 186 360 245 322 428 231 136 541 535 — 518 148 — 83 289 48 44 246 122 475 248 149 286 212 277 256
— 551 535 531 441 546 378 461 512 333 215 586 677 1,000 643 281 800 308 452 372 320 418 524 570 456 293 743 519 427 435
— 504 294 327 240 400 300 360 467 266 172 556 566 — 555 200 — 177 339 141 616 300 310 509 312 200 400 285 329 298
278 121 123 225 45 254 194 274 475 123 — 619 304 — 282 145 — — — — — 160 88 — — 72 — — 120 230
472 284 534 492 396 479 345 430 577 234 — 696 517 — 478 331 — — — — — 409 540 — — 266 — — 310 395
333 151 225 326 118 285 220 310 562 160 — 719 351 — 315 229 — — — — — 230 312 — — 138 — — 150 266
40
Supplementary Table 9 | In-hospital mortality prediction performance in the eICU Collaborative Research Database. Per-analyte precision, sensitivity, specificity, and balanced accuracy for PerRI and NORMARI abnormality flags, restricted to measurements classified as normal by PopRI . Analytes with fewer than 100 measurements are omitted (—). PerRI
NORMARI
Analyte
N
Events
Precision
Sensitivity
Specificity
Accuracy
Precision
Sensitivity
Specificity
Accuracy
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
— 891 8,843 7,391 8,035 19,881 16,083 26,608 29,891 27,847 289 9,534 2,919 — 3,383 57,801 — 32,416 22,846 43,792 34,479 40,769 33,546 4,595 16,275 8,985 11 176 4,859 25,827
— 49 1,143 913 793 1,058 1,214 2,149 2,393 1,883 49 696 254 — 264 5,680 — 3,272 1,998 4,335 3,564 3,828 2,580 355 889 1,041 2 34 486 1,714
— 0.05 0.13 0.12 0.10 0.06 0.08 0.08 0.06 0.08 0.06 0.04 0.10 — 0.09 0.11 — 0.11 0.09 0.11 0.12 0.10 0.07 0.10 0.06 0.14 0.29 0.20 0.11 0.07
— 0.20 0.60 0.56 0.53 0.48 0.33 0.37 0.23 0.31 0.06 0.15 0.49 — 0.42 0.38 — 0.49 0.41 0.61 0.59 0.44 0.65 0.34 0.49 0.35 1.00 0.56 0.38 0.42
— 0.75 0.42 0.43 0.48 0.54 0.67 0.62 0.69 0.75 0.80 0.74 0.58 — 0.64 0.66 — 0.54 0.61 0.43 0.48 0.57 0.31 0.73 0.57 0.70 0.44 0.46 0.67 0.58
— 0.48 0.51 0.50 0.50 0.51 0.50 0.49 0.46 0.53 0.43 0.45 0.53 — 0.53 0.52 — 0.51 0.51 0.52 0.53 0.50 0.48 0.54 0.53 0.53 0.72 0.51 0.52 0.50
— 0.08 0.16 0.12 0.11 0.06 0.08 0.10 0.08 0.10 0.04 0.08 0.13 — 0.12 0.12 — 0.12 0.10 0.14 0.11 0.10 0.07 0.10 0.06 0.15 0.67 0.22 0.11 0.07
— 0.18 0.21 0.20 0.17 0.15 0.17 0.16 0.16 0.15 0.02 0.10 0.18 — 0.20 0.20 — 0.29 0.18 0.30 0.91 0.65 0.37 0.15 0.17 0.16 1.00 0.18 0.16 0.15
— 0.88 0.83 0.80 0.85 0.87 0.83 0.88 0.83 0.90 0.90 0.91 0.89 — 0.87 0.83 — 0.76 0.85 0.79 0.13 0.37 0.59 0.88 0.84 0.88 0.89 0.85 0.86 0.87
— 0.53 0.52 0.50 0.51 0.51 0.50 0.52 0.50 0.53 0.46 0.51 0.53 — 0.54 0.52 — 0.53 0.51 0.54 0.52 0.51 0.48 0.52 0.51 0.52 0.94 0.51 0.51 0.51
41
Supplementary Table 10 | Acute kidney injury prediction performance in the eICU Collaborative Research Database. Per-analyte precision, sensitivity, specificity, and balanced accuracy for PerRI and NORMARI abnormality flags, restricted to measurements classified as normal by PopRI . Analytes with fewer than 100 measurements are omitted (—). PerRI
NORMARI
Analyte
N
Events
Precision
Sensitivity
Specificity
Accuracy
Precision
Sensitivity
Specificity
Accuracy
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
— 891 8,843 7,391 8,035 19,881 16,083 26,608 29,891 27,847 289 9,534 2,919 — 3,383 57,801 — 32,416 22,846 43,792 34,479 40,769 33,546 4,595 16,275 8,985 11 176 4,859 25,827
— 130 1,712 1,408 1,384 1,164 2,079 3,839 3,601 2,062 78 1,257 266 — 255 8,390 — 4,932 3,339 6,744 5,045 6,154 4,508 429 1,681 1,660 3 42 864 3,491
— 0.14 0.19 0.17 0.17 0.05 0.13 0.13 0.09 0.07 0.36 0.11 0.10 — 0.09 0.14 — 0.16 0.14 0.15 0.15 0.14 0.13 0.11 0.10 0.19 0.29 0.21 0.19 0.12
— 0.23 0.58 0.51 0.50 0.38 0.34 0.34 0.23 0.23 0.23 0.20 0.47 — 0.43 0.34 — 0.48 0.37 0.56 0.53 0.40 0.65 0.31 0.42 0.31 0.67 0.48 0.36 0.39
— 0.75 0.41 0.42 0.47 0.53 0.67 0.62 0.68 0.74 0.85 0.74 0.58 — 0.64 0.66 — 0.54 0.61 0.43 0.47 0.56 0.30 0.73 0.57 0.70 0.38 0.43 0.67 0.57
— 0.49 0.50 0.47 0.49 0.46 0.51 0.48 0.45 0.48 0.54 0.47 0.53 — 0.54 0.50 — 0.51 0.49 0.49 0.50 0.48 0.48 0.52 0.50 0.51 0.52 0.46 0.52 0.48
— 0.17 0.21 0.18 0.16 0.05 0.15 0.16 0.11 0.05 0.33 0.15 0.10 — 0.08 0.17 — 0.17 0.14 0.16 0.15 0.15 0.13 0.11 0.11 0.20 0.33 0.22 0.21 0.14
— 0.15 0.18 0.19 0.14 0.12 0.20 0.14 0.16 0.06 0.10 0.10 0.13 — 0.14 0.20 — 0.28 0.15 0.23 0.88 0.61 0.39 0.14 0.17 0.13 0.33 0.14 0.17 0.13
— 0.88 0.83 0.80 0.85 0.87 0.83 0.88 0.83 0.89 0.92 0.92 0.89 — 0.87 0.83 — 0.76 0.84 0.78 0.13 0.37 0.59 0.88 0.84 0.88 0.75 0.84 0.87 0.87
— 0.51 0.51 0.49 0.49 0.49 0.52 0.51 0.49 0.48 0.51 0.51 0.51 — 0.50 0.52 — 0.52 0.50 0.50 0.51 0.49 0.49 0.51 0.51 0.51 0.54 0.49 0.52 0.50
42
Supplementary Table 11 | Sepsis prediction performance in the eICU Collaborative Research Database. Peranalyte precision, sensitivity, specificity, and balanced accuracy for PerRI and NORMARI abnormality flags, restricted to measurements classified as normal by PopRI . Analytes with fewer than 100 measurements are omitted (—). PerRI
NORMARI
Analyte
N
Events
Precision
Sensitivity
Specificity
Accuracy
Precision
Sensitivity
Specificity
Accuracy
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
— 891 8,843 7,391 8,035 19,881 16,083 26,608 29,891 27,847 289 9,534 2,919 — 3,383 57,801 — 32,416 22,846 43,792 34,479 40,769 33,546 4,595 16,275 8,985 11 176 4,859 25,827
— 87 2,106 1,925 2,036 2,759 2,034 4,164 4,190 4,045 72 1,557 332 — 363 9,268 — 5,623 3,540 7,707 5,542 6,847 5,676 594 2,064 2,286 2 56 1,118 3,531
— 0.08 0.23 0.25 0.24 0.12 0.13 0.14 0.11 0.14 0.32 0.13 0.11 — 0.10 0.16 — 0.17 0.15 0.17 0.16 0.15 0.17 0.13 0.12 0.26 0.29 0.35 0.24 0.12
— 0.20 0.57 0.54 0.49 0.40 0.34 0.34 0.25 0.25 0.22 0.20 0.40 — 0.35 0.34 — 0.47 0.36 0.57 0.52 0.40 0.67 0.27 0.41 0.30 1.00 0.61 0.35 0.36
— 0.75 0.41 0.42 0.47 0.53 0.67 0.62 0.68 0.74 0.84 0.74 0.57 — 0.63 0.65 — 0.53 0.61 0.43 0.47 0.56 0.31 0.73 0.57 0.70 0.44 0.48 0.67 0.57
— 0.47 0.49 0.48 0.48 0.47 0.51 0.48 0.47 0.50 0.53 0.47 0.49 — 0.49 0.49 — 0.50 0.49 0.50 0.50 0.48 0.49 0.50 0.49 0.50 0.72 0.55 0.51 0.47
— 0.13 0.23 0.25 0.23 0.12 0.14 0.17 0.12 0.13 0.33 0.18 0.12 — 0.12 0.17 — 0.19 0.15 0.19 0.16 0.17 0.17 0.10 0.12 0.25 0.00 0.30 0.26 0.12
— 0.16 0.16 0.19 0.14 0.11 0.19 0.13 0.14 0.10 0.11 0.10 0.12 — 0.15 0.18 — 0.27 0.15 0.23 0.88 0.62 0.41 0.09 0.16 0.12 0.00 0.14 0.16 0.11
— 0.88 0.83 0.80 0.85 0.87 0.83 0.87 0.83 0.90 0.93 0.92 0.89 — 0.87 0.83 — 0.76 0.84 0.79 0.13 0.37 0.59 0.88 0.84 0.88 0.67 0.84 0.87 0.87
— 0.52 0.50 0.49 0.49 0.49 0.51 0.50 0.49 0.49 0.52 0.51 0.50 — 0.51 0.51 — 0.51 0.50 0.51 0.50 0.50 0.50 0.48 0.50 0.50 0.33 0.49 0.51 0.49
43
Supplementary Table 12 | Prolonged ICU stay prediction performance in the eICU Collaborative Research Database. Per-analyte precision, sensitivity, specificity, and balanced accuracy for PerRI and NORMARI abnormality flags, restricted to measurements classified as normal by PopRI . Analytes with fewer than 100 measurements are omitted (—). PerRI
NORMARI
Analyte
N
Events
Precision
Sensitivity
Specificity
Accuracy
Precision
Sensitivity
Specificity
Accuracy
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
— 891 8,843 7,391 8,035 19,881 16,083 26,608 29,891 27,847 289 9,534 2,919 — 3,383 57,801 — 32,416 22,846 43,792 34,479 40,769 33,546 4,595 16,275 8,985 11 176 4,859 25,827
— 158 2,882 2,412 2,561 3,352 2,941 5,167 5,495 5,026 137 718 467 — 486 11,154 — 7,342 4,454 9,765 7,956 8,509 7,568 778 2,917 2,983 1 117 1,495 5,040
— 0.17 0.32 0.32 0.30 0.15 0.20 0.15 0.12 0.19 0.56 0.06 0.16 — 0.13 0.20 — 0.25 0.18 0.22 0.24 0.17 0.22 0.18 0.19 0.35 0.00 0.66 0.33 0.17
— 0.24 0.57 0.55 0.48 0.41 0.35 0.28 0.20 0.27 0.20 0.18 0.41 — 0.33 0.35 — 0.51 0.35 0.57 0.55 0.36 0.67 0.29 0.45 0.32 0.00 0.54 0.36 0.38
— 0.75 0.41 0.42 0.46 0.53 0.68 0.60 0.67 0.75 0.85 0.75 0.57 — 0.63 0.66 — 0.55 0.60 0.43 0.48 0.55 0.30 0.73 0.57 0.71 0.30 0.44 0.67 0.57
— 0.50 0.49 0.48 0.47 0.47 0.51 0.44 0.43 0.51 0.53 0.46 0.49 — 0.48 0.50 — 0.53 0.48 0.50 0.51 0.45 0.49 0.51 0.51 0.51 0.15 0.49 0.52 0.47
— 0.23 0.35 0.33 0.31 0.19 0.23 0.23 0.23 0.19 0.50 0.13 0.17 — 0.15 0.26 — 0.28 0.19 0.27 0.24 0.21 0.27 0.19 0.20 0.35 0.33 0.59 0.37 0.20
— 0.16 0.19 0.20 0.14 0.15 0.21 0.15 0.20 0.11 0.09 0.15 0.12 — 0.14 0.23 — 0.31 0.15 0.27 0.90 0.64 0.48 0.14 0.18 0.13 1.00 0.14 0.17 0.13
— 0.88 0.84 0.80 0.85 0.87 0.84 0.88 0.84 0.90 0.92 0.92 0.89 — 0.87 0.84 — 0.77 0.84 0.80 0.14 0.37 0.62 0.89 0.84 0.88 0.80 0.81 0.87 0.87
— 0.52 0.51 0.50 0.50 0.51 0.53 0.52 0.52 0.50 0.50 0.54 0.51 — 0.50 0.54 — 0.54 0.50 0.53 0.52 0.51 0.55 0.51 0.51 0.51 0.90 0.47 0.52 0.50
44
Supplementary Table 13 | In-hospital mortality prediction performance in the INSPIRE cohort. Per-analyte precision, sensitivity, specificity, and balanced accuracy for PerRI and NORMARI abnormality flags, restricted to measurements classified as normal by PopRI . Analytes with fewer than 100 measurements are omitted (—). PerRI
NORMARI
Analyte
N
Events
Precision
Sensitivity
Specificity
Accuracy
Precision
Sensitivity
Specificity
Accuracy
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
26 15,833 23,817 17,229 24,951 18,100 12,367 19,435 3,345 32,702 — 3,293 9,823 — 11,389 31,429 — — — — — 26,515 27,257 — — 23,415 — — 16,931 23,905
0 62 387 383 513 225 197 310 188 355 — 58 45 — 61 580 — — — — — 367 260 — — 397 — — 178 281
0.00 0.00 0.02 0.02 0.03 0.01 0.02 0.01 0.04 0.03 — 0.02 0.01 — 0.01 0.03 — — — — — 0.02 0.01 — — 0.03 — — 0.01 0.01
— 0.19 0.50 0.31 0.45 0.32 0.28 0.20 0.21 0.35 — 0.31 0.40 — 0.43 0.39 — — — — — 0.50 0.43 — — 0.38 — — 0.21 0.23
0.73 0.81 0.53 0.65 0.63 0.69 0.82 0.77 0.68 0.89 — 0.78 0.69 — 0.73 0.76 — — — — — 0.69 0.49 — — 0.79 — — 0.78 0.78
— 0.50 0.51 0.48 0.54 0.51 0.55 0.48 0.44 0.62 — 0.54 0.55 — 0.58 0.57 — — — — — 0.60 0.46 — — 0.59 — — 0.49 0.51
0.00 0.01 0.02 0.02 0.03 0.03 0.07 0.04 0.08 0.07 — 0.02 0.01 — 0.01 0.07 — — — — — 0.07 0.01 — — 0.02 — — 0.01 0.02
— 0.07 0.17 0.14 0.12 0.10 0.12 0.12 0.35 0.18 — 0.28 0.20 — 0.07 0.30 — — — — — 0.33 0.28 — — 0.09 — — 0.03 0.09
0.92 0.96 0.89 0.87 0.92 0.96 0.97 0.95 0.76 0.97 — 0.75 0.94 — 0.96 0.92 — — — — — 0.94 0.75 — — 0.93 — — 0.96 0.95
— 0.52 0.53 0.51 0.52 0.53 0.55 0.54 0.55 0.58 — 0.51 0.57 — 0.51 0.61 — — — — — 0.63 0.52 — — 0.51 — — 0.50 0.52
45
Supplementary Table 14 | Perioperative infection prediction performance in the INSPIRE cohort. Per-analyte precision, sensitivity, specificity, and balanced accuracy for PerRI and NORMARI abnormality flags, restricted to measurements classified as normal by PopRI . Analytes with fewer than 100 measurements are omitted (—). PerRI
NORMARI
Analyte
N
Events
Precision
Sensitivity
Specificity
Accuracy
Precision
Sensitivity
Specificity
Accuracy
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
26 15,833 23,817 17,229 24,951 18,100 12,367 19,435 3,345 32,702 — 3,293 9,823 — 11,389 31,429 — — — — — 26,515 27,257 — — 23,415 — — 16,931 23,905
2 437 1,174 1,005 1,439 736 481 854 421 1,124 — 161 201 — 277 1,374 — — — — — 975 1,137 — — 1,292 — — 763 1,043
0.14 0.03 0.04 0.05 0.05 0.03 0.03 0.03 0.09 0.04 — 0.04 0.02 — 0.02 0.04 — — — — — 0.03 0.03 — — 0.07 — — 0.04 0.03
0.50 0.17 0.37 0.29 0.33 0.23 0.15 0.14 0.22 0.14 — 0.19 0.31 — 0.24 0.22 — — — — — 0.28 0.36 — — 0.25 — — 0.19 0.17
0.75 0.81 0.52 0.64 0.62 0.68 0.81 0.76 0.67 0.89 — 0.77 0.69 — 0.73 0.76 — — — — — 0.69 0.49 — — 0.79 — — 0.78 0.78
0.62 0.49 0.44 0.46 0.48 0.46 0.48 0.45 0.44 0.51 — 0.48 0.50 — 0.49 0.49 — — — — — 0.49 0.42 — — 0.52 — — 0.48 0.48
0.00 0.03 0.04 0.05 0.06 0.02 0.06 0.04 0.12 0.05 — 0.06 0.03 — 0.03 0.07 — — — — — 0.07 0.03 — — 0.05 — — 0.08 0.04
0.00 0.04 0.10 0.12 0.09 0.02 0.04 0.04 0.24 0.04 — 0.30 0.08 — 0.05 0.13 — — — — — 0.12 0.20 — — 0.07 — — 0.06 0.04
0.92 0.96 0.89 0.87 0.92 0.96 0.97 0.95 0.75 0.97 — 0.75 0.94 — 0.96 0.92 — — — — — 0.94 0.75 — — 0.93 — — 0.97 0.95
0.46 0.50 0.49 0.49 0.50 0.49 0.51 0.50 0.50 0.51 — 0.53 0.51 — 0.51 0.53 — — — — — 0.53 0.47 — — 0.50 — — 0.51 0.50
46
Supplementary Table 15 | Prolonged hospital stay prediction performance in the INSPIRE cohort. Per-analyte precision, sensitivity, specificity, and balanced accuracy for PerRI and NORMARI abnormality flags, restricted to measurements classified as normal by PopRI . Analytes with fewer than 100 measurements are omitted (—). PerRI
NORMARI
Analyte
N
Events
Precision
Sensitivity
Specificity
Accuracy
Precision
Sensitivity
Specificity
Accuracy
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
26 15,833 23,817 17,229 24,951 18,100 12,367 19,435 3,345 32,702 — 3,293 9,823 — 11,389 31,429 — — — — — 26,515 27,257 — — 23,415 — — 16,931 23,905
19 10,599 18,548 13,384 19,532 13,840 8,496 15,298 2,885 22,812 — 2,352 5,503 — 6,945 23,052 — — — — — 18,858 19,942 — — 17,682 — — 12,395 17,474
0.71 0.68 0.78 0.76 0.76 0.74 0.67 0.74 0.83 0.70 — 0.67 0.52 — 0.60 0.71 — — — — — 0.69 0.71 — — 0.75 — — 0.73 0.69
0.26 0.19 0.48 0.34 0.36 0.30 0.18 0.22 0.31 0.12 — 0.21 0.29 — 0.27 0.24 — — — — — 0.30 0.50 — — 0.21 — — 0.22 0.21
0.71 0.81 0.53 0.61 0.59 0.65 0.81 0.72 0.62 0.89 — 0.74 0.66 — 0.72 0.73 — — — — — 0.67 0.46 — — 0.78 — — 0.78 0.75
0.49 0.50 0.50 0.48 0.48 0.48 0.49 0.47 0.46 0.50 — 0.48 0.47 — 0.49 0.48 — — — — — 0.48 0.48 — — 0.50 — — 0.50 0.48
1.00 0.66 0.70 0.74 0.72 0.70 0.65 0.73 0.89 0.67 — 0.74 0.49 — 0.53 0.72 — — — — — 0.68 0.68 — — 0.70 — — 0.74 0.76
0.10 0.03 0.10 0.13 0.07 0.04 0.03 0.04 0.25 0.03 — 0.26 0.06 — 0.04 0.08 — — — — — 0.06 0.23 — — 0.07 — — 0.04 0.05
1.00 0.96 0.85 0.84 0.90 0.95 0.97 0.94 0.80 0.97 — 0.78 0.93 — 0.95 0.92 — — — — — 0.93 0.71 — — 0.91 — — 0.97 0.96
0.55 0.50 0.47 0.48 0.48 0.49 0.50 0.49 0.53 0.50 — 0.52 0.49 — 0.49 0.50 — — — — — 0.49 0.47 — — 0.49 — — 0.50 0.50
47
Supplementary Table 16 | Unplanned ICU admission prediction performance in the INSPIRE cohort. Per-analyte precision, sensitivity, specificity, and balanced accuracy for PerRI and NORMARI abnormality flags, restricted to measurements classified as normal by PopRI . Analytes with fewer than 100 measurements are omitted (—). PerRI
NORMARI
Analyte
N
Events
Precision
Sensitivity
Specificity
Accuracy
Precision
Sensitivity
Specificity
Accuracy
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
26 15,833 23,817 17,229 24,951 18,100 12,367 19,435 3,345 32,702 — 3,293 9,823 — 11,389 31,429 — — — — — 26,515 27,257 — — 23,415 — — 16,931 23,905
15 4,067 7,129 5,695 7,856 5,436 3,418 5,038 1,870 6,625 — 622 1,745 — 2,448 8,519 — — — — — 6,473 6,866 — — 6,578 — — 4,750 6,052
0.43 0.22 0.28 0.31 0.27 0.29 0.23 0.19 0.51 0.21 — 0.14 0.15 — 0.21 0.25 — — — — — 0.20 0.20 — — 0.28 — — 0.23 0.19
0.20 0.16 0.45 0.34 0.32 0.30 0.15 0.17 0.29 0.12 — 0.16 0.27 — 0.27 0.22 — — — — — 0.26 0.41 — — 0.21 — — 0.18 0.17
0.64 0.80 0.52 0.64 0.60 0.68 0.81 0.74 0.65 0.89 — 0.76 0.68 — 0.73 0.75 — — — — — 0.68 0.46 — — 0.79 — — 0.76 0.76
0.42 0.48 0.48 0.49 0.46 0.49 0.48 0.46 0.47 0.50 — 0.46 0.47 — 0.50 0.48 — — — — — 0.47 0.44 — — 0.50 — — 0.47 0.47
1.00 0.20 0.26 0.35 0.25 0.33 0.20 0.34 0.57 0.19 — 0.22 0.25 — 0.24 0.32 — — — — — 0.32 0.22 — — 0.27 — — 0.22 0.23
0.13 0.03 0.10 0.14 0.06 0.04 0.02 0.06 0.25 0.03 — 0.29 0.09 — 0.05 0.09 — — — — — 0.09 0.21 — — 0.07 — — 0.03 0.04
1.00 0.96 0.88 0.87 0.92 0.96 0.97 0.96 0.76 0.97 — 0.76 0.94 — 0.96 0.93 — — — — — 0.94 0.74 — — 0.93 — — 0.96 0.95
0.57 0.49 0.49 0.51 0.49 0.50 0.49 0.51 0.51 0.50 — 0.52 0.52 — 0.50 0.51 — — — — — 0.51 0.48 — — 0.50 — — 0.49 0.50
48
Supplementary Table 17 | Positive predictive value across outcomes in Clalit Health Services. Measurements classified as normal by PopRI are excluded. Per 100 patients flagged abnormal by PerRI or NORMARI , the number who experienced each outcome (anemia, chronic kidney disease, all-cause mortality, and type 2 diabetes), by analyte. Analytes with fewer than 100 measurements are omitted (—). Anemia
Chronic Kidney Disease
Mortality
Type 2 Diabetes
Analyte
PerRI
NORMARI
PerRI
NORMARI
PerRI
NORMARI
PerRI
NORMARI
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
21 21 19 19 19 14 21 30 100 16 — 16 10 17 9 21 20 16 14 17 19 19 18 13 14 24 21 — 23 18
30 25 24 22 23 20 27 38 100 24 — 23 12 21 12 25 24 18 17 20 18 26 25 18 15 22 26 — 27 22
39 44 40 40 40 26 45 52 75 34 — 30 28 38 28 41 43 38 30 37 37 41 35 30 32 46 43 — 47 37
48 51 47 43 45 34 53 62 100 57 — 38 34 42 33 48 50 41 34 40 36 51 45 36 33 44 50 — 54 41
23 26 25 24 26 18 28 36 50 21 — 18 15 22 15 27 26 23 18 23 24 26 23 17 18 30 27 — 30 23
34 33 30 27 29 24 35 46 80 32 — 26 19 26 19 34 30 27 22 26 24 34 32 22 19 29 31 — 37 27
70 78 74 74 74 73 77 93 100 75 — 53 70 71 70 78 79 71 73 73 74 76 73 71 71 79 80 — 78 73
78 81 79 78 79 80 82 97 100 84 — 64 75 74 76 83 86 74 77 75 73 84 79 78 72 77 87 — 81 78
Overall
21
25
39
46
24
30
75
80
49
Supplementary Table 18 | Positive predictive value across outcomes in the eICU Collaborative Research Database. Measurements classified as normal by PopRI are excluded. Per 100 patients flagged abnormal by PerRI or NORMARI , the number who experienced each outcome (acute kidney injury, in-hospital mortality, prolonged ICU stay, and sepsis), by analyte. Analytes with fewer than 100 measurements are omitted (—). Acute Kidney Injury
Mortality
Prolonged LOS (>7d)
Sepsis
Analyte
PerRI
NORMARI
PerRI
NORMARI
PerRI
NORMARI
PerRI
NORMARI
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
— 14 19 17 16 5 13 13 9 7 36 11 10 — 9 14 — 16 14 15 15 14 13 11 10 19 29 21 19 12
— 17 21 18 16 5 15 16 11 5 33 16 10 — 8 17 — 17 14 16 15 15 13 11 11 20 33 22 22 14
— 5 13 12 10 6 8 8 6 8 6 4 10 — 9 11 — 11 9 11 12 10 7 10 6 14 29 20 11 7
— 8 16 12 11 6 8 10 8 10 4 8 13 — 12 12 — 12 10 14 11 10 7 10 6 15 67 22 11 8
— 17 32 32 30 15 20 15 12 19 56 6 16 — 13 20 — 25 18 22 24 17 22 18 19 35 0 66 33 17
— 23 35 33 31 19 23 23 23 19 50 13 17 — 15 26 — 28 19 27 24 21 27 19 20 35 33 59 37 20
— 8 23 25 24 12 13 14 11 14 32 13 11 — 10 16 — 18 15 18 16 15 16 13 12 26 29 35 24 12
— 13 23 25 23 12 14 16 12 13 33 18 12 — 12 17 — 19 15 19 16 17 17 10 12 25 0 30 26 12
Overall
15
16
10
13
23
27
18
17
50
Supplementary Table 19 | Positive predictive value across outcomes in the INSPIRE cohort. Measurements classified as normal by PopRI are excluded. Per 100 patients flagged abnormal by PerRI or NORMARI , the number who experienced each outcome (in-hospital mortality, perioperative infection, prolonged hospital stay, and unplanned ICU admission), by analyte. Analytes with fewer than 100 measurements are omitted (—). Mortality
Perioperative Infection
Prolonged LOS (>7d)
Unplanned ICU Admission
Analyte
PerRI
NORMARI
PerRI
NORMARI
PerRI
NORMARI
PerRI
NORMARI
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
0 0 2 2 2 1 2 1 4 3 — 2 1 — 1 3 — — — — — 2 1 — — 3 — — 1 1
0 1 2 2 3 3 7 4 8 7 — 2 1 — 1 7 — — — — — 7 1 — — 2 — — 1 2
14 2 4 5 5 3 3 3 9 4 — 4 2 — 2 4 — — — — — 3 3 — — 7 — — 4 3
0 3 4 5 6 2 6 4 12 5 — 6 2 — 3 7 — — — — — 7 3 — — 5 — — 8 4
71 68 78 76 76 74 67 74 83 70 — 67 52 — 60 71 — — — — — 69 71 — — 75 — — 73 69
100 66 70 74 72 70 65 73 89 67 — 74 49 — 52 72 — — — — — 68 68 — — 70 — — 74 76
43 22 28 31 27 29 23 19 51 21 — 14 15 — 21 25 — — — — — 20 20 — — 28 — — 22 19
100 20 26 35 25 33 20 34 57 19 — 22 25 — 24 32 — — — — — 32 22 — — 27 — — 22 23
Overall
2
3
4
5
71
71
25
31
51
Supplementary Table 20 | Lead time for NORMARI abnormality detection relative to PopRI . Per analyte and cohort: number of tests flagged by NORMARI alone (NORMARI -only) and the subset later confirmed by PopRI (later population), with median lead time and interquartile range (hours for eICU-CRD and INSPIRE, months for CHS). eICU-CRD
CHS
INSPIRE
Analyte
NORMA-only
Later PopRI
Median (h)
IQR (h)
NORMA-only
Later PopRI
Median (mo)
IQR (mo)
NORMA-only
Later PopRI
Median (h)
IQR (h)
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC Overall
— 2,185 4,362 6,749 7,720 25,578 45,543 24,313 52,641 40,464 425 20,333 2,406 2 4,408 62,211 3 40,943 38,761 19,132 101,702 115,529 46,110 6,592 18,706 11,136 28 317 11,311 31,253 740,863
— 865 757 1,235 1,317 7,039 13,222 8,477 26,365 8,744 72 17,921 998 0 1,526 16,970 2 5,361 11,733 1,789 4,805 31,882 5,482 2,129 3,308 1,721 7 38 2,592 8,381 184,738
— 30.0 49.6 47.6 47.4 28.0 25.1 34.7 25.6 26.0 25.2 5.6 24.2 — 23.7 40.9 35.1 45.5 25.1 48.2 37.5 26.6 49.7 23.9 26.3 47.4 70.3 60.0 34.8 35.3 34.8
— 23.1–73.9 24.2–120.2 23.9–99.4 23.9–95.5 23.1–71.3 19.6–69.2 23.0–72.8 19.3–64.5 21.6–67.7 20.8–49.3 3.2–12.1 17.7–48.1 — 14.3–47.1 23.0–78.6 29.2–41.0 23.4–93.1 22.5–49.9 23.7–120.7 23.2–72.5 21.1–69.9 24.0–120.8 16.5–47.6 22.8–70.0 23.8–95.5 42.0–120.9 25.4–145.5 22.9–74.1 23.3–72.7 23.1–72.6
14,227 807,619 1,196,400 1,199,702 1,195,620 568,660 574,875 44,929 96 1,070,918 — 358,479 1,083,585 515,071 1,065,707 1,337,595 764,300 2,300,580 840,255 4,464,421 20,300,446 805,446 1,356,487 1,202,775 1,580,050 2,966,137 675,432 — 962,840 2,183,898 51,436,550
5,064 436,471 543,146 703,509 551,419 402,995 356,694 36,351 80 687,844 — 248,307 762,204 120,646 585,178 881,518 172,816 1,420,584 776,047 1,941,232 8,370,452 506,896 660,368 769,663 933,310 1,268,891 330,110 — 530,403 1,402,700 25,404,898
12.4 7.9 15.5 11.9 10.0 7.6 6.4 4.9 2.9 7.5 — 5.5 8.5 21.9 16.1 4.5 17.6 11.2 7.6 18.8 20.0 8.0 8.5 8.9 16.2 26.4 14.8 — 6.9 7.4 8.7
5.7–27.5 1.2–29.9 3.1–50.6 2.4–41.0 1.6–38.0 1.1–29.2 0.8–29.0 0.5–22.7 0.3–8.7 1.4–27.1 — 0.9–20.6 1.8–30.2 7.3–53.3 3.4–44.3 0.4–25.8 6.4–44.5 2.8–37.1 1.6–29.9 4.1–62.8 4.8–53.3 0.9–34.7 1.1–40.3 1.8–32.9 3.1–52.3 6.0–68.5 5.0–39.6 — 0.9–28.6 1.0–33.0 1.7–33.8
24 1,493 3,442 3,210 2,204 1,983 1,419 3,124 3,787 2,321 — 14,216 2,240 — 1,576 5,722 — — — — — 4,825 8,493 — — 2,742 — — 1,227 2,389 66,437
0 29 201 187 29 409 342 932 2,418 826 — 8,944 484 — 301 2,611 — — — — — 1,308 277 — — 71 — — 18 348 19,735
— 23.5 60.4 47.6 16.6 24.9 39.8 24.0 8.4 43.2 — 4.6 10.4 — 14.8 15.7 — — — — — 29.8 23.3 — — 23.8 — — 13.0 26.0 23.7
— 13.3–25.7 24.7–102.3 19.2–109.0 8.2–24.0 13.0–55.2 16.3–77.0 11.4–54.2 3.5–19.2 16.1–121.8 — 2.5–8.7 3.4–21.5 — 5.7–32.6 5.8–37.2 — — — — — 11.3–75.0 8.9–64.1 — — 9.7–83.0 — — 11.9–24.0 13.8–57.4 11.4–54.7
52
Supplementary Table 21 | Proportional hazards analysis for all-cause mortality in Clalit Health Services. Per-analyte hazard ratios (with 95% confidence intervals) and concordance indices for PopRI , PerRI , and NORMARI abnormality flags, adjusted for age and sex. PopRI
PerRI
NORMARI
Analyte
N train (events)
N test (events)
HR [95% CI]
C
HR [95% CI]
C
HR [95% CI]
C
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
9517 (3118) 13612 (4221) 20262 (5615) 25254 (6667) 25296 (6662) 26980 (7187) 15946 (4688) 846 (307) — 29331 (7624) — 28928 (7343) 31026 (7576) 20800 (5296) 31020 (7578) 25819 (7121) 23706 (6093) 28147 (6984) 31009 (7580) 30848 (7546) 28679 (7094) 25977 (7133) 30976 (7564) 30946 (7556) 30874 (7528) 13506 (3888) 25805 (6534) — 13638 (4217) 31053 (7583)
6345 (2079) 9076 (2815) 13508 (3743) 16836 (4445) 16864 (4442) 17987 (4792) 10631 (3125) 564 (205) — 19555 (5083) — 19286 (4895) 20685 (5051) 13867 (3530) 20680 (5052) 17213 (4748) 15804 (4062) 18765 (4656) 20673 (5053) 20566 (5031) 19120 (4729) 17318 (4756) 20652 (5043) 20631 (5038) 20583 (5018) 9005 (2593) 17204 (4357) — 9092 (2812) 20702 (5055)
0.84 [0.65, 1.09] 4.15 [3.83, 4.49] 2.05 [1.94, 2.17] 2.08 [1.95, 2.22] 1.85 [1.76, 1.94] 2.82 [2.53, 3.15] 1.93 [1.81, 2.07] 3.57 [1.67, 7.61] — 3.00 [2.82, 3.18] — 1.60 [1.39, 1.85] 3.60 [3.15, 4.12] 1.68 [1.56, 1.81] 3.70 [3.32, 4.12] 1.63 [1.53, 1.73] 0.81 [0.77, 0.85] 2.04 [1.89, 2.20] 1.20 [0.89, 1.61] 1.70 [1.62, 1.78] 1.28 [1.22, 1.35] 2.17 [2.04, 2.31] 1.98 [1.89, 2.08] — 3.09 [2.82, 3.39] 1.61 [1.50, 1.72] 0.93 [0.87, 0.99] — 2.23 [2.08, 2.38] 1.86 [1.75, 1.97]
0.689 0.785 0.760 0.761 0.758 0.747 0.754 0.767 — 0.770 — 0.749 0.774 0.733 0.779 0.748 0.735 0.767 0.767 0.774 0.765 0.753 0.779 — 0.774 0.744 0.739 — 0.756 0.773
0.58 [0.54, 0.63] 1.21 [1.10, 1.33] 0.96 [0.87, 1.05] 0.94 [0.86, 1.03] 1.02 [0.94, 1.12] 0.81 [0.75, 0.88] 1.08 [0.99, 1.17] 1.90 [0.97, 3.71] — 1.04 [0.97, 1.12] — 0.77 [0.70, 0.84] 0.79 [0.73, 0.86] 0.73 [0.67, 0.81] 0.83 [0.77, 0.90] 1.07 [0.99, 1.15] 0.74 [0.69, 0.79] 0.98 [0.91, 1.07] 0.96 [0.90, 1.03] 0.90 [0.83, 0.97] 0.93 [0.87, 0.99] 1.03 [0.95, 1.12] 0.92 [0.83, 1.01] 0.78 [0.71, 0.85] 0.90 [0.83, 0.97] 1.13 [1.04, 1.22] 0.80 [0.75, 0.86] — 1.07 [0.98, 1.16] 0.98 [0.89, 1.07]
0.699 0.740 0.741 0.749 0.744 0.739 0.740 0.760 — 0.744 — 0.748 0.765 0.726 0.764 0.742 0.736 0.760 0.767 0.762 0.762 0.738 0.765 0.764 0.762 0.737 0.739 — 0.736 0.763
0.82 [0.65, 1.03] 3.11 [2.80, 3.46] 2.01 [1.86, 2.18] 1.35 [1.27, 1.43] 1.55 [1.47, 1.65] 2.58 [2.31, 2.87] 1.88 [1.73, 2.05] 4.10 [1.82, 9.27] — 3.30 [3.07, 3.54] — 1.61 [1.38, 1.86] 3.36 [2.86, 3.94] 1.72 [1.57, 1.89] 3.72 [3.24, 4.26] 1.69 [1.56, 1.83] 0.87 [0.82, 0.92] 2.18 [1.90, 2.50] 0.99 [0.76, 1.28] 1.95 [1.77, 2.16] 0.10 [0.04, 0.23] 2.02 [1.88, 2.16] 1.88 [1.77, 1.98] — 2.23 [1.96, 2.52] 1.09 [0.98, 1.21] 0.96 [0.90, 1.03] — 2.73 [2.46, 3.04] 1.90 [1.75, 2.07]
0.688 0.757 0.749 0.752 0.750 0.746 0.749 0.755 — 0.768 — 0.749 0.771 0.730 0.772 0.746 0.734 0.763 0.767 0.766 0.762 0.750 0.775 — 0.765 0.738 0.739 — 0.745 0.768
53
Supplementary Table 22 | Proportional hazards analysis for in-hospital mortality in the eICU Collaborative Research Database. Per-analyte hazard ratios (with 95% confidence intervals) and concordance indices for PopRI , PerRI , and NORMARI abnormality flags, adjusted for age and sex. PopRI
PerRI
NORMARI
Analyte
N train (events)
N test (events)
HR [95% CI]
C
HR [95% CI]
C
HR [95% CI]
C
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
— 11151 (1560) 7966 (1118) 7288 (976) 7446 (912) 33379 (3378) 33080 (3474) 34320 (3560) 40679 (4792) 31723 (3152) 684 (122) 52615 (5182) 34809 (3553) — 35018 (3622) 38603 (3943) — 28067 (2889) 29987 (3063) 29488 (3017) 21532 (2191) 36226 (3839) 30955 (3214) 31071 (3221) 28255 (2852) 7465 (930) — 186 (37) 8881 (1307) 30250 (2917)
— 7435 (1040) 5311 (746) 4859 (650) 4964 (608) 22253 (2252) 22054 (2316) 22881 (2374) 27120 (3194) 21149 (2102) 456 (82) 35077 (3455) 23206 (2368) — 23346 (2415) 25736 (2628) — 18712 (1926) 19992 (2042) 19659 (2011) 14355 (1461) 24152 (2559) 20638 (2142) 20714 (2148) 18838 (1901) 4978 (620) — 124 (24) 5921 (871) 20167 (1945)
— 1.04 [0.75, 1.43] 0.97 [0.86, 1.09] 0.98 [0.86, 1.11] 1.68 [1.47, 1.93] 1.74 [1.58, 1.92] 1.81 [1.64, 1.98] 1.52 [1.40, 1.65] 1.94 [1.76, 2.13] 1.89 [1.74, 2.05] 1.23 [0.81, 1.87] 2.14 [1.73, 2.65] 1.11 [0.89, 1.38] — 1.01 [0.82, 1.23] 1.34 [1.26, 1.43] — 0.99 [0.92, 1.07] 0.85 [0.78, 0.92] 1.26 [1.16, 1.37] 0.95 [0.84, 1.07] 1.22 [1.14, 1.31] 2.50 [2.32, 2.68] 1.05 [0.88, 1.25] 1.85 [1.65, 2.07] 1.39 [1.22, 1.59] — 1.09 [0.54, 2.18] 1.67 [1.44, 1.93] 2.39 [2.17, 2.63]
— 0.544 0.534 0.538 0.571 0.573 0.577 0.572 0.586 0.602 0.546 0.565 0.558 — 0.561 0.563 — 0.558 0.563 0.559 0.568 0.551 0.654 0.562 0.601 0.552 — 0.594 0.576 0.619
— 1.01 [0.65, 1.57] 1.09 [0.93, 1.28] 1.05 [0.89, 1.24] 1.82 [1.52, 2.19] 1.84 [1.59, 2.14] 1.68 [1.48, 1.91] 1.69 [1.51, 1.89] 1.82 [1.60, 2.08] 1.97 [1.78, 2.17] 1.07 [0.69, 1.65] 1.65 [1.26, 2.15] 1.45 [0.98, 2.15] — 1.11 [0.83, 1.50] 1.31 [1.22, 1.41] — 0.95 [0.88, 1.03] 0.87 [0.78, 0.97] 1.35 [1.25, 1.47] 1.17 [1.08, 1.28] 1.37 [1.25, 1.50] 1.77 [1.58, 1.99] 1.07 [0.85, 1.34] 2.23 [1.89, 2.63] 1.28 [1.12, 1.47] — 1.03 [0.43, 2.48] 1.68 [1.37, 2.06] 2.48 [2.16, 2.85]
— 0.544 0.534 0.539 0.562 0.559 0.564 0.568 0.566 0.599 0.512 0.563 0.556 — 0.561 0.563 — 0.558 0.558 0.564 0.575 0.557 0.570 0.561 0.589 0.550 — 0.590 0.554 0.598
— 1.14 [0.78, 1.65] 1.12 [0.99, 1.26] 0.99 [0.87, 1.13] 1.61 [1.40, 1.86] 1.77 [1.59, 1.97] 1.75 [1.58, 1.94] 1.58 [1.44, 1.72] 1.90 [1.71, 2.11] 1.98 [1.81, 2.16] 1.14 [0.75, 1.75] 2.19 [1.75, 2.74] 1.31 [1.01, 1.71] — 1.07 [0.85, 1.35] 1.36 [1.27, 1.45] — 0.99 [0.92, 1.06] 0.89 [0.82, 0.98] 1.35 [1.25, 1.45] 1.20 [1.02, 1.40] 1.22 [1.11, 1.35] 1.67 [1.54, 1.82] 1.10 [0.90, 1.33] 1.91 [1.68, 2.16] 1.38 [1.22, 1.58] — 1.22 [0.59, 2.49] 1.66 [1.42, 1.95] 2.41 [2.17, 2.67]
— 0.544 0.535 0.537 0.575 0.571 0.571 0.574 0.576 0.604 0.537 0.565 0.558 — 0.561 0.564 — 0.558 0.561 0.569 0.571 0.552 0.579 0.561 0.595 0.552 — 0.598 0.562 0.617
54
Supplementary Table 23 | Proportional hazards analysis for acute kidney injury in the eICU Collaborative Research Database. Per-analyte hazard ratios (with 95% confidence intervals) and concordance indices for PopRI , PerRI , and NORMARI abnormality flags, adjusted for age and sex. PopRI
PerRI
NORMARI
Analyte
N train (events)
N test (events)
HR [95% CI]
C
HR [95% CI]
C
HR [95% CI]
C
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
— 11121 (2490) 7949 (1702) 7267 (1472) 7429 (1464) 33279 (4964) 32980 (5407) 34213 (5471) 40566 (6164) 31638 (4454) 682 (211) 52476 (6866) 34716 (5015) — 34923 (5009) 38496 (5647) — 27981 (4313) 29898 (4621) 29400 (4528) 21459 (3144) 36123 (5599) 30866 (4735) 30981 (4736) 28174 (4340) 7447 (1482) 20 (4) 186 (54) 8860 (1890) 30162 (4532)
— 7414 (1660) 5300 (1134) 4846 (982) 4954 (976) 22187 (3310) 21987 (3604) 22809 (3647) 27044 (4109) 21092 (2969) 456 (141) 34984 (4577) 23144 (3343) — 23282 (3339) 25664 (3764) — 18654 (2875) 19933 (3080) 19601 (3018) 14307 (2096) 24082 (3733) 20578 (3156) 20654 (3158) 18783 (2894) 4966 (988) 14 (2) 124 (36) 5908 (1260) 20108 (3021)
— 1.28 [1.01, 1.63] 1.25 [1.14, 1.37] 1.08 [0.97, 1.19] 1.19 [1.08, 1.32] 1.89 [1.76, 2.04] 1.43 [1.34, 1.53] 1.43 [1.34, 1.52] 1.40 [1.31, 1.50] 1.77 [1.66, 1.89] 1.32 [0.97, 1.79] 1.06 [0.94, 1.19] 1.91 [1.55, 2.36] — 1.97 [1.62, 2.40] 1.40 [1.33, 1.48] — 1.12 [1.05, 1.19] 1.17 [1.10, 1.25] 1.12 [1.04, 1.21] 1.30 [1.18, 1.43] 1.13 [1.07, 1.19] 1.38 [1.30, 1.46] 1.75 [1.50, 2.04] 1.73 [1.60, 1.87] 1.15 [1.04, 1.28] 0.52 [0.07, 3.77] 1.18 [0.67, 2.09] 1.28 [1.15, 1.43] 1.27 [1.19, 1.35]
— 0.516 0.544 0.504 0.529 0.566 0.541 0.540 0.534 0.576 0.520 0.519 0.524 — 0.527 0.557 — 0.528 0.531 0.523 0.535 0.527 0.543 0.520 0.571 0.526 0.640 0.551 0.517 0.537
— 1.45 [1.03, 2.03] 1.12 [0.99, 1.26] 0.94 [0.82, 1.07] 1.07 [0.95, 1.21] 1.63 [1.47, 1.82] 1.40 [1.28, 1.52] 1.24 [1.14, 1.34] 1.35 [1.23, 1.49] 1.55 [1.44, 1.66] 1.32 [0.94, 1.86] 0.96 [0.82, 1.13] 2.59 [1.78, 3.79] — 2.18 [1.61, 2.96] 1.33 [1.26, 1.41] — 1.08 [1.01, 1.16] 1.14 [1.05, 1.25] 1.02 [0.96, 1.08] 1.07 [1.00, 1.15] 1.06 [0.99, 1.13] 1.03 [0.95, 1.12] 1.71 [1.40, 2.09] 1.65 [1.50, 1.83] 1.11 [1.00, 1.23] — 0.97 [0.47, 2.01] 1.34 [1.15, 1.55] 1.17 [1.08, 1.26]
— 0.515 0.516 0.505 0.504 0.535 0.533 0.525 0.524 0.555 0.520 0.518 0.521 — 0.524 0.544 — 0.527 0.529 0.522 0.523 0.528 0.517 0.521 0.553 0.527 0.600 0.515 0.510 0.524
— 1.36 [1.04, 1.79] 1.24 [1.12, 1.36] 1.06 [0.95, 1.18] 1.14 [1.03, 1.27] 1.73 [1.60, 1.87] 1.39 [1.29, 1.49] 1.43 [1.34, 1.52] 1.33 [1.24, 1.44] 1.59 [1.49, 1.69] 1.40 [1.01, 1.93] 1.04 [0.92, 1.17] 1.97 [1.56, 2.49] — 2.01 [1.61, 2.50] 1.39 [1.32, 1.47] — 1.14 [1.07, 1.21] 1.13 [1.05, 1.21] 1.06 [1.00, 1.13] 1.10 [0.98, 1.24] 1.01 [0.94, 1.08] 1.15 [1.08, 1.22] 1.78 [1.50, 2.11] 1.71 [1.58, 1.86] 1.12 [1.01, 1.24] 0.89 [0.08, 10.23] 1.02 [0.58, 1.82] 1.31 [1.16, 1.48] 1.28 [1.20, 1.36]
— 0.515 0.534 0.504 0.522 0.555 0.539 0.538 0.528 0.560 0.523 0.519 0.524 — 0.526 0.554 — 0.528 0.530 0.523 0.526 0.525 0.524 0.521 0.568 0.532 0.640 0.525 0.518 0.531
55
Supplementary Table 24 | Proportional hazards analysis for sepsis in the eICU Collaborative Research Database. Per-analyte hazard ratios (with 95% confidence intervals) and concordance indices for PopRI , PerRI , and NORMARI abnormality flags, adjusted for age and sex. PopRI
PerRI
NORMARI
Analyte
N train (events)
N test (events)
HR [95% CI]
C
HR [95% CI]
C
HR [95% CI]
C
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
— 11106 (2754) 7938 (2020) 7256 (1877) 7415 (1904) 33240 (5690) 32941 (5802) 34176 (5909) 40521 (6798) 31595 (5314) 682 (186) 52429 (7603) 34671 (5744) — 34876 (5787) 38461 (6288) — 27943 (5010) 29860 (5363) 29362 (5220) 21432 (3423) 36085 (6030) 30826 (5403) 30939 (5515) 28135 (5022) 7437 (1878) 20 (6) 186 (63) 8848 (2278) 30127 (5170)
— 7405 (1836) 5292 (1347) 4838 (1252) 4944 (1270) 22161 (3793) 21962 (3868) 22785 (3939) 27015 (4532) 21064 (3542) 456 (125) 34953 (5068) 23114 (3829) — 23252 (3858) 25642 (4192) — 18630 (3340) 19908 (3575) 19575 (3480) 14289 (2282) 24058 (4020) 20552 (3603) 20626 (3677) 18758 (3349) 4959 (1252) 14 (5) 124 (42) 5899 (1518) 20086 (3447)
— 2.27 [1.72, 3.00] 1.17 [1.07, 1.28] 1.01 [0.92, 1.11] 1.08 [0.98, 1.18] 1.22 [1.15, 1.30] 1.53 [1.44, 1.63] 1.34 [1.26, 1.41] 1.26 [1.18, 1.34] 1.28 [1.21, 1.35] 0.98 [0.72, 1.33] 0.98 [0.88, 1.09] 1.50 [1.26, 1.79] — 1.50 [1.28, 1.76] 1.48 [1.40, 1.55] — 1.20 [1.14, 1.27] 1.32 [1.24, 1.41] 1.24 [1.16, 1.33] 1.27 [1.16, 1.39] 1.05 [1.00, 1.11] 1.16 [1.10, 1.23] 1.38 [1.22, 1.57] 1.83 [1.70, 1.96] 1.02 [0.93, 1.12] 0.71 [0.13, 3.89] 0.95 [0.56, 1.59] 1.07 [0.97, 1.18] 1.48 [1.39, 1.57]
— 0.519 0.544 0.510 0.502 0.533 0.549 0.539 0.540 0.547 0.513 0.527 0.528 — 0.531 0.560 — 0.533 0.547 0.529 0.520 0.519 0.532 0.522 0.557 0.517 0.425 0.409 0.529 0.545
— 2.42 [1.64, 3.56] 0.98 [0.88, 1.10] 1.01 [0.90, 1.13] 0.96 [0.87, 1.07] 1.08 [0.99, 1.17] 1.56 [1.43, 1.71] 1.29 [1.20, 1.39] 1.17 [1.07, 1.27] 1.21 [1.14, 1.29] 1.01 [0.72, 1.41] 0.93 [0.80, 1.08] 1.75 [1.31, 2.34] — 1.53 [1.22, 1.93] 1.32 [1.25, 1.40] — 1.11 [1.04, 1.18] 1.30 [1.20, 1.41] 1.12 [1.05, 1.18] 1.05 [0.98, 1.12] 1.00 [0.93, 1.06] 1.03 [0.95, 1.11] 1.47 [1.24, 1.74] 1.67 [1.52, 1.83] 0.98 [0.89, 1.08] — 1.28 [0.63, 2.63] 1.08 [0.95, 1.22] 1.28 [1.19, 1.38]
— 0.518 0.503 0.510 0.517 0.527 0.538 0.528 0.532 0.536 0.515 0.527 0.527 — 0.527 0.541 — 0.525 0.529 0.510 0.513 0.519 0.526 0.518 0.535 0.511 0.375 0.439 0.527 0.523
— 2.24 [1.66, 3.02] 1.12 [1.02, 1.22] 0.99 [0.90, 1.09] 1.08 [0.98, 1.18] 1.16 [1.09, 1.24] 1.51 [1.40, 1.62] 1.32 [1.24, 1.40] 1.19 [1.12, 1.28] 1.21 [1.15, 1.28] 1.09 [0.79, 1.50] 0.96 [0.86, 1.08] 1.47 [1.22, 1.78] — 1.50 [1.26, 1.79] 1.41 [1.34, 1.48] — 1.19 [1.13, 1.26] 1.31 [1.23, 1.40] 1.22 [1.15, 1.29] 1.04 [0.93, 1.16] 0.98 [0.92, 1.05] 1.07 [1.01, 1.13] 1.36 [1.19, 1.56] 1.77 [1.64, 1.91] 1.01 [0.92, 1.11] 0.76 [0.08, 7.71] 0.89 [0.53, 1.50] 1.11 [1.00, 1.24] 1.39 [1.31, 1.48]
— 0.518 0.532 0.512 0.495 0.527 0.544 0.536 0.534 0.537 0.522 0.527 0.527 — 0.529 0.551 — 0.533 0.541 0.520 0.514 0.519 0.530 0.519 0.545 0.515 0.425 0.405 0.531 0.536
56
Supplementary Table 25 | Proportional hazards analysis for prolonged ICU stay in the eICU Collaborative Research Database. Prolonged ICU stay is defined as ICU stay greater than 7 days. Per-analyte hazard ratios (with 95% confidence intervals) and concordance indices for PopRI , PerRI , and NORMARI abnormality flags, adjusted for age and sex. PopRI
PerRI
NORMARI
Analyte
N train (events)
N test (events)
HR [95% CI]
C
HR [95% CI]
C
HR [95% CI]
C
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
— 11140 (3587) 7956 (2821) 7276 (2596) 7437 (2658) 33336 (7040) 33039 (6996) 34278 (7202) 40636 (7193) 31684 (6787) 684 (328) 52561 (7373) 34765 (7128) — 34972 (7128) 38559 (7280) — 28026 (6451) 29944 (6867) 29445 (6783) 21492 (4995) 36186 (7220) 30913 (7030) 31027 (7105) 28216 (6496) 7456 (2547) 20 (8) 186 (134) 8871 (3098) 30209 (7003)
— 7428 (2392) 5305 (1881) 4851 (1731) 4958 (1772) 22225 (4694) 22026 (4664) 22853 (4801) 27092 (4795) 21123 (4525) 456 (218) 35041 (4915) 23178 (4753) — 23316 (4752) 25707 (4854) — 18684 (4301) 19964 (4578) 19631 (4523) 14328 (3330) 24125 (4814) 20609 (4686) 20686 (4737) 18812 (4331) 4972 (1699) 14 (5) 124 (90) 5914 (2065) 20140 (4668)
— 0.75 [0.61, 0.93] 0.68 [0.63, 0.73] 0.82 [0.76, 0.88] 0.86 [0.80, 0.93] 0.66 [0.62, 0.71] 0.79 [0.75, 0.84] 0.74 [0.70, 0.78] 0.63 [0.58, 0.68] 0.68 [0.65, 0.72] 1.02 [0.81, 1.29] 0.44 [0.36, 0.53] 0.53 [0.46, 0.61] — 0.50 [0.43, 0.58] 0.77 [0.73, 0.80] — 0.92 [0.87, 0.97] 0.67 [0.63, 0.71] 0.88 [0.83, 0.94] 0.85 [0.78, 0.92] 0.74 [0.71, 0.78] 0.90 [0.86, 0.94] 0.55 [0.49, 0.62] 0.67 [0.63, 0.71] 0.86 [0.79, 0.93] 0.67 [0.11, 4.11] 0.82 [0.57, 1.18] 1.03 [0.95, 1.12] 0.75 [0.71, 0.79]
— 0.535 0.551 0.533 0.545 0.550 0.541 0.552 0.547 0.566 0.512 0.537 0.541 — 0.543 0.553 — 0.537 0.559 0.539 0.535 0.560 0.534 0.549 0.557 0.537 0.600 0.527 0.524 0.557
— 0.70 [0.51, 0.96] 0.84 [0.76, 0.92] 0.87 [0.78, 0.96] 0.98 [0.89, 1.07] 0.76 [0.70, 0.83] 0.77 [0.71, 0.83] 0.84 [0.79, 0.90] 0.72 [0.66, 0.80] 0.66 [0.62, 0.70] 0.97 [0.75, 1.25] 0.53 [0.41, 0.70] 0.64 [0.51, 0.81] — 0.51 [0.42, 0.62] 0.81 [0.77, 0.85] — 0.86 [0.81, 0.91] 0.67 [0.62, 0.72] 0.97 [0.92, 1.02] 0.98 [0.92, 1.04] 0.80 [0.75, 0.85] 1.22 [1.14, 1.30] 0.53 [0.45, 0.63] 0.70 [0.65, 0.76] 0.90 [0.83, 0.98] — 0.74 [0.46, 1.19] 0.97 [0.87, 1.09] 0.79 [0.74, 0.85]
— 0.536 0.533 0.520 0.534 0.535 0.537 0.538 0.540 0.562 0.482 0.535 0.537 — 0.537 0.543 — 0.544 0.552 0.540 0.529 0.543 0.542 0.541 0.546 0.542 — 0.533 0.527 0.545
— 0.77 [0.60, 0.97] 0.73 [0.68, 0.79] 0.85 [0.78, 0.92] 0.87 [0.81, 0.94] 0.67 [0.63, 0.72] 0.75 [0.71, 0.80] 0.75 [0.71, 0.80] 0.61 [0.56, 0.66] 0.69 [0.66, 0.73] 0.95 [0.75, 1.22] 0.47 [0.38, 0.57] 0.53 [0.45, 0.62] — 0.52 [0.44, 0.61] 0.76 [0.73, 0.80] — 0.83 [0.79, 0.87] 0.65 [0.61, 0.69] 0.83 [0.80, 0.88] 0.93 [0.84, 1.03] 0.76 [0.71, 0.82] 0.96 [0.91, 1.01] 0.54 [0.47, 0.62] 0.67 [0.63, 0.72] 0.87 [0.81, 0.95] 0.20 [0.00, 12.22] 0.91 [0.63, 1.31] 0.98 [0.90, 1.07] 0.74 [0.69, 0.78]
— 0.535 0.548 0.527 0.545 0.548 0.544 0.548 0.545 0.566 0.461 0.536 0.539 — 0.542 0.554 — 0.541 0.559 0.546 0.532 0.549 0.537 0.547 0.554 0.534 0.500 0.528 0.527 0.552
57
Supplementary Table 26 | Proportional hazards analysis for in-hospital mortality in the INSPIRE cohort. Per-analyte hazard ratios (with 95% confidence intervals) and concordance indices for PopRI , PerRI , and NORMARI abnormality flags, adjusted for age and sex. PopRI
PerRI
NORMARI
Analyte
N train (events)
N test (events)
HR [95% CI]
C
HR [95% CI]
C
HR [95% CI]
C
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
233 (13) 18427 (470) 16668 (437) 15676 (441) 15814 (415) 17938 (447) 17863 (456) 18561 (442) 3742 (331) 24427 (421) — 15690 (419) 23831 (460) — 21702 (466) 21804 (440) — — — — — 21663 (455) 20394 (466) — — 17133 (424) — — 17152 (470) 20554 (458)
156 (8) 12285 (314) 11113 (292) 10452 (294) 10544 (277) 11960 (298) 11909 (304) 12374 (294) 2495 (221) 16285 (281) — 10461 (279) 15888 (307) — 14469 (310) 14537 (294) — — — — — 14442 (303) 13597 (310) — — 11423 (282) — — 11436 (313) 13703 (305)
0.13 [0.01, 2.06] 6.58 [4.90, 8.83] 2.10 [1.72, 2.56] 1.66 [1.37, 2.02] 2.17 [1.76, 2.68] 2.40 [1.90, 3.02] 3.08 [2.48, 3.83] 2.79 [2.21, 3.53] 3.71 [2.58, 5.34] 2.63 [2.12, 3.26] — 1.87 [1.41, 2.49] 4.33 [3.01, 6.23] — 4.38 [3.06, 6.29] 2.98 [2.43, 3.67] — — — — — 2.37 [1.95, 2.88] 4.55 [3.66, 5.65] — — 3.36 [2.76, 4.10] — — 4.44 [3.57, 5.52] 3.80 [3.03, 4.78]
0.446 0.805 0.745 0.632 0.742 0.724 0.764 0.791 0.643 0.731 — 0.681 0.730 — 0.719 0.821 — — — — — 0.741 0.846 — — 0.758 — — 0.822 0.755
0.13 [0.01, 2.06] 4.28 [3.06, 6.00] 2.40 [1.89, 3.04] 1.48 [1.20, 1.82] 1.99 [1.62, 2.45] 2.21 [1.71, 2.86] 2.89 [2.25, 3.71] 2.61 [2.02, 3.39] 3.25 [2.10, 5.03] 2.47 [1.95, 3.13] — 1.91 [1.37, 2.65] 3.96 [2.49, 6.29] — 3.82 [2.48, 5.89] 2.61 [2.10, 3.24] — — — — — 2.90 [2.31, 3.65] 3.75 [2.89, 4.86] — — 2.78 [2.25, 3.43] — — 3.20 [2.54, 4.04] 3.42 [2.68, 4.37]
0.446 0.719 0.727 0.636 0.713 0.670 0.713 0.718 0.616 0.745 — 0.658 0.676 — 0.683 0.747 — — — — — 0.738 0.752 — — 0.761 — — 0.733 0.751
— 5.87 [4.30, 8.02] 2.43 [1.98, 2.98] 1.51 [1.24, 1.83] 2.21 [1.81, 2.70] 2.41 [1.91, 3.06] 3.24 [2.58, 4.07] 2.81 [2.20, 3.60] 3.20 [2.15, 4.76] 2.70 [2.16, 3.38] — 2.07 [1.39, 3.09] 4.07 [2.76, 6.00] — 4.31 [2.94, 6.31] 2.99 [2.42, 3.69] — — — — — 2.64 [2.15, 3.24] 3.96 [3.15, 4.96] — — 2.87 [2.36, 3.51] — — 3.93 [3.15, 4.90] 3.83 [3.03, 4.84]
0.482 0.785 0.754 0.637 0.732 0.707 0.776 0.773 0.627 0.750 — 0.655 0.724 — 0.717 0.806 — — — — — 0.765 0.816 — — 0.765 — — 0.799 0.799
58
Supplementary Table 27 | Proportional hazards analysis for perioperative infection in the INSPIRE cohort. Peranalyte hazard ratios (with 95% confidence intervals) and concordance indices for PopRI , PerRI , and NORMARI abnormality flags, adjusted for age and sex. PopRI
PerRI
NORMARI
Analyte
N train (events)
N test (events)
HR [95% CI]
C
HR [95% CI]
C
HR [95% CI]
C
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
226 (32) 18112 (691) 16373 (656) 15382 (665) 15519 (665) 17637 (682) 17550 (687) 18256 (689) 3583 (413) 24107 (694) — 15420 (629) 23504 (722) — 21375 (716) 21492 (704) — — — — — 21349 (703) 20078 (706) — — 16829 (676) — — 16841 (680) 20235 (707)
152 (21) 12076 (461) 10916 (438) 10255 (443) 10347 (444) 11758 (455) 11700 (458) 12172 (460) 2390 (275) 16072 (462) — 10281 (419) 15670 (482) — 14251 (478) 14329 (469) — — — — — 14234 (469) 13386 (471) — — 11220 (450) — — 11228 (454) 13491 (471)
0.82 [0.35, 1.97] 1.22 [1.05, 1.43] 1.16 [0.98, 1.38] 1.03 [0.88, 1.21] 1.45 [1.15, 1.84] 1.09 [0.94, 1.27] 1.17 [1.00, 1.36] 1.09 [0.94, 1.27] 1.00 [0.81, 1.22] 1.08 [0.92, 1.26] — 0.97 [0.81, 1.16] 1.09 [0.92, 1.28] — 1.06 [0.90, 1.25] 1.04 [0.88, 1.24] — — — — — 1.12 [0.95, 1.33] 1.06 [0.90, 1.26] — — 1.07 [0.87, 1.31] — — 1.27 [1.08, 1.49] 1.18 [1.01, 1.37]
0.559 0.624 0.593 0.610 0.642 0.557 0.582 0.603 0.575 0.573 — 0.547 0.616 — 0.602 0.610 — — — — — 0.604 0.611 — — 0.576 — — 0.601 0.558
1.30 [0.43, 4.00] 1.14 [0.97, 1.34] 1.12 [0.96, 1.30] 1.02 [0.87, 1.19] 1.20 [1.02, 1.40] 1.06 [0.90, 1.24] 1.21 [1.04, 1.42] 1.03 [0.88, 1.20] 0.98 [0.78, 1.23] 0.97 [0.84, 1.13] — 1.00 [0.81, 1.22] 1.10 [0.90, 1.35] — 1.00 [0.82, 1.20] 1.05 [0.90, 1.22] — — — — — 1.03 [0.88, 1.19] 1.15 [0.99, 1.33] — — 1.10 [0.94, 1.29] — — 1.41 [1.21, 1.64] 1.21 [1.05, 1.41]
0.572 0.626 0.598 0.610 0.625 0.548 0.593 0.595 0.577 0.571 — 0.547 0.612 — 0.592 0.609 — — — — — 0.600 0.603 — — 0.584 — — 0.582 0.564
1.80 [0.58, 5.56] 1.15 [0.98, 1.34] 1.16 [0.99, 1.36] 1.08 [0.93, 1.26] 1.26 [1.03, 1.54] 1.12 [0.96, 1.31] 1.15 [0.99, 1.34] 1.14 [0.98, 1.32] 1.02 [0.82, 1.27] 1.07 [0.91, 1.25] — 1.00 [0.79, 1.27] 1.05 [0.88, 1.26] — 1.06 [0.89, 1.26] 1.02 [0.87, 1.19] — — — — — 0.97 [0.83, 1.13] 1.22 [1.05, 1.41] — — 1.02 [0.85, 1.22] — — 1.25 [1.08, 1.46] 1.26 [1.09, 1.47]
0.572 0.631 0.590 0.614 0.644 0.556 0.585 0.605 0.577 0.571 — 0.547 0.615 — 0.601 0.611 — — — — — 0.593 0.638 — — 0.578 — — 0.607 0.559
59
Supplementary Table 28 | Proportional hazards analysis for prolonged hospital stay in the INSPIRE cohort. Prolonged hospital stay is defined as hospital stay greater than 7 days. Per-analyte hazard ratios (with 95% confidence intervals) and concordance indices for PopRI , PerRI , and NORMARI abnormality flags, adjusted for age and sex. PopRI
PerRI
NORMARI
Analyte
N train (events)
N test (events)
HR [95% CI]
C
HR [95% CI]
C
HR [95% CI]
C
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
233 (164) 18427 (14140) 16668 (13163) 15676 (12343) 15814 (12427) 17938 (14031) 17863 (13403) 18561 (14572) 3742 (3239) 24427 (17097) — 15690 (11167) 23831 (16876) — 21702 (15789) 21804 (15982) — — — — — 21663 (15923) 20394 (15193) — — 17133 (13177) — — 17152 (13544) 20554 (15316)
156 (109) 12285 (9427) 11113 (8776) 10452 (8230) 10544 (8286) 11960 (9355) 11909 (8936) 12374 (9714) 2495 (2160) 16285 (11399) — 10461 (7446) 15888 (11251) — 14469 (10527) 14537 (10655) — — — — — 14442 (10615) 13597 (10129) — — 11423 (8786) — — 11436 (9031) 13703 (10211)
0.95 [0.62, 1.45] 0.70 [0.67, 0.73] 0.67 [0.64, 0.71] 0.85 [0.82, 0.88] 0.71 [0.65, 0.77] 0.75 [0.72, 0.77] 0.89 [0.86, 0.92] 0.77 [0.74, 0.79] 0.83 [0.77, 0.89] 0.65 [0.63, 0.68] — 1.03 [0.99, 1.07] 0.81 [0.78, 0.83] — 0.74 [0.71, 0.76] 0.73 [0.69, 0.76] — — — — — 0.71 [0.68, 0.74] 0.69 [0.66, 0.73] — — 0.87 [0.82, 0.92] — — 0.79 [0.76, 0.83] 0.80 [0.78, 0.83]
0.511 0.566 0.558 0.545 0.539 0.553 0.539 0.552 0.516 0.580 — 0.536 0.560 — 0.566 0.538 — — — — — 0.546 0.553 — — 0.534 — — 0.542 0.545
0.82 [0.46, 1.45] 0.73 [0.71, 0.76] 1.10 [1.06, 1.13] 1.02 [0.99, 1.06] 1.12 [1.08, 1.16] 0.96 [0.92, 0.99] 0.95 [0.92, 0.98] 0.97 [0.93, 1.00] 0.98 [0.91, 1.06] 0.71 [0.68, 0.73] — 1.07 [1.02, 1.11] 0.85 [0.82, 0.88] — 0.75 [0.72, 0.78] 1.02 [0.99, 1.05] — — — — — 0.95 [0.92, 0.98] 1.22 [1.18, 1.26] — — 0.95 [0.92, 0.99] — — 0.96 [0.93, 0.99] 0.97 [0.94, 1.00]
0.500 0.568 0.532 0.533 0.539 0.538 0.538 0.525 0.507 0.574 — 0.538 0.555 — 0.557 0.533 — — — — — 0.528 0.553 — — 0.532 — — 0.535 0.537
0.88 [0.52, 1.48] 0.66 [0.64, 0.69] 0.82 [0.79, 0.85] 0.92 [0.89, 0.95] 0.95 [0.90, 1.00] 0.76 [0.74, 0.79] 0.88 [0.85, 0.91] 0.83 [0.81, 0.86] 0.85 [0.79, 0.92] 0.68 [0.66, 0.71] — 1.02 [0.97, 1.07] 0.80 [0.78, 0.83] — 0.71 [0.69, 0.74] 0.85 [0.82, 0.89] — — — — — 0.76 [0.73, 0.79] 1.04 [1.00, 1.07] — — 0.95 [0.91, 0.99] — — 0.78 [0.75, 0.81] 0.84 [0.81, 0.87]
0.523 0.572 0.538 0.536 0.530 0.551 0.540 0.540 0.515 0.578 — 0.536 0.560 — 0.568 0.532 — — — — — 0.545 0.542 — — 0.532 — — 0.545 0.543
60
Supplementary Table 29 | Proportional hazards analysis for unplanned ICU admission in the INSPIRE cohort. Per-analyte hazard ratios (with 95% confidence intervals) and concordance indices for PopRI , PerRI , and NORMARI abnormality flags, adjusted for age and sex. PopRI
PerRI
NORMARI
Analyte
N train (events)
N test (events)
HR [95% CI]
C
HR [95% CI]
C
HR [95% CI]
C
A1C ALB ALP ALT AST BUN CA CL CO2 CRE DBIL GLU HCT HDL HGB K LDL MCH MCHC MCV MPV NA PLT RBC RDW TBIL TC TGL TP WBC
233 (105) 18427 (5221) 16668 (5096) 15676 (5066) 15814 (5026) 17938 (5376) 17863 (5265) 18561 (5536) 3742 (2059) 24427 (5246) — 15690 (4539) 23831 (6037) — 21702 (5792) 21804 (5871) — — — — — 21663 (5875) 20394 (5724) — — 17133 (5043) — — 17152 (5154) 20554 (5663)
156 (70) 12285 (3481) 11113 (3397) 10452 (3378) 10544 (3351) 11960 (3585) 11909 (3510) 12374 (3691) 2495 (1373) 16285 (3498) — 10461 (3026) 15888 (4025) — 14469 (3862) 14537 (3915) — — — — — 14442 (3916) 13597 (3817) — — 11423 (3363) — — 11436 (3437) 13703 (3775)
1.09 [0.68, 1.74] 0.85 [0.80, 0.91] 1.07 [0.99, 1.15] 0.95 [0.90, 1.02] 1.14 [1.01, 1.28] 1.00 [0.94, 1.06] 0.85 [0.80, 0.90] 1.29 [1.22, 1.37] 0.93 [0.86, 1.02] 1.26 [1.18, 1.35] — 1.03 [0.97, 1.10] 0.94 [0.89, 0.99] — 1.01 [0.96, 1.07] 1.11 [1.02, 1.20] — — — — — 1.31 [1.22, 1.40] 1.35 [1.25, 1.45] — — 1.11 [1.02, 1.21] — — 0.87 [0.81, 0.94] 1.16 [1.10, 1.23]
0.495 0.541 0.535 0.535 0.530 0.523 0.534 0.530 0.521 0.563 — 0.529 0.529 — 0.532 0.520 — — — — — 0.532 0.537 — — 0.534 — — 0.539 0.540
1.00 [0.59, 1.69] 0.90 [0.86, 0.96] 0.96 [0.91, 1.01] 0.94 [0.88, 0.99] 0.84 [0.79, 0.89] 0.94 [0.89, 0.99] 0.87 [0.82, 0.91] 1.10 [1.04, 1.15] 0.93 [0.84, 1.02] 1.19 [1.12, 1.26] — 0.93 [0.87, 1.00] 0.98 [0.92, 1.04] — 1.10 [1.04, 1.17] 0.93 [0.88, 0.98] — — — — — 0.98 [0.93, 1.03] 0.83 [0.79, 0.87] — — 1.07 [1.01, 1.14] — — 0.87 [0.82, 0.92] 1.00 [0.95, 1.05]
0.499 0.542 0.539 0.537 0.540 0.524 0.536 0.521 0.520 0.562 — 0.531 0.529 — 0.532 0.522 — — — — — 0.511 0.540 — — 0.535 — — 0.542 0.537
1.17 [0.66, 2.07] 0.94 [0.89, 1.00] 1.01 [0.95, 1.08] 1.01 [0.96, 1.07] 0.90 [0.82, 0.98] 1.01 [0.96, 1.07] 0.84 [0.79, 0.89] 1.42 [1.34, 1.49] 0.90 [0.82, 0.99] 1.22 [1.15, 1.30] — 1.26 [1.15, 1.37] 1.05 [0.99, 1.10] — 1.05 [0.99, 1.11] 1.13 [1.05, 1.20] — — — — — 1.37 [1.29, 1.45] 0.99 [0.93, 1.04] — — 1.09 [1.02, 1.17] — — 0.89 [0.83, 0.95] 1.15 [1.09, 1.22]
0.528 0.539 0.536 0.534 0.532 0.523 0.536 0.544 0.518 0.562 — 0.529 0.528 — 0.533 0.518 — — — — — 0.539 0.531 — — 0.535 — — 0.538 0.541
61