ConceptioArchivearXiv CS
arXiv CSopen access

Benchmarking Machine Learning Architectures for Antimicrobial Stewardship in Pediatric ICUs

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
neuralnetworks
machine learning, deep learning, neural networks

Preprint: Under Review 1–33

Benchmarking Machine Learning Architectures for Antimicrobial Stewardship in Pediatric ICUs

arXiv:2605.22611v1 [cs.LG] 21 May 2026

Niklas Raehse1,2 Luregn J. Schlapbach1 Daphné Chopard1,3

[email protected] [email protected] [email protected] 1 Department of Intensive Care and Neonatology and Children’s Research Center, University of Zurich, University Children’s Hospital Zurich, Zurich, Switzerland 2 Department of Health Sciences and Technology, ETH Zurich, Zurich, Switzerland 3 Department of Computer Science, ETH Zurich, Zurich, Switzerland

Abstract Antimicrobial stewardship (AMS) is critical for reducing unnecessary antibiotic exposure, particularly in pediatric intensive care units (PICUs), where clinical uncertainty leads to frequent broad-spectrum use, and overuse amplifies antimicrobial resistance and may have long-term consequences. Machine learning has been proposed to support AMS by identifying patient-level opportunities for intervention from electronic health record data. However, prior work has focused on adult populations and has predominantly relied on static tabular representations, leaving open questions about target design, temporal modeling, and generalizability in pediatric settings. In this work, we present a systematic benchmarking study of AMS intervention prediction in the PICU. Using both a public dataset and a private institutional cohort, we define four clinically relevant proxy targets for reducing antibiotic use (intravenous-to-oral switching, de-escalation, discontinuation, and short-course therapy) and compare tabular, sequence-based, and graph-based temporal models under a unified evaluation framework. We find that model performance is primarily driven by target prevalence and data characteristics rather than model complexity. Sequence models provide improvements in precision-recall trade-off over tabular approaches at coarse (24-hour) resolution, with limited additional gains when finer temporal structure is incorporated. However, these gains come at the cost of poorer calibration, with simpler tabular models producing more reliable probability estimates. Multi-task learning yields only marginal improvements, suggesting limited shared structure across stewardship targets. Our results highlight the importance of target selection, temporal representation, and calibration in clinical machine learning, and provide practical guidance for developing reliable decision support systems for AMS in pediatric care. The code will be made publicly available1 .

1. Introduction Antimicrobial resistance is a growing global threat, and optimizing antibiotic use through antimicrobial stewardship (AMS) is essential to mitigate its impact (World Health Organization, 2015; Barlam et al., 2016). In the pediatric intensive care unit (PICU), antibiotics are among the most frequently administered therapies (Willems et al., 2021) and are often initiated by default due to non-specific signs of systemic inflammation in critically ill 1. https://anonymous.4open.science/r/AMS_intervention_prediction-C024

© N. Raehse1,2 , L.J. Schlapbach1 & D. Chopard1,3 .

PICU stewardship benchmarking

children (Weiss et al., 2020; Evans et al., 2018). This leads to substantial overuse, much of which may be unnecessary or overly broad (Chiotos et al., 2022; Abbas et al., 2016), contributing to toxicity and antimicrobial resistance, with amplified long-term consequences in children (Romandini et al., 2021; Neuman et al., 2018; Sivasankar et al., 2023; Fanelli et al., 2020a). AMS programs aim to identify opportunities to modify therapy, including de-escalation, discontinuation, or switching from intravenous to oral treatment (Pollack and Srinivasan, 2014; Khdour et al., 2018; Deresinski, 2007). However, these processes are resource-intensive and may miss timely intervention opportunities (World Health Organization, 2019), motivating the development of clinical decision support (CDS) tools to assist clinicians in identifying when antibiotic therapy could be safely adjusted (Giacobbe et al., 2024). Machine learning (ML) methods using electronic health record (EHR) data have been proposed to support AMS by predicting when stewardship interventions are likely (Giacobbe et al., 2024; Tran-The et al., 2024). These approaches typically rely on retrospective targets derived from historical clinical decisions, which reflect observed practice rather than optimal care, but remain useful for capturing real-world stewardship behavior at scale. Identifying contexts in which therapy modification is likely to be appropriate is particularly valuable, as clinicians may be hesitant to adjust treatment (Chiotos and Tamma, 2020). Despite these advances, AMS intervention prediction in the PICU remains an open challenge (Fanelli et al., 2020b). Compared to adult settings, pediatric data are more heterogeneous and sparse (Shah et al., 2023), and little work has addressed stewardship prediction in this population. Consequently, it remains unclear which intervention targets are most informative and which modeling approaches best capture the temporal and clinical structure of stewardship decisions. In particular, commonly used targets such as intravenous-to-oral switching are underutilized in pediatric settings (Berthe-Aucejo et al., 2012; McMullan et al., 2016), highlighting the need to reassess target definitions. From a modeling perspective, patient data can be represented as static snapshots or temporal sequences, but the impact of these choices has not been systematically studied in AMS, and model selection is often driven by convention rather than evidence. In this work, we address these gaps by systematically benchmarking ML architectures for AMS intervention prediction in the PICU. We evaluate models across two cohorts, a publicly available PICU dataset (Zeng et al., 2020) and a private institutional dataset, to obtain a broader understanding of findings across pediatric populations, and define four stewardship proxy outcomes clinically relevant in the PICU settings: intravenous-to-oral switching, de-escalation, discontinuation, and short-course therapy. We compare tabular, sequence-based, and graph-based temporal models under a consistent framework, providing insights into how temporal representation and model choice affect predictive performance and offering guidance for the development of reliable CDS tools for pediatric stewardship. Generalizable Insights about Machine Learning in the Context of Healthcare We derive several insights for applying machine learning in clinical settings beyond antimicrobial stewardship. First, evaluation metrics can lead to different conclusions about model utility: while AUROC is often comparable across models, meaningful differences emerge in AUPRC and F1 score, where temporal models show improvements over tabular ones. These metrics better reflect performance at clinically relevant decision thresholds and are therefore more informative for developing CDS tools. Second, prediction target definition strongly in2

PICU stewardship benchmarking

fluences both learnability and practical relevance. Commonly used targets may be too rare or poorly aligned with practice in specific populations, limiting their usefulness. Finally, jointly modeling related clinical tasks does not necessarily yield performance gains, suggesting that shared representations are only beneficial when tasks exhibit sufficiently strong common structure. Together, these findings emphasize that evaluation metrics, representation, target design, and validation strategy are key determinants of model effectiveness in clinical machine learning, and highlight the importance of studying underrepresented populations such as pediatric cohorts, where data characteristics and clinical practices differ substantially from adult settings.

2. Related Work AMS intervention prediction targets. Several studies have proposed predicting AMS interventions from EHR data (Beaudoin et al., 2014; Goodman et al., 2022; Bystritsky et al., 2020; Tran-The et al., 2024), operationalizing stewardship actions such as intravenous-tooral switching, de-escalation, and discontinuation as prediction targets. However, most work focuses on a single intervention type (Bolton et al., 2024, 2025) or collapses multiple actions into a single binary outcome (Bystritsky et al., 2020; Goodman et al., 2022), potentially obscuring differences between stewardship decisions. Moreover, while these targets are clinically related and may share common underlying signals, prior work has not explicitly modeled these relationships. Multi-task learning (MTL) has shown benefits in predicting related tasks (Harutyunyan et al., 2019), but has not been explored for AMS prediction. Furthermore, existing studies are primarily conducted in adult ICUs, and it remains unclear how well these targets transfer to pediatric settings, where intervention patterns and data characteristics differ (Fay and Bryant, 2020; Shah et al., 2023) Modeling approaches for AMS prediction. Prior work spans a range of modeling approaches. Early studies rely on tabular representations with classical models such as logistic regression or gradient boosting (Beaudoin et al., 2014; Tran-The et al., 2024), emphasizing interpretability and ease of deployment. These approaches represent patients as snapshots by aggregating dynamic features into summary statistics, thereby discarding temporal structure for simplicity (Beaudoin et al., 2014; Goodman et al., 2022; Bystritsky et al., 2020; Tran-The et al., 2024). More recent work has explored deep neural networks with post-hoc explanation (Bolton et al., 2024), but these models operate on similar aggregated inputs and therefore do not explicitly model temporal dynamics. While sequence-based models can capture longitudinal information, they have not been systematically evaluated for AMS prediction. Existing studies remain fragmented across private and public datasets, limiting reproducibility and cross-site comparison, and do not jointly assess the impact of temporal representations and model choice. Causal outcome estimation for AMS. A related line of work estimates outcomes under alternative treatment decisions rather than predicting observed interventions. For example, Bolton et al. (2022) use a recurrent neural network and synthetic controls to model counterfactual antibiotic cessation. While this approach targets decision support more directly, it relies on strong assumptions and unobserved confounding. In contrast, predicting observed interventions provides a simpler and more reproducible supervised learning setting, albeit as a proxy for optimal decisions. 3

PICU stewardship benchmarking

PICU-specific challenges. ML for pediatric AMS remains underexplored (Fanelli et al., 2020b). Compared to adult settings, PICU data are more heterogeneous, with greater variability in physiology and treatment patterns, as well as smaller cohort sizes, limiting the applicability of adult-derived models (Shah et al., 2023). While ML has been applied to related pediatric tasks such as infection detection (Martin et al., 2022; Lee et al., 2022; Fenta et al., 2024; Clarke et al., 2022), antibiotic susceptibility prediction (Oonsivilai et al., 2018), and clinical deterioration (Velez et al., 2025), no prior work has studied AMS intervention prediction models in the PICU. To address these gaps, we provide a systematic, cross-cohort evaluation of AMS intervention prediction in PICU using both a public and a private dataset. We jointly examine target formulation, temporal representation, and modeling paradigm, enabling a direct comparison of tabular, sequence, and graph-based models across clinically relevant stewardship tasks.

3. Methods We formulate AMS intervention prediction as a patient-day classification task using EHR data. Our goal is to identify, at the start of each antibiotic day, whether a stewardshiprelevant intervention is likely to occur during that day. This section describes the prediction setting, data representation pipeline, model classes, and evaluation protocol (also see Figure 1), whereas cohort-specific preprocessing details are provided in Section 4. The code is made publicly available for reproducibility2 . 3.1. Prediction setting Predictions are made once per day at a fixed time, using only information available up to that point. Each prediction is defined for a patient-day with active antibiotic therapy and aims to predict whether an intervention occurs within the following 24 hours. For tabular models, predictions are based on a single patient-day feature vector summarizing the most recent patient state. Temporal models use the same patient-day representation at each time step, but operate on sequences of consecutive patient-days, allowing them to incorporate information from prior days. This unified formulation ensures that all models receive the same per-day input, while differing only in their ability to leverage longitudinal context. 3.2. Feature space We extract a consistent set of clinical variables across cohorts, including (i) patient and admission context, such as age and admission type, (ii) physiological measurements, namely vital signs and laboratory values, and (iii) antibiotic exposure features (e.g., ATC code, route, concurrent therapies). Full listing of all input features is available in Table 2 in Appendix A alongside more detailed descriptions. 2. https://anonymous.4open.science/r/AMS_intervention_prediction-C024

4

PICU stewardship benchmarking

PIC 37 raw features

12,514 admissions

Admission

Antibiotics

• Age, sex • Weight, height • Length of stay

• ATC code • Treatment day •…

Laboratory

1h resolution bins of last 24h

Vitals

• pH • Lactate •…

static/semi-static features

• Blood pressure • Temperature •… dynamic features

SpO2 --

Bin 24

105

97

1D CNN

Learned representation of the last 24 hours from 1-hour bins

Pt.

Day

Age (years)

Primary ATC code

Concurrent ATC day count

sum

min

max

mean

dim_300

A

1

4

J01CA04

1

48

4320

70

110

90

0.966

0.796

A

2

4

J01CA04

2

52

4680

72

108

90

0.904

0.734

A

3

4

J02AC01

1

50

4450

68

105

89

0.842

0.672

BP (mmHg) [last 24h]

dim_1

Augmented sub-day representation [1h]

Patient-day representation [24h]

Tabular [24h]

Tabular [1h]

LR, LGBM, MLP, TabPFN

LR, LGBM, MLP

Sequential [24h]

Sequential [1h]

Graph

GRU, LSTM, TCN, Transformer, SSM

GRU, LSTM, TCN, Transformer, SSM

RAINDROP

STL

IV-to-oral

Irregular temporal representation

Learnable features [per day]

+

modeling strategies

HR 100

last value per hour

Patient-day dataset

targets

Bin Bin 1

aggregation over the last 24 h

64 features per patient per day

cohorts and features

Private

or

13,449 admissions

MTL

de-escalation

discontinuation

short-course

Figure 1: Benchmarking overview. EHR are represented at two temporal resolutions: (i) a 24-hour representation based on summary statistics of dynamic variables over the preceding day, and (ii) augmented by a 1-hour representation capturing intraday dynamics through a learned embedding of hourly bins using a CNN trained jointly with the prediction task. We benchmark three classes of models under a unified framework. All models are evaluated on four antimicrobial stewardship intervention targets. 3.3. Patient-day representation and temporal representations We formulate the task at the patient-day level, where each instance corresponds to a 24hour interval ending at a fixed daily prediction time. This defines a consistent prediction unit across all models. We consider multiple temporal resolutions of this patient-day representation to disentangle the effect of temporal granularity and modeling assumptions (see Figure 1). These representations provide different levels of temporal detail over the same prediction unit, enabling a controlled comparison of modeling approaches. Patient-day representation (24h resolution) Physiological measurements are aggregated into summary statistics over the preceding 24 hours (e.g., minimum, maximum, mean)

5

PICU stewardship benchmarking

and combined with static and semi-static features, yielding a single feature vector per patient-day of the recent state. This corresponds to the standard tabular representation. Augmented sub-day representation (1h resolution) To capture intra-day dynamics, we construct hourly representations within the same 24-hour window. Observations are binned into hourly intervals, with the last value retained per bin and missing values forwardfilled. A convolutional encoder summarizes this sequence into a compact embedding, which is concatenated with the 24h-resolution feature vector. As a result, the 1h representation extends rather than replaces the 24h representation, allowing models to jointly leverage coarse summary statistics and fine-grained temporal structure. The embedding is learned jointly with the prediction task. This representation allows tabular models to incorporate temporal information, and enables direct comparison across model classes. Irregular temporal representation. We additionally consider a representation that preserves the native irregular sampling of EHR data within each patient-day. This is used by graph-based models, which directly operate on irregular observations and capture temporal and inter-variable relationships without discretization. 3.4. Model classes We evaluate three classes of models, each corresponding to a different inductive bias about how clinical decisions arise from patient data. Implementations details can be found in Appendix A.4. Tabular models Tabular models operate on independent patient-day feature vectors and assume that relevant information is captured by the current state (i.e., previous patientday). In line with previous work on AMS intervention prediction (Goodman et al., 2022; Bystritsky et al., 2020; Tran-The et al., 2024; Bolton et al., 2024), we include logistic regression (Nigam et al., 1999), LightGBM (Ke et al., 2017), a MLP, and the transformerbased tabular foundation model TabPFN (Hollmann et al., 2025). Sequence-based temporal models Sequence models explicitly capture temporal dependencies across patient trajectories. They process patient data as sequences of observations and incorporate information from previous timesteps (here, patient days) to inform predictions, either through recurrent state updates or through architectures that aggregate context across the sequence. We evaluate recurrent architectures (GRU, (Cho et al., 2014), LSTM (Hochreiter and Schmidhuber, 1997)), a Temporal Convolutional Network (TCN (Bai et al., 2018)), a Transformer (Vaswani et al., 2017), and a State Space Model (SSM) inspired by Mamba (Gu and Dao, 2024). We consider both single-task learning (STL) and multi-task learning (MTL) settings. In MTL, a shared encoder is trained jointly across all targets with task-specific output heads, allowing the model to exploit potential relationships between stewardship actions. Graph-based temporal models To model irregularly sampled data without discretization, we include RAINDROP (Zhang et al., 2022), a graph neural network that represents multivariate time series as a graph over variables. Temporal dynamics and inter-variable relationships are captured through message passing and attention mechanisms. Unlike

6

PICU stewardship benchmarking

sequence models, RAINDROP processes each patient-day independently while modeling intra-day temporal structure. 3.5. Targets We define four AMS proxy targets derived from antibiotic treatment trajectories (see Fig. 2): (i) IV-to-oral, transition from intravenous to non-intravenous formulation within a course; (ii) De-escalation, switch to an antibiotic with strictly lower spectrum (based on ASI); (iii) Discontinuation termination of antibiotic therapy prior to discharge with no immediate restart; (iv) Short-course, total antibiotic duration longer than 24h and shorter than 96h (excluding prophylactic cefazolin). Each target is defined at the patient-day level and labeled positive only on the day the event occurs. IV-to-oral and de-escalation correspond to established stewardship actions aimed at reducing treatment invasiveness and narrowing the antimicrobial spectrum. Discontinuation and short-course targets capture opportunities to limit overall antibiotic exposure. Antibiotics course Event

IV-to-oral

Mode: Intravenous

Mode: Oral < 1 day

Time (days)

Event

De-escalation

Broad, ASI = 8 <1 day

Narrow, ASI = 2 Time (days) Event

Event

Discontinuation

Discharge

≥3 days

Time (days) Event

Short-course Time (days)

<96 hours total duration

Figure 2: AMS intervention targets Target design. We build on prior work that operationalizes stewardship interventions from retrospective prescription data (Tran-The et al., 2024), but adapt these definitions to the PICU setting and extend them in two important ways. First, we introduce shortcourse therapy as an additional proxy outcome. In pediatric critical care, clinically indicated antibiotic treatments typically require durations longer than 96 hours to achieve therapeutic effect (Chiotos et al., 2022; Abbas et al., 2016). Short courses may therefore reflect situations where antibiotic therapy was initiated under uncertainty and subsequently discontinued when deemed unnecessary, providing a complementary signal for potential over-treatment. Second, we define de-escalation strictly in terms of a reduction in antibiotic spectrum rather than also a reduction in the number of agents. This distinction is important: reducing the number of antibiotics does not necessarily correspond to narrower antimicrobial coverage and may in some cases reflect escalation (e.g., consolidation to a broad-spectrum agent). 7

PICU stewardship benchmarking

By relying on spectrum-based definitions, we better align the target with stewardship intent and avoid conflating treatment simplification with true de-escalation. Interpretation as intervention opportunities. These targets correspond to observed stewardship interventions and reflect real-world clinical decision-making rather than optimal treatment decisions. Because they are derived from retrospective clinical practice, they reflect when clinicians chose to modify or discontinue therapy, not necessarily when they should have done so. Despite this limitation, such proxies remain valuable for CDS. They provide a scalable and reproducible way to approximate points in the care trajectory where antimicrobial therapy may be unnecessary, overly broad, or suitable for modification. In practice, they enable models to identify patient-days that resemble past situations where clinicians judged that treatment could be safely reduced or adjusted, thereby supporting prospective audit and feedback workflows. Note that by modeling multiple intervention types separately, rather than collapsing them into a single outcome, we aim to capture the diversity of stewardship decisions and provide more clinically actionable signals. 3.6. Evaluation Datasets are split per patient-day into training and test sets (80-20) using a patient-wise split, ensuring that all stays from a given patient are assigned to the same split. Models are trained using the same evaluation protocol across all model classes to ensure fair comparison. Performance is evaluated at the patient-day level using standard classification metrics, including Area Under the Receiver Operating Characteristic (AUROC) (Hanley and McNeil, 1982), F1 Score (van Rijsbergen, 1979), and Area Under the Precision Recall Curve (AUPRC) (Davis and Goadrich, 2006). All experiments are run with a fixed random seed and all results are reported on the held-out test set.

4. Cohorts 4.1. Cohort Selection To evaluate AMS intervention prediction across heterogeneous pediatric settings, we use two complementary PICU cohorts: a publicly available dataset and a private institutional dataset. This dual-cohort design enables both reproducibility and assessment of cross-site comparison, addressing a key limitation of prior work which typically relies on a single type of data source. In both cohorts, we extract antibiotic treatment trajectories and construct patient-day observations corresponding to days on which antibiotic therapy is active, following the prediction setting described in Section 3. PIC The Chinese Pediatric Intensive Care (PIC) database (Li et al., 2020) is a publicly available dataset containing de-identified EHR data from patients admitted to the PICU of the Children’s Hospital of Zhejiang University School of Medicine between 2010 and 2018. It includes demographics, vital signs, laboratory measurements, and medication records, with physiological variables primarily recorded manually. We include all patients with at least one antibiotic prescription during their stay, resulting in 5,671 admissions for 5,515 distinct patients. The cohort is characterized by a young population (median age 0.83 years,

8

PICU stewardship benchmarking

IQR 0.16-3.81) and relatively long ICU stays (median 10.16 days, IQR 4.59-19.52). As a public dataset, PIC provides a reproducible benchmark for method comparison. Private PICU dataset The private cohort consists of de-identified EHR data from a single PICU3 with ethical approval, covering admissions between 2015 and 2023. In contrast to PIC, this dataset reflects a different clinical environment and data collection process, including both sensor-derived and nurse-validated physiological measurements. After applying the same inclusion criteria, the cohort contains 2,789 admissions for 2,059 distinct patients, with a younger and shorter-stay population (median age 0.40 years, IQR 0.04-2.91; median length of stay 5.18 days, IQR 1.77-15.78). This cohort enables evaluation of model robustness and generalization across institutions. 4.2. Cohort-specific characteristics of AMS targets The distribution of AMS intervention targets differs substantially across cohorts and highlights important pediatric-specific considerations. In particular, intravenous-to-oral switching, commonly used as a proxy for stewardship interventions in adult settings, is extremely rare in both PICU cohorts (0.34% in PIC and 0.98% in the private cohort, see Table 4 in the Appendix). In contrast, short-course therapy is substantially more frequent, especially in the private cohort, and often co-occurs with discontinuation, suggesting that it captures a meaningful and prevalent pattern of early treatment cessation. Other combinations of intervention types are rare, indicating that stewardship actions are typically isolated events (see Figure 7 in the Appendix). These differences motivate the inclusion of short-course therapy as an additional proxy outcome tailored to the pediatric setting, complementing established intervention targets and enabling a broader characterization of opportunities for reducing unnecessary antibiotic exposure. 4.3. Data Extraction We extract a consistent set of 37 clinical variables from both cohorts, covering demographics, admission context, antibiotic prescriptions, vital signs, and laboratory measurements (See Table A in the Appendix for full list). To enable consistent modeling and fair comparison, both datasets are mapped to a shared feature space covering demographics, clinical measurements, and antibiotic exposure. Antibiotics records are limited to an overlapping set and harmonized using ATC codes (see Table 3 in the Appendix). We assume that prescription timestamps approximate administered therapy in PIC. For the Private cohort, the available timestamps correspond to actual antibiotic administrations. Laboratory measurements are aligned using LOINC codes (Regenstrief Institute, 2024). 4.4. Feature construction We apply the unified feature construction pipeline described in Section 3 to both cohorts, ensuring that differences in performance reflect underlying data characteristics rather than preprocessing choices. Here, we highlight cohort-specific implementation details. Antibiotic prescriptions are mapped to systemic antibiotics and merged into clinically contiguous 3. institution anonymized for review

9

PICU stewardship benchmarking

courses based on temporal proximity (within 24 hours). This step is particularly important for the Private cohort, where prescriptions were mostly recorded as individual bolus administrations rather than continuous treatments. Very short courses (< 24 hours), typically reflecting prophylactic use, are excluded. When multiple antibiotic courses overlap within a patient-day, all active agents are retained and jointly contribute to feature construction and target labeling, rather than collapsing to a single representative treatment. For the Private cohort, antibiotic prescription timestamps correspond to actual administrations, while in PIC they approximate administered therapy. This difference does not affect feature construction but is relevant for interpreting temporal alignment. All remaining preprocessing and feature aggregation steps follow the shared pipeline described in Section 3.

5. Results 5.1. Comparison to the adult setting and previous work We compare tabular models trained on 24-hour aggregated pediatric data (LightGBM and MLP) to previously reported results in adult settings. Results are shown in Table 1. Table 1: Comparison of tabular models trained on 24-hour aggregated data with results reported by Tran-The et al. (2024) on a private Korean cohort and Bolton et al. (2024) on MIMIC-IV and eICU. For comparison, the early- and late-de-escalation target metrics of Tran-The et al. (2024) were averaged. We report AUROC, AUPRC, the true positive rate (TPR), true negative rate (TNR), positive predictive value (PPV), negative predictive value (NPV). Task

Cohort

Model

Prev.

AUROC

AUPRC

IV-to-oral

Private

LGBM MLP LGBM MLP LGBM MLP MLP

0.010 0.010 0.003 0.003 0.030 – –

0.841 0.566 0.906 0.898 0.810 0.80 0.77

0.072 0.013 0.000 0.144 0.160 0.37 0.33

0.000 1.000 0.000 0.990 0.000 1.000 0.000 0.990 0.167 1.000 0.563 0.997 0.000 1.000 0.000 0.997 0.780 0.710 0.070 0.990 0.85 0.25 – – 0.90 0.35 – –

LGBM MLP LGBM MLP LGBM

0.052 0.052 0.040 0.040 0.035

0.936 0.930 0.939 0.908 0.750

0.429 0.364 0.390 0.270 0.130

0.199 0.991 0.552 0.958 0.177 0.985 0.399 0.956 0.164 0.995 0.593 0.966 0.067 0.996 0.391 0.962 0.640 0.760 0.080 0.985

LGBM MLP LGBM MLP LGBM

0.057 0.057 0.017 0.017 0.090

0.718 0.643 0.832 0.765 0.800

0.153 0.085 0.106 0.061 0.360

0.005 1.000 0.500 0.944 0.000 1.000 – 0.943 0.007 1.000 0.286 0.983 0.000 1.000 – 0.983 0.710 0.720 0.210 0.960

PIC Tran The MIMIC-IV† eICU† De-escalation

Private PIC Tran The

Discontinuation

Private PIC Tran The

TPR

TNR

PPV

NPV

Not directly comparable to other cohorts due to differences in cohort selection and label definition (e.g., restricted to patients receiving both IV and oral antibiotics and evaluated at the per-day level).

Intervention prevalence differs substantially between adult and pediatric settings. IV-tooral switching is much rarer in PICU cohorts (0.3-1.0%) than in adult data (3.0%), and discontinuation is also less frequent (1.7-5.7% vs 9%). We observe high AUROC values

10

PICU stewardship benchmarking

in pediatric cohorts (e.g., ∼0.93-0.94 for de-escalation), but precision-recall performance is markedly lower for rare targets. For IV-to-oral for example, LightGBM shows low AUPRC and TPR scores compared to Tran-The et al. (2024), reflecting extreme class imbalance. This highlights that AUROC alone can overestimate performance in such settings. For more prevalent targets, strong AUROC does not consistently translate to improved AUPRC, indicating differences in label distribution and clinical practice. Overall, these results show that performance does not directly transfer from adult to pediatric cohorts. Model behavior is strongly influenced by target prevalence and cohort characteristics, motivating the need for PICU-specific evaluation and the inclusion of additional targets such as short-course therapy. 5.2. Effect of sequence modeling at 24-hour temporal resolution We next investigate whether sequence-based models can better exploit longitudinal structure in patient trajectories compared to tabular models when using the same 24-hour aggregated input representation. Specifically, we compare sequence-based models (recurrent, convolutional, and attention-based) to tabular baselines using the same 24-hour aggregated representation across PIC and Private cohorts. Results are illustrated in Figure 3 with full details available in Appendix C.1. Across tasks and cohorts, sequence models achieve simiIV-to-oral

De-escalation

1.0

Cohort PIC Private

0.8

Score

0.6

Models Tabular [24h]

0.4

Logistic regression LightGBM 0.2

MLP TabPFN

0.0

AUROC

F1

AUPRC

AUROC

Discontinuation

F1

AUPRC

Sequential [24h] GRU

Short-course

LSTM

1.0

TCN Transformer

0.8

SSM

Score

0.6

0.4

0.2

0.0

AUROC

F1

AUPRC

AUROC

F1

AUPRC

Figure 3: Comparison of tabular and sequence models using 24-hour aggregated representations on PIC and Private cohorts across four antimicrobial stewardship targets. lar AUROC to tabular models (e.g., ∼0.85-0.95 for IV-to-oral and de-escalation), indicating that most predictive signal is already captured by the current patient-day representation. Differences are more apparent in AUPRC and F1. For rare targets (IV-to-oral, discon-

11

PICU stewardship benchmarking

tinuation), sequence models provide only modest and inconsistent gains, while for more prevalent targets, sequence models provide substantial gains. Similar trends are observed across cohorts. We observe that for de-escalation and short-course on the Private cohort, the sequence models improved in F1 ∼0.25 to 0.5 and ∼0.18 to 0.45 on average. Smaller but similar performance gains were found for PIC. The results suggest that the explicit temporal modeling can improve the prediction of interventions over tabular approaches, and that exploring finer-grained temporal representations may offer additional improvements. 5.3. Effect of finer temporal resolution and irregular time series modeling We next evaluate whether increasing temporal resolution improves performance. We augment 24-hour features with embeddings derived from 1-hour data and compare tabular and sequence models to a graph-based model (RAINDROP) operating on irregular observations on the PIC cohort (Figure 4). Corresponding results for the private cohort are available in Figure 8 of the Appendix. Incorporating finer-grained temporal information improves performance slightly across tasks. For de-escalation, AUPRC increases from approximately 0.45-0.50 at 24h to 0.50-0.55 with 1h representations. IV-to-oral and discontinuation show limited and rather inconsistent differences (AUPRC < 0.15). Sequence models seem to benefit somewhat from higher-resolution data, with small gains compared to their 24-hour counterparts. Tabular models augmented with temporal embeddings show comparable performance compared to their standard tabular baselines but generally remain below sequence-based approaches. RAINDROP shows subpar performance across tasks and cohorts, with F1 and AUPRC scores falling between those of tabular and sequence models and often lower AUROC, suggesting that its approach to modeling irregular sampling may not fully capture day-level and temporal patterns. We also observe differences in variability across model classes. Sequence models show stable performance across architectures, with differences typically within ∼0.02-0.05 AUPRC. In contrast, tabular models vary more substantially depending on the algorithm. This variability is task-dependent: LightGBM performs best for the most imbalanced target (IV-to-oral, AUPRC ∼0.20-0.30), while TabPFN achieves the strongest results for more prevalent targets such as de-escalation and short-course, followed by LightGBM. Overall, these results indicate that temporal resolution and structure do matter, but they provide only provide incremental benefits, and fully leveraging them requires models that can naturally handle sequences rather than static snapshots. 5.4. Evaluating reliability across model classes Trust is a critical factor in the adoption of clinical decision support systems (CDSS) for antimicrobial stewardship (AMS) (Bolton et al., 2025). In this context, well-calibrated predictions are essential, as they determine how reliably predicted probabilities reflect true intervention likelihood. We therefore evaluate calibration across the different models and temporal representations on the PIC cohort. Results are shown in Figure 5. Corresponding results for the private cohort are available in Appendix Figure 9. Calibration varies substantially across model classes and tasks. Tabular models, particularly LightGBM, show the best calibration, with predicted probabilities closely aligned with observed event rates for more prevalent targets such as de-escalation and short-course. In contrast, sequence mod12

PICU stewardship benchmarking

IV-to-oral

De-escalation

1.0

0.8

Models 0.6

Score

Tabular [24h] Logistic regression LightGBM MLP

0.4

TabPFN Tabular [1h]

0.2

Logistic regression LightGBM MLP

0.0 AUROC

F1

AUPRC

AUROC

F1

AUPRC

Sequential [24h] GRU LSTM

Discontinuation

Short-course

TCN Transformer

1.0

SSM Sequential [1h]

0.8

GRU LSTM TCN

0.6

Score

Transformer SSM Graph-based [~1h]

0.4

RAINDROP

0.2

0.0 AUROC

F1

AUPRC

AUROC

F1

AUPRC

Figure 4: Comparison of models using augmented 1-hour representations on the PIC cohort. Tabular and sequence models are enhanced with learned embeddings derived from hourly data, and compared to a graph-based model (RAINDROP) operating on irregular time series. els are over-confident, with predicted probabilities consistently higher than the observed event rates across bins. The graph-based model RAINDROP exhibits similar behavior, with weak probability separation and poor calibration. Importantly, increasing temporal resolution does not substantially improve calibration. Models using 1-hour representations exhibit similar over-confidence as their 24-hour counterparts, indicating that calibration is largely independent of temporal resolution in this setting. For rare targets (IV-to-oral, discontinuation), all models show unstable calibration, with noisy estimates and unreliable behavior at higher probability ranges due to extreme class imbalance. Overall, these results highlight a clear trade-off between discrimination and calibration. While more complex temporal models can improve ranking performance, simpler tabular models provide more reliable probability estimates, which is critical for building a trustworthy CDSS.

13

PICU stewardship benchmarking

1.0

IV-to-oral

1.0

Test prevalence=0.0034

0.8 Observed event rate

Observed event rate

0.8 0.6 0.4 0.2 0.0 0.0 1.0

0.2

0.4 0.6 Mean predicted probability

0.8

0.4

Tabular [1h] Logistic regression LightGBM MLP

0.0 0.0

1.0

Discontinuation

1.0

Test prevalence=0.0174

0.2

0.4 0.6 Mean predicted probability

0.8

1.0

Short-course

Test prevalence=0.0572

0.6 0.4 0.2

Sequential [24h] GRU LSTM TCN Transformer SSM

0.6

Sequential [1h] GRU LSTM TCN Transformer SSM

0.4

Graph-based [~1h] RAINDROP

0.8 Observed event rate

Observed event rate

Models

Perfect calibration Tabular [24h] Logistic regression LightGBM MLP

0.6

0.2

0.8

0.0 0.0

De-escalation Test prevalence=0.0403

0.2

0.2

0.4 0.6 Mean predicted probability

0.8

1.0

0.0 0.0

0.2

0.4 0.6 Mean predicted probability

0.8

1.0

Figure 5: Calibration plots for the different models on PIC across four AMS targets. 5.5. Multi-task for AMS intervention prediction We finally evaluate whether MTL can improve performance by jointly modeling multiple stewardship targets. Since these interventions are clinically related and sometimes considered jointly (e.g., intervention vs no intervention), we hypothesize that sharing representations across tasks may improve learning. We therefore compare STL and MTL on targets de-escalation, discontinuation, and short-course as these co-occur (Appendix Figure 7) as we expect them to be more related to each other than to IV-to-oral. For sequence models using 24-hour representations. AUROC performance on the PIC cohort across both settings is shown in Figure 6. Appendix Figure 10 shows the same figure for the Private cohort and Appendix section C.3 describes the performance results for each outcome across cohorts in detail. Overall, MTL does not provide consistent improvements over STL. Performance differences are small across architectures, with most tasks showing similar or slightly worse AUROC under MTL. A modest benefit is observed for discontinuation, where MTL yields small but consistent gains across models (on the order of ∼0.01–0.02 AUROC, except for the Transformer). In contrast, no clear improvement is observed for de-escalation or short-course, where performance remains comparable between STL and MTL. These results suggest that, while AMS targets are conceptually related, they may not share sufficiently strong predictive structure to benefit from joint learning in this setting. Instead, task-specific modeling appears to remain important, particularly for more prevalent or heterogeneous targets.

14

PICU stewardship benchmarking

De-escalation 1.0

0.97 0.96

0.96 0.96

Discontinuation 0.96 0.96

0.95 0.95

Short-course

0.96 0.95

0.9 0.84

0.87 0.87

0.85

0.84 0.84

AUROC

0.82

0.83

0.84

0.83

0.81

0.84

0.83

0.84

0.85

0.84

0.82

0.81

0.80

0.8 0.76

0.7

0.6

0.5 U

GR

TM LS

N er TC orm nsf Tra

M SS

U GR

TM

LS

N er TC orm nsf Tra

M

SS

U GR

TM

LS

N er TC orm nsf Tra

M

SS

Backbone architecture Single-task learning (STL)

Multi-task learning (MTL)

Figure 6: Comparison of single-task and multi-task learning for sequence models across three AMS targets on the PIC cohort.

6. Discussion In this work, we systematically evaluate multiple modeling paradigms for AMS intervention prediction in the PICU across cohorts, targets, and temporal representations, yielding several key insights. First, performance is driven by data characteristics and sequential memory rather than model complexity. Outcome prevalence and temporal structure largely determine both absolute performance and model ranking. In particular, targets such as IV-to-oral switching are underutilized in pediatric settings (Berthe-Aucejo et al., 2012; McMullan et al., 2016) and rarely observed in our PICU cohorts, which results in poor precision-recall performance despite reasonable AUROC. These findings indicate limited transferability from adult cohorts and the importance of specific model development for the pediatric setting. The similarities in data characteristics and model performance between our PICU cohorts suggest feasibility of generalizable pediatric AMS prediction models. Second, more fine-grained temporal modeling only slightly benefits sequential models in terms of F1 and AUPRC compared to tabular approaches, suggesting most signal is captured by aggregated features. This would support simpler time-agnostic model development such as done by Tran-The et al. (2024); Bolton et al. (2024), but we find benefits in temporal sequence modeling for F1 and AUPRC scores. Third, we observe a trade-off between discrimination and reliability. While sequence and graph-based models improve ranking performance, they are consistently less well calibrated than tabular models. Trust and accuracy have been found to be critical for antibiotic CDS systems (Laka et al., 2021; Bolton et al., 2025), and temporal resolution does not improve calibration, the need for explicit calibration strategies are underlined. Lastly, multi-task learning provides limited benefit. Despite conceptual overlap between a subset of stewardship actions, it yields only marginal improvements for some tasks and none for others, suggesting weak shared predictive structure and supporting task-specific modeling. Overall, these findings suggest that for PICU AMS intervention prediction, target definition, temporal representation, and calibration are

15

PICU stewardship benchmarking

as critical as model choice, and increasing model complexity alone is insufficient for reliable clinical decision support. Limitations This work has several limitations. First, as discussed above, prediction targets are derived from retrospective clinical decisions and reflect historical practice rather than optimal care, limiting their interpretation as ground truth. Second, performance is strongly affected by class imbalance, particularly for rare targets such as IV-to-oral switching, where low precision and unstable calibration limit clinical utility. Third, temporal representations rely on design choices such as fixed aggregation windows, which may obscure clinically relevant patterns. Finally, this study is retrospective and does not assess real-world clinical impact, which will require prospective evaluation.

References Qalab Abbas, Aysha Ul Haq, Ritika Kumar, Syed Asad Ali, Karim Hussain, and Sadia Shakoor. Evaluation of antibiotic use in pediatric intensive care unit of a developing country. Indian Journal of Critical Care Medicine, 20(5):291–294, 2016. doi: 10.4103/ 0972-5229.182197. 2, 7 Shaojie Bai, J Zico Kolter, and Vladlen Koltun. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv preprint arXiv:1803.01271, 2018. 6, 29, 30 Tamar F Barlam, Sara E Cosgrove, Lilly M Abbo, et al. Implementing an antibiotic stewardship program: guidelines by the infectious diseases society of america and the society for healthcare epidemiology of america. Clinical Infectious Diseases, 62(10):e51–e77, 2016. 1 Mathieu Beaudoin, Froduald Kabanza, Vincent Nault, and Louis Valiquette. An antimicrobial prescription surveillance system that learns from experience. AI Magazine, 35(1): 15–15, 2014. doi: 10.1609/aimag.v35i1.2500. 3 A Berthe-Aucejo, M Postaire, A Cheikhlard, JR Zahar, and P Bourget. Antibiotic treatment of appendicular peritonitis in children: is the oral route done? Archives de Pediatrie: Organe Officiel de la Societe Francaise de Pediatrie, 19(12):1303–1307, 2012. 2, 15 William J. Bolton, Richard Wilson, Mark Gilchrist, et al. Machine learning and synthetic outcome estimation for individualised antimicrobial cessation. Frontiers in Digital Health, 4:997219, 2022. doi: 10.3389/fdgth.2022.997219. 3 William J. Bolton, Richard Wilson, Mark Gilchrist, Pantelis Georgiou, Alison Holmes, and Timothy M. Rawson. Personalising intravenous to oral antibiotic switch decision making through fair interpretable machine learning. Nature Communications, 15(1):506, 2024. doi: 10.1038/s41467-024-44740-2. 3, 6, 10, 15 William J. Bolton, Richard Wilson, Mark Gilchrist, Pantelis Georgiou, Alison Holmes, and Timothy M. Rawson. The impact of artificial intelligence-driven decision support on uncertain antimicrobial prescribing: A randomised, multimethod study. The Lancet Digital Health, 7(11):100912, 2025. doi: 10.1016/j.landig.2025.100912. 3, 12, 15 16

PICU stewardship benchmarking

Rachel J. Bystritsky, Alex Beltran, Albert T. Young, Andrew Wong, Xiao Hu, and Sarah B. Doernberg. Machine learning for the prediction of antimicrobial stewardship intervention in hospitalized patients receiving broad-spectrum agents. Infection Control & Hospital Epidemiology, 41(9):1022–1027, 2020. doi: 10.1017/ice.2020.213. 3, 6 Kathleen Chiotos and Pranita D. Tamma. Antibiotics: How can we make it as easy to stop as it is to start? Clinical Microbiology and Infection, 26(12):1600–1601, December 2020. ISSN 1198-743X. doi: 10.1016/j.cmi.2020.08.029. 2 Kathleen Chiotos, Jennifer Blumenthal, Juri Boguniewicz, Debra L. Palazzi, Erika L. Stalets, Jessica H. Rubens, Pranita D. Tamma, Stephanie S. Cabler, Jason Newland, Hillary Crandall, Emily Berkman, Robert P. Kavanagh, Hannah R. Stinson, and Jeffrey S. Gerber. Antibiotic indications and appropriateness in the pediatric intensive care unit: A 10-center point prevalence study. Clinical Infectious Diseases, 76(3):e1021–e1030, 2022. doi: 10.1093/cid/ciac698. 2, 7 Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations using RNN encoder–decoder for statistical machine translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1724– 1734, 2014. doi: 10.3115/v1/D14-1179. 6, 29, 30 Sarah LN Clarke, Kevon Parmesar, Moin A Saleem, and Athimalaipet V Ramanan. Future of machine learning in paediatrics. Archives of Disease in Childhood, 107(3):223–228, 2022. 4 Jesse Davis and Mark Goadrich. The relationship between Precision-Recall and ROC curves. In Proceedings of the 23rd International Conference on Machine Learning (ICML), pages 233–240, 2006. doi: 10.1145/1143844.1143874. 8 Stan Deresinski. Principles of Antibiotic Therapy in Severe Infections: Optimizing the Therapeutic Approach by Use of Laboratory and Clinical Data. Clinical Infectious Diseases, 45(Supplement 3):S177–S183, September 2007. ISSN 1058-4838. doi: 10.1086/519472. 2 Idris V. R. Evans, Gary S. Phillips, Elizabeth R. Alpern, Derek C. Angus, Marcus E. Friedrich, Niranjan Kissoon, Stanley Lemeshow, Mitchell M. Levy, Margaret M. Parker, Kathleen M. Terry, R. Scott Watson, Scott L. Weiss, Jerry Zimmerman, and Christopher W. Seymour. Association Between the New York Sepsis Care Mandate and InHospital Mortality for Pediatric Sepsis. JAMA, 320(4):358–367, July 2018. ISSN 00987484. doi: 10.1001/jama.2018.9071. 2 Umberto Fanelli, Vincenzo Chiné, Marco Pappalardo, Pierpacifico Gismondi, and Susanna Esposito. Improving the Quality of Hospital Antibiotic Use: Impact on MultidrugResistant Bacterial Infections in Children. Frontiers in Pharmacology, 11, May 2020a. ISSN 1663-9812. doi: 10.3389/fphar.2020.00745. 2 Umberto Fanelli, Marco Pappalardo, Vincenzo Chinè, Pierpacifico Gismondi, Cosimo Neglia, Alberto Argentiero, Adriana Calderaro, Andrea Prati, and Susanna Esposito. 17

PICU stewardship benchmarking

Role of Artificial Intelligence in Fighting Antimicrobial Resistance in Pediatrics. Antibiotics, 9(11):767, November 2020b. ISSN 2079-6382. doi: 10.3390/antibiotics9110767. 2, 4 Michael-John Fay and Penelope A Bryant. Antimicrobial stewardship in children: Where to from here? Journal of Paediatrics and Child Health, 56(10):1504–1507, 2020. 3 Haile Mekonnen Fenta, Temesgen T Zewotir, Saloshni Naidoo, Rajen N Naidoo, and Henry Mwambi. Factors of acute respiratory infection among under-five children across subsaharan african countries using machine learning approaches. Scientific Reports, 14(1): 15801, 2024. 4 Jeffrey S. Gerber, Adam L. Hersh, Matthew P. Kronman, Jason G. Newland, Rachael K. Ross, and Talene A. Metjian. Development and application of an antibiotic spectrum index for benchmarking antibiotic selection patterns across hospitals. Infection Control & Hospital Epidemiology, 38(8):993–997, 2017. doi: 10.1017/ice.2017.114. 24, 31 Daniele Roberto Giacobbe, Cristina Marelli, Sabrina Guastavino, Sara Mora, Nicola Rosso, Alessio Signori, Cristina Campi, Mauro Giacomini, and Matteo Bassetti. Explainable and interpretable machine learning for antimicrobial stewardship: Opportunities and challenges. Clinical Therapeutics, 46(6):474–480, 2024. doi: 10.1016/j.clinthera.2024.02. 010. 2 A. L. Goldberger, L. A. N. Amaral, L. Glass, J. M. Hausdorff, P. Ch. Ivanov, R. G. Mark, J. E. Mietus, G. B. Moody, C.-K. Peng, and H. E. Stanley. PhysioBank, PhysioToolkit, and PhysioNet: Components of a new research resource for complex physiologic signals. Circulation, 101(23):e215–e220, 2000 (June 13). Circulation Electronic Pages: http://circ.ahajournals.org/content/101/23/e215.full PMID:1085218; doi: 10.1161/01.CIR.101.23.e215. 32 Katherine E. Goodman, Emily L. Heil, Kimberly C. Claeys, Mary Banoub, and Jacqueline T. Bork. Real-world antimicrobial stewardship experience in a large academic medical center: Using statistical and machine learning approaches to identify intervention “hotspots” in an antibiotic audit and feedback program. Open Forum Infectious Diseases, 9(7):ofac289, 2022. doi: 10.1093/ofid/ofac289. 3, 6 Albert Gu and Tri Dao. Mamba: Linear-Time Sequence Modeling with Selective State Spaces, May 2024. 6, 30 James A. Hanley and Barbara J. McNeil. The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology, 143(1):29–36, 1982. doi: 10.1148/ radiology.143.1.7063747. 8 Hrayr Harutyunyan, Hrant Khachatrian, David C. Kale, Greg Ver Steeg, and Aram Galstyan. Multitask learning and benchmarking with clinical time series data. Scientific Data, 6(1):96, June 2019. ISSN 2052-4463. doi: 10.1038/s41597-019-0103-9. 3 Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural Computation, 9(8):1735–1780, 1997. doi: 10.1162/neco.1997.9.8.1735. 6, 29, 30 18

PICU stewardship benchmarking

Noah Hollmann, Samuel Müller, Lennart Purucker, Arjun Krishnakumar, Max Körfer, Shi Bin Hoo, Robin Tibor Schirrmeister, and Frank Hutter. Accurate predictions on small data with a tabular foundation model. Nature, 01 2025. doi: 10.1038/s41586-024-08328-6. URL https://www.nature.com/articles/s41586-024-08328-6. 6, 29 Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. Lightgbm: A highly efficient gradient boosting decision tree. In Advances in Neural Information Processing Systems, volume 30, 2017. 6, 29 Alex Kendall, Yarin Gal, and Roberto Cipolla. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7482–7491, 2018. doi: 10.1109/CVPR.2018.00781. 30 Maher R. Khdour, Hussein O. Hallak, Mamoon A. Aldeyab, Mowaffaq A. Nasif, Aliaa M. Khalili, Ahamad A. Dallashi, Mohammad B. Khofash, and Michael G. Scott. Impact of antimicrobial stewardship programme on hospitalized patients at the intensive care unit: A prospective audit and feedback study. British Journal of Clinical Pharmacology, 84(4): 708–715, 2018. ISSN 1365-2125. doi: 10.1111/bcp.13486. 2 Diederik P. Kingma and Jimmy Ba. Adam: A Method for Stochastic Optimization, January 2017. 27 Mah Laka, Adriana Milazzo, and Tracy Merlin. Factors That Impact the Adoption of Clinical Decision Support Systems (CDSS) for Antibiotic Management. International Journal of Environmental Research and Public Health, 18(4):1901, January 2021. ISSN 1660-4601. doi: 10.3390/ijerph18041901. 15 Bongjin Lee, Hyun Jung Chung, Hyun Mi Kang, Do Kyun Kim, and Young Ho Kwak. Development and validation of machine learning-driven prediction model for serious bacterial infection among febrile children in emergency departments. PLoS One, 17(3):e0265500, 2022. 4 Haomin Li, Xian Zeng, and Gang Yu. Paediatric Intensive Care database. PhysioNet, November 2020. doi: 10.13026/32x9-wv38. URL https://doi.org/10.13026/ 32x9-wv38. Version 1.1.0. 8 Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H. Chi. Modeling task relationships in multi-task learning with multi-gate mixture-of-experts. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD ’18, pages 1930–1939, New York, NY, USA, July 2018. Association for Computing Machinery. ISBN 978-1-4503-5552-0. doi: 10.1145/3219819.3220007. 30 Blake Martin, Peter E DeWitt, Halden F Scott, Sarah Parker, and Tellen D Bennett. Machine learning approach to predicting absence of serious bacterial infection at picu admission. Hospital pediatrics, 12(6):590–603, 2022. 4 B. J. McMullan, D. Andresen, C. C. Blyth, et al. Antibiotic duration and timing of the switch from intravenous to oral route for bacterial infections in children: systematic 19

PICU stewardship benchmarking

review and guidelines. The Lancet Infectious Diseases, 16(8):e139–e152, 2016. doi: 10. 1016/S1473-3099(16)30024-X. 2, 15 Hadar Neuman, Paul Forsythe, Atara Uzan, Orly Avni, and Omry Koren. Antibiotics in early life: dysbiosis and the damage done. FEMS Microbiology Reviews, 42(4):489–499, 2018. doi: 10.1093/femsre/fuy018. 2 Kamal Nigam, John Lafferty, and Andrew McCallum. Using maximum entropy for text classification. In IJCAI-99 workshop on machine learning for information filtering, volume 1, pages 61–67. Stockholom, Sweden, 1999. 6 Mathupanee Oonsivilai, Yin Mo, Nantasit Luangasanatip, Yoel Lubell, Thyl Miliya, Pisey Tan, Lorn Loeuk, Paul Turner, and Ben S. Cooper. Using machine learning to guide targeted and locally-tailored empiric antibiotic prescribing in a children’s hospital in Cambodia. Wellcome Open Research, 3:131, October 2018. ISSN 2398-502X. doi: 10. 12688/wellcomeopenres.14847.1. 4 F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830, 2011. 28, 29 Loria A. Pollack and Arjun Srinivasan. Core Elements of Hospital Antibiotic Stewardship Programs From the Centers for Disease Control and Prevention. Clinical Infectious Diseases, 59(suppl 3):S97–S100, October 2014. ISSN 1058-4838. doi: 10.1093/cid/ciu542. 2 Regenstrief Institute. LOINC. https://loinc.org/, 2024. Logical Observation Identifiers Names and Codes; used for laboratory concept alignment. 9 Alessandra Romandini, Arianna Pani, Paolo Andrea Schenardi, Giulia Angela Carla Pattarino, Costantino De Giacomo, and Francesco Scaglione. Antibiotic resistance in pediatric infections: Global emerging threats, predicting the near future. Antibiotics, 10(4): 393, 2021. doi: 10.3390/antibiotics10040393. 2 Neel Shah, Ahmed Arshad, Monty B. Mazer, Christopher L. Carroll, Steven L. Shein, and Kenneth E. Remy. The use of machine learning and artificial intelligence within pediatric critical care. Pediatric Research, 93(2):405–412, January 2023. ISSN 1530-0447. doi: 10.1038/s41390-022-02380-6. 2, 3, 4 Shivani Sivasankar, Jennifer L Goldman, and Mark A Hoffman. Variation in antibiotic resistance patterns for children and adults treated at 166 non-affiliated US facilities using EHR data. JAC-Antimicrobial Resistance, 5(1):dlac128, 2023. doi: 10.1093/jacamr/ dlac128. 2 Tam Tran-The, Eunjeong Heo, Sanghee Lim, Yewon Suh, Kyu-Nam Heo, Eunkyung Euni Lee, Ho-Young Lee, Eu Suk Kim, Ju-Yeun Lee, and Se Young Jung. Development of machine learning algorithms for scaling-up antibiotic stewardship. International Journal

20

PICU stewardship benchmarking

of Medical Informatics, 181:105300, 2024. doi: 10.1016/j.ijmedinf.2023.105300. 2, 3, 6, 7, 10, 11, 15 C. J. van Rijsbergen. Information Retrieval. Butterworths, London, 2nd edition, 1979. 8 Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30, 2017. 6, 29, 30 Tom Velez, Oluwakemi Badaki-Makun, Danielle Hirsch, Danielle Claire Mercurio, Holly Depinet, Maya Dewan, Rishikesan Kamaleswaran, Jocelyn Grunwell, Maria Triantafyllou, Fehima Abdelrahman, Charles Macias, and Ioannis Koutroulis. Early prediction of antibiotic need and bacteremia risk in non-immunocompromised pediatric emergency patients using machine learning. Pediatric Research, pages 1–9, December 2025. ISSN 1530-0447. doi: 10.1038/s41390-025-04656-z. 4 Scott L. Weiss, Mark J. Peters, Waleed Alhazzani, Michael S. D. Agus, Heidi R. Flori, David P. Inwald, Simon Nadel, Luregn J. Schlapbach, Robert C. Tasker, Andrew C. Argent, Joe Brierley, Joseph Carcillo, Enitan D. Carrol, Christopher L. Carroll, Ira M. Cheifetz, Karen Choong, Jeffry J. Cies, Andrea T. Cruz, Daniele De Luca, Akash Deep, Saul N. Faust, Claudio Flauzino De Oliveira, Mark W. Hall, Paul Ishimine, Etienne Javouhey, Koen F. M. Joosten, Poonam Joshi, Oliver Karam, Martin C. J. Kneyber, Joris Lemson, Graeme MacLaren, Nilesh M. Mehta, Morten Hylander Møller, Christopher J. L. Newth, Trung C. Nguyen, Akira Nishisaki, Mark E. Nunnally, Margaret M. Parker, Raina M. Paul, Adrienne G. Randolph, Suchitra Ranjit, Lewis H. Romer, Halden F. Scott, Lyvonne N. Tume, Judy T. Verger, Eric A. Williams, Joshua Wolf, Hector R. Wong, Jerry J. Zimmerman, Niranjan Kissoon, and Pierre Tissieres. Surviving sepsis campaign international guidelines for the management of septic shock and sepsis-associated organ dysfunction in children. Intensive Care Medicine, 46(1):10–67, February 2020. ISSN 1432-1238. doi: 10.1007/s00134-019-05878-6. 2 WHO. Anatomical therapeutic chemical (ATC) classification, 2024. URL https://www. whocc.no/atc_ddd_index/. Accessed: 2026-04-14. 31 Rand R. Wilcox. Introduction to Robust Estimation and Hypothesis Testing. Academic Press, December 2011. ISBN 978-0-12-387015-5. 28 Jef Willems, Eline Hermans, Petra Schelstraete, Pieter Depuydt, and Pieter De Cock. Optimizing the Use of Antibiotic Agents in the Pediatric Intensive Care Unit: A Narrative Review. Paediatric Drugs, 23(1):39–53, 2021. ISSN 1174-5878. doi: 10.1007/ s40272-020-00426-y. 1 World Health Organization. Global action plan on antimicrobial resistance. https://www. who.int/antimicrobial-resistance/global-action-plan/en/, 2015. 1 World Health Organization. Antimicrobial stewardship programmes in health-care facilities in low- and middle-income countries: A practical toolkit. https://www.who.int/ publications/i/item/9789241515488, 2019. 2 21

PICU stewardship benchmarking

Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn. Gradient surgery for multi-task learning. In Advances in Neural Information Processing Systems, volume 33, pages 5824–5836, 2020. 30 Xian Zeng, Gang Yu, Yang Lu, Linhua Tan, Xiujing Wu, Shanshan Shi, Huilong Duan, Qiang Shu, and Haomin Li. PIC, a paediatric-specific intensive care database. Scientific Data, 7(1):14, 2020. doi: 10.1038/s41597-020-0355-4. 2 Xiang Zhang, Marko Zeman, Theodoros Tsiligkaridis, and Marinka Zitnik. Graph-guided network for irregularly sampled multivariate time series. In International Conference on Learning Representations, ICLR, 2022. 6, 30

22

PICU stewardship benchmarking

Appendix A. Data Table 2: List of variables included in the patient-day dataset. Raw variable

Source Table

Patient and admission context Age at admission PATIENTS (years) Gender PATIENTS Ethnicity ADMISSIONS Admission type ADMISSIONS Insurance ADMISSIONS Current LOS hours at ADMISSIONS day start Cumulative ICU hours ADMISSIONS Physiological measurements

Aggregated features (Previous day, unless stated otherwise)

Notes

One-Hot Encoding (OHE) OHE OHE OHE Visit trajectory context icu hours cum

Visit trajectory context

Vital signs Temperature Heart Rate Respiratory Rate Oxygen Saturation Diastolic Pressure Systolic Pressure Height Weight

CHARTEVENTS CHARTEVENTS CHARTEVENTS CHARTEVENTS CHARTEVENTS CHARTEVENTS CHARTEVENTS CHARTEVENTS

sum, mean, n, min, max sum, mean, n, min, max sum, mean, n, min, max sum, mean, n, min, max sum, mean, n, min, max sum, mean, n, min, max sum, mean, n, min, max sum, mean, n, min, max

PICDB ITEMID: 1001 PICDB ITEMID: 1003 PICDB ITEMID: 1004 PICDB ITEMID: 1006 PICDB ITEMID: 1015 PICDB ITEMID: 1016 PICDB ITEMID: 1013 PICDB ITEMID: 1014

Vital signs trend features Vital trend features

CHARTEVENTS

delta prev day max, delta abx start max, var 3d

Built for TEMP, RR, HR, SBP, DBP, SPO2. Day-to-day differences, deviations from antibiotic start, and 3-day variance.

sum, mean, n sum, mean, n

LOINC: 1751-1 LOINC: 1742-6

sum, mean, n sum, mean, n sum, mean, n sum, mean, n sum, mean, n sum, mean, n sum, mean, n sum, mean, n sum, mean, n sum, mean, n, recent 90d

LOINC: 2532-0 LOINC: 718-7 LOINC: 789-8 LOINC: 2532-0 LOINC: 2947-0 LOINC: 11557-6 LOINC: 11558-4 LOINC: 11556-8 LOINC: 20564-1 LOINC: 26464-8

sum, mean, n, recent 90d

LOINC: 32264

sum, mean, n, recent 90d

LOINC: 33959-8

distinct atc ever, concurrent atc day, high priority atc ever, any oral rx last 3d

Primary antibiotic class Antibiotic exposure dynamics: current and historical classes, concurrent count, high-priority prior exposure, recent oral use.

Laboratory values Albumin LABEVENTS Alanine LABEVENTS Aminotransferase (ALT) Lactate Dehydrogenase LABEVENTS Hemoglobin LABEVENTS Red Blood Cells LABEVENTS Lactate LABEVENTS Sodium Whole Blood LABEVENTS PCO2 LABEVENTS Ph LABEVENTS PO2 LABEVENTS Oxygen Saturation LABEVENTS White Blood Count LABEVENTS (WBC) C-Reactive Protein LABEVENTS (CRP) Procalcitonin LABEVENTS Antibiotic exposure features Primary ATC code PRESCRIPTIONS Antibiotic medication PRESCRIPTIONS features

23

PICU stewardship benchmarking

A.1. Description of patient-day data Patient and admission context Static and semi-static features include age at admission, sex, ethnicity, admission type, and insurance status (See Table 2). These features capture baseline patient characteristics and care context. These are encoded using one-hot representations where applicable. We also include visit-level context such as length of stay at the start of the day and cumulative ICU time to capture disease progression and care trajectory. Physiological measurements Time-varying clinical variables include vital signs (e.g., temperature, heart rate, respiratory rate, blood pressure, and oxygen saturation) and laboratory measurements (e.g., C-reactive protein (CRP), white blood cell count (WBC), and lactate). These features capture the evolving physiological state of the patient. For the 24h aggregated data, we compute summary statistics over the pre-day window, including sum, mean, count, and where appropriate, minimum and maximum. For sparsely recorded inflammatory markers (e.g. CRP, procalcitonin), we additionally include longer-term summaries (e.g., most recent value in last 90 days) to capture baseline trends. To incorporate short-term dynamics within the tabular representation, we also compute engineered trend features for key vital signs (e.g. temperature, heart rate, respiratory rate, blood pressure, oxygen saturation), including day-to-day differences, deviations from antibiotic start, and short-window variability (e.g. 3-day (lookback) variance). Antibiotic exposure features. We include features describing antibiotic treatment patterns, including current and prior antibiotic classes, route of administration, and concurrent therapies. These features provide context for the AMS interventions. We derive features describing antibiotic usage patterns from prescription data, including current and historical antibiotic classes (ATC codes), number of concurrent antibiotics, prior exposure to highpriority antibiotics, and recent oral antibiotic use. These features aim to capture treatment dynamics relevant for stewardship decisions. To support the definition of the de-escalation outcome, we map antibiotics to their ASI scores (Gerber et al., 2017) (see Appendix Table 3). A.2. Antibiotics mapping The de-escalation target definition relies on definition of the antibiotics spectrum. For this we rely use the ASI score. Details are provided in Table A.2. Antibiotics tracked for creating patient-days dataset. Mappings to the Anatomical Therapeutic Chemical (ATC) standard were used for extraction Table 3: from the hospital information system of the private cohort. A mapping to the Antibiotic Spectrum Index (ASI) score (Gerber et al., 2017) and the corresponding product names in the PRESCRIPTIONS table of PICDB is included. ATC

Label

ASI

J02AX01 J01FA09

5-fluorocytosine acetylspiramycin

4

J01GB06

amikacin

5

PIC product name — Clarithromycin Granule for Oral Suspension; Clarithromycin Tablets Amikacin Sulfate Injection Continued on next page

24

PICU stewardship benchmarking

Table 3 — continued ATC

Label

ASI

PIC product name

J01CA04

amoxicillin

2

J01CR02 J02AA01 J01CA01

amoxicillin/clavulanic acid amphotericin b ampicillin

6

J01CR01 J01FA10

ampicillin/sulbactam azithromycin

6 4

J01DF01 J01DC04 J01DB04 J01DE01 J01DD62

aztreonam cefaclor cefazolin cefepime cefoperazone/sulbactam

5 4 3 6 6

J01DD01 J01DC05 J01DC01

5 4 4

J01DD02 J01DD04 J01DC02

cefotaxime cefotetan cefoxitin cefoxitin screen ceftazidime ceftriaxone cefuroxime

4 5 4

J01BA01 J01MA02 J01FA09

chloramphenicol ciprofloxacin clarithromycin

5 8 4

J01FF01 J01EE01

clindamycin compound sulfamethoxazole

4 4

J01AA02 J01DH03 J01FA01

doxycycline erbepenem erythromycin

5 9 2

beta-

J02AC01

extended spectrum lactamase detection fluconazole

Amoxicillin Sodium and Clavulanate Potassium for Injection; Amoxicillin and Clavulanate Potassium Granules; Amoxicillin Sodium and Sulbactam Sodium for Injection; Amoxicillin and Clavulanate Potassium Tablets — Amphotericin B Liposome for Injection Ampicillin Sodium and Sulbactam Sodium for Injection; Ampicillin Sodium for Injection — Azithromycin for Injection; Azithromycin for Suspension; Azithromycin Lactobionate for Injection; Azithromycin Tablets Aztreonam for Injection Cefaclor for Suspension; Cefaclor Capsules — Cefepime Hydrochloride For Injection Cefoperazone Sodium and Sulbactam Sodium for Injection Cefotaxime Sodium for Injection — — — Ceftazidime for Injection Ceftriaxone Sodium for Injection Cefuroxime Sodium For Injection; Cefuroxime Sodium for Injection; Cefuroxime Axetil Tablets Chloramphenicol Injection Ciprofloxacin and Sodium Chloride Injection Clarithromycin Granule for Oral Suspension; Clarithromycin Tablets Clindamycin Phosphate for Injection Compound Sulfamethoxazole Tablets; Compound Sulfamethoxazole Injection — Ertapenem for Injection Erythromycin Lactobionate for Injection; Erythromycin Ointment; Erythromycin Eye Ointment —

J01MA16 J01GB03

gatifloxacin gentamicin

9 5

J01GB03

gentamicin 500

5

J01DH51

imipenem induced clindamycin tance itraconazole josamycin koalaranin levofloxacin

resis-

Fluconazole Capsules; Fluconazole and Sodium Chloride Injection; Fluconazole Injection — Gentamycin Sulfate and Procaine Hydrochloride Capsules; Gentamycin Sulfate,Procaine Hydrochloride and Vitamin B Gentamycin Sulfate and Procaine Hydrochloride Capsules; Gentamycin Sulfate,Procaine Hydrochloride and Vitamin B Imipenem and Cilastatin Sodium for Injection —

J02AC02 J01FA07 J01MA12

2

10

— — — Levofloxacin Eye Drops; Levofloxacin and Sodium Chloride Injection; Levofloxacin Hydrochloride Ear Drops

4 9

Continued on next page

25

PICU stewardship benchmarking

Table 3 — continued ATC

Label

ASI

PIC product name

J01XX08 J01DH02

5 10

Linezolid Injection; Linezolid Tablets Meropenem for Injection —

J01AA08 J01MA14 J01XE01 J01MA01 J01CF04 J01CE01

linezolid meropenem methicillin-resistant staphylococcus minocycline moxifloxacin nitrofurantoin ofloxacin oxacillin penicillin

5 9 2 9 1 2

J01CA12

piperacillin

5

J01CR05 J01XB02 J01FG02 J04AB02 J01FA06 J01MA09 J01GA01 J01FA15 J01AA07 J01CA13 J01AA12 J01GB01

piperacillin/tazobactam polymyxin b quinupristin/dalfopristin rifampin roxithromycin sparfloxacin streptomycin 2000 telithromycin tetracycline ticarcillin tigecycline tobramycin

8 4 5 3 4 8 5 4 5 5 13 5

J01XA01

vancomycin

5

J02AC03

voriconazole

— — Nitrofurantoin Enteric-coated Tablets Ofloxacin Eye Ointment; Ofloxacin Gel Oxacillin Sodium for Iniection Benzylpenicillin Sodium for Injection; Benzathine Benzylpenicillin for Injection Piperacillin Sodium and Tazobactam Sodium Injection; Piperacillin Sodium and Tazobactam Sodium for Injection; Piperacillin Sodium And Sulbactam Sodium For Injection — Polymyxin B Sulfate for injection — Rifampicin Capsules; Rifampicin for Eye Use — — — — — — Tigecycline for Injection Tobramycin and Dexamethasone Ophthalmic Ointment; Tobramycin Sulfate Injection; Tobramycin Dexamethasone Eye Drops; Tobramycin Eye Drops Vancomycin Hydrochloride for Intra Venous; Norvancomycin Hydrochloride for Injection Voriconazole for Injection; Voriconazole Tablets —

beta-lactamase

A.3. Construction of patient-day data The datasets are organized at the patient-day level: each row is one calendar antibiotic day on the PICU. Antibiotic prescriptions are gap-merged (if within 24h) into clinically contiguous courses for each admission and antibiotic. Patient-day rows are enumerated only from those merged intervals. Multi-day antibiotic coverage therefore yields multiple rows; days without qualifying overlap are excluded. Merged ’same-day’ antibiotic courses shorter than 24 hours are dropped as these are mostly prophylactic and the duration is too short for PAF. If several courses overlap the same day, all are considered when attaching flags (or across overlapping courses) rather than collapsing to a single agent. Prediction is defined at the start of each patient-day (DAY START), default 8:00AM, and all features are constructed using only information available up to this time point to prevent information leakage (DAY START-W,DAY START), with a default 24h aggregation window. A.3.1. 24h resolution Time-varying clinical variables, including vital signs, laboratory measurements, and treatment features, are aggregated within a fixed temporal windows anchored to the start of each 26

PICU stewardship benchmarking

patient-day. We compute clinically interpretable summary statistics (e.g., mean, extrema, and most recent value) to represent the current patient state. Treatment-related features capture antibiotic regimen characteristics such as route, number of concurrent agents, and days since antibiotic initiation. The final feature representation combines these components using one-hot encoding and engineered summary features, resulting in a total of 194 features (Figure 2 above). A.3.2. 1 hour bins We augment each patient-day timestep with a fixed-resolution grid of pre-anchor chart and laboratory measurements. For day index d with anchor time DAY START(d), numeric vitals and labs with timestamps in [DAY START(d) - W, DAY START(d))—and not at or after the anchor—are assigned to B contiguous sub-intervals of equal duration within that window (default W = 24 h). Within each item vocabulary column, values are forward-filled across bins so later bins carry the most recent observation carried forward, yielding a dense matrix of shape (B × S) with S timeseries variables, as a simple multivariate representation of irregularly sampled vitals and labs on a common time base. Encoding the temporal matrix This grid is encoded by a small EventGridEncoder module comprising a linear projection per bin followed by one-dimensional convolutions with global pooling over the bin axis, yielding a fixed-size embedding vector event embed dim per patient-day timestep. The architecture reflects a deliberate hierarchical decomposition of clinical time: the convolutional encoder captures intra-day physiological dynamics within the W -hour pre-anchor window, while the downstream sequence backbone—GRU, LSTM, TCN, Transformer, or SSM—models inter-day clinical trajectory across the admission. This separation avoids the sequence-length explosion that would result from unrolling bins directly into the admission timeline (expanding sequence length by a factor of B), and keeps the per-day representation backbone-agnostic. Fusion of 1h embeddings to aggregated data for sequential models The resulting event embed dim-dimensional day vector is concatenated with the scaled tabular feature vector for that day to form a fused representation of dimension tab dim + event embed dim (e.g. 300 + 64 = 364), which is then passed through a LayerNorm before entering the sequence trunk and task heads. Where applicable, pre-anchor tabular columns (Table 2 in Appendix) whose information overlaps the longitudinal window were omitted from the tabular branch so coarse window-level summaries are not duplicated. The EventGridEncoder is not pre-trained or trained as a standalone classifier; it is optimised end-to-end within the full prediction network comprising the tabular pathway, event encoder, sequence trunk, and prediction head. The training objective is masked timestep binary cross-entropy with logits over valid sequence positions, with per-task positive-class reweighting (pos weight), minimised with AdamW (Kingma and Ba, 2017). Gradients backpropagate through the prediction head, sequence trunk, fusion concatenation, and the Conv1d layers of the EventGridEncoder, so the convolutional filters learn whatever within-day temporal structure improves the final supervised prediction objective, subject to gradient clipping (clip grad norm = 2.0) and early stopping on held-out validation loss.

27

PICU stewardship benchmarking

Fusion of 1h embeddings and aggregated data for tabular models For the tabular model [1h] benchmark, we adopt a gated recurrent unit (GRU) as the default sequence backbone to encode the 1h temporal representations. Specifically, the concatenated vector of dimension tab dim+event embed dim at each timestep is passed through a unidirectional GRU, producing a sequence of hidden states that summarise the evolving clinical trajectory across the admission. The GRU is chosen as a strong and computationally efficient baseline with inductive bias toward temporal dependencies, providing a consistent reference point across experiments. The intra-day convolutional encoder described above captures shortrange physiological dynamics within each W -hour window, while the GRU aggregates these embeddings over time into a longitudinal patient representation. All tabular benchmark results are reported using this CNN + GRU configuration, with the embeddings exported from the CNN encoder and appended to the tabular aggregated feature representation. With this setup we can give the tabular models a fair temporal dimension, and understand if the difference between tabular and sequential models is not just because of access to more data. A.3.3. Sub-day irregular sequence representations Finally, graph-based temporal models represent multivariate clinical time series as a graph over variables rather than as an explicit sequence, and therefore do not maintain a conventional notion of sequence “memory.” Instead, they capture temporal dynamics alongside inter-variable relationships through structured message passing, making them particularly well suited to irregularly sampled data and providing a complementary inductive bias to sequence-based approaches. The Raindrop framework closely aligns with the longitudinal fusion setup in terms of input construction: both operate on the same leak-safe pre-anchor window (DAY STARTW,DAY START) and transform irregular chart and laboratory events into binned sensor trajectories. The key distinction lies in where temporal modeling is performed. Raindrop treats each patient-day as an independent sample, learning temporal structure within that day (approximately ∼1h ’bins’) via graph-based observation propagation combined with causal Transformer-style pooling, before fusing the resulting representation with static tabular features. This representation therefore preserves the native, irregular sampling of EHR data without discretization and allows us to gauge the information loss from binning. A.4. Implementation details A.4.1. Data preprocessing All numerical data is Winsorized (Wilcox, 2011) during cleaning to reduce the effect of spurious outliers. Data is further standardized using sklearn’s (Pedregosa et al., 2011) StandardScaler unless stated otherwise. For sequence models, rows are grouped by admission id, ordered by timestamp of day start, truncated to at most 512 timesteps, and trained with masked per-timestep logistic losses so padded days do not contribute to optimization.

28

PICU stewardship benchmarking

A.4.2. Tabular models Tabular models operate on a flattened patient-day representation, where each instance corresponds to a single day of a patient’s stay. These models do not explicitly model temporal dependencies across timesteps. Instead, temporal information is incorporated indirectly through engineered features, such as length-of-stay variables and short-term trend summaries. As a result, each prediction is made independently based on the current patient state and aggregated context of the previous day, without access to the full temporal trajectory. We do not specifically handle class imbalance for the tabular models. Logistic regression Logistic regression uses L2 regularization with standard solver lbfgs and max iter=3000 and default inverse of regularization strength C=1.0) LightGBM LightGBM uses 500 trees with learning rate=0.05 and num leaves=31 (Ke et al., 2017). We don’t standardize nor scale the features for this model. MLP classifier The Sci-kit learn MLPClassifier (Pedregosa et al., 2011) was used with hidden layer size=(64,0), ReLu activation function, batch size of 256, initial learning rate of 0.001, and 200 max iteration with early stopping. TabPFN TabPFN-v2.6 weights were downloaded locally from the official GitHub https: //github.com/PriorLabs/TabPFN (Hollmann et al., 2025). Feature values were converted to floating point and non-finite values replaced with zeros before inference. No standardization or scaling is applied to the input data. For the Private cohort we trained/tested on the full sets (14,573/3,573 patient-days), for PIC we trained on the full training set (64,051), and tested on a stratified subset of 500 patient because of compute limitations. A.4.3. Sequence-based temporal models Sequence-based temporal models train one network per target. In all sequence-based temporal models runs, class imbalance is handled with a scalar pos weight for the target-specific loss, and the readout is one logit per patient-day rather than a pooled admission-level prediction. Recurrent architectures The recurrent models are Gated Recurrent Unit (GRU) (Cho et al., 2014) and Long Short-Term Memory (LSTM) (Hochreiter and Schmidhuber, 1997) trunks with two stacked layers, hidden size 128, dropout 0.2 between recurrent layers, and a linear per-timestep head (unidirectional recurrent trunk). Temporal Convolutional Network The Temporal Convolutional Network (TCN) (Bai et al., 2018) is implemented as a residual stack of left-padded causal dilated temporal convolutions over the admission trajectory. The causal structure ensures that predictions at timestep t depend only on observations up to and including t. By using convolutions, the TCN is able to capture longer-range temporal dependencies while maintaining a fixed computational cost per timestep. The output at each timestep is passed to a linear prediction head. Transformer We use a Transformer encoder (Vaswani et al., 2017) with learned positional embeddings to model temporal dependencies. The architecture employs multi-head selfattention with GELU activations and a norm first configuration. To handle variable-length 29

PICU stewardship benchmarking

sequences, key-padding masks are applied to ignore padded timesteps. A causal attention mask is used to ensure that predictions at timestep t depend only on observations up to and including t, preventing information leakage from future timesteps. State Space Model The state space model is a lightweight Mamba-inspired selective SSM (Gu and Dao, 2024), implemented as a residual stack of gated, input-conditioned causal state-update blocks over the admission trajectory. Four input-dependent projections are computed from the hidden representation x ∈ RB×T ×H at each timestep t. This gating mechanism allows the model to interpolate between propagating the updated latent state and passing the input through directly. Residual connections with dropout and a final layer normalisation are applied across the block stack, and padding safety is enforced by masking state updates and zeroing outputs beyond each sequence’s valid length. The causal structure ensures that predictions at timestep t depend only on observations up to and including t. The output at each timestep is passed to a linear prediction head. This implementation is intentionally lightweight: it does not require a CUDA scan kernel and omits the convolutional mixer branch and full SSD parameterisation of the canonical Mamba release (Gu and Dao, 2024), prioritising ease of integration and fair comparison to the other simple sequence-based temporal architectures. A.4.4. Multi-task learning setting In the MTL setting, all four targets are predicted jointly using a shared encoder, with taskspecific output heads. The models use the same per-timestep inputs as in the single task learning setting, i.e., tabular patient-day features alone, or fused with learnable features when applicable. We consider several MTL strategies. In hard parameter sharing, a single encoder is shared across tasks with separate linear heads. In uncertainty-weighted MTL (Kendall et al., 2018), task losses are combined using learned homoscedastic uncertainty parameters. We also evaluate a multi-gate mixture-of-experts (MMoE) architecture (Ma et al., 2018) with PCGrad (Yu et al., 2020, projecting conflicting gradients), where the shared sequence representation is passed through a small set of expert networks with taskspecific gating, and gradient updates are adjusted to reduce destructive interference between tasks. All MTL variants can be paired with GRU (Cho et al., 2014), LSTM (Hochreiter and Schmidhuber, 1997), TCN (Bai et al., 2018), Transformer (Vaswani et al., 2017), or SSM sequence encoders. A.4.5. Graph-based temporal models RAINDROP We use RAINDROP (Zhang et al., 2022), a graph-based model designed for irregularly sampled multivariate time series. Similar to the sequence-based temporal models, the model operates on pre-anchor chart and laboratory measurements, but without fixed binning, thus preserving the original temporal structure. Input observations are represented as a feature-wise time series, where each variable corresponds to a node in a sensor graph. Temporal dependencies are modeled using attention mechanisms with a causal mask, ensuring that predictions at timestep t depend only on past observations. We use a standard RAINDROP configuration with 2 layers, 2 attention heads and a dropout of 0.2, along with 2 observation propagation layers over the sensor graph. Temporal representations are aggregated using mean pooling and concatenated with tabular patient context features. 30

PICU stewardship benchmarking

Table 4: Cross-cohort summary for PIC and the Private patient-day datasets. Test-set prevalences and antibiotics course duration metrics split by outcomes are also included. Quantity

PIC

Private

Patient-days (all) Patient-days (train) Patient-days (test) Distinct patients (all) Distinct admissions (all)

79,831 64,051 15,780 5,515 5,671

18,148 14,573 3,575 2,059 2,789

Test-set prevalence IV-to-oral de-escalation discontinuation short-course

0.0035 0.0392 0.0169 0.0577

0.0098 0.0520 0.0568 0.1860

Test-set positive-course duration, days (median [Q1, Q3]) IV-to-oral 6.62 [5.62, 9.25] 3.19 [2.67, 5.07] de-escalation 6.62 [3.68, 10.02] 1.99 [1.00, 3.30] discontinuation 6.76 [3.98, 9.71] 2.00 [1.33, 4.33] short-course 2.60 [1.81, 3.07] 1.88 [1.01, 2.50]

The resulting representation is passed to a classification head mlp static to produce binary predictions.

Appendix B. Targets For the outcomes we first derive course-level stewardship events from merged antibiotic prescription courses, then attach them to the patient-day unit. IV-to-oral An IV-to-oral event is defined when a course contains an earlier intravenous prescription followed by a non-intravenous formulation. The timestamp of the first qualifying non-intravenous prescription is recorded as the event time. The corresponding patientday is labeled positive only on the calendar day of the switch. Clinically, this target is best interpreted as a proxy for route-change review opportunity rather than a recommendation that a switch was appropriate. De-escalation A de-escalation event is defined when a subsequent antibiotic prescription within the same admission has a different Anatomic Therapeutic Classification (ATC) (WHO, 2024) code and a strictly lower antibiotic spectrum index (ASI) (Gerber et al., 2017) than the current course (Table 3). Only the patient-day corresponding to the first qualifying de-escalation event is labeled positive. This target captures a prescription-based spectrum-narrowing proxy rather than adjudicated stewardship intent.

31

PICU stewardship benchmarking

Figure 7: Intersections of AMS intervention targets across cohorts. Discontinuation A discontinuation event is defined when an antibiotic course ends before discharge, with at least 72 hours remaining in the admission and no overlapping antibiotic prescriptions in the subsequent interval. The label is assigned to the final day of the course. The label therefore depends on an operational drug-free-window threshold rather than direct clinician annotation. Short-course A short-course event is defined when the duration of an antibiotic course is less than 96 hours. Note that since we exclude stays ¡24h, this means that a short-course antibiotic target is actually a course with a length between 24 and 96 hours. Additionally, Cefazolin (ATC: J01DB04) is excluded from this definition due to its common prophylactic use in PICU. As with discontinuation, only the final day of the course is labeled as positive. This is an auxiliary duration-based target and should likewise be viewed as a stewardshiprelevant proxy rather than a gold-standard decision label. B.1. Datasets B.1.1. PIC PIC version 1.1.0 was retrieved from Physionet (Goldberger et al., 2000 (June 13). B.2. Data preprocessing Missing data are encoded using indicator variables in the sequence representations. Numeric variables are clipped at predefined percentile thresholds to reduce the impact of recording artifacts while preserving clinically plausible ranges. Undefined aggregates (e.g., no measurements within a window) are assigned neutral default values (e.g., zero counts), consistent with feature design. All preprocessing steps are fitted on the training cohort and applied unchanged to the test cohorts to prevent information leakage. Additional implementation details are provided in Appendix A.4.

32

PICU stewardship benchmarking

IV-to-oral

De-escalation

1.0

0.8

Models 0.6

Score

Tabular [24h] Logistic regression LightGBM MLP

0.4

TabPFN Tabular [1h]

0.2

Logistic regression LightGBM MLP

0.0 AUROC

F1

AUPRC

AUROC

F1

AUPRC

Sequential [24h] GRU LSTM

Discontinuation

Short-course

TCN Transformer

1.0

SSM Sequential [1h]

0.8

GRU LSTM TCN

0.6

Score

Transformer SSM Graph-based [~1h]

0.4

RAINDROP

0.2

0.0 AUROC

F1

AUPRC

AUROC

F1

AUPRC

Figure 8: Comparison of models using augmented 1-hour representations on the Private cohort. Tabular and sequence models are enhanced with learned embeddings derived from hourly data, and compared to a graph-based model (RAINDROP) operating on irregular time series.

Appendix C. Additional Results C.1. Detailed results per target C.2. Calibration plots Private C.3. MTL modeling detailed results C.4. Comparing STL and MTL models on Private

33

PICU stewardship benchmarking

Table 5: Full results for target Short-course Short-course Private Model

PIC

AUROC

F1

AUPRC

AUROC

F1

AUPRC

0.719 0.761 0.733 0.779

0.092 0.243 0.126 0.211

0.324 0.403 0.339 0.486

0.743 0.819 0.758 0.829

0.035 0.208 0.048 0.328

0.169 0.335 0.204 0.473

0.719 0.761 0.735

0.092 0.243 0.155

0.324 0.403 0.345

0.741 0.819 0.767

0.039 0.193 0.081

0.167 0.337 0.219

0.793 0.773 0.787 0.760 0.720

0.462 0.495 0.517 0.487 0.462

0.437 0.469 0.476 0.444 0.400

0.857 0.835 0.848 0.821 0.763

0.275 0.240 0.274 0.250 0.211

0.389 0.352 0.348 0.288 0.198

0.793 0.789 0.791 0.742 0.743

0.477 0.467 0.474 0.437 0.435

0.431 0.424 0.419 0.363 0.356

0.865 0.863 0.845 0.821 0.806

0.261 0.289 0.268 0.257 0.246

0.404 0.395 0.351 0.318 0.280

0.704

0.450

0.349

0.750

0.242

0.186

Tabular [24h] Logistic regression LightGBM MLP TabPFN Tabular [1h] Logistic regression LightGBM baseline MLP Sequential [24h] GRU LSTM TCN Transformer SSM Sequential [1h] GRU LSTM TCN Transformer SSM Graph-based [∼1h] RAINDROP

34

PICU stewardship benchmarking

Table 6: Full results for target IV-to-oral IV-to-oral Private Model

PIC

AUROC

F1

AUPRC

AUROC

F1

AUPRC

0.865 0.841 0.566 0.875

0.000 0.000 0.000 0.000

0.049 0.072 0.013 0.089

0.906 0.906 0.898 0.990

0.000 0.257 0.000 0.000

0.100 0.223 0.144 0.236

0.865 0.841 0.605

0.000 0.000 0.000

0.049 0.072 0.014

0.898 0.894 0.906

0.000 0.267 0.000

0.106 0.273 0.098

0.791 0.792 0.846 0.802 0.840

0.076 0.055 0.047 0.049 0.053

0.073 0.071 0.043 0.037 0.070

0.915 0.905 0.902 0.932 0.935

0.088 0.090 0.135 0.164 0.077

0.102 0.107 0.100 0.143 0.100

0.857 0.821 0.881 0.838 0.832

0.072 0.061 0.097 0.101 0.063

0.064 0.061 0.074 0.046 0.050

0.926 0.916 0.931 0.926 0.944

0.067 0.075 0.068 0.099 0.083

0.167 0.198 0.127 0.153 0.116

0.869

0.000

0.078

0.873

0.000

0.133

Tabular [24h] Logistic regression LightGBM MLP TabPFN Tabular [1h] Logistic regression LightGBM baseline MLP Sequential [24] GRU LSTM TCN Transformer SSM Sequential [1h] GRU LSTM TCN Transformer SSM Graph-based [∼1h] RAINDROP

35

PICU stewardship benchmarking

Table 7: Full results for target De-escalation De-escalation Private Model

PIC

AUROC

F1

AUPRC

AUROC

F1

AUPRC

0.901 0.936 0.930 0.955

0.124 0.292 0.245 0.393

0.278 0.429 0.364 0.529

0.907 0.939 0.908 0.934

0.140 0.256 0.114 0.333

0.262 0.390 0.270 0.427

0.901 0.936 0.925

0.124 0.292 0.226

0.278 0.429 0.334

0.906 0.936 0.922

0.145 0.229 0.155

0.267 0.382 0.295

0.920 0.955 0.966 0.923 0.962

0.466 0.542 0.564 0.399 0.512

0.453 0.605 0.602 0.360 0.611

0.966 0.964 0.959 0.950 0.955

0.451 0.438 0.372 0.338 0.360

0.533 0.542 0.491 0.405 0.446

0.941 0.937 0.949 0.939 0.949

0.445 0.467 0.489 0.398 0.413

0.475 0.504 0.509 0.368 0.438

0.966 0.966 0.961 0.947 0.964

0.436 0.425 0.391 0.342 0.378

0.562 0.549 0.512 0.440 0.536

0.460

0.329

0.894

0.289

0.229

Tabular [24h] Logistic regression LightGBM MLP TabPFN Tabular [1h] Logistic regression LightGBM MLP Sequential [24h] GRU LSTM TCN Transformer SSM Sequential [∼1h] GRU LSTM TCN Transformer SSM

Irregular Time Series Models RAINDROP

0.925

36

PICU stewardship benchmarking

Table 8: Full results for target Discontinuation Discontinuation Private Model

PIC

AUROC

F1

AUPRC

AUROC

F1

AUPRC

0.679 0.718 0.643 0.741

0.000 0.010 0.000 0.000

0.109 0.153 0.085 0.218

0.797 0.832 0.765 0.851

0.000 0.014 0.000 0.222

0.074 0.106 0.061 0.219

0.679 0.718 0.631

0.000 0.010 0.000

0.109 0.153 0.085

0.800 0.827 0.789

0.000 0.021 0.000

0.076 0.096 0.064

0.712 0.732 0.741 0.720 0.713

0.168 0.236 0.245 0.228 0.249

0.137 0.213 0.204 0.174 0.172

0.843 0.836 0.815 0.838 0.815

0.099 0.108 0.095 0.100 0.095

0.092 0.111 0.085 0.093 0.082

0.726 0.713 0.724 0.675 0.679

0.193 0.178 0.189 0.139 0.156

0.129 0.126 0.133 0.110 0.109

0.839 0.828 0.830 0.821 0.820

0.089 0.091 0.092 0.090 0.093

0.085 0.094 0.081 0.080 0.089

0.226

0.169

0.745

0.000

0.058

Tabular [24h] Logistic regression LightGBM baseline MLP TabPFN Tabular [1h] Logistic regression LightGBM MLP Sequential [24h] GRU LSTM TCN Transformer STL (SSM) Sequential [∼1h] GRU LSTM TCN Transformer SSM

Irregular Time Series Models RAINDROP

0.696

37

PICU stewardship benchmarking

1.0

IV-to-oral

1.0

Test prevalence=0.0098

0.8 Observed event rate

Observed event rate

0.8 0.6 0.4 0.2 0.0 0.0 1.0

0.2

0.4 0.6 Mean predicted probability

0.8

0.4

Tabular [1h] Logistic regression LightGBM MLP

0.0 0.0

1.0

Discontinuation

1.0

Test prevalence=0.0568

0.2

0.4 0.6 Mean predicted probability

0.8

Sequential [24h] GRU LSTM TCN Transformer SSM

1.0

Short-course

Test prevalence=0.1860

0.6

Sequential [1h] GRU LSTM TCN Transformer SSM

0.4

Graph-based [~1h] RAINDROP

0.8 Observed event rate

Observed event rate

Models

Perfect calibration Tabular [24h] Logistic regression LightGBM MLP

0.6

0.2

0.8 0.6 0.4 0.2 0.0 0.0

De-escalation Test prevalence=0.0520

0.2

0.2

0.4 0.6 Mean predicted probability

0.8

1.0

0.0 0.0

0.2

0.4 0.6 Mean predicted probability

0.8

1.0

Figure 9: Calibration plots for the different models on Private across four AMS targets.

Table 9: Comparing MTL models to tabular baselines for Short-course. SC Private Architecture

Approach

Sequential [24h] GRU hard-sharing GRU uncertainty-weighted GRU MMoE+PCGrad LSTM hard-sharing LSTM uncertainty-weighted LSTM MMoE+PCGrad TCN hard-sharing TCN uncertainty-weighted TCN MMoE+PCGrad Transformer hard-sharing Transformer uncertainty-weighted Transformer MMoE+PCGrad SSM hard-sharing SSM uncertainty-weighted SSM MMoE+PCGrad

PIC

AUROC

F1

AUPRC

AUROC

F1

AUPRC

0.774 0.775 0.779 0.771 0.767 0.780 0.781 0.776 0.768 0.776 0.763 0.762 0.737 0.739 0.724

0.501 0.504 0.501 0.497 0.482 0.511 0.519 0.516 0.497 0.496 0.494 0.491 0.481 0.482 0.465

0.457 0.459 0.474 0.454 0.458 0.458 0.460 0.450 0.459 0.451 0.456 0.446 0.417 0.421 0.393

0.866 0.865 0.867 0.845 0.845 0.864 0.837 0.848 0.855 0.813 0.827 0.810 0.801 0.790 0.788

0.287 0.285 0.280 0.274 0.271 0.260 0.264 0.267 0.254 0.229 0.223 0.227 0.235 0.227 0.232

0.423 0.422 0.427 0.367 0.366 0.402 0.322 0.343 0.360 0.263 0.305 0.253 0.282 0.248 0.249

38

PICU stewardship benchmarking

Table 10: Comparing MTL models to tabular baselines for De-escalation. De-escalation Private Architecture

Approach

Sequential [24h] GRU hard-sharing GRU uncertainty-weighted GRU MMoE+PCGrad LSTM hard-sharing LSTM uncertainty-weighted LSTM MMoE+PCGrad TCN hard-sharing TCN uncertainty-weighted TCN MMoE+PCGrad Transformer hard-sharing Transformer uncertainty-weighted Transformer MMoE+PCGrad SSM hard-sharing SSM uncertainty-weighted SSM MMoE+PCGrad

PIC

AUROC

F1

AUPRC

AUROC

F1

AUPRC

0.955 0.956 0.962 0.951 0.949 0.958 0.962 0.959 0.960 0.959 0.955 0.956 0.957 0.958 0.953

0.464 0.469 0.458 0.447 0.423 0.456 0.499 0.511 0.482 0.465 0.463 0.423 0.469 0.477 0.465

0.520 0.525 0.595 0.540 0.516 0.558 0.525 0.510 0.515 0.548 0.507 0.545 0.511 0.536 0.493

0.964 0.965 0.965 0.961 0.962 0.964 0.961 0.962 0.964 0.949 0.951 0.948 0.954 0.951 0.943

0.433 0.447 0.443 0.404 0.403 0.410 0.395 0.434 0.424 0.331 0.351 0.349 0.352 0.328 0.304

0.549 0.550 0.560 0.511 0.511 0.537 0.512 0.536 0.538 0.397 0.432 0.389 0.443 0.416 0.391

Table 11: Comparing MTL models to tabular baselines for Discontinuation. Discontinuation Private Architecture

Approach

Sequential [24h] GRU hard-sharing GRU uncertainty-weighted GRU MMoE+PCGrad LSTM hard-sharing LSTM uncertainty-weighted LSTM MMoE+PCGrad TCN hard-sharing TCN uncertainty-weighted TCN MMoE+PCGrad Transformer hard-sharing Transformer uncertainty-weighted Transformer MMoE+PCGrad SSM hard-sharing SSM uncertainty-weighted SSM MMoE+PCGrad

PIC

AUROC

F1

AUPRC

AUROC

F1

AUPRC

0.742 0.742 0.743 0.743 0.730 0.743 0.731 0.743 0.740 0.747 0.732 0.722 0.713 0.717 0.699

0.234 0.233 0.229 0.248 0.226 0.243 0.238 0.248 0.226 0.239 0.242 0.226 0.238 0.246 0.217

0.234 0.237 0.212 0.199 0.187 0.180 0.228 0.209 0.188 0.208 0.201 0.208 0.173 0.177 0.164

0.855 0.855 0.860 0.843 0.844 0.854 0.828 0.841 0.849 0.828 0.844 0.822 0.830 0.815 0.824

0.099 0.102 0.102 0.106 0.105 0.105 0.091 0.096 0.092 0.096 0.086 0.083 0.101 0.094 0.087

0.106 0.104 0.115 0.100 0.100 0.121 0.075 0.095 0.101 0.092 0.111 0.081 0.092 0.086 0.093

39

PICU stewardship benchmarking

De-escalation

Discontinuation

Short-course

1.0 0.96

0.95 0.95

0.94

0.97 0.96

0.96

0.96 0.96

0.92

AUROC

0.9

0.8

0.79 0.74

0.73

0.73

0.74

0.74

0.78

0.77 0.77

0.75 0.73

0.79 0.78 0.76

0.78 0.74

0.72

0.72

0.71 0.71

0.7

0.6

0.5 U

GR

TM

LS

N er TC orm nsf Tra

M

SS

U

GR

TM

LS

N er TC orm nsf Tra

SS

M

U

GR

TM

LS

N er TC orm nsf Tra

SS

M

Backbone architecture Single-task learning (STL)

Multi-task learning (MTL)

Figure 10: Comparing AUROC scores across STL and MTL (hard-sharing) architectures on the Private cohort.

40

Record · ID 216850 · SHA-256 dc286e4d685b8d39
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.