Conceptio › Archive › arXiv CS
arXiv CSopen access

ECG Mirage: Revealing and Mitigating the Underutilisation of ECGs in Vision-Language Models for Clinical Prediction

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

ECG MIRAGE: REVEALING AND MITIGATING THE UNDERUTILISATION OF ECGS IN VISION-LANGUAGE MODELS FOR CLINICAL PREDICTION Jinning Liang

Mingcheng Zhu

Tingting Zhu

University of Oxford, Oxford, United Kingdom [email protected]

arXiv:2609.21755v1 [cs.AI] 18 Sep 2026

ABSTRACT Emergency department (ED) decision-making relies on heterogeneous clinical information, including patient history, vital signs, laboratory results, and electrocardiograms (ECGs). Vision– language models (VLMs) can jointly process these modalities, but strong predictive performance does not necessarily imply meaningful use of the correct patient’s ECG. We term this failure mode ECG Mirage: apparent multimodal capability without useful dependence on patient-specific ECG information. We distinguish two forms: ECG neglect, where ECGs provide little predictive benefit, and ECG confusion, where matched ECGs outperform no-image inputs but not mismatched ECGs. To evaluate these behaviours, we compare predictions obtained with matched ECGs, outcome-discordant mismatched ECGs, and no-image inputs while holding the clinical text and prediction targets fixed. Across four VLMs on MDS-ED, matched ECGs provide no consistent advantage for either ICU admission or clinical deterioration prediction. We then train four restricted visual prompts using supervised learning followed by conditional direct preference optimisation, while keeping the VLM backbone frozen. The resulting models achieve balanced accuracies of 70.6% for ICU admission and 67.5% for deterioration and increase the matched-versusmismatched performance gap to approximately 16.5 and 5.5 percentage points, respectively. Overall, our study identifies ECG Mirage in multimodal clinical prediction and introduces visual prompt tuning as an efficient mitigation strategy. Code is available at https://github.com/JasonZuu/ECG-Mirage. Index Terms— Electrocardiography, vision-language models, multimodal mirage, efficient optimisation, AI for healthcare. 1. INTRODUCTION Predicting clinical outcomes in the emergency department (ED) presents an opportunity for multimodal signal processing from patient history, vital signs, laboratory measurements and electrocardiograms (ECGs) [1, 2]. Vision–language models (VLMs) offer a flexible approach to combining these sources, processing textual electronic health records (EHRs) alongside rendered ECG images [3, 4]. Related EHR-based work has investigated efficient prompt compression for clinical prediction [5] and adaptation across clinical conditions [6]. Recent work has demonstrated progress in ECG interpretation and in linking diagnostic predictions to waveform evidence [7, 8]. These capabilities motivate the use of VLMs for outcome prediction, where ECGs must be interpreted alongside the broader clinical context. However, processing both modalities does not necessarily imply that both contribute to the prediction. This leads to a central question: do VLMs benefit from patientspecific ECG information beyond the EHR context?

Case 1: EHR + Correct ECG EHR (same for all cases)

ECG (correct patient)

Age: 62 Sex: Male Chief complaint: Chest pain History: Hypertension, Meds: Aspirin, Metformin ...

Case 2: EHR + Mismatched Patient’s ECG EHR (same for all cases)

Frozen VLM

ECG (mismatched patient)

Age: 62 Sex: Male Chief complaint: Chest pain History: Hypertension, Meds: Aspirin, Metformin ...

Case 3: EHR + No ECG

ECG information is ignored.

ECG (missing)

EHR (same for all cases)

Same prediction

Age: 62 Sex: Male Chief complaint: Chest pain History: Hypertension, Meds: Aspirin, Metformin ...

A:A: ECG ECGMirage Mirage

Objective:Objective: maximise maximise p(y = y | x, I, q) Visible Visible Trainable VPT Tokens

Trainable VPT Tokens

Visible Visible ECG Image Tokens

ECG Image Tokens Mask

⋯...

Textual Tokens

Textual Tokens

...

⋯

Mask Frozen VLM

Loss: L task

y Frozen VLM

B: Visual Prompt Tuning

Fig. 1. Evaluation and training share a controlled ECG substitution. The reference answer remains fixed in conditional preference training. The no-image input contains only EHR text and the task Loss: B: Visual Prompt Tuning question and is used only for evaluation.

Prior work shows that VLM benchmark performance can obscure limited visual dependence, including strong reliance on accompanying text in medical image prediction [9, 10]. For clinical outcome prediction, an improvement from adding an ECG does not necessarily indicate that the model benefits from patient-specific information. A model could perform similarly when given another patient’s ECG, despite outperforming a no-image control. We refer to this discrepancy as ECG Mirage, where apparent multimodal predictive capability is not accompanied by reliable benefit from the ECG information. ECG Mirage operationalises this broader multimodal failure mode for clinical outcome prediction through patient-level ECG substitution controls. To examine this phenomenon, we compare matched ECGs, outcome-discordant mismatched ECGs and a no-image control while keeping the EHR, prediction question and target unchanged. Across the evaluated zero-shot VLMs and clin-

ical tasks, matched ECGs offer no consistent predictive advantage over the controls. These findings motivate training that encourages models to distinguish matched from mismatched ECG evidence. To mitigate ECG Mirage, we propose visual prompt tuning (VPT), a lightweight training approach that combines prompt tuning with restricted visual prompts and conditional preference optimisation. The approach encourages VLMs to use patient-matched ECG information when predicting clinical outcomes from multimodal clinical data. Specifically, we train a small set of prompts while keeping the VLM backbone frozen. An attention mask prevents textual tokens from attending directly to these prompts, allowing them to influence predictions through only image tokens. We first train these prompts to predict clinical outcomes from matched EHR–ECG pairs. We then use conditional direct preference optimisation [11,12] to encourage the model to favour the correct outcome when given the patient’s own ECG over a mismatched ECG from another patient. Experiments on ICU admission and clinical deterioration prediction show that our approach mitigates ECG Mirage, improving matchedECG balanced accuracy and widening the performance gap between matched- and mismatched-ECG cases. Our contributions are threefold. (1) We identify and characterise ECG Mirage across four zero-shot VLMs, showing that multimodal prediction performance does not necessarily reflect benefit from patient-matched ECGs. (2) We propose a lightweight approach that combines restricted visual prompt tuning with supervised learning and conditional preference optimisation to mitigate ECG Mirage. (3) Through experiments on two clinical prediction tasks, we demonstrate improvements in matched-ECG balanced accuracy and greater separation from mismatched ECGs and no-image inputs. 2. METHODOLOGY 2.1. Clinical Prediction with VLMs We consider clinical outcome prediction from electronic health records (EHRs) and electrocardiogram (ECG) images. For each encounter, let x denote the EHR input available at prediction time, I the matched ECG image, and q the task question. The target answer y encodes either the binary ICU-admission outcome or the six clinical deterioration labels. We represent this answer as a token sequence y = (y1 , . . . , yT ). Given the EHR input x, ECG image I, and task question q, a VLM with parameters θ models pθ (y | x, I, q) =

T Y

pθ (yt | y<t , x, I, q).

(1)

t=1

We denote the resulting outcome prediction by ŷ = fθ (x, I, q), using a fixed decoding procedure across input conditions. 2.2. ECG Mirage We define ECG Mirage as apparent multimodal predictive capability without reliable benefit from patient-matched ECG information. Let I + denote the matched ECG and I − an ECG from another patient with discordant outcomes. In the no-image condition (∅), the VLM receives only the EHR input and task question. The EHR input x, task question q, and target y remain fixed across conditions. For a task loss ℓ, the expected prediction losses are   R+ = E ℓ fθ (x, I + , q), y ,   (2) R− = E ℓ fθ (x, I − , q), y , R0 = E[ℓ(fθ (x, ∅, q), y)] .

We assess whether patient-matched ECGs provide predictive benefit by requiring lower expected loss than both mismatched ECGs and EHR input alone: R+ < R −

and

R+ < R0 .

(3)

The first comparison assesses the benefit of the correct patient’s ECG over an outcome-discordant substitute. The second assesses the added benefit of the ECG over the EHR text alone. We distinguish two illustrative ECG Mirage patterns. ECG neglect occurs when R+ ≥ R− and R+ ≥ R0 , indicating no predictive advantage from the matched ECG over either control. ECG confusion occurs when R− ≤ R+ < R0 , indicating an advantage over the no-image condition without an advantage over the mismatched ECG. Thus, improvement from including an ECG does not necessarily establish benefit from the correct patient’s ECG. These conditions describe predictive behaviour rather than internal model processing. Poorer control performance can inflate the matched-ECG advantage without improving matched predictions. 2.3. Visual Prompt Tuning (VPT) and Optimisation We adapt visual prompt tuning [13] through a restricted attention pathway and two-stage optimisation. Specifically, let HI ∈ RNI ×d denote the image token embeddings and Hx,q ∈ RNT ×d the textual token embeddings, where d is the hidden dimension. We introduce K = 4 learnable visual prompt embeddings Hv ∈ RK×d before the image and textual tokens, yielding   H(0) = Hv , HI , Hx,q , (4) where the brackets denote concatenation along the token dimension. The backbone parameters θ remain frozen, and only Hv is optimised. To restrict direct access to the prompts, let P, I, and T denote the positions of visual prompts, image tokens, and textual tokens, respectively. The textual positions include both input text and answer tokens. We define an additive attention mask ( −∞, i ∈ T and j ∈ P, Mij = (5) 0, otherwise, where i and j index queries and keys, respectively. For each attention head, we compute the masked attention as   QK⊤ Attention(Q, K, V) = softmax √ + A + M V, (6) dk where Q, K, and V are the query, key, and value matrices, dk is the key dimension, and A is the existing attention mask. In our implementation, image-token queries attend to preceding visual-prompt keys and values in full-attention layers, whereas textual queries are masked from attending directly to the prompts. The resulting prompt-conditioned image representations influence answer prediction through subsequent full-attention and recurrent linear-attention layers. To restrict direct prompt contributions through the recurrent pathway, we additionally zero the visual-prompt hidden states at the input to every recurrent token mixer of Qwen3.5. In the first training stage, we optimise the visual prompts on matched encounters. LSFT = − log pθ,Hv (y | x, I + , q).

(7)

Only target answer tokens contribute to this loss. This stage learns task-specific prompts through the restricted visual pathway, but provides no explicit comparison between matched and mismatched

(b) Clinical deterioration

60

60

40

40

(%)

(%)

(a) ICU admission

20

0

20

Matched Mismatched No image Qwen3.5 4B

0

Gemma4 MedGemma1.5 Qwen3.8 4B 4B 27B

Matched Mismatched No image Qwen3.5 4B

Gemma4 MedGemma1.5 Qwen3.8 4B 4B 27B

Fig. 2. Zero-shot performance on ICU admission and clinical deterioration prediction under matched ECG, mismatched ECG and no-image conditions. Bars show full-test balanced accuracy for ICU admission and macro balanced accuracy for deterioration (%). Error bars indicate one standard deviation from 100 case-bootstrap resamples.

ECGs. In the second stage, we initialise the policy and a frozen reference from the same supervised checkpoint. Each pair contains a matched ECG I + and a mismatched ECG I − , with the EHR input x, task question q, and target answer y held fixed. We define the policy’s corresponding log-probability margin as mHv = log pθ,Hv (y | x, I + , q) − log pθ,Hv (y | x, I − , q).

(8)

Let mref denote the same margin under the frozen reference. The second-stage training objective is LDPO = − log σ(β[mHv − mref ]) + λLSFT ,

(9)

where σ is the logistic sigmoid, β = 0.1 scales the preference margin, and λ = 0.1 weights the supervised term. DPO favours the target answer under matched over mismatched ECGs. 3. RESULTS We investigate whether zero-shot VLMs benefit from patientmatched ECGs in clinical outcome prediction (RQ1), whether our VPT approach improves predictive performance while mitigating ECG Mirage (RQ2), and how conditional preference optimisation, the supervised training stage and the visual attention mask contribute to these outcomes (RQ3). 3.1. Experimental setup We use MIMIC-IV-ED [14] with linked MIMIC-IV [15] and MIMIC-IV-ECG [16] following MDS-ED cohort construction and patient-level partitioning [2]. For ICU admission, the training, validation and test sets comprise 108,877/5,802/6,048 visits from 63,676/3,513/3,626 patients, respectively. For deterioration, the corresponding sets comprise 109,299/5,819/6,077 visits from 63,929/ 3,525/3,644 patients. ICU admission is assessed over the entire hospital stay, while deterioration outcomes are assessed within 24 hours of ED arrival, using the records from the first 90 minutes of observation in the ED admission [2]. The clinical deterioration task comprises six outcomes: severe hypoxaemia, vasopressor use, mechanical ventilation, extracorporeal membrane oxygenation, inotrope use, and in-hospital cardiac arrest. Prediction is performed 90 minutes after ED arrival. All EHR information available within this window is converted into textual input following [17], with task-specific leakage exclusions applied.

The matched ECG is the first recording from the ED encounter acquired by prediction time. We compare matched, mismatched, and no-image conditions while keeping the EHR input, task question, and target fixed. The mismatched condition randomly selects an ECG from another patient with a different task label. The no-image condition supplies only EHR text and the task question. For the deep learning baseline, we use the S4–MLP model [2], which combines an ECG waveform encoder with an MLP for fusion with tabular clinical features. For VLM adaptation baselines, we include low-rank adaptation (LoRA) [18] and prompt tuning [19]. All methods use the same cohort splits, prediction-time cutoff, and leakage exclusions. Models are trained separately for the two prediction tasks. For S4–MLP, we follow the original training configuration reported in MDS-ED [2]. For VLM adaptation, we use AdamW [20] with a learning rate of 10−3 and a batch size of 32. The learning rate increases linearly over the first 10% of training steps and follows a cosine decay schedule over the remaining 90%. Training runs for a maximum of one epoch, with validation every 500 optimisation steps. We select the best checkpoint by task-specific validation F1 and stop after three consecutive checks without improvement. For ICU admission, we report balanced accuracy and macro F1. For clinical deterioration, we report macro balanced accuracy and macro positive-class F1 across the six outcomes, excluding unavailable labels from the corresponding calculations. We report scores computed on the complete test set together with standard deviations estimated from 100 bootstrap resamples [21]. Each resample draws test encounters with replacement, retaining their predictions and reference labels, and the same resampling indices are used across methods and input conditions. These standard deviations describe test-set sampling variability rather than variation across training runs. 3.2. RQ1: Examining ECG Mirage in zero-shot prediction In this study, we examined whether zero-shot VLMs derive predictive benefit from ECGs when EHR is also available. We evaluated Qwen3.5-4B [22], Qwen3.8-27B [22], Gemma4-4B [23] and MedGemma1.5-4B [24] on ICU admission and clinical deterioration prediction under three conditions: matched ECGs, mismatched ECGs from other patients, and no image. Figure 2 shows evidence of ECG Mirage. For ICU admission, all four models had slightly lower balanced-accuracy point estimates with matched ECGs than with mismatched ECGs. Qwen3.5-4B achieved 52.3%, 52.5% and 53.4% under matched, mismatched and

Table 1. Performance on ICU admission and clinical deterioration. All VLM methods use Qwen3.5-4B. S4-MLP follows the MDS-ED architecture and main training hyperparameters. Subscripts indicate case-bootstrap standard deviations over 100 resamples. ICU Admission Matched Method

BAcc

S4-MLP

67.3±0.8 49.6±1.6

F1

Mismatched

Clinical Deterioration No image

Matched

Mismatched

BAcc

F1

BAcc

F1

BAcc

—

—

—

—

55.0±0.8 12.8±1.7

F1

No image

BAcc

F1

BAcc

F1

—

—

—

—

Zero-shot 52.3±0.3 50.7±0.7 52.5±0.4 51.2±0.7 53.4±0.4 52.9±0.8 58.1±1.6 12.8±2.0 57.2±1.5 11.8±2.0 56.2±1.5 13.1±2.8 LoRA 74.7±1.0 78.8±0.9 73.8±1.0 78.0±0.9 74.7±1.0 78.5±0.9 69.2±2.4 34.4±3.3 68.4±2.3 33.5±3.2 68.4±2.4 30.3±2.4 Prompt Tuning 71.2±0.9 75.4±0.9 70.6±0.9 74.7±0.9 54.3±0.4 54.3±0.8 67.1±2.0 22.8±2.3 66.7±2.1 23.0±2.3 63.8±1.6 17.2±2.0 VPT (Ours)

70.6±0.9 62.8±0.7 54.1±1.0 47.7±0.6 55.5±0.5 56.5±0.9 67.5±2.1 21.7±2.8 62.1±1.6 17.4±1.7 55.6±1.2 11.0±1.9

no-image conditions, respectively. MedGemma performed better with matched ECGs than without an image, but not better than with mismatched ECGs, indicating that gains from including an ECG do not necessarily reflect a benefit from patient-specific information. For clinical deterioration, Gemma4-4B also performed worse with matched ECGs than under either control, whereas both Qwen models achieved higher matched-ECG point estimates than under either control. Overall, matched ECGs provided no consistent predictive advantage across the evaluated models and tasks, supporting the existence of ECG Mirage. These models may rely primarily on EHR text, consistent with prior findings that medical images add little predictive value when clinical text is informative [10]. 3.3. RQ2: Mitigating ECG Mirage We investigated whether our VPT approach mitigates ECG Mirage while improving clinical outcome prediction. Table 1 compares VPT with zero-shot inference, LoRA and supervised prompt tuning, using Qwen3.5-4B as the VLM backbone. S4–MLP served as a deeplearning baseline for predictive performance. VPT improved matched-ECG performance over zero-shot inference on both tasks, achieving balanced accuracies of 70.6% for ICU admission and 67.5% for clinical deterioration, with corresponding F1 scores of 62.8% and 21.7% under the task-specific definitions. Matched ECGs outperformed mismatched ECGs by 16.5 and 5.5 percentage points, respectively, and no-image inputs by 15.1 and 11.9 percentage points. VPT had the largest matched–mismatched balanced-accuracy gaps among the evaluated VLM methods. Although LoRA achieved higher absolute predictive performance, its performance remained similar across ECG conditions. These comparisons show that improvements in task performance need not be accompanied by greater benefit from patient-matched ECGs. LoRA substantially improved prediction even without an image, while the prompt-tuning results indicate that a benefit from adding an ECG need not depend on patient matching. Our VPT approach combined improved matched performance over zero-shot inference with higher scores for matched ECGs than for either control, supporting mitigation of ECG Mirage in the evaluated setting. Its main benefit was therefore improved prediction over the unadapted model combined with a clearer distinction between matched and mismatched ECGs. 3.4. RQ3: Contributions of the adaptation components We conducted ablation experiments on clinical deterioration prediction to assess the contributions of conditional preference optimisation, supervised initialisation and the visual attention mask. Table 2 compares VPT with three variants: w/o DPO, which uses

Table 2. Ablation study on clinical deterioration prediction. Performance is measured with balanced accuracy (%). Method

Matched Mismatched No image

VPT

67.5±2.1

62.1±1.6

55.6±1.2

w/o DPO 63.4±1.6 w/o SFT Stage 56.5±1.4 w/o Visual Mask 64.3±1.8

63.7±1.8 56.3±1.4 64.1±1.7

55.4±1.1 55.5±1.2 62.2±1.6

supervised training alone; w/o SFT stage, which removes the separate supervised training stage and jointly optimises the SFT and DPO losses from the outset; and w/o Visual Mask, which allows text tokens to attend directly to the learned prompts. VPT achieved the highest matched balanced accuracy among the evaluated configurations (67.5%). The w/o DPO variant achieved similar scores with matched (63.4%) and mismatched (63.7%) ECGs, although both exceeded its no-image performance (55.4%). Relative to this variant, VPT improved matched performance and reduced mismatched performance to 62.1%, with both changes contributing to the larger separation. The Joint SFT+DPO variant achieved 56.5% with matched ECGs and similar scores under mismatched (56.3%) and no-image (55.5%) conditions, favouring a separate supervised initialisation stage over joint optimisation from the outset in this setting. The w/o Visual Mask variant retained sequential training but achieved similar matched (64.3%) and mismatched (64.1%) scores.

4. CONCLUSION In this study, we identified ECG Mirage in multimodal clinical outcome prediction, where VLM performance can obscure benefit from patient-matched ECG information. Across four zero-shot VLMs and two ED prediction tasks, matched ECGs offered no consistent advantage over mismatched ECGs or EHR text alone. We introduced a lightweight approach combining restricted visual prompts and conditional preference optimisation with a frozen backbone. VPT improved matched-ECG performance over zero-shot inference and increased matched–mismatched performance gaps on both tasks. These findings support evaluating multimodal clinical models through both predictive performance and patient-matched modality benefit. However, LoRA achieved higher predictive performance with matched ECGs than VPT on both tasks. Future work aims to improve predictive performance while ensuring that models benefit from patient-specific ECG information.

5. COMPLIANCE WITH ETHICAL STANDARDS This study retrospectively analysed de-identified data from MIMICIV, accessed through PhysioNet under the respective data use agreements. The Beth Israel Deaconess Medical Centre Institutional Review Board approved the original MIMIC-IV data sharing with a waiver of informed consent. Our study involved no patient recruitment or clinical intervention. 6. ACKNOWLEDGEMENTS The authors declare no conflicts of interest. 7. REFERENCES [1] Emma Chen et al., “Multimodal clinical benchmark for emergency care (MC-BEC): A comprehensive benchmark for evaluating foundation models in emergency medicine,” Advances in Neural Information Processing Systems, vol. 36, pp. 45794– 45811, 2023. [2] Juan Miguel Lopez Alcaraz, Hjalmar Bouma, and Nils Strodthoff, “Enhancing clinical decision support with physiological waveforms—a multimodal benchmark in emergency care,” Computers in Biology and Medicine, vol. 192, pp. 110196, 2025. [3] Taha Razzaq, Murtaza Taj, and Asim Iqbal, “Multimodal ai in healthcare: Review of vision-language foundation models for real-world medical applications,” Journal of Biomedical Informatics, p. 105075, 2026. [4] Mingcheng Zhu, Yu Liu, Zhiyao Luo, and Tingting Zhu, “The taxonomies, training, and applications of event stream modelling for electronic health records,” arXiv preprint arXiv:2603.14003, 2026. [5] Mingcheng Zhu, Zhiyao Luo, Yu Liu, and Tingting Zhu, “From token to token pair: Efficient prompt compression for large language models in clinical prediction,” in Proceedings of the 43rd International Conference on Machine Learning, 2026. [6] Mingcheng Zhu, Yu Liu, Zhiyao Luo, and Tingting Zhu, “Bridging data gaps of rare conditions in ICU: a multi-disease adaptation approach for clinical prediction,” npj Digital Medicine, vol. 9, no. 1, pp. 7, 2026. [7] Ruoqi Liu, Yuelin Bai, Xiang Yue, and Ping Zhang, “Teaching multimodal LLMs to comprehend 12-lead electrocardiographic images,” npj Digital Medicine, vol. 9, no. 1, pp. 349, 2026. [8] Xiang Lan, Feng Wu, Kai He, Qinghao Zhao, Shenda Hong, and Mengling Feng, “GEM: Empowering MLLM for grounded ECG understanding with time series and images,” Advances in Neural Information Processing Systems, vol. 38, pp. 94421–94455, 2025. [9] Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang, Yu Qiao, Dahua Lin, et al., “Are we on the right way for evaluating large visionlanguage models?,” Advances in Neural Information Processing Systems, vol. 37, pp. 27056–27087, 2024. [10] Thomas A Buckley, James A Diao, Cam N Srivastava, Peter G Brodeur, Pranav Rajpurkar, Adam Rodman, and Arjun K Manrai, “Multimodal foundation models exploit text to make medical image predictions,” Nature Communications, vol. 17, pp. 7475, 2026.

[11] Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn, “Direct preference optimization: Your language model is secretly a reward model,” in Advances in Neural Information Processing Systems, 2023, vol. 36, pp. 53728–53741. [12] Fei Wang, Wenxuan Zhou, James Y Huang, Nan Xu, Sheng Zhang, Hoifung Poon, and Muhao Chen, “mDPO: Conditional preference optimization for multimodal large language models,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 8078– 8088. [13] Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim, “Visual prompt tuning,” in European conference on computer vision. Springer, 2022, pp. 709–727. [14] Alistair Johnson, Lucas Bulgarelli, Tom Pollard, Leo Anthony Celi, Roger Mark, and Steven Horng, “MIMIC-IV-ED,” PhysioNet, 2023, Version 2.2. [15] Alistair Johnson, Lucas Bulgarelli, Tom Pollard, Steven Horng, Leo Anthony Celi, and Roger Mark, “MIMIC-IV,” PhysioNet, 2023, Version 2.2. [16] Brian Gow, Tom Pollard, Larry A Nathanson, Alistair Johnson, Benjamin Moody, Chrystinne Fernandes, Nathaniel Greenbaum, Jonathan W Waks, Parastou Eslami, Tanner Carbonati, et al., “MIMIC-IV-ECG: Diagnostic electrocardiogram matched subset,” PhysioNet, 2023, Version 1.0. [17] Tianyi Chen, Mingcheng Zhu, Zhiyao Luo, and Tingting Zhu, “Cross-representation benchmarking in time-series electronic health records for clinical outcome prediction,” in ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2026, pp. 7076–7080. [18] Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen, “LoRA: Low-rank adaptation of large language models,” in International Conference on Learning Representations, 2022. [19] Brian Lester, Rami Al-Rfou, and Noah Constant, “The power of scale for parameter-efficient prompt tuning,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021, pp. 3045–3059, Association for Computational Linguistics. [20] Ilya Loshchilov and Frank Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations, 2019. [21] Bradley Efron, “Bootstrap methods: Another look at the jackknife,” The Annals of Statistics, vol. 7, no. 1, pp. 1–26, 1979. [22] Qwen Team, “Qwen3.5: Towards native multimodal agents,” https://qwen.ai/blog?id=qwen3.5, 2026. [23] Gemma Team, Sherif El Abd, Vaibhav Aggarwal, Robin Algayres, Alek Andreev, Olivier Bachem, Ian Ballantyne, Cormac Brick, Victor Cărbune, Michelle Casbon, et al., “Gemma 4 technical report,” arXiv preprint arXiv:2607.02770, 2026. [24] Andrew Sellergren et al., “MedGemma 1.5 technical report,” arXiv preprint arXiv:2604.05081, 2026.

Record · ID 1006898 · SHA-256 27de3bab6e730ead
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.