ConceptioArchivearXiv CS
arXiv CSopen access

EchoBridge: Long-Tail-Aware ECG-Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

EchoBridge: Long-Tail-Aware ECG–Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings Xiaocheng Fang∗

School of Intelligence University of the Chinese Science and Technology, Academy of Sciences Peking University Beijing, China Beijing, China [email protected] [email protected]

arXiv:2607.24553v1 [cs.LG] 27 Jul 2026

Guangkun Nie∗

Jieyi Cai

Haoyu Wang

School of Intelligence University of the Chinese Science and Technology, Academy of Sciences Peking University Beijing, China Beijing, China [email protected] [email protected]

Jiarui Jin∗

Yujie Xiao

Bo Liu∗

Chenyang He

School of Intelligence Science and Technology, Peking University Beijing, China [email protected]

National Institute of Health Data Science, Peking University Beijing, China [email protected]

School of Intelligence Science and Technology, Peking University Beijing, China [email protected]

National Institute of Health Data Science, Peking University Beijing, China [email protected]

Qinghao Zhao

Gaofeng Cheng

Hongyan Li∗†

Shenda Hong†

Department of Cardiology, Peking University People’s Hospital Beijing, China [email protected]

University of the Chinese Academy of Sciences Beijing, China [email protected]

School of Intelligence Science and Technology, Peking University Beijing, China [email protected]

National Institute of Health Data Science, Peking University Beijing, China [email protected]

Abstract Standardized echocardiography conclusions provide meaningful supervision for learning ECG representations of echocardiographyderived cardiac findings. Global ECG–text alignment may entangle modality-specific factors, while long-tailed finding distributions provide sparse positive supervision for low-prevalence conditions. We propose EchoBridge with Complementary Shared–Private Projection (CSPP) and Adaptive Prototype Boundary Calibration (APBC). CSPP maps each modality into shared and auxiliary private projections, reduces directional redundancy via within-modality orthogonality, and bidirectionally aligns normalized shared projections. APBC organizes the shared hypersphere with class-specific prototypes, training-frequency-adaptive angular margins, and spherical Riesz repulsion. We evaluate EchoBridge on EchoNext-Mini and independent PKUPH and SHTMU cohorts under four protocols: prompt-based inference without downstream classifier training, indomain frozen linear probing, target-domain cross-center frozen linear probing, and source-only cross-center transfer, supplemented by finding-specific analyses. EchoBridge improves classifier-free ∗ Also with State Key Laboratory of General Artificial Intelligence, Peking University. † Corresponding authors.

Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. KDD ’27, San Jose, CA, USA © 2018 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-XXXX-X/2018/06 https://doi.org/XXXXXXX.XXXXXXX

AUROC, AUPRC, and F1 over the strongest baselines by 7.88, 5.61, and 4.54 points, respectively, and achieves the highest point estimates across all in-domain and target-domain probing budgets and both source-only transfer cohorts. Finding-specific analyses show gains for most conditions, including several low-prevalence valvular findings. Code: https://github.com/PKUDigitalHealth/EchoBridge.

CCS Concepts • Applied computing → Health informatics; • Computing methodologies → Artificial intelligence. ACM Reference Format: Xiaocheng Fang, Jieyi Cai, Guangkun Nie, Haoyu Wang, Jiarui Jin, Yujie Xiao, Bo Liu, Chenyang He, Qinghao Zhao, Gaofeng Cheng, Hongyan Li, and Shenda Hong. 2018. EchoBridge: Long-Tail-Aware ECG–Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings. In Proceedings of The 33rd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD ’27). ACM, New York, NY, USA, 16 pages. https: //doi.org/XXXXXXX.XXXXXXX

1

Introduction

Echocardiography is a central imaging modality for evaluating cardiac structure, function, and valvular abnormalities [7, 14, 31]. Direct ECG–echocardiography learning from raw images or videos can be difficult to conduct at scale in retrospective multicenter studies because of storage, accessibility, and data-governance constraints [20, 35]. In routine practice, echocardiographic examinations are accompanied by diagnostic reports containing detailed findings and a concise conclusion that aggregates the principal structural, functional, and valvular abnormalities [3, 22]. We use

KDD ’27, August, 2027, San Jose, CA, USA

Fang et al.

12-lead ECG

Echocardiography Text ECG Auxiliary Private

Confounder

Shared

Pathology-Related Information

Leads to

Global Alignment

Echo-Text Auxiliary Private

Confounder

Chamber structure: Left ventricular hypertrophy.

Class 1

Updated More Frequently Discriminative

Myocardial function: Left ventricular systolic dysfunction.

Class 4

Valve findings: Moderate or severe tricuspid regurgitation, Moderate or severe mitral regurgitation.

Error-Prone Class 2

Class 3

Majority Class Data Minority Class Data Previous Boundary Updated Boundary

Updated Less Frequently

Figure 1: Two challenges in ECG–echocardiography text alignment. Left: Global alignment entangles modality-specific factors and weakens clinically relevant cross-modal correspondence. Right: Prevalence imbalance provides fewer positive constraints and less reliable sample–prototype organization for low-prevalence findings. these conclusion-level summaries as cross-modal supervision because they preserve clinically salient finding compositions while reducing measurement-, acquisition-, and template-specific variation. When paired with temporally matched ECG recordings, these summaries connect echocardiography-derived clinical semantics with cardiac electrical activity and provide a practical source of cross-modal supervision. Recent AI–ECG studies show that ECGs encode echocardiographyconfirmed abnormalities, including reduced ejection fraction and broader structural heart disease [2, 32, 37]. MERL and ECG-CLIP use natural-language supervision for transferable representation learning [28, 40], while MERL-ECHO and Wearable-Echo-FM align ECGs with echocardiography reports for structural cardiac findings [21, 36]. Building on these methods, we learn a class-aware normalized space from paired ECGs and standardized conclusions, capturing paired-sample ECG–text alignment and organizing both modalities by predefined findings; evaluation covers prompt-based classifier-free inference for pretraining-seen findings, in-domain and target-domain cross-center frozen linear probing across label budgets, and source-only cross-center transfer. Technically, aligning ECGs with standardized echocardiography conclusions poses two challenges. 1) Cross-modal representation interference. ECG representations may encode continuous variation in rhythm, conduction, waveform morphology, and acquisition-related factors, whereas standardized conclusions provide compact compositional semantics for predefined structural, functional, and valvular findings. Global alignment may force heterogeneous ECG factors into a single conclusion-oriented embedding space, reducing the separation between finding-relevant and modality-dependent variation. 2) Long-tailed representation bias. Echocardiography-derived findings often have highly imbalanced prevalence. Frequent findings contribute more positive samples and repeated optimization signals, potentially biasing the space toward head classes. Low-prevalence findings receive fewer positive constraints, reducing intra-class compactness and prototype separation and yielding less reliable decision boundaries and uneven representation quality across prevalence levels. To address these challenges, we propose EchoBridge, an ECG– echocardiography text alignment framework for learning transferable ECG representations of echocardiography-derived findings. EchoBridge comprises two components. First, Complementary Shared–Private Projection maps ECG and text representations into alignment-oriented shared and auxiliary private projections. Crossmodal alignment and class-prototype supervision operate on the shared projections, while within-modality orthogonality reduces

directional redundancy between branches and a symmetric contrastive objective aligns the ℓ2 -normalized shared representations. Second, Adaptive Prototype Boundary Calibration organizes the shared hypersphere using class-specific prototypes for predefined findings. Training-frequency-adaptive angular margins strengthen positive sample–prototype constraints for low-prevalence findings, while spherical Riesz repulsion discourages prototype concentration. Together, these components integrate instance-level ECG– text correspondence with class-aware geometric organization under imbalanced finding distributions. We evaluate EchoBridge on EchoNext-Mini and two independent hospital cohorts from Peking University People’s Hospital (PKUPH) and the Second Hospital of Tianjin Medical University (SHTMU) under four protocols: promptbased classifier-free inference for pretraining-seen findings, indomain frozen linear probing across multiple label budgets, targetdomain cross-center linear probing, and source-only cross-center transfer, complemented by finding-specific analyses. The main contributions are as follows: • We propose EchoBridge, an ECG–echocardiography text alignment framework for transferable ECG representations of echocardiography-derived findings under cross-modal interference and long-tailed distributions. • We develop Complementary Shared–Private Projection, which maps each modality into shared and auxiliary private projections, reduces directional redundancy via within-modality orthogonality, and bidirectionally aligns the normalized shared space. • We propose Adaptive Prototype Boundary Calibration, which structures the shared hypersphere with cross-modal class prototypes, training-frequency-adaptive angular margins, and spherical Riesz repulsion to improve class compactness and prototype separation under imbalance. • We evaluate EchoBridge on EchoNext-Mini and two independent hospital cohorts using prompt-based classifier-free inference, in-domain and target-domain cross-center frozen linear probing across label budgets, source-only cross-center transfer, and finding-specific analyses.

2 Methodology 2.1 Overview of EchoBridge EchoBridge learns transferable echocardiography-related ECG representations from paired 12-lead ECGs and standardized echocardiography conclusions. As shown in Figure 2, it combines Complementary Shared–Private Projection and Adaptive Prototype Boundary

EchoBridge: Long-Tail-Aware ECG–Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings

A. Complementary Shared–Private Projection

Paired

Echo-Text Encoder

Echo-Text Private Echo-Text Feature

Echo-Text

Class 1

Push

Pull Class 2

Echo-Text Shared Feature

Pull

Shared

Push

Class 4

Paired

ECG

Norm

ECG Encoder

ECG Private ECG Feature

B. Adaptive Prototype Boundary Calibration

Hyperspherical Contrast

Norm

Chamber structure: Left ventricular hypertrophy. Myocardial function: Left ventricular systolic dysfunction. Valve findings: Moderate or severe tricuspid regurgitation, Moderate or severe mitral regurgitation.

KDD ’27, August, 2027, San Jose, CA, USA

ECG Shared Feature

Class 3

Echo-Text Shared Feature

Class Prototype Class Boundary Direction of Boundary Update

ECG Shared Feature Positive Contrast Negative Contrast

Figure 2: Overview of EchoBridge. Given paired ECGs and echocardiography summaries, (A) Complementary Shared–Private Projection maps each modality into alignment-oriented shared and auxiliary private branches, with within-modality orthogonality reducing directional redundancy between them. (B) Adaptive Prototype Boundary Calibration structures the shared space using class-specific prototypes, frequency-adaptive angular margins, and spherical Riesz repulsion. Calibration through normalized bidirectional alignment. Given an ECG 𝑥𝑖 and paired summary 𝑡𝑖 , encoders 𝑓𝑒 (·) and 𝑓𝑡 (·) produce representations ℎ𝑒,𝑖 and ℎ𝑡,𝑖 , which independent heads map into shared and auxiliary private projections. Within-modality orthogonality reduces directional redundancy between branches, while symmetric contrastive learning aligns the ℓ2 -normalized shared representations by increasing paired-sample similarity over unpaired samples. Adaptive Prototype Boundary Calibration structures the shared space with class-specific prototypes for predefined cardiac findings. Training-frequency-adaptive angular margins strengthen positive alignment for low-prevalence findings, while spherical Riesz repulsion discourages prototype concentration and promotes pairwise separation on the unit hypersphere.

2.2

Complementary Shared–Private Projection

ECG waveforms encode rich continuous variation in rhythm, conduction, morphology, and acquisition conditions, whereas standardized echocardiography conclusions provide compact compositional semantics for predefined cardiac findings. Their global representations therefore contain factors with different relevance to cross-modal correspondence. Direct global alignment may entangle finding-relevant and modality-specific information within a single embedding space. We propose Complementary Shared–Private Projection, which maps each modality into alignment-oriented shared and auxiliary private projections. Cross-modal alignment and class-prototype supervision operate on the shared projections, while within-modality orthogonality reduces sample-wise directional redundancy between the two branches. The terms shared and private specify their optimization roles, while their detailed semantic contents remain unconstrained. Shared–Private Projection. Given an ECG 𝑥𝑖 and paired echocardiography conclusion 𝑡𝑖 , the ECG and text encoders produce global representations: ℎ𝑒,𝑖 = 𝑓𝑒 (𝑥𝑖 ),

ℎ𝑡,𝑖 = 𝑓𝑡 (𝑡𝑖 ).

(1)

Independent projection heads map each representation into shared and private projections: 𝑠 𝑧𝑒,𝑖 = 𝜙𝑒𝑠 (ℎ𝑒,𝑖 ),

𝑧𝑒,𝑖 = 𝜙𝑒 (ℎ𝑒,𝑖 ),

𝑝

𝑝

(2)

𝑠 𝑧𝑡,𝑖 = 𝜙𝑡𝑠 (ℎ𝑡,𝑖 ),

𝑝 𝑝 𝑧𝑡,𝑖 = 𝜙𝑡 (ℎ𝑡,𝑖 ).

(3)

Cross-modal alignment and class-prototype supervision operate 𝑠 and 𝑧𝑠 . The independently parameon the shared projections 𝑧𝑒,𝑖 𝑡,𝑖 𝑝 𝑝 terized private projections 𝑧𝑒,𝑖 and 𝑧𝑡,𝑖 provide auxiliary capacity without direct cross-modal alignment, allowing the shared branches to emphasize ECG–text correspondence. To reduce sample-wise directional redundancy between the shared and private projections, we introduce a within-modality orthogonality objective: Lorth =

𝑁 i 1 ∑︁ h 𝑠 𝑝 2 𝑝 2 𝑠 , , 𝑧¯𝑡,𝑖 𝑧¯𝑒,𝑖 , 𝑧¯𝑒,𝑖 + 𝑧¯𝑡,𝑖 𝑁 𝑖=1

(4)

where 𝑁 is the batch size, 𝑧¯ = 𝑧/∥𝑧 ∥ 2 denotes an ℓ2 -normalized representation, and ⟨·, ·⟩ denotes the inner product. Because the normalized inner product corresponds to cosine similarity, minimizing Lorth drives the shared and private projections toward orthogonality within each modality. This geometric regularization reduces collinearity while leaving the semantic content of the shared and private projections unconstrained. Shared-Space Alignment. The shared ECG and text projections are normalized onto the unit hypersphere: 𝑠 𝑧ˆ𝑒,𝑖 =

𝑠 𝑧𝑒,𝑖 𝑠 ∥ ∥𝑧𝑒,𝑖 2

,

𝑠 𝑧ˆ𝑡,𝑖 =

𝑠 𝑧𝑡,𝑖 𝑠 ∥ ∥𝑧𝑡,𝑖 2

.

(5)

For a mini-batch of 𝑁 paired samples, the cross-modal similarity matrix is defined as: D E 𝑠 , 𝑧ˆ𝑠 𝑧ˆ𝑒,𝑖 𝑡,𝑗 𝑆𝑖 𝑗 = , (6) 𝜏 where 𝜏 is the temperature parameter. Because both representations are ℓ2 -normalized, their inner product equals cosine similarity.

KDD ’27, August, 2027, San Jose, CA, USA

Fang et al.

We optimize bidirectional ECG–text correspondence using a symmetric contrastive objective: " # 𝑁 exp(𝑆𝑖𝑖 ) exp(𝑆𝑖𝑖 ) 1 ∑︁ − log Í𝑁 Lalign = − log Í𝑁 . (7) 2𝑁 𝑖=1 𝑗=1 exp(𝑆𝑖 𝑗 ) 𝑗=1 exp(𝑆 𝑗𝑖 ) The first term retrieves the paired echocardiography conclusion from an ECG query, while the second performs reverse retrieval. This objective forms a common normalized space for paired ECG– text correspondence, which the following class-prototype objective organizes by echocardiography-derived findings.

2.3

Adaptive Prototype Boundary Calibration

Pairwise ECG–text alignment provides instance-level supervision, whereas fine-grained cardiac findings require explicit class-level organization. Long-tailed distributions may produce weak boundaries for infrequent classes, while unconstrained prototypes may cluster geometrically. We therefore propose Adaptive Prototype Boundary Calibration (APBC), combining frequency-adaptive angular margins calibrated by training-set label frequencies with spherical repulsion that separates normalized class prototypes. We maintain a learnable prototype matrix: 𝑃 = [𝑝 1, . . . , 𝑝𝐶 ] ⊤ ∈ R𝐶 ×𝑑 ,

(8)

where 𝐶 denotes the number of fine-grained cardiac findings and 𝑑 the shared-space dimension. Each prototype is normalized as: 𝑝𝑐 𝑝ˆ𝑐 = . (9) ∥𝑝𝑐 ∥ 2 𝑠 and 𝑧ˆ𝑠 , For normalized shared ECG and text representations 𝑧ˆ𝑒,𝑖 𝑡,𝑖 the prototype similarities are: 𝑒 𝑠 ˆ 𝑢𝑖,𝑐 = 𝑧ˆ𝑒,𝑖 , 𝑝𝑐 ,

𝑡 𝑠 ˆ 𝑢𝑖,𝑐 = 𝑧ˆ𝑡,𝑖 , 𝑝𝑐 .

(10)

𝑡 𝑡 ℓ𝑖,𝑐 = 𝛾𝑢𝑖,𝑐 ,

(11)

The corresponding logits are: 𝑒 𝑒 ℓ𝑖,𝑐 = 𝛾𝑢𝑖,𝑐 ,

where 𝛾 = exp(𝑠) is a learnable positive scale. Sharing 𝑃 across modalities provides a common category-level reference in the normalized space. Frequency-Adaptive Angular Margin. To address label-frequency imbalance, we assign each cardiac finding a class-specific angular margin based on its training-set positive rate. Let 𝑟𝑐 denote the positive rate of class 𝑐. Its margin is:   median(𝑟 ) 1 · , 𝑚 min, 𝑚 max , (12) 𝑚𝑐 = clip 𝑚 0 · 𝑟𝑐 + 𝜖 𝜅 where 𝑚 0 is the base margin, 𝜖 ensures numerical stability, and 𝑚 min and 𝑚 max bound the margin. The data-dependent normalization factor is: 𝐶 1 ∑︁ median(𝑟 ) 𝜅= , (13) 𝐶 𝑐=1 𝑟𝑐 + 𝜖 which normalizes the mean pre-clipping margin to 𝑚 0 . For each sample–class pair, the adjusted logits are:   𝑒 𝑒 𝑡 𝑡 ℓ˜𝑖,𝑐 = 𝛾 cos 𝜃 𝑖,𝑐 + 𝑦𝑖,𝑐 𝑚𝑐 , ℓ˜𝑖,𝑐 = 𝛾 cos 𝜃 𝑖,𝑐 + 𝑦𝑖,𝑐 𝑚𝑐 , (14)     𝑒 = arccos 𝑢 𝑒 𝑡 𝑡 where 𝜃 𝑖,𝑐 𝑖,𝑐 and 𝜃 𝑖,𝑐 = arccos 𝑢𝑖,𝑐 denote the ECG– prototype and text–prototype angles, respectively, and 𝑦𝑖,𝑐 ∈ {0, 1} indicates whether sample 𝑖 is positive for class 𝑐. Thus, the margin

Margin

Before

After

Error-Prone

Discriminative

Class Prototype

Class Boundary

Adaptive Angular Margin

Data Sample Pull/Push Contrast

Figure 3: Adaptive Prototype Boundary Calibration is designed to improve positive-class compactness and prototype separation on the hypersphere. applies only to positive sample–prototype pairs, leaving negativepair logits unchanged. Lower-prevalence findings receive larger bounded margins, encouraging more discriminative representations around tail-class prototypes. The adjusted logits supervise the shared ECG and text branches:   Lproto = BCE ℓ˜𝑒 , 𝑦 + BCE ℓ˜𝑡 , 𝑦 , (15) where 𝑦 is the multi-label target matrix, and ℓ˜𝑒 and ℓ˜𝑡 are the adjusted ECG and text logits. This objective draws positive representations toward corresponding prototypes and suppresses responses to prototypes of negative findings, linking instance-level ECG–text alignment with class-level multi-label supervision. Spherical Riesz Repulsion. Prototype supervision constrains sample–prototype relationships, while class prototypes may still cluster on the hypersphere. We therefore apply a spherical Riesz repulsion loss to the normalized prototypes: ∑︁ 2 −𝑞 ∥ 𝑝ˆ𝑎 − 𝑝ˆ𝑏 ∥ 2 , Lriesz = 𝑞 > 0. (16) 𝐶 (𝐶 − 1) 1≤𝑎<𝑏 ≤𝐶

Because the loss increases as pairwise distance decreases, minimizing Lriesz discourages prototype concentration and promotes separation on the unit hypersphere. The Adaptive Prototype Boundary Calibration objective is: Lapbc = Lproto + 𝜆𝑟 Lriesz,

(17)

where 𝜆𝑟 controls the contribution of spherical prototype repulsion.

2.4

Overall Training Objective

EchoBridge jointly optimizes Complementary Shared–Private Projection and Adaptive Prototype Boundary Calibration with the objective: L = Lalign + Lorth + Lapbc . (18)

3 Experiments 3.1 Datasets and Splits We evaluate EchoBridge on the open-source EchoNext-Mini [15] and two private real-world cohorts from Peking University People’s Hospital (PKUPH) and the Second Hospital of Tianjin Medical University (SHTMU). Each sample pairs an ECG with echocardiographyderived findings from a temporally matched examination under a dataset-specific matching window. We partition each dataset into patient-disjoint training, validation, and test sets using a 7:1:2 ratio

EchoBridge: Long-Tail-Aware ECG–Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings EchoNext-Mini

PKUPH 24,220

LVH

23,892

1,090 616

MR 257 RVE 210

1,264

1,878

LVSD

LVWMA 273

4,054

AS AR

LVSD

8,451

MR

3,216

LVH

5,108

LVH LVE

10,651

TR

Table 1: Prompt-based classifier-free inference performance on EchoNext-Mini for pretraining-seen cardiac findings.

10,928

LAE

7,623

LAE LVSD

SHTMU 10,707

LVDD

TR

601

MR

598

AR 283

RAE 207 PR 821 0.00

0.05

ECGs = 100,000 train 70,000 val 10,000 test 20,000

TR 162

0.10

0.0

0.15

0.20

Positive prevalence

0.25

ECGs = 27,158 train 19,010 val 2,716 test 5,432 0.1

0.2

0.3

Positive prevalence

0.4

ECGs = 18,588 train 13,011 val 1,859 test 3,718

RAE 108 0.0

0.1

0.2

0.3

0.4

Positive prevalence

0.5

0.6

0.7

Figure 4: Dataset statistics and label distributions across EchoNext-Mini, PKUPH, and SHTMU. Each panel also reports the ECG cohort size and the train/validation/test split. to reduce information leakage. All datasets exhibit long-tailed label distributions; Figure 4 reports their statistics, label distributions, and split sizes.

3.2

Data Preprocessing

ECG Signal Processing. We applied record-level quality control and excluded unreadable ECGs, records with substantial missingness, and samples that could not be reliably linked to patient identifiers or labels. All signals were resampled to 500 Hz by linear interpolation and processed using a fixed denoising pipeline comprising a 0.5-Hz high-pass filter, a second-order Butterworth low-pass filter with a 50-Hz cutoff, and a 50/60-Hz notch filter. Recordings were standardized to 10-second segments, with longer recordings divided into consecutive temporal windows. Each segment was Z-score normalized, and missing leads or signal values were zero-filled to preserve a consistent input shape. Echocardiography Conclusion-Style Text Construction. Clinical echocardiography reports typically contain detailed findings and a concise conclusion or impression. Findings may include quantitative measurements, image-quality descriptions, and contextual observations with variable relevance to predefined cardiac findings, limiting their consistency as supervision for finding-specific representation learning. We therefore use conclusion-level semantics. For each examination, we deterministically verbalize the structured multi-label finding vector into a standardized conclusion-style summary organized by chamber structure, myocardial function, and valvular abnormalities. This preserves multi-label co-occurrence, mirrors the concise compositional form of clinical conclusions, and controls unrelated lexical, formatting, and template variation. As an informal quality check, five senior cardiologists inspected a random subset of the generated conclusion-style summaries and provided qualitative feedback that the sampled summaries were broadly clinically plausible and semantically consistent with the corresponding structured findings. This review was not designed as a formal annotation or inter-rater agreement study. The resulting text supports ECG–text alignment and class-specific prototype learning under controlled conclusion-level semantics.

3.3

KDD ’27, August, 2027, San Jose, CA, USA

Baselines and Implementation Details

We compare EchoBridge with representative ECG-only self-supervised and ECG–text pretraining methods. Prompt-based classifier-free inference uses ECG–text dual encoders with fixed definition-only textual prototypes for pretraining-seen findings. In all linear-probing

Methods

Ref.

AUROC

AUPRC

F1

CLIP [33] SigLIP [39] PCME++ [8] MERL-ECHO [36] ECG-CLIP [40] D-BETA [16] SGERA [5]

ICML’21 ICCV’23 ICLR’24 medRxiv’25 npj DM’25 ICML’25 ICML’26

62.57 [61.70, 63.41] 65.07 [64.23, 65.91] 62.02 [61.14, 62.87] 64.02 [63.17, 64.86] 60.47 [59.55, 61.36] 63.99 [63.15, 64.83] 67.85 [67.03, 68.67]

17.03 [16.42, 17.82] 17.35 [16.73, 18.14] 17.59 [16.95, 18.38] 18.63 [17.98, 19.48] 15.47 [14.87, 16.22] 20.12 [19.46, 20.96] 21.18 [20.51, 22.03]

24.13 [23.49, 25.17] 25.00 [24.36, 26.04] 23.07 [22.39, 24.06] 23.92 [23.26, 24.87] 22.91 [22.28, 23.89] 27.29 [26.58, 28.36] 26.95 [26.26, 27.99]

EchoBridge

Ours

75.73 [75.00, 76.54]

26.79 [25.95, 27.98]

31.83 [31.15, 33.30]

experiments, the pretrained ECG representation extractor is frozen, and an identical linear multi-label classifier is trained with 1%, 10%, or 100% of the available labels. All methods share patient-level partitions, ECG preprocessing, label subsets, metrics, and validationbased threshold selection. We use official implementations when available; otherwise, we reproduce methods under matched backbone, optimization, and training-budget settings where applicable. Appendix F.4 further compares EchoBridge with task-specific ResNet-18 classifiers trained end-to-end on EchoNext-Mini. EchoBridge was implemented in PyTorch and trained on one NVIDIA A100 GPU using a one-dimensional ResNet-18 ECG encoder and MedCPT-Article-Encoder text encoder [19]. AdamW optimization uses a learning rate of 1 × 10−5 , weight decay of 1 × 10−8 , batch size 64, and up to 15 epochs. Further architecture and optimization details are provided in the appendix.

3.4

Evaluation Protocols and Metrics

We evaluate pretrained ECG representations under four protocols. Prompt-based classifier-free inference predicts pretrainingseen findings by comparing frozen ECG representations with fixed definition-only textual prototypes. In-domain frozen linear probing trains a linear multi-label classifier on EchoNext-Mini using 1%, 10%, or 100% of the available labels. Target-domain cross-center probing freezes the EchoNext-Mini-pretrained extractor and trains a new linear classifier on PKUPH or SHTMU with the same label ratios. Source-only transfer trains the classifier and selects thresholds exclusively on EchoNext-Mini, then evaluates findings shared with each external cohort. All experiments use patient-disjoint test sets and report macro-averaged AUROC, AUPRC, and F1 as percentage values. Class-specific thresholds maximize validation-set F1 and remain fixed for testing; source-only transfer retains EchoNext-Mini thresholds. We estimate 95% confidence intervals using 1,000 nonparametric bootstrap repetitions. Appendix E.1 provides detailed protocols and evaluation procedures.

4 Results 4.1 Prompt-Based Classifier-Free Inference Table 1 evaluates whether the frozen shared space supports classifierfree prediction using fixed definition-only textual prototypes. EchoBridge achieves 75.73 AUROC, 26.79 AUPRC, and 31.83 F1, exceeding the strongest competing method for each metric by 7.88, 5.61, and 4.54 percentage points, respectively. These gains indicate improved score ranking and positive-finding discrimination

KDD ’27, August, 2027, San Jose, CA, USA

Fang et al.

Table 2: Frozen linear-probing performance on EchoNext-Mini using 1%, 10%, and 100% of the available training labels. Each entry reports the macro-averaged metric value with its 95% bootstrap confidence interval in brackets. Methods

1% Linear Probing

Ref. AUROC

10% Linear Probing

100% Linear Probing

AUPRC

F1

AUROC

AUPRC

F1

AUROC

AUPRC

F1

ECG-only Self-Supervised Learning SimCLR [6] ICML’20 60.38 [59.42, 61.39] ST-MEM [29] ICLR’24 62.44 [61.41, 63.40] HeartLang [17] ICLR’25 60.15 [59.17, 61.18]

13.78 [13.20, 14.58] 17.09 [16.47, 17.86] 16.98 [16.39, 17.79]

20.55 [19.85, 21.62] 22.74 [22.00, 23.78] 21.86 [21.15, 22.94]

62.19 [61.33, 63.10] 68.42 [67.49, 69.28] 70.78 [69.90, 71.71]

16.81 [16.15, 17.75] 19.92 [19.22, 20.83] 22.58 [21.91, 23.53]

23.17 [22.43, 24.37] 25.92 [25.14, 27.09] 27.41 [26.66, 28.62]

68.70 [67.92, 69.53] 71.16 [70.31, 71.94] 74.59 [73.79, 75.44]

19.96 [19.10, 21.06] 21.72 [20.82, 22.79] 25.37 [24.50, 26.48]

25.65 [24.85, 27.02] 27.57 [26.73, 28.91] 30.69 [29.88, 32.07]

ECG–Text Pretraining CLIP [33] ICML’21 60.11 [59.06, 61.09] SigLIP [39] ICCV’23 61.59 [60.60, 62.53] PCME++ [8] ICLR’24 57.08 [56.04, 58.12] MERL-ECHO [36]medRxiv’25 61.90 [60.95, 62.87] ECG-CLIP [40] npj DM’25 61.56 [60.54, 62.58] D-BETA [16] ICML’25 68.99 [67.99, 69.94] SGERA [5] ICML’26 68.76 [67.79, 69.76]

14.30 [13.67, 15.08] 15.35 [14.76, 16.11] 12.74 [12.12, 13.56] 16.04 [15.47, 16.81] 17.71 [17.10, 18.51] 23.02 [22.42, 23.78] 24.49 [23.91, 25.28]

22.12 [21.37, 23.17] 22.63 [21.92, 23.66] 20.04 [19.30, 21.13] 22.80 [22.11, 23.84] 23.10 [22.37, 24.17] 28.08 [27.36, 29.11] 27.73 [27.03, 28.79]

64.99 [64.04, 65.87] 66.24 [65.35, 67.08] 63.04 [62.10, 63.98] 66.36 [65.51, 67.23] 69.57 [68.65, 70.49] 75.93 [75.03, 76.78] 74.97 [74.10, 75.87]

18.97 [18.26, 19.89] 19.43 [18.76, 20.33] 17.62 [16.92, 18.58] 19.40 [18.75, 20.31] 21.80 [21.11, 22.74] 26.84 [26.16, 27.74] 27.69 [27.03, 28.62]

23.65 [22.86, 24.83] 23.91 [23.16, 25.07] 22.77 [21.99, 23.99] 25.01 [24.28, 26.18] 27.01 [26.24, 28.21] 30.41 [29.65, 31.57] 31.74 [31.00, 32.93]

70.46 [69.59, 71.26] 70.97 [70.16, 71.73] 69.98 [69.12, 70.84] 71.90 [71.13, 72.69] 73.89 [73.05, 74.73] 76.43 [75.61, 77.20] 77.19 [76.40, 78.01]

22.14 [21.23, 23.22] 21.87 [21.00, 22.93] 20.95 [20.05, 22.07] 23.32 [22.47, 24.39] 24.27 [23.38, 25.37] 29.13 [28.25, 30.19] 30.25 [29.39, 31.34]

26.86 [26.01, 28.21] 27.11 [26.30, 28.44] 26.09 [25.25, 27.48] 29.16 [28.37, 30.50] 29.86 [29.03, 31.23] 32.82 [32.00, 34.15] 33.57 [32.77, 34.93]

EchoBridge

Ours

72.77 [71.84, 73.79] 26.53 [25.67, 27.56] 31.32 [30.49, 32.59] 76.94 [76.10, 77.84] 30.45 [29.55, 31.52] 34.13 [33.18, 35.52] 78.79 [77.98, 79.55] 32.66 [31.65, 33.88] 35.82 [35.02, 37.36]

Table 3: Target-domain cross-center frozen linear-probing performance on PKUPH using 1%, 10%, and 100% of the available target-cohort training labels. All ECG representation extractors are pretrained on EchoNext-Mini and frozen; only a linear multi-label classifier is trained on PKUPH. Methods

1% Linear Probing

Ref. AUROC

10% Linear Probing

100% Linear Probing

AUPRC

F1

AUROC

AUPRC

F1

AUROC

AUPRC

F1

ECG-only Self-Supervised Learning SimCLR [6] ICML’20 55.34 [52.99, 57.57] ST-MEM [29] ICLR’24 58.12 [55.73, 60.46] HeartLang [17] ICLR’25 56.29 [53.92, 58.64]

11.92 [11.56, 12.44] 12.54 [12.03, 13.26] 11.99 [11.64, 12.67]

17.92 [17.43, 18.85] 18.73 [18.06, 20.42] 18.18 [17.64, 20.38]

59.47 [56.94, 61.89] 69.84 [67.41, 72.18] 71.18 [68.85, 73.40]

12.92 [12.52, 13.47] 16.48 [15.57, 18.07] 17.25 [16.45, 18.95]

18.29 [17.87, 19.35] 23.46 [22.31, 26.05] 24.50 [23.51, 27.64]

72.61 [70.40, 74.89] 76.84 [74.72, 78.93] 76.95 [74.88, 78.86]

18.51 [17.40, 20.57] 22.91 [21.34, 24.76] 22.90 [21.28, 25.94]

27.30 [25.90, 30.01] 30.92 [29.36, 33.18] 31.31 [30.80, 36.02]

ECG–Text Pretraining CLIP [33] ICML’21 65.44 [62.74, 68.11] SigLIP [39] ICCV’23 68.51 [66.26, 71.03] PCME++ [8] ICLR’24 63.29 [60.39, 66.07] MERL-ECHO [36]medRxiv’25 67.79 [64.92, 70.24] ECG-CLIP [40] npj DM’25 68.47 [65.83, 70.94] D-BETA [16] ICML’25 69.63 [67.06, 71.84] SGERA [5] ICML’26 70.91 [68.42, 72.47]

16.27 [14.93, 18.52] 17.70 [16.29, 20.13] 14.61 [13.77, 16.17] 15.83 [15.02, 17.37] 18.34 [16.72, 20.05] 18.91 [17.23, 20.19] 18.62 [16.91, 20.10]

25.07 [23.38, 28.33] 26.65 [25.07, 30.34] 22.55 [21.42, 25.20] 24.21 [22.90, 26.96] 28.05 [26.41, 29.20] 27.42 [25.68, 29.08] 27.88 [26.03, 29.16]

68.92 [66.26, 71.35] 68.80 [66.02, 71.56] 66.98 [64.51, 69.54] 71.55 [68.85, 73.80] 72.66 [70.13, 74.86] 73.58 [71.12, 75.18] 73.21 [70.75, 75.12]

18.24 [16.78, 20.61] 18.97 [17.45, 21.49] 15.73 [14.90, 17.33] 17.54 [16.71, 19.17] 19.72 [18.03, 21.24] 20.10 [18.39, 21.52] 20.51 [18.72, 21.61]

26.11 [24.68, 29.63] 27.54 [25.98, 31.13] 22.90 [21.82, 25.44] 25.08 [23.87, 28.18] 29.18 [27.38, 30.42] 28.76 [26.91, 30.31] 28.96 [27.02, 30.37]

74.53 [72.14, 76.92] 73.01 [70.68, 75.49] 75.06 [73.17, 76.86] 77.88 [76.88, 80.71] 78.02 [76.01, 79.12] 77.48 [75.41, 79.03] 77.66 [75.62, 79.17]

21.25 [19.64, 23.73] 20.18 [18.65, 22.67] 18.77 [17.63, 20.75] 21.24 [20.11, 23.29] 23.62 [22.01, 25.12] 23.84 [22.21, 25.18] 24.11 [22.47, 25.29]

28.87 [27.38, 32.18] 28.56 [26.67, 32.17] 26.39 [24.89, 29.49] 29.90 [28.54, 32.94] 31.76 [30.18, 33.16] 32.06 [30.45, 33.22] 31.84 [30.26, 33.19]

EchoBridge

Ours

72.68 [70.26, 75.00] 20.32 [18.31, 23.88] 29.31 [27.44, 33.90] 75.42 [72.92, 77.72] 21.79 [19.82, 25.28] 30.60 [28.55, 35.02] 79.48 [77.32, 81.46] 25.43 [23.70, 28.62] 33.40 [31.86, 37.62]

under imbalanced multi-label evaluation. Because all evaluated findings are incorporated during pretraining through standardized conclusion-style summaries and class-specific prototype supervision, this protocol measures the accessibility of pretraining-seen finding semantics through textual prototypes. The results support stronger correspondence between ECG representations and predefined echocardiography-derived findings.

4.2

In-Domain Frozen Linear Probing under Different Label Ratios

Table 2 evaluates frozen ECG representations on EchoNext-Mini using linear multi-label classifiers trained with 1%, 10%, or 100% of the available labels. EchoBridge achieves the highest AUROC, AUPRC, and F1 among the evaluated ECG-only and ECG–echocardiography text methods at every budget. With 1% labels, it obtains 72.77 AUROC, 26.53 AUPRC, and 31.32 F1, exceeding the strongest baseline for each metric by 3.78, 2.04, and 3.24 points, respectively. These gains indicate that finding-related information remains linearly accessible under limited supervision. At 10%, the corresponding improvements are 1.01, 2.76, and 2.39 points; at 100%, they are

1.60, 2.41, and 2.25 points. Consistent AUPRC and F1 gains indicate improved positive-finding discrimination under class imbalance. ECG–summary alignment with class-aware prototype supervision therefore yields frozen representations that remain linearly accessible across in-domain label budgets.

4.3

Target-Domain Cross-Center Frozen Linear Probing under Different Label Ratios

Tables 3 and 4 evaluate cross-center transferability of ECG representations pretrained on EchoNext-Mini. For each method, the ECG representation extractor, including the encoder and output projection where applicable, is frozen, and a new linear multi-label classifier is trained independently on PKUPH or SHTMU using 1%, 10%, or 100% of the target-cohort labels. Model selection and class-specific thresholds use only the corresponding validation split. This protocol measures linear accessibility under institutional and label-distribution shifts across target-domain supervision levels. On PKUPH, EchoBridge achieves the highest point estimates for all metrics at every label ratio. Relative to the strongest baseline for each metric, it improves AUROC, AUPRC, and F1 by 1.77, 1.41,

EchoBridge: Long-Tail-Aware ECG–Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings

KDD ’27, August, 2027, San Jose, CA, USA

Table 4: Target-domain cross-center frozen linear-probing performance on SHTMU using 1%, 10%, and 100% of the available target-cohort training labels. All ECG representation extractors are pretrained on EchoNext-Mini and frozen; only a linear multi-label classifier is trained on SHTMU. Methods

1% Linear Probing

Ref. AUROC

10% Linear Probing

100% Linear Probing

AUPRC

F1

AUROC

AUPRC

F1

AUROC

AUPRC

F1

ECG-only Self-Supervised Learning SimCLR [6] ICML’20 56.20 [53.85, 58.46] ST-MEM [29] ICLR’24 57.02 [54.61, 59.31] HeartLang [17] ICLR’25 55.41 [52.94, 57.59]

15.95 [15.46, 16.78] 16.71 [15.96, 17.83] 16.38 [15.62, 17.49]

22.93 [22.20, 24.50] 23.86 [22.98, 25.96] 23.20 [22.46, 25.09]

63.43 [61.00, 65.61] 65.91 [63.53, 68.17] 67.48 [65.52, 69.65]

18.85 [18.20, 19.96] 21.38 [20.46, 22.86] 23.07 [22.10, 24.54]

25.05 [24.57, 26.96] 27.44 [26.52, 29.72] 28.56 [27.88, 30.34]

69.91 [67.50, 72.05] 73.88 [71.81, 75.81] 73.39 [71.60, 75.29]

23.48 [22.62, 24.99] 25.92 [24.88, 27.65] 25.12 [24.15, 26.89]

29.33 [28.42, 31.36] 31.68 [30.61, 32.87] 31.10 [30.28, 33.56]

ECG–Text Pretraining CLIP [33] ICML’21 59.06 [56.56, 61.72] SigLIP [39] ICCV’23 60.55 [58.06, 63.07] PCME++ [8] ICLR’24 60.01 [57.56, 62.44] MERL-ECHO [36]medRxiv’25 64.11 [61.47, 66.72] ECG-CLIP [40] npj DM’25 65.22 [62.69, 67.68] D-BETA [16] ICML’25 66.43 [63.91, 68.81] SGERA [5] ICML’26 67.74 [65.36, 69.12]

17.95 [17.09, 19.80] 18.19 [17.25, 20.23] 17.28 [16.57, 18.36] 18.58 [17.75, 20.03] 20.04 [19.01, 21.82] 21.28 [20.16, 22.73] 21.05 [19.98, 22.64]

25.28 [24.09, 27.98] 26.22 [24.71, 28.97] 24.47 [23.66, 26.50] 26.34 [25.28, 28.80] 28.35 [26.93, 30.42] 28.12 [26.74, 30.32] 29.21 [27.85, 30.57]

64.25 [61.72, 66.79] 63.23 [60.78, 65.62] 61.80 [59.37, 64.32] 66.56 [63.89, 69.09] 67.88 [65.41, 70.22] 69.46 [67.08, 71.68] 70.33 [68.05, 71.72]

19.93 [19.10, 21.65] 19.60 [18.74, 21.18] 18.33 [17.61, 19.51] 20.26 [19.34, 21.74] 22.85 [21.76, 24.21] 22.77 [21.72, 24.31] 22.58 [21.49, 24.16]

26.66 [25.67, 29.20] 26.80 [25.29, 29.60] 25.18 [24.35, 27.57] 27.57 [26.54, 29.94] 28.63 [27.51, 30.54] 29.42 [28.23, 30.76] 29.40 [28.26, 30.98]

69.11 [66.98, 71.26] 66.47 [64.22, 68.74] 68.58 [66.46, 70.63] 71.79 [69.61, 73.89] 73.92 [71.83, 75.72] 73.68 [71.59, 75.52] 74.37 [72.42, 75.78]

24.04 [22.88, 26.12] 22.11 [20.96, 24.26] 21.90 [21.00, 23.37] 23.73 [22.73, 25.36] 25.47 [24.34, 27.23] 26.72 [25.58, 27.96] 26.41 [25.31, 27.82]

29.97 [28.57, 32.79] 28.58 [26.93, 31.81] 27.97 [27.00, 30.52] 30.22 [29.20, 32.86] 31.72 [30.54, 32.79] 31.36 [30.25, 32.78] 31.60 [30.53, 32.79]

EchoBridge

Ours

69.40 [67.03, 71.57] 23.06 [22.07, 24.48] 30.78 [29.82, 33.10] 72.06 [69.91, 74.15] 24.76 [23.73, 26.24] 31.00 [30.07, 33.00] 75.99 [74.02, 77.85] 28.11 [26.96, 29.89] 32.98 [32.25, 35.53]

Table 5: Source-only cross-center transfer from EchoNext-Mini to PKUPH and SHTMU. The ECG representation extractor and linear multi-label classifier are trained on EchoNext-Mini, and all model parameters and class-specific decision thresholds are fixed during target-cohort evaluation. Methods

PKUPH

Ref.

SHTMU

AUROC

AUPRC

F1

AUROC

AUPRC

F1

ECG-only Self-Supervised Learning SimCLR [6] ICML’20 ST-MEM [29] ICLR’24 HeartLang [17] ICLR’25

70.78 [69.96, 71.61] 73.30 [72.51, 74.08] 73.22 [72.43, 74.01]

13.46 [12.85, 14.19] 14.47 [13.84, 15.23] 14.59 [13.95, 15.36]

22.26 [21.48, 23.31] 25.23 [24.42, 26.31] 26.50 [25.67, 27.61]

70.28 [69.39, 71.16] 70.43 [69.55, 71.32] 71.44 [70.57, 72.31]

17.33 [16.64, 18.18] 17.86 [17.18, 18.72] 18.76 [18.05, 19.64]

24.21 [23.34, 25.38] 24.10 [23.24, 25.27] 24.52 [23.63, 25.70]

ECG–Text Pretraining CLIP [33] ICML’21 SigLIP [39] ICCV’23 PCME++ [8] ICLR’24 MERL-ECHO [36] medRxiv’25 ECG-CLIP [40] npj DM’25 D-BETA [16] ICML’25 SGERA [5] ICML’26

70.42 [69.57, 71.25] 66.53 [65.64, 67.42] 66.92 [66.04, 67.81] 76.30 [75.55, 77.05] 74.09 [73.31, 74.87] 76.78 [76.04, 77.52] 72.36 [71.56, 73.16]

16.59 [15.92, 17.40] 13.98 [13.34, 14.76] 14.59 [13.94, 15.38] 16.54 [15.87, 17.36] 14.60 [13.96, 15.38] 17.45 [16.77, 18.27] 15.94 [15.28, 16.74]

25.67 [24.84, 26.77] 22.03 [21.23, 23.12] 23.58 [22.77, 24.68] 26.40 [25.58, 27.50] 24.39 [23.58, 25.48] 27.13 [26.31, 28.21] 25.82 [25.00, 26.91]

67.24 [66.31, 68.17] 60.62 [59.61, 61.62] 64.51 [63.55, 65.47] 72.47 [71.62, 73.32] 72.41 [71.56, 73.27] 73.58 [72.75, 74.41] 69.83 [68.92, 70.74]

19.09 [18.35, 19.98] 16.20 [15.50, 17.08] 15.73 [15.02, 16.61] 18.75 [18.04, 19.62] 19.20 [18.47, 20.10] 20.00 [19.26, 20.90] 18.31 [17.59, 19.19]

23.68 [22.78, 24.87] 22.11 [21.20, 23.32] 21.69 [20.80, 22.88] 25.69 [24.80, 26.86] 24.74 [23.86, 25.91] 25.35 [24.46, 26.52] 24.69 [23.80, 25.87]

EchoBridge

77.92 [77.20, 78.64]

18.98 [18.28, 19.82]

27.39 [26.57, 28.48]

74.11 [73.29, 74.93]

22.08 [21.31, 23.01]

28.16 [27.24, 29.36]

Ours

and 1.26 points with 1% labels; 1.84, 1.28, and 1.42 points with 10%; and 1.46, 1.32, and 1.34 points with 100%, respectively. The gains at 1% and 10% indicate adaptability under limited target-center supervision. On SHTMU, EchoBridge also achieves the highest point estimates across all metric–budget combinations. Its AUROC, AUPRC, and F1 gains are 1.66, 1.78, and 1.57 points with 1% labels; 1.73, 1.69, and 1.58 points with 10%; and 1.62, 1.39, and 1.26 points with 100%, respectively. Consistent gains across both cohorts support target-domain adaptability and label efficiency under institutional and prevalence shifts.

4.4

Source-Only Cross-Center Transfer

Table 5 evaluates source-only cross-center transfer by training the linear classifier and selecting class-specific thresholds exclusively on EchoNext-Mini, then applying the fixed representation extractor, classifier, and thresholds to findings shared with each target cohort. EchoBridge exceeds the strongest competing method for each metric by 1.14 AUROC, 1.53 AUPRC, and 0.26 F1 points on PKUPH, and by 0.53, 2.08, and 2.47 points on SHTMU, respectively. On PKUPH,

the larger AUROC and AUPRC gains relative to F1 suggest improved score ranking with limited change under the transferred sourcedomain thresholds. On SHTMU, the larger AUPRC and F1 gains indicate improved positive-finding discrimination under the shifted prevalence distribution. These results complement target-domain linear probing and demonstrate direct cross-center transfer without target-domain training, calibration, or threshold adjustment.

4.5

Ablation Study

Table 6 reports a matched label-only control and ablations of CSPP and the APBC refinements FAAM and SRR. All ECG–text variants combine bidirectional alignment with class-prototype supervision, using Align.+Proto. as the reference. The label-only BCE control tests whether structured-label supervision alone explains the gains. It uses the same ECG encoder, 256-dimensional projection, data split, optimizer, batch size, epoch budget, and schedule as the ECG– text variants, with pretraining driven solely by structured multilabel BCE. Its encoder and projection are then frozen, and a new linear classifier is trained with 1%, 10%, or 100% of the probe labels.

KDD ’27, August, 2027, San Jose, CA, USA

Fang et al.

Table 6: Matched label-only control and component-wise ablation of CSPP, FAAM, and SRR on EchoNext-Mini. Performance is evaluated using prompt-based classifier-free inference and frozen linear probing.

Label-only BCE

×

Align. + Proto. + CSPP + APBC + CSPP + FAAM + CSPP + SRR EchoBridge

APBC

Prompt-based

FAAM

SRR

AUROC

×

×

×

✓ ✓ ✓ ✓ ✓

× ✓ × ✓ ✓

× × ✓ ✓ ×

× × ✓ × ✓

67.18 71.05 73.36 74.02 73.91

18.83 22.39 24.71 25.18 25.04

75.73

26.79

1% Linear Probing

AUPRC

F1

F1

AUROC

AUPRC

F1

AUROC

AUPRC

F1

64.79

18.06

24.53

73.79

26.16

31.07

76.87

29.89

33.96

24.09 27.74 30.02 30.46 30.31

68.48 70.08 70.88 71.69 71.43

21.90 23.58 24.51 25.40 25.12

28.00 29.27 29.94 30.52 30.28

74.11 75.16 75.69 76.23 76.06

27.20 28.36 29.03 29.66 29.47

31.80 32.70 33.18 33.58 33.40

76.61 77.42 77.83 78.24 78.11

30.18 31.06 31.58 32.07 31.92

34.08 34.76 35.13 35.42 35.27

31.83

72.77

26.53

31.32

76.94

30.45

34.13

78.79

32.66

35.82

Finding-Specific Performance across Prevalence Levels

As shown in Figure 5, under 100% frozen linear probing on EchoNextMini, EchoBridge achieves the highest AUROC point estimate for all seven findings, with the largest gains for moderate-or-severe aortic stenosis and tricuspid regurgitation at 1.59 and 1.57 points, respectively. It also obtains the highest AUPRC for six findings, improving pulmonic and tricuspid regurgitation by 8.15 and 2.36 points, respectively, while its LVH AUPRC is 0.67 points lower. Appendix F reports finding-specific AUROC and AUPRC on EchoNext-Mini, PKUPH, and SHTMU under the corresponding 100% frozen linear-probing protocols. On both target cohorts, EchoBridge achieves the highest point estimates for every finding across ventricular function, chamber abnormalities, and valvular disease. These cross-institutional gains include several low-prevalence valvular abnormalities. Their varying magnitudes indicate conditionspecific benefits, while estimates based on few positive test cases require cautious interpretation under severe class imbalance.

5

100% Linear Probing

AUPRC

Prompt-based inference is reported only for ECG–text variants because it requires a text-aligned space. Align.+Proto. outperforms the label-only control under 1% and 10% probing, indicating greater linear accessibility with limited supervision, while their 100% results are comparable. CSPP improves prompt-based inference and frozen probing across all budgets, and complete APBC without CSPP improves every metric. With CSPP, both FAAM and SRR outperform CSPP alone, with slightly larger gains from FAAM. EchoBridge performs best across all protocols, supporting their complementary contributions beyond Align.+Proto. and matched label supervision.

4.6

10% Linear Probing

AUROC

Limitations and Ethical Considerations

The retrospective study covers a limited number of institutions; prospective multicenter validation across populations, acquisition systems, and clinical workflows remains necessary, particularly for low-prevalence valvular abnormalities. Performance may also vary across demographic groups, ECG devices, acquisition protocols, and care settings. All private data were de-identified and processed in access-controlled institutional environments. The PKUPH and SHTMU cohorts were retrospectively collected under approvals from the Institutional Review Boards of Peking University People’s Hospital (Approval No. 2024PHB428-001) and the Second Hospital

HeartLang

D-BETA

(a)

SGERA

Δ AUPRC (percentage points)

CSPP

90 85

AUROC (%)

ECG–Text Align. + Proto.

Configuration

80 75 70 65 60 LVSD

LVH

AR

AS

TR

MR

PR

EchoBridge

(b)

10 9 8 7 6 5 4 3 2 1 0 −1 −2

+8.15

+2.36 +2.26 +1.07

+0.94 +0.72

-0.67

LVSD

LVH

AR

AS

TR

MR

PR

Figure 5: Finding-specific performance on EchoNext-Mini under 100% frozen linear probing. (a) AUROC by method for each echocardiography-derived finding. (b) EchoBridge’s AUPRC difference from the strongest baseline for each finding in percentage points; positive values indicate gains.

of Tianjin Medical University (Approval No. KY2025K386), respectively. EchoBridge is intended for research and clinician-assisted screening, with predictions interpreted by qualified professionals alongside other clinical evidence. False-negative predictions, particularly for rare abnormalities, may delay further assessment. Clinical deployment requires site-specific validation, calibration, monitoring, regulatory review, and explicit human oversight.

6

Conclusion

We presented EchoBridge, an ECG–echocardiography text alignment framework for transferable representations of echocardiographyderived cardiac findings. It combines Complementary Shared–Private Projection and Adaptive Prototype Boundary Calibration to address cross-modal interference and long-tailed distributions. Across EchoNext-Mini and two independent cohorts, EchoBridge outperforms representative ECG-only and ECG–text baselines across classifier-free inference, in-domain and target-domain frozen probing, source-only transfer, and finding-specific evaluation.

Generative AI Usage LLMs assisted with auxiliary coding, translation, editing, and drafting definition-only prompts for echocardiography-derived findings. GPT-5.5 Thinking generated initial candidates, which senior cardiologists reviewed and standardized against predefined label definitions before testing. The authors independently developed the methodology, conducted experiments, interpreted results, and assume full responsibility for the manuscript.

EchoBridge: Long-Tail-Aware ECG–Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings

References [1] Arya Aminorroaya, Lovedeep S Dhingra, Aline F Pedroso, Sumukh Vasisht Shankar, Andreas Coppi, Akshay Khunte, Murilo Foppa, Luisa CC Brant, Sandhi M Barreto, Antonio Luiz P Ribeiro, et al. 2025. Development and multinational validation of an ensemble deep learning algorithm for detecting and predicting structural heart disease using noisy single-lead electrocardiograms. European Heart Journal-Digital Health 6, 4 (2025), 554–566. [2] Zachi I Attia, Suraj Kapa, Francisco Lopez-Jimenez, Paul M McKie, Dorothy J Ladewig, Gaurav Satam, Patricia A Pellikka, Maurice Enriquez-Sarano, Peter A Noseworthy, Thomas M Munger, et al. 2019. Screening for cardiac contractile dysfunction using an artificial intelligence–enabled electrocardiogram. Nature medicine 25, 1 (2019), 70–74. [3] Chieh-Ju Chao, Jean-Benoit Delbrouck, Mohammad Asadi, Imon Banerjee, Juan M Farina, Francesca Galasso, Ahmed K Mahmoud, Mohammed Tiseer Abbas, YuChiang Wang, Reza Arsanjani, et al. 2025. EchoGraph system for automated quality assessment of echocardiography reports. NPJ Digital Medicine (2025). [4] Jian Chen, Xiaoru Dong, Wei Wang, Shaorui Zhou, Lequan Yu, and Xiping Hu. 2025. DERI: Cross-Modal ECG Representation Learning with Deep ECG-Report Interaction. In 34th International Joint Conference on Artificial Intelligence, IJCAI 2025. International Joint Conferences on Artificial Intelligence, 4824–4832. [5] Jian Chen, Yipeng Du, Wenhao Yuan, Shuai Wang, Jinfeng Xu, Zewei Liu, Running Zhao, and Edith C. H. Ngai. 2026. SGERA: Stein-Guided ECG-Report Alignment for ECG Representation Learning. In Forty-third International Conference on Machine Learning. [6] Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In International conference on machine learning. PmLR, 1597–1607. [7] Matthew Christensen, Milos Vukadinovic, Neal Yuan, and David Ouyang. 2024. Vision–language foundation model for echocardiogram interpretation. Nature Medicine 30, 5 (2024), 1481–1488. [8] Sanghyuk Chun. 2023. Improved probabilistic image-text representations. arXiv preprint arXiv:2305.18171 (2023). [9] Lovedeep S Dhingra, Arya Aminorroaya, Veer Sangha, Aline F Pedroso, Sumukh Vasisht Shankar, Andreas Coppi, Murilo Foppa, Luisa CC Brant, Sandhi M Barreto, Antonio Luiz P Ribeiro, et al. 2025. Ensemble deep learning algorithm for structural heart disease screening using electrocardiographic images: PRESENT SHD. Journal of the American College of Cardiology 85, 12 (2025), 1302–1313. [10] Pierre Elias, Timothy J Poterucha, Vijay Rajaram, Luca Matos Moller, Victor Rodriguez, Shreyas Bhave, Rebecca T Hahn, Geoffrey Tison, Sean A Abreau, Joshua Barrios, et al. 2022. Deep learning electrocardiographic analysis for detection of left-sided valvular heart disease. Journal of the American College of Cardiology 80, 6 (2022), 613–626. [11] Xiaocheng Fang, Zhengyao Ding, Jieyi Cai, Yujie Xiao, Bo Liu, Jiarui Jin, Haoyu Wang, Guangkun Nie, Shun Huang, Ting Chen, et al. 2026. ECGFlowCMR: Pretraining with ECG-Generated Cine CMR Improves Cardiac Disease Classification and Phenotype Prediction. arXiv preprint arXiv:2601.20904 (2026). [12] Xiaocheng Fang, Jiarui Jin, Haoyu Wang, Che Liu, Jieyi Cai, Yujie Xiao, Guangkun Nie, Bo Liu, Shun Huang, Hongyan Li, et al. 2025. PPGFlowECG: Latent Rectified Flow with Cross-Modal Encoding for PPG-Guided ECG Generation and Cardiovascular Disease Detection. arXiv preprint arXiv:2509.19774 (2025). [13] Goro Fujiki, Satoshi Kodera, Naoto Setoguchi, Kengo Tanabe, Kotaro Miyaji, Shunichi Kushida, Mike Saji, Mamoru Nanasato, Hisataka Maki, Hideo Fujita, et al. 2025. Deep learning-based identification of echocardiographic abnormalities from electrocardiograms. JACC: Asia 5, 1_Part_1 (2025), 88–98. [14] Amirata Ghorbani, David Ouyang, Abubakar Abid, Bryan He, Jonathan H Chen, Robert A Harrington, David H Liang, Euan A Ashley, and James Y Zou. 2020. Deep learning interpretation of echocardiograms. NPJ digital medicine 3, 1 (2020), 10. [15] John Weston Hughes, Linyuan Jing, Joshua Finer, Dustin Hartzel, Christopher Kelsey, Aaron Long, Daniel Rocha, Jeffrey Ruhl, Timothy Poterucha, and Pierre Elias. 2026. EchoNext-mini: A dataset and baseline AI model for detecting structural heart disease from electrocardiograms. NEJM AI 3, 5 (2026), AIdbp2500516. [16] Manh Pham Hung, Aaqib Saeed, and Dong Ma. 2025. Boosting Masked ECGText Auto-Encoders as Discriminative Learners. In Forty-second International Conference on Machine Learning. [17] Jiarui Jin, Haoyu Wang, Hongyan Li, Jun Li, Jiahui Pan, and Shenda Hong. 2025. Reading your heart: Learning ecg words and sentences via pre-training ecg language model. arXiv preprint arXiv:2502.10707 (2025). [18] Jiarui Jin, Haoyu Wang, Xingliang Wu, Xiaocheng Fang, Xiang Lan, Zihan Wang, Deyun Zhang, Bo Liu, Yingying Zhang, Xian Wu, et al. 2026. ECG-R1: ProtocolGuided and Modality-Agnostic MLLM for Reliable ECG Interpretation. arXiv preprint arXiv:2602.04279 (2026). [19] Qiao Jin, Won Kim, Qingyu Chen, Donald C Comeau, Lana Yeganova, W John Wilbur, and Zhiyong Lu. 2023. Medcpt: Contrastive pre-trained transformers with large-scale pubmed search logs for zero-shot biomedical information retrieval. Bioinformatics 39, 11 (2023), btad651.

KDD ’27, August, 2027, San Jose, CA, USA

[20] Georgios A Kaissis, Marcus R Makowski, Daniel Rückert, and Rickmer F Braren. 2020. Secure, privacy-preserving and federated machine learning in medical imaging. Nature Machine Intelligence 2, 6 (2020), 305–311. [21] Elizabeth Knight, Evangelos K Oikonomou, Arya Aminorroaya, Aline F Pedroso, and Rohan Khera. 2026. Wearable-Echo-FM: an ECG echo foundation model for 1-lead electrocardiography. European Heart Journal-Digital Health 7, 4 (2026), ztag049. [22] Gloria Hyunjung Kwak, Dana Moukheiber, Mira Moukheiber, Lama Moukheiber, Sulaiman Moukheiber, Neel M Butala, Leo A Celi, and Christina W Chen. 2025. Large open access database of echocardiogram reports in intensive care unit patients. Scientific Data 12, 1 (2025), 1153. [23] Joon-Myoung Kwon, Soo Youn Lee, Ki-Hyun Jeon, Yeha Lee, Kyung-Hee Kim, Jinsik Park, Byung-Hee Oh, and Myong-Mook Lee. 2020. Deep learning–based algorithm for detecting aortic stenosis using electrocardiography. Journal of the American Heart Association 9, 7 (2020), e014717. [24] Sravan Kumar Lalam, Hari Krishna Kunderu, Shayan Ghosh, Harish Kumar, Samir Awasthi, Ashim Prasad, Francisco Lopez-Jimenez, Zachi I Attia, Samuel Asirvatham, Paul Friedman, et al. 2023. Ecg representation learning with multimodal ehr data. Transactions on Machine Learning Research (2023). [25] Jun Li, Che Liu, Sibo Cheng, Rossella Arcucci, and Shenda Hong. 2024. Frozen language model helps ecg zero-shot learning. In Medical Imaging with Deep Learning. PMLR, 402–415. [26] Che Liu, Cheng Ouyang, Zhongwei Wan, Haozhe Wang, Wenjia Bai, and Rossella Arcucci. 2025. Knowledge-enhanced multimodal ecg representation learning with arbitrary-lead inputs. arXiv preprint arXiv:2502.17900 (2025). [27] Che Liu, Zhongwei Wan, Sibo Cheng, Mi Zhang, and Rossella Arcucci. 2024. Etp: Learning transferable ecg representations via ecg-text pre-training. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 8230–8234. [28] Che Liu, Zhongwei Wan, Cheng Ouyang, Anand Shah, Wenjia Bai, and Rossella Arcucci. 2024. Zero-shot ecg classification with multimodal learning and test-time clinical knowledge enhancement. arXiv preprint arXiv:2403.06659 (2024). [29] Yeongyeon Na, Minje Park, Yunwon Tae, and Sunghoon Joo. 2024. Guiding masked representation learning to capture spatio-temporal relationship of electrocardiogram. arXiv preprint arXiv:2402.09450 (2024). [30] Guangkun Nie, Gongzheng Tang, Yujie Xiao, Jun Li, Shun Huang, Deyun Zhang, Qinghao Zhao, and Shenda Hong. 2025. Anyppg: An ecg-guided ppg foundation model trained on over 100,000 hours of recordings for holistic health profiling. arXiv preprint arXiv:2511.01747 (2025). [31] David Ouyang, Bryan He, Amirata Ghorbani, Neal Yuan, Joseph Ebinger, Curtis P Langlotz, Paul A Heidenreich, Robert A Harrington, David H Liang, Euan A Ashley, et al. 2020. Video-based AI for beat-to-beat assessment of cardiac function. Nature 580, 7802 (2020), 252–256. [32] Timothy J Poterucha, Linyuan Jing, Ramon Pimentel Ricart, Michael Adjei-Mosi, Joshua Finer, Dustin Hartzel, Christopher Kelsey, Aaron Long, Daniel Rocha, Jeffrey A Ruhl, et al. 2025. Detecting structural heart disease from electrocardiograms using AI. Nature 644, 8075 (2025), 221–230. [33] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning. PmLR, 8748–8763. [34] Alvaro E Ulloa-Cerna, Linyuan Jing, John M Pfeifer, Sushravya Raghunath, Jeffrey A Ruhl, Daniel B Rocha, Joseph B Leader, Noah Zimmerman, Greg Lee, Steven R Steinhubl, et al. 2022. rECHOmmend: an ECG-based machine learning approach for identifying patients at increased risk of undiagnosed structural heart disease detectable by echocardiography. Circulation 146, 1 (2022), 36–47. [35] Milos Vukadinovic, I-Min Chiu, Xiu Tang, Neal Yuan, Tien-Yu Chen, Paul Cheng, Debiao Li, Susan Cheng, Bryan He, and David Ouyang. 2026. Comprehensive echocardiogram evaluation with view primed vision language AI. Nature 650, 8103 (2026), 970–977. [36] Wai-Chak Wong, Che Liu, Pierre Elias, John Weston Hughes, Chun-Yu Leung, Xiao-Yan Qian, Hang-Long Li, Yuk-Ming Lau, Chao-Fan Tao, Ali Choo, et al. 2025. Contrastive Multi-modal Training with Electrocardiography and Natural Language Echocardiography Reports for Zero-shot Prediction of Structural Heart Disease. medRxiv (2025), 2025–09. [37] Xiaoxi Yao, David R Rushlow, Jonathan W Inselman, Rozalina G McCoy, Thomas D Thacher, Emma M Behnken, Matthew E Bernard, Steven L Rosas, Abdulla Akfaly, Artika Misra, et al. 2021. Artificial intelligence–enabled electrocardiograms for identification of patients with low ejection fraction: a pragmatic, randomized clinical trial. Nature medicine 27, 5 (2021), 815–819. [38] Han Yu, Peikun Guo, and Akane Sano. 2024. Ecg semantic integrator (esi): A foundation ecg model pretrained with llm-enhanced cardiological text. arXiv preprint arXiv:2405.19366 (2024). [39] Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. 2023. Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF international conference on computer vision. 11975–11986. [40] Xue Zhou, Tianhui Li, Hiromasa Hayama, Keijiro Nakamura, Shing-Hong Liu, Wenxi Chen, and Xin Zhu. 2025. Diagnosis of cardiac conditions from 12-lead

KDD ’27, August, 2027, San Jose, CA, USA

electrocardiogram through natural language supervision. npj Digital Medicine 8, 1 (2025), 697.

A

Acknowledgments

As an informal quality check, five senior cardiologists inspected a randomly sampled subset of mappings from structured cardiac findings to standardized conclusion-style summaries. Each sampled summary was presented with its corresponding finding vector. The reviewers provided qualitative feedback that the sampled summaries were broadly clinically plausible and semantically consistent with the encoded findings. This review was not designed as a formal annotation, quantitative validation, or inter-rater agreement study. The expert panel comprised Qinghao Zhao (Peking University People’s Hospital, Beijing, China), Guanyu Mu (The Second Hospital of Tianjin Medical University, Tianjin, China), Xingliang Wu (Tianjin Institute of Cardiology, Tianjin, China), Xinxin Di (The First Affiliated Hospital of USTC, Hefei, China), and Jing Zhao (The First Affiliated Hospital of Anhui Medical University, Hefei, China).

B Related Work B.1 Direct Echocardiography Supervision Echocardiography-derived supervision is widely used for ECGbased structural heart disease assessment. Early studies paired ECG and echocardiography data to detect ventricular systolic dysfunction or reduced ejection fraction [2, 37], followed by work on specific abnormalities such as aortic stenosis and left-sided valvular disease [10, 23]. These studies established supervised ECG-to-echo prediction using echocardiographic measurements or diagnoses as screening targets. More recent work expanded to broader structural heart disease assessment. rECHOmmend [34] used a composite endpoint to identify patients at risk of undiagnosed echocardiographydetectable disease, while Fujiki et al. [13] studied multi-label prediction across ventricular function, chamber structure, and valvular abnormalities. EchoNext scaled this direction with paired examinations and public resources for evaluating echo-confirmed structural heart disease detection [15, 32]. Related studies also examined ECG images and single-lead recordings [1, 9]. These methods use predefined echocardiography-derived targets, motivating compositional ECG–echocardiography alignment organized around predefined findings and their co-occurrence patterns.

B.2

ECG–Report Representation Learning

ECG–text pretraining is a major approach to language-supervised ECG representation learning. METS [25] and ETP [27] learn shared embeddings from paired ECGs and machine-generated or clinical reports for zero-shot classification and label-efficient evaluation. MERL [28] uses clinical knowledge-enhanced prompts at inference, while ECG-CLIP [40] scales CLIP-style supervision to large ECG–report datasets for zero-shot diagnosis across cardiac conditions. Multimodal ECG pretraining further combines reports, EHR records, PPG, and cardiac imaging [11, 12, 24, 30]. Recent methods introduce richer interaction and structured supervision [18]. DERI [4] combines multiple alignment objectives with mutual feature reconstruction, ESI [38] enriches reports with LLM-generated cardiological descriptions, and K-MERL [26] extracts structured knowledge from free-text reports while supporting arbitrary-lead

Fang et al.

inputs. D-BETA [16] integrates masked ECG–text autoencoding with discriminative contrastive learning, while SGERA [5] uses Stein-guided alignment to address structural and statistical modality discrepancies. These studies establish ECG–report alignment for transferable ECG representation learning. Their supervision primarily comes from ECG reports, diagnostic statements, and broader clinical records, emphasizing rhythm, conduction, waveform morphology, and general ECG diagnoses.

B.3

ECG–Echocardiography Text Alignment

Recent studies use paired ECGs and echocardiography reports as cross-modal supervision for structural cardiac representation learning. MERL-ECHO applies CLIP-style contrastive pretraining to encode 12-lead ECGs and echocardiography reports in a shared space for zero-shot structural heart disease prediction [36]. WearableEcho-FM extends this paradigm to single-lead ECGs and evaluates label-efficient fine-tuning for left ventricular systolic dysfunction, diastolic dysfunction, and composite structural heart disease [21]. These studies establish ECG–echocardiography text alignment as a viable approach for transferring structural cardiac information to ECG encoders and improving label efficiency. However, global crossmodal alignment may entangle modality-specific factors, while imbalanced finding distributions provide uneven class-level supervision. EchoBridge addresses these limitations through shared– private projection and class-aware geometric organization of the normalized representation space.

C Implementation Details of EchoBridge C.1 Algorithm of EchoBridge Algorithm 1 summarizes EchoBridge’s end-to-end training, including shared–private projection, within-modality orthogonality, bidirectional ECG–text alignment, frequency-adaptive prototype calibration, spherical Riesz repulsion, and joint optimization.

C.2

Encoder and Projection Architectures

ECG Encoder. We use a one-dimensional ResNet-18 initialized from scratch to encode each 10-second, 12-lead ECG recording with input shape 12 × 5000. The encoder follows the four-stage ResNet18 architecture, with all two-dimensional operations replaced by one-dimensional counterparts. We remove the max-pooling layer after the stem convolution to preserve temporal resolution. The stem comprises a bias-free Conv1D(12, 64, 7, 2, 3) layer followed by batch normalization and ReLU. Each basic residual block contains two 3-tap convolutions: Conv1D(𝑘 = 3) → BN → ReLU → Conv1D(𝑘 = 3) → BN . (19) The first block of each stage except the first uses stride 2 to reduce temporal resolution and increase channels. Dimension-changing shortcuts use a bias-free 1 × 1 convolution followed by batch normalization. A ReLU is applied after adding the residual and shortcut paths. Adaptive average pooling produces the 512-dimensional ECG representation ℎ𝑒 . Table 7 summarizes the architecture. Text Encoder. We use the pretrained MedCPT-Article-Encoder [19], a BERT-base model with 12 Transformer layers, a hidden size of 768, 12 self-attention heads, and a 3,072-dimensional feed-forward layer.

EchoBridge: Long-Tail-Aware ECG–Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings

KDD ’27, August, 2027, San Jose, CA, USA

Algorithm 1 End-to-End Training of EchoBridge 𝐶 ; alignment Require: Labeled paired data D = { (𝑥𝑖 , 𝑡𝑖 , y𝑖 ) }; training epochs 𝐸; number of cardiac findings 𝐶; training-set positive rates r = {𝑟𝑐 }𝑐=1 temperature 𝜏; margin parameters (𝑚 0 , 𝑚 min , 𝑚 max ); numerical constant 𝜖; Riesz exponent 𝑞; repulsion weight 𝜆𝑟 . Ensure: Optimized model parameters Θ. 1: Initialize the learnable prototype matrix 𝑃 = [𝑝 1 , . . . , 𝑝𝐶 ] ⊤ ∈ R𝐶 ×𝑑 and learnable logit-scale parameter 𝑠. Í median(r) 2: Compute the normalization factor 𝜅 ← 𝐶1 𝐶 . 𝑐=1 𝑟𝑐 +𝜖   median(r) 3: Compute the fixed class-specific margins 𝑚𝑐 ← clip 𝑚 0 𝑟 +𝜖 𝜅1 , 𝑚 min , 𝑚 max for 𝑐 = 1, . . . , 𝐶. 𝑐 4: for 𝑒 = 1, . . . , 𝐸 do 𝐵 ∼ D do 5: for each mini-batch { (𝑥𝑖 , 𝑡𝑖 , y𝑖 ) }𝑖=1 6: Encode the paired inputs: ℎ𝑒,𝑖 ← 𝑓𝑒 (𝑥𝑖 ) and ℎ𝑡,𝑖 ← 𝑓𝑡 (𝑡𝑖 ). 𝑠 ← 𝜙 𝑠 (ℎ ) and 𝑧 𝑝 ← 𝜙 𝑝 (ℎ ). 7: Obtain the shared and private ECG projections: 𝑧𝑒,𝑖 𝑒,𝑖 𝑒 𝑒 𝑒,𝑖 𝑒,𝑖 𝑝 𝑝 𝑠 8: Obtain the shared and private text projections: 𝑧𝑡,𝑖 ← 𝜙𝑡𝑠 (ℎ𝑡,𝑖 ) and 𝑧𝑡,𝑖 ← 𝜙𝑡 (ℎ𝑡,𝑖 ). 9: Construct normalized shared and private copies and compute the within-modality orthogonality loss Lorth using Eq. (4). 𝑠 ← 𝑧 𝑠 /∥𝑧 𝑠 ∥ , 𝑧ˆ𝑠 ← 𝑧 𝑠 /∥𝑧 𝑠 ∥ , and 𝑝ˆ ← 𝑝 /∥𝑝 ∥ . 10: Normalize the shared representations and class prototypes: 𝑧ˆ𝑒,𝑖 𝑐 𝑐 𝑐 2 𝑒,𝑖 𝑒,𝑖 2 𝑡,𝑖 𝑡,𝑖 𝑡,𝑖 2 11: Compute the cross-modal similarity matrix and symmetric alignment loss Lalign using Eqs. (6)–(7). 𝑒 = ⟨𝑧ˆ𝑠 , 𝑝ˆ ⟩ and 𝑢 𝑡 = ⟨𝑧ˆ𝑠 , 𝑝ˆ ⟩ using Eq. (10). 12: Compute ECG–prototype and text–prototype cosine similarities 𝑢𝑖,𝑐 𝑒,𝑖 𝑐 𝑖,𝑐 𝑡,𝑖 𝑐 𝑒 and ℓ˜𝑡 by applying 𝑚 only to positive sample–prototype pairs according to Eq. (14). 13: Set 𝛾 ← exp(𝑠 ) and compute the margin-adjusted logits ℓ˜𝑖,𝑐 𝑐 𝑖,𝑐 14: Compute the multi-label prototype loss Lproto ← BCE( ℓ˜𝑒 , y) + BCE( ℓ˜𝑡 , y) using Eq. (15). 15: Compute the spherical Riesz repulsion loss Lriesz over all normalized class-prototype pairs using Eq. (16). 16: Compute Lapbc ← Lproto + 𝜆𝑟 Lriesz and L ← Lalign + Lorth + Lapbc . 𝑝 𝑝 17: Update 𝑓𝑒 , 𝑓𝑡 , 𝜙𝑒𝑠 , 𝜙𝑒 , 𝜙𝑡𝑠 , 𝜙𝑡 , 𝑃 , and 𝑠 using AdamW and ∇Θ L. 18: end for 19: end for 20: return ECG encoder 𝑓𝑒 and shared ECG projection head 𝜙𝑒𝑠 .

Table 7: Layer-wise configuration of the one-dimensional ResNet-18 ECG encoder for an input ECG of shape 12 × 5000. Each residual stage contains two basic blocks, and the reported stride applies to the first block of each stage. Component Stem convolution Residual stage 1 Residual stage 2 Residual stage 3 Residual stage 4 Adaptive average pooling

Blocks

Input channels

Output channels

Kernel size

First-block stride

Temporal length

1 2 2 2 2 1

12 64 64 128 256 512

64 64 128 256 512 512

7 3 3 3 3 –

2 1 2 2 2 –

2500 2500 1250 625 313 1

It uses GELU, 0.1 dropout, a 30,522-token vocabulary, and a maximum sequence length of 512. The 768-dimensional pooler_output serves as the global text representation ℎ𝑡 . All parameters are jointly fine-tuned during cross-modal pretraining. Shared and Private Projections. Four independent two-layer heads map the ECG and text representations into 256-dimensional shared and auxiliary private projections:

𝑝

𝜙𝑒𝑠 , 𝜙𝑒 : Linear(512, 256) → GELU → Linear(256, 256),

(20)

𝑝 𝜙𝑡𝑠 , 𝜙𝑡 : Linear(768, 256) → GELU → Linear(256, 256).

(21)

Shared projections are ℓ2 -normalized before cross-modal alignment and prototype supervision. Private projections remain unnormalized, while normalized copies of both branches compute the withinmodality orthogonality loss.

D

Additional Dataset Details

EchoNext-Mini labels follow the published structured definitions of the original dataset. For PKUPH and SHTMU, echocardiographyderived findings were extracted from physician-authored reports using regular-expression-based natural language processing pipelines. The pipelines normalized synonymous clinical terms, handled negation and uncertainty expressions, parsed severity modifiers, and mapped report statements to predefined binary findings. Senior cardiologist Qinghao Zhao subsequently reviewed all extracted labels from both cohorts for clinical consistency.

D.1

EchoNext-Mini Dataset

EchoNext-Mini [15] is a de-identified public subset of EchoNext derived from routine clinical data at Columbia University Irving Medical Center, including Columbia and Allen hospitals. It contains 100,000 10-second, 12-lead ECGs from 36,286 adults aged at least 18 years. Recordings were acquired between 2008 and 2022 at 250 Hz and linked to transthoracic echocardiograms performed within one year. Each record includes demographic, acquisition, and automated

KDD ’27, August, 2027, San Jose, CA, USA

ECG measurement metadata; ages above 90 were capped at 90 for de-identification. Echocardiography-derived labels were constructed from structured report fields and quantitative measurements, including left ventricular ejection fraction (LVEF), ventricular wall thickness, pulmonary artery systolic pressure, tricuspid regurgitation velocity, valvular severity, right ventricular systolic function, and pericardial effusion. The dataset provides 11 component abnormalities and a composite moderate-or-greater structural heart disease label. Binary criteria include LVEF ≤ 45% for left ventricular systolic dysfunction, maximum septal or posterior wall thickness ≥ 13 mm for increased wall thickness, and moderate-or-severe grading for aortic stenosis and aortic, mitral, tricuspid, and pulmonic regurgitation. Positive labels require the ECG to precede the echocardiogram by at most one year. Echocardiograms with prosthetic valves, missing LVEF, or unavailable wall-thickness measurements were excluded. We evaluate seven findings: left ventricular systolic dysfunction, left ventricular hypertrophy, and moderate-or-severe aortic regurgitation, aortic stenosis, tricuspid regurgitation, mitral regurgitation, and pulmonic regurgitation. Prevalence is strongly imbalanced: left ventricular systolic dysfunction and hypertrophy each occur in approximately 24% of samples, whereas pulmonic and aortic regurgitation occur in approximately 0.8%–1.3%.

D.2

PKUPH Dataset

The PKUPH cohort was retrospectively assembled from longitudinal ECG and transthoracic echocardiography data collected at Peking University People’s Hospital from June 2015 to May 2023. Of 74,220 screened individuals, the final cohort included 20,768 patients and 27,158 ECG recordings. ECG–echocardiography pairs were matched by prioritizing same-day examinations; otherwise, the temporally closest ECG within a ±10-day window was selected. Pairs outside this window and ECGs with corrupted waveforms, incomplete leads, or inconsistent metadata were excluded. The mean age was 61.6 ± 14.4 years, and 9,193 patients (44.3%) were male. Cardiac findings were extracted from physician-authored echocardiography reports using rule-based natural language processing. The cohort includes broader diastolic-function- and wallmotion-related findings than EchoNext-Mini, including left ventricular diastolic dysfunction and wall-motion abnormality, and exhibits a distinct prevalence distribution, enabling cross-center evaluation under institutional and label-prevalence shifts.

D.3

SHTMU Dataset

The SHTMU cohort was retrospectively assembled from longitudinal ECG and transthoracic echocardiography data collected at the Second Hospital of Tianjin Medical University between January and December 2024. Among 479,089 individuals screened from the institutional repository, 462,468 lacked available echocardiography data. The resulting cohort included 16,621 patients and 18,588 ECG recordings. ECG–echocardiography pairs were constructed using the same temporal matching strategy as for PKUPH. Same-day examinations were prioritized; when no same-day ECG was available, the temporally closest ECG within ±10 days of the echocardiography report was selected. Pairs outside this window were excluded,

Fang et al.

together with ECGs containing corrupted waveforms, incomplete lead data, or inconsistent metadata. The patients had a mean age of 65.3 ± 13.7 years, and 9,654 were male. Compared with PKUPH, SHTMU showed a substantially different finding-prevalence profile, characterized by a high prevalence of left atrial enlargement and a low prevalence of findings such as right atrial enlargement. The cohort therefore supports geographically independent cross-center evaluation under concurrent institutional and label-prevalence shifts.

E Additional Experimental Details E.1 Detailed Evaluation Protocols Prompt-Based Classifier-Free Inference. We freeze the pretrained ECG encoder, shared ECG projection head, text encoder, and shared text projection head, and perform inference without training a downstream classifier. Each cardiac finding is represented by a fixed definition-only prompt, which the text branch encodes into an ℓ2 -normalized textual prototype. Each ECG is mapped into the normalized shared space, and its cosine similarity with each textual prototype serves as the corresponding class-wise prediction score. Class-specific thresholds are selected on the EchoNext-Mini validation set and fixed before test evaluation. Because all evaluated findings are incorporated during pretraining through standardized summaries and class-prototype supervision, this protocol evaluates the semantic accessibility of pretraining-seen findings. Generalization to unseen disease categories lies outside its scope. In-Domain Frozen Linear Probing. We freeze the pretrained ECG encoder and output projection and train a linear multi-label classifier on EchoNext-Mini using 1%, 10%, or 100% of the available training labels. This protocol evaluates the linear accessibility of the frozen representations across different label budgets. Label subsets are sampled using fixed random seeds shared across methods. Model selection and class-specific threshold determination use only the EchoNext-Mini validation split, and performance is reported on the independent patient-disjoint test set. All methods use identical label subsets, classifier architectures, optimization settings, and evaluation procedures. Target-Domain Cross-Center Frozen Linear Probing. We evaluate the adaptability of EchoNext-Mini-pretrained representations to the independent PKUPH and SHTMU cohorts. For each method, the pretrained ECG representation extractor, including the ECG encoder and output projection where applicable, is frozen. A new linear multi-label classifier is trained independently on each target cohort using 1%, 10%, or 100% of its available training labels. Label subsets are sampled using fixed random seeds shared across methods. Model selection and class-specific threshold determination use only the validation split of the corresponding target cohort, and performance is reported on its patient-disjoint test set. This protocol evaluates the linear accessibility of the frozen representations under institutional and label-distribution shifts across different levels of target-domain supervision. Source-Only Cross-Center Transfer. We evaluate direct crosscenter transfer without target-domain training or calibration. A

EchoBridge: Long-Tail-Aware ECG–Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings

KDD ’27, August, 2027, San Jose, CA, USA

Table 8: Definition-only prompts used to construct textual prototypes for prompt-based classifier-free inference on EchoNextMini. Each prompt describes the corresponding echocardiography-derived finding through target-defining anatomical, functional, hemodynamic, severity-related, or quantitative characteristics. Cardiac finding

Text prompt

Left ventricular systolic dysfunction

The echocardiographic examination demonstrates impaired global left ventricular contractile function, typically characterized by a left ventricular ejection fraction of 45% or lower. The echocardiographic examination demonstrates increased left ventricular myocardial wall thickness or mass, typically involving thickening of the interventricular septum, the left ventricular posterior wall, or both. The echocardiographic examination demonstrates moderate or severe diastolic regurgitant blood flow across the aortic valve from the aorta into the left ventricle, reflecting clinically significant aortic valve incompetence. The echocardiographic examination demonstrates moderate or severe narrowing and restricted opening of the aortic valve, resulting in hemodynamically significant obstruction of systolic blood flow from the left ventricle into the aorta, typically accompanied by increased transvalvular velocity or pressure gradient and reduced valve area. The echocardiographic examination demonstrates moderate or severe systolic regurgitant blood flow across the tricuspid valve from the right ventricle into the right atrium, reflecting clinically significant tricuspid valve incompetence. The echocardiographic examination demonstrates moderate or severe systolic regurgitant blood flow across the mitral valve from the left ventricle into the left atrium, reflecting clinically significant mitral valve incompetence. The echocardiographic examination demonstrates moderate or severe diastolic regurgitant blood flow across the pulmonic valve from the pulmonary artery into the right ventricle, reflecting clinically significant pulmonic valve incompetence.

Left ventricular hypertrophy Moderate or severe aortic regurgitation Moderate or severe aortic stenosis

Moderate or severe tricuspid regurgitation Moderate or severe mitral regurgitation Moderate or severe pulmonic regurgitation

Table 9: Hyperparameters used for EchoBridge pretraining.

E.2

Hyperparameter

Table 9 summarizes the optimization and alignment hyperparameters used for EchoBridge pretraining.

Frequency-adaptive angular margin Base margin 𝑚 0 Minimum margin 𝑚 min Maximum margin 𝑚 max Spherical Riesz regularization Riesz exponent 𝑞 Regularization weight 𝜆𝑟 Training configuration Learning rate Batch size Max epochs Random seed Data-loader workers Checkpoint interval

Value 0.20 rad 0.05 rad 0.50 rad 2.0 0.05 1 × 10 −5 64 15 42 8 5 epochs

linear multi-label classifier is trained on frozen ECG representations using the complete EchoNext-Mini training set, while model selection and class-specific threshold determination use only the EchoNext-Mini validation split. The ECG encoder, output projection, linear classifier, and decision thresholds are then fixed and applied unchanged to PKUPH and SHTMU. Evaluation is restricted to cardiac findings shared between EchoNext-Mini and each target cohort. Target-domain samples are excluded from representation learning, classifier training, model selection, calibration, and threshold determination.

E.3

Pretraining Hyperparameters

Definition-Only Prompt Construction for Classifier-Free Inference

For prompt-based classifier-free inference, each echocardiographyderived finding is represented by a fixed definition-only prompt encoded by the frozen text branch as an ℓ2 -normalized prototype. Because standardized pretraining summaries derive from structured labels, direct label-name prompts may create lexical overlap with training text. We instead describe each phenotype through targetdefining functional, anatomical, hemodynamic, severity-related, and quantitative characteristics, excluding etiologies, associated abnormalities, and explicit ECG manifestations. GPT-5.5 Thinking generated initial candidates, which senior cardiologists reviewed and standardized according to predefined EchoNext-Mini label definitions. The final prompts in Table 8 were fixed before testing and constructed independently of validation and test performance. During inference, normalized ECG representations and textual prototypes are compared by cosine similarity to obtain class-wise scores. This protocol evaluates the semantic accessibility of pretrainingseen findings without downstream classifier training.

E.4

Downstream Classifier Training

For all linear-probing experiments, the pretrained ECG encoder and output projection are frozen. We train a linear multi-label classifier with sigmoid outputs using binary cross-entropy with logits. Optimization uses AdamW with a learning rate of 1 × 10−3 , weight decay of 1 × 10−6 , a batch size of 128, and a maximum of 30 epochs. Gradients are restricted to the classifier parameters. All methods

KDD ’27, August, 2027, San Jose, CA, USA

Fang et al.

Table 10: Finding-specific AUROC and AUPRC on EchoNext-Mini under 100% frozen linear probing. P/N denotes the numbers of positive and negative test samples for each echocardiography-derived finding. Mod./Sev. denotes moderate-or-severe disease. LVSD

LVH

Mod./Sev. AR

Mod./Sev. AS

Mod./Sev. TR

Mod./Sev. MR

Mod./Sev. PR

P/N=4833/15167

P/N=4954/15046

P/N=268/19732

P/N=826/19174

P/N=2172/17828

P/N=1715/18285

P/N=154/19846

ECG-only Self-Supervised Learning SimCLR [6] ICML’20 75.18/50.06 ST-MEM [29] ICLR’24 78.73/54.72 HeartLang [17] ICLR’25 82.97/63.16

65.64/35.39 68.53/39.32 70.76/41.92

62.00/2.19 59.17/1.91 61.73/2.31

65.60/7.87 74.00/11.96 76.99/14.33

67.95/20.33 70.77/22.42 74.34/28.05

69.65/17.91 72.41/18.96 75.94/22.36

74.87/5.99 74.54/2.73 79.43/5.44

ECG-Text Pretraining CLIP [33] ICML’21 SigLIP [39] ICCV’23 PCME++ [8] ICLR’24 MERL-ECHO [36] medRxiv’25 ECG-CLIP [40] npj DM’25 D-BETA [16] ICML’25 SGERA [5] ICML’26

79.20/57.87 78.86/57.46 76.63/53.10 80.34/59.32 81.71/61.14 84.32/69.67 85.28/69.56

67.01/38.22 67.14/38.24 67.61/39.14 68.32/38.55 70.08/41.82 72.62/51.65 73.51/49.82

63.34/2.19 64.18/2.13 64.96/2.33 64.13/2.14 64.28/2.15 68.20/2.54 66.84/2.44

67.19/8.42 67.08/8.30 68.33/8.61 71.92/10.39 73.81/10.95 75.62/14.05 76.73/15.77

69.65/22.41 70.40/22.64 68.15/20.37 71.54/24.42 73.06/25.61 76.31/34.74 77.92/35.58

72.64/20.83 71.89/19.89 71.80/19.89 74.84/21.74 75.06/21.34 77.90/26.68 78.31/27.82

74.23/5.04 77.25/4.41 72.34/3.17 72.18/6.69 79.23/6.90 80.04/4.57 81.72/10.74

EchoBridge

86.15/70.74

74.74/50.98

69.53/3.48

78.58/16.49

79.49/37.94

79.83/30.08

83.24/18.89

Methods

Ref.

Ours

Table 11: Finding-specific AUROC and AUPRC on PKUPH under 100% target-domain cross-center frozen linear probing. P/N denotes the numbers of positive and negative test samples for each echocardiography-derived finding. Mod./Sev. denotes moderate-or-severe disease. LVSD

LVDD

LVWMA

LVH

LAE

LVE

RAE

RVE

P/N=129/5303

P/N=2150/3282

P/N=52/5380

P/N=1053/4379

P/N=1554/3878

P/N=211/5221

P/N=48/5384

P/N=35/5397

P/N=26/5406

P/N=49/5383

ECG-only Self-Supervised Learning SimCLR [6] ICML’20 84.68/19.40 63.84/53.78 90.26/13.22 60.60/29.82 59.91/37.88 77.42/19.48 72.47/2.98 71.51/1.24 ST-MEM [29] ICLR’24 88.78/30.41 65.78/57.34 91.57/17.18 63.06/31.50 62.14/40.83 81.08/26.59 77.48/4.78 79.69/9.59 HeartLang [17] ICLR’25 88.95/29.73 66.46/56.71 91.53/17.37 63.76/31.31 62.48/40.49 81.62/26.45 77.88/4.70 78.11/10.55

69.04/1.78 79.03/3.94 79.11/3.79

76.37/5.52 79.79/6.94 79.60/7.90

ECG-Text Pretraining CLIP [33] ICML’21 SigLIP [39] ICCV’23 PCME++ [8] ICLR’24 MERL-ECHO [36] medRxiv’25 ECG-CLIP [40] npj DM’25 D-BETA [16] ICML’25 SGERA [5] ICML’26

87.13/27.84 84.99/22.53 87.62/19.32 91.06/25.45 90.48/33.74 89.55/34.33 90.92/34.09

74.09/3.64 76.45/4.88 73.05/3.51 73.88/4.95 74.51/2.98 76.22/1.72 77.77/4.26 80.00/6.86 78.97/4.90 81.42/12.88 78.06/4.95 80.38/11.58 77.52/4.62 79.91/13.39

72.89/2.69 69.82/1.47 73.39/1.84 79.56/3.13 79.85/3.62 80.35/4.13 80.38/4.28

77.42/7.13 75.37/6.93 78.08/5.29 81.30/6.38 80.75/8.11 81.00/8.25 80.61/8.06

EchoBridge

92.91/37.73 67.47/58.39 93.30/19.41 65.18/32.09 63.72/41.99 84.12/29.95 79.86/5.30 82.88/16.10

82.93/4.68

82.38/8.70

Methods

Ref.

Ours

65.35/55.96 64.25/55.15 66.02/55.12 66.98/56.46 66.81/56.47 66.22/56.78 66.59/57.18

90.44/16.97 89.88/15.70 90.75/13.43 92.45/15.39 92.27/18.25 91.65/18.57 92.22/18.54

61.94/31.03 60.41/29.91 63.43/30.83 64.45/31.16 64.23/31.15 63.72/31.36 64.06/31.65

use the same frozen-representation protocol, classifier architecture, optimizer, hyperparameters, and validation-based model selection.

F Additional Results F.1 In-Domain Finding-Specific Performance Table 10 evaluates findings with positive prevalence ranging from approximately 25% to below 1%. Under 100% frozen linear probing, EchoBridge achieves the highest AUROC for all seven findings and the highest AUPRC for six, extending its gains beyond frequent LVSD and LVH. Relative to the strongest baseline for each finding, its largest AUROC gains are 1.59 and 1.57 points for moderate-orsevere AS and TR, while its largest AUPRC gain is 8.15 points for PR. These results support improved discrimination across prevalence levels, particularly for low-prevalence valvular findings.

60.51/39.50 60.16/38.93 61.29/38.22 62.65/39.63 62.68/39.63 62.61/40.55 62.49/40.84

F.2

79.08/22.86 78.29/22.72 79.29/18.95 82.58/23.68 82.74/27.45 81.26/27.90 81.90/28.45

Mod./Sev. TR Mod./Sev. MR

Cross-Center Finding-Specific Performance

Finding-Specific Performance on PKUPH. Table 11 reports finding-specific performance on PKUPH under 100% target-domain cross-center frozen linear probing. EchoBridge achieves the highest AUROC and AUPRC point estimates for all ten findings. The largest gains occur for LVSD, exceeding the strongest baseline by 1.85 AUROC and 3.40 AUPRC points, followed by LVE with gains of 1.38 and 1.50 points. Improvements also cover chamber enlargement, left ventricular wall-motion abnormality, ventricular hypertrophy, and valvular findings. Moderate-or-severe TR, RVE, and RAE contain only 26, 35, and 48 positive test cases, respectively, so their AUPRC values require cautious interpretation. The consistently higher point estimates support cross-center transfer across frequent and low-prevalence findings under institutional and label-distribution shifts.

EchoBridge: Long-Tail-Aware ECG–Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings

KDD ’27, August, 2027, San Jose, CA, USA

Table 12: Finding-specific AUROC and AUPRC on SHTMU under 100% target-domain cross-center frozen linear probing. P/N denotes the numbers of positive and negative test samples for each echocardiography-derived finding. Mod./Sev. denotes moderate-or-severe disease. LVSD

LVH

LAE

RAE

Mod./Sev. AR

Mod./Sev. TR

Mod./Sev. MR

P/N=369/3349

P/N=663/3055

P/N=2196/1522

P/N=28/3690

P/N=56/3662

P/N=121/3597

P/N=124/3594

ECG-only Self-Supervised Learning SimCLR [6] ICML’20 80.91/48.78 ST-MEM [29] ICLR’24 88.69/53.43 HeartLang [17] ICLR’25 87.71/51.76

65.13/29.08 68.44/31.20 69.79/33.81

59.61/61.76 62.32/66.96 61.83/63.46

77.99/1.09 79.35/2.30 78.85/1.32

59.52/1.12 61.04/1.43 60.01/1.41

71.92/10.14 76.52/11.63 76.63/10.61

74.29/12.39 80.80/14.49 78.91/13.47

ECG-Text Pretraining CLIP [33] ICML’21 SigLIP [39] ICCV’23 PCME++ [8] ICLR’24 MERL-ECHO [36] medRxiv’25 ECG-CLIP [40] npj DM’25 D-BETA [16] ICML’25 SGERA [5] ICML’26

79.88/49.91 74.67/46.04 79.62/46.12 84.67/49.19 87.91/52.08 86.79/54.20 89.01/53.99

64.32/30.05 61.47/27.56 63.69/27.65 66.65/28.90 68.73/30.26 67.51/33.74 69.97/32.68

58.78/63.26 56.95/58.20 58.35/58.70 60.75/61.71 63.09/66.02 62.70/68.38 63.33/69.46

77.41/2.15 76.94/2.29 76.47/0.85 78.80/1.84 79.31/2.43 78.85/2.00 78.48/1.15

58.82/1.16 58.75/1.24 58.16/1.04 60.41/1.03 60.43/1.66 60.96/1.44 59.92/1.83

71.21/9.66 67.36/8.55 70.81/8.36 73.95/10.55 78.07/11.88 78.90/12.82 79.28/11.73

73.35/12.09 69.15/10.89 72.96/10.58 77.30/12.89 79.90/13.96 80.05/14.46 80.60/14.03

EchoBridge

91.23/56.40

71.05/35.07

64.01/70.79

80.23/2.81

61.87/2.27

80.40/13.38

83.16/16.02

Methods

Ref.

Ours

Table 13: Comparison of global, shared, and private representations on EchoNext-Mini. Prompt-based classifier-free inference uses fixed definition-only textual prototypes, whereas 100% frozen linear probing trains a linear multi-label classifier using the complete labeled training set. Representation

Prompt-based

100% Linear Probing

AUROC AUPRC F1 AUROC AUPRC F1 Global representation w/o CSPP Private representation Shared representation

73.36 – 75.73

24.71 30.02 – – 26.79 31.83

77.83 75.67 78.79

31.58 35.13 26.39 31.30 32.66 35.82

Finding-Specific Performance on SHTMU. Table 12 reports finding-specific performance on SHTMU under 100% target-domain cross-center frozen linear probing. EchoBridge achieves the highest AUROC and AUPRC point estimates for all seven findings, with gains over the strongest baseline of 0.68–2.36 AUROC points and 0.38–2.20 AUPRC points. Improvements span frequent LAE and low-prevalence RAE and moderate-or-severe aortic, tricuspid, and mitral regurgitation, indicating linear accessibility across chamber, ventricular, and valvular abnormalities at a second independent institution. RAE and moderate-or-severe aortic regurgitation contain only 28 and 56 positive test cases, respectively, so these estimates require cautious interpretation under severe class imbalance.

F.3

Shared–Private Representation Analysis

Table 13 compares the global representation from the no-CSPP variant with EchoBridge’s shared and private representations. The shared representation performs best under both protocols. Relative to the global representation, it improves prompt-based AUROC, AUPRC, and F1 by 2.37, 2.08, and 1.81 points, respectively, and 100% frozen linear-probing performance by 0.96, 1.08, and 0.69 points. These gains indicate greater accessibility of echocardiographyderived finding information to textual prototypes and linear classifiers in the shared space.

Table 14: Comparison with fully supervised task-specific models on EchoNext-Mini using the complete training set. The supervised baselines are trained end-to-end, whereas EchoBridge uses a frozen pretrained ECG encoder and shared projection with a linear multi-label classifier. Methods

AUPRC

F1

ResNet-18 + BCE 77.95 [77.25, 78.70] ResNet-18 + ASL 79.46 [78.73, 80.17] ResNet-18 + Cosine BCE 79.78 [79.04, 80.48]

AUROC

30.77 [29.67, 32.02] 32.75 [31.61, 34.11] 33.46 [32.39, 34.72]

34.66 [33.76, 36.14] 36.39 [35.51, 37.94] 36.60 [35.77, 38.01]

EchoBridge

32.66 [31.65, 33.88]

35.82 [35.02, 37.36]

78.79 [77.98, 79.55]

The private representation remains predictive under linear probing, showing that the auxiliary branch retains task-relevant ECG information. The shared branch exceeds it by 3.12 AUROC, 6.27 AUPRC, and 4.52 F1 points, with larger AUPRC and F1 gaps indicating better positive-finding discrimination under class imbalance. Prompt-based inference is evaluated only in the shared space because the private branch is unaligned with textual prototypes. These results support the complementary roles of both branches, while characterizing private-branch information requires further modality-specific analysis.

F.4

Comparison with Fully Supervised Models

Table 14 compares EchoBridge’s frozen representation with taskspecific ECG classifiers trained end-to-end using all EchoNext-Mini training labels. All supervised baselines use the same ResNet-18 backbone and differ only in their classification heads or training objectives. ResNet-18 + BCE and ResNet-18 + ASL use linear heads optimized with binary cross-entropy and asymmetric loss, respectively. ResNet-18 + Cosine BCE uses an ℓ2 -normalized cosine classifier with binary cross-entropy, providing a supervised reference for EchoBridge’s prototype-based formulation.

KDD ’27, August, 2027, San Jose, CA, USA

Among the fully supervised models, ResNet-18 + Cosine BCE achieves the highest performance, with 79.78 AUROC, 33.46 AUPRC, and 36.60 F1. With the pretrained ECG encoder and shared projection frozen, EchoBridge obtains 78.79 AUROC, 32.66 AUPRC, and 35.82 F1 using a linear classifier trained with all probe labels. These values exceed ResNet-18 + BCE by 0.84 AUROC, 1.89 AUPRC, and 1.16 F1 points, while remaining 0.99, 0.80, and 0.78 points below the strongest fully supervised baseline, respectively.

Fang et al.

This comparison contextualizes EchoBridge’s source-domain performance under complete downstream supervision. EchoBridge provides a competitive frozen cross-modal representation while additionally supporting prompt-based classifier-free inference and source-only cross-center transfer through the same pretrained representation space.

Record · ID 405657 · SHA-256 7f51a0be72313e52
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.