Learning Cardiac Features: ECG Biometrics Across Time and Exercise Luca Thiebaud1 , Paul Chauchat1 , Mustapha Ouladsine1 , and Stéphane Delliaux2,3 1
arXiv:2609.21962v1 [cs.AI] 18 Sep 2026
2
Aix-Marseille Univ, CNRS, LIS, Marseille, France Aix-Marseille Univ, Inserm-INRAE, C2VN, Marseille, France 3 University Hospitals of Marseille, Marseille, France
Abstract. Electrocardiograms (ECGs) carry subject-specific patterns enabling reliable individual discrimination, forming the basis of ECG biometrics. Beyond authentication, this paradigm holds significant potential to secure sensitive cardiac data and to serve as a pretext task in self-supervised learning. Yet, most studies remain confined to singlesession, resting data, leaving robustness to temporal and physiological variations largely untested. We address this gap by evaluating ECG biometrics under realistic conditions involving exercise-induced stress and cross-session variability. A Siamese ResNet with late multi-lead fusion strategy is trained on a large ECG dataset extracted from cardiopulmonary exercise tests and evaluated with a exercise- and time-aware protocol, as well as on public benchmarks. This first extensive assessment of ECG biometrics under combined physiological and temporal variability achieves an intra-session rest-to-peak EER of 1.7% and stateof-the-art 3.9% on the CYBHi dataset. Findings support the presence of an intrinsic cardiac signature resilient to physiological and temporal drift.
Keywords: ECG biometrics · Siamese networks · Deep learning · Exercise stress · CPET
1
Introduction
The electrocardiogram (ECG) is a fundamental tool for assessing cardiac electrical activity [15]. While its analysis has traditionally depended on expert interpretation, machine learning—and particularly deep learning—has greatly advanced automatic ECG-based diagnosis and classification [24, 18], notably through ResNet architectures [7, 26]. To overcome the reliance on labeled data, recent work has turned to self-supervised learning (SSL) based approaches, enabling pre-training large architectures without the need for large amounts of annotated data [32]. In this context, ECG biometrics has emerged as a ECG SSL pretext task [10, 7], enhancing downstream clinical tasks such as pathology classification.
2
L. Thiebaud et al.
Beyond model training, ECG-based biometrics also stand as a secure and physiology-rooted alternative to conventional face or fingerprint recognition methods [23]. This raises critical privacy implications at the same time, as even anonymized ECG traces can reveal identifying information across large datasets [20]. Therefore, ECG biometrics appears both as a compelling pretext task for downstream medical purposes, and as a crucial privacy-related approach in itself which should be thoroughly analyzed. Deep learning methods have become the state-of-the-art in ECG-based biometrics, outperforming classical fiducial and handcrafted feature approaches [23, 22]. However, methodological heterogeneity—differences in tasks (identification vs. authentication), dataset scale, acquisition, or evaluation—still impedes objective benchmarking [23]. Identification systems (one-to-many, 1:N) typically require retraining when new users are added, limiting scalability, while authentication protocols (one-to-one, 1:1) enable verification without retraining and better reflect real-world use. Yet, most existing works still rely on identification settings with small populations, often under a few hundred subjects [22]. Another persistent weakness lies in temporal variability. Many studies rely on intra-session evaluation (training and testing within the same recording period), reporting near-perfect accuracies that reflect session consistency more than biometric permanence [23]. In contrast, the few cross-session analyses (i.e., training and testing on recordings acquired days to years apart) [13, 6, 23, 14] reveal substantial performance degradation over time, questioning both the permanence of ECG as a biometric trait and the quality of the learned representations. Finally, the impact of physiological ECG changes on identifiability remains largely unexplored. Most datasets are recorded at rest, and the few studies examining exercise or emotional stress [33, 31, 16, 8] report substantial drops in recognition accuracy, highlighting limited robustness to physiological variability. These limitations highlight the need for large, diverse datasets covering temporal and physiological variability, and for learning strategies able to extract invariant ECG representations resilient to context-dependent changes. In this study, we propose a deep learning–based ECG biometric authentication method and assess its performance under cross-session and exercise-induced variability using a large-scale dataset. Our main contributions are as follows: – A Siamese ResNet with late multi-lead fusion that improves robustness to cross-session variability compared to conventional early fusion schemes – The first extensive evaluation of ECG biometric stability under exercise stress, showing consistent performance from rest to exertion – State-of-the-art authentication results, including cross-dataset evaluations, surpassing existing methods The paper is organized as follows: Section 2 reviews related work; Section 3 describes the CPET dataset; Section 4 details the methodology; Section 5 presents experiments and results; and Section 6 concludes the study.
Learning Cardiac Features: ECG Biometrics Across Time and Exercise
2
Related Works
2.1
Evolution of ECG Biometric Methodologies: From Fiducial to Deep Models
3
Early ECG biometrics employed fiducial approaches, extracting waveform landmarks (P-QRS-T peaks, RR intervals) [11] and non-fiducial methods using autocorrelation or wavelet transforms [22]. Both suffer from noise sensitivity, imprecise peak detection, and limited robustness to cross-session variability and physiological changes. Recent years have seen deep learning methods dominate ECG biometrics, leveraging raw signals to overcome fiducial limitations [23]. Identification-driven works use standard CNN architectures. The ECG can be turned into an image to allow the use of standard image processing tools: Depthwise separable convolutions showed good performance in intra-session arrhythmic beat identification [2], and an ensemble of fine-tuned pretrained ResNets and DenseNets showed very good performance in a cross-session setting [30]. While a transformer-based approach have produced good authentication performances [6], almost all methods rely on Siamese networks [22]. The backbone can be pretrained on identification [13] or autoencoding tasks [23]. 2.2
Methodological Limitations
Cross- vs. Intra-session Several recent reviews show that most ECG biometric studies are still conducted in intra-session settings and report excellent performance. However, such protocols may conflate subject-specific cardiac traits with session-dependent acquisition factors, including electrode placement, posture, or psychophysiological state. This raises the question of whether intra-session results fully reflect stable biometric properties or instead partly capture sessionspecific effects [22, 23]. In contrast, more recent works [23, 13, 6, 14] explicitly address this limitation by enforcing strict cross-session evaluation splits and consistently reporting the associated performance degradation. Moreover, [23] introduced a benchmarking framework encompassing both mono- and multi-session settings across multiple databases, advancing the standardization of ECG biometric evaluation. Identification vs. Authentication Finally, many works rely on identification rather than authentication protocols [30, 2]. Identification (1:N) aims at determining the identity of a subject among a fixed set of enrolled users, whereas authentication (1:1) verifies a claimed identity by comparing two biometric samples. Identification thus generally requires retraining when new individuals are added, limiting scalability, while authentication approaches allow verification without retraining. Moreover, authentication enables more realistic evaluation protocols such as patient-level splits, where testing is performed on individuals not seen during training. This clearly transpires in cross-session studies, where identification results easily reach 99% or higher. However, in this setting, the associated networks
4
L. Thiebaud et al.
are trained on a first session of each patient, therefore patients whose second session lies in the test set were already seen in the training phase, inducing data leakage [30]. Although more robust, cross-session authentication is still rarely explored in the literature. Notably, [30] reported that among 21 studies (yielding 32 results on authentication and identification tasks across various ECG databases), only five results correspond to cross-session authentication protocols. 2.3
Dataset Limitations
Population Size Most ECG-biometrics studies rely on limited sample sizes. As emphasized in [22], “existing works on ECG for user authentication do not consider a population size close to a real application.” Typically, datasets include only a few dozen participants and seldom exceed a few hundred individuals. Impact of Physiological Variability on Biometric Performance Most studies on ECG biometrics rely on resting recordings, the standard condition in available databases. Only a few works have investigated performance under varying physiological conditions [33, 31, 16, 8]. These studies are generally based on small datasets (20–70 subjects), are limited to identification tasks, and consistently report marked degradation when comparing rest and non-resting conditions. In [31], recognition accuracy dropped to approximately 60% when enrollment was performed at rest and verification immediately after exertion, before recovering above 90% within one minute and 96% within five minutes, highlighting the impact of post-exercise physiological recovery. Similarly, [8] reported a decrease from 99.7% (rest–rest) to 88% (exercise–exercise), and only 17.5% under cross-condition evaluation (rest–exercise). Overall, the current state of the art highlights the lack of robustness of ECG biometrics under physiological stress, compounded by methodological limitations such as small dataset sizes and the predominant use of identification rather than authentication protocols.
3
CPET Dataset
3.1
Study Population
The data were collected from the Pulmonary Function Testing Laboratory of the Hôpital Nord, Assistance Publique – Hôpitaux de Marseille, between 2010 and 2024, with approval from the Clinical Ethics Committee. All participants underwent symptom-limited incremental exercise tests on an ergocycle, performed as part of routine cardiopulmonary evaluation. A total of 1651 adult patients were included in the study (mean age = 58.5 ± 14.8 years; 59.4% males). Among them, 1523 participants underwent a single CPET (mean age = 58.9 ± 14.7 years; 58.9% males), while 128 participants completed multiple sessions during the study period (mean age = 53.8 ± 15.2 years; 65.6% males).
Learning Cardiac Features: ECG Biometrics Across Time and Exercise
CPET Protocol and Data Acquisition
60 (a) 45 30 15 Rest Warm-up Exercise 0 00 03 07 10
Time (min)
Power Extracted ECG strips Recovery
14
18
Amplitude (µV)
Power (W)
3.2
5
150 (b) 0 I 150 II III 300 0
aVR aVL aVF
2
V1 V2 V3
4
V4 V5 V6
6
Time (s)
8
10
Fig. 1. CPET protocol and ECG extraction. (a) Incremental workload with 10s ECG strips (red) sampled during the test. (b) Example 12-lead ECG strip during the incremental phase.
Exercise Protocol CPET is performed on a cycle ergometer following a standard incremental protocol [4], comprising four phases (Fig. 1(a)): rest (baseline recording, no cycling), warm-up (low constant workload), incremental exercise (progressive ramp to exhaustion or clinical stop), and recovery (post-exercise monitoring until return toward baseline). ECG Acquisition During CPET For medical and patient security purposes, ECG is continuously monitored during CPET using CardioSoft v7.0 system (GE Healthcare). For each session, continuous 10-second 12-lead ECG strips are extracted throughout the test to provide balanced coverage of the CPET phases (see Fig. 1(a)). The signal sampling rate is 500 Hz. Due to the nature of the exercise test, the signals can be noisy with motion artifacts (Fig. 1(b)). 3.3
ECG Preprocessing and Beat Extraction
Signal Filtering and Segmentation All ECG signals are processed using tools from the NeuroKit2 Python library [21]. Noise removal is performed using a 5th-order Butterworth high-pass filter at 0.5 Hz combined with 50 Hz powerline filtering to correct baseline drift and reduce electrical interference. R-peaks are then detected using the NeuroKit2 detection algorithm. Individual heartbeats are segmented around each detected R-peak using a fixed-length window of 0.8 s (0.32 s before and 0.48 s after the peak), as in [23]. This segmentation strategy has been shown to provide more stable morphological alignment than random cropping [17]. All extracted beats are finally standardized using z-score normalization prior to model training. Quality-based Beat Exclusion Due to the nature of the exercise test, signals can be too noisy or exhibit motion artifacts. We assessed beat quality using the template-matching quality index [21], discarding beats below 85%. This threshold, selected based on visual inspection and kept fixed without further optimization, allowed retaining more than 70% of the beats.
6
L. Thiebaud et al.
Table 1. ECG datasets used in this study. ∗ MISI: Median Inter-Session Interval. † Finger-to-finger. Unlike public datasets, mostly recorded at rest, CPET includes varying physiological conditions. CPET PTB CYBHi Heartprint ECG-ID Total indiv. 1651 290 Indiv. with ≥2 sess. 128 113 MISI* (days) 346 4 ECG leads 12 12 Recording condition Varying Rest
3.4
63 63 104 1† Rest
78 78 1572 1† Rest
[9]
90 22 90 0 Same day 1† 12 Rest Exercise
MIT-BIH 423 0 12 Rest
Comparison with Existing ECG Databases
Public ECG datasets suitable for cross-session biometric evaluation remain scarce and are mostly recorded at rest, notably PTB [5], CYBHi [28] and Heartprint [14], which we use for comparison. PTB provides 12-lead recordings with a short median inter-session interval (MISI) of 4 days, whereas CYBHi and Heartprint use single-lead finger acquisitions with MISIs of about 3 months and 4 years, respectively. For Heartprint, only the S1–S3L configuration was considered to maximize the inter-session interval. ECG-ID [19] also enables crosssession analysis but was excluded due to its shorter MISI, while MIT-BIH [25] and [9] contain only one session per subject. As shown in Table 1, the CPET dataset differs markedly from public ECG databases by combining a larger cohort, repeated exercise recordings, and a MISI of 346 days, whereas most public datasets are limited to rest. For all datasets, only the first two recordings were retained when more than two were available, following [23]. 3.5
Exercise-Induced ECG Morphological Variability
During incremental exercise, cardiac electrical activity evolves according to physiological adaptation, mainly driven by the autonomic regulation. The muscularly and neurally induced cardiovascular and hemodynamic changes associated with exercise primarily lead to heart rate increase. (Table 2) and signal deformation. Figure 2 illustrates intra-session variability for one patient: HR rises progressively (top), while ECG morphology across leads (bottom) shows P-wave amplification, R-wave attenuation with QRS-axis shift, and ST–T alterations partially recovering post-exercise [29]. With a fixed 0.8 s window, cycle shortening at high intensity leads to multiple beats per segment, introducing additional intra-session variability that challenges biometric recognition. Table 2. Mean heart rate (HR, bpm) across subjects during CPET phases.
HR (µ ±σ)
Rest
Warmup
VT
Peak
Recovery
85 ±15
93 ±15
111 ±20
136 ±25
108 ±19
Learning Cardiac Features: ECG Biometrics Across Time and Exercise
150
(a)
100 0 min
Lead I
HR
200
7
5 min
10 min
15 min
20 min
0 5
(b) 0.8 s
Rest
0.8 s
Warm-up
0.8 s
Peak
Recovery
0.8 s
Fig. 2. For one patient. (a) Heart rate (HR, bpm) over the test. (b) Lead-I ECG morphology across phases (0.8 s, Z-score normalized), showing exercise-induced changes.
4
Methods
4.1
Siamese ResNet Architecture
We adopt a Siamese architecture, given its state-of-the-art performance in authentication and its ability to reuse the backbone as a generic feature extractor for downstream medical tasks, a key motivation of this work. The backbone is a 1D ResNet-18 [12], a strong lightweight baseline for ECG analysis [7, 26]. As shown in Fig. 3, two beats bA , bB ∈ R(l,400) (with l leads, of length 0.8 s at 500 Hz) are processed independently by a shared-weight backbone. Feature maps are aggregated via adaptive concatenated pooling (max + average) into fixed-length vectors rA , rB ∈ R1024 , then projected through a shared head (FC 1024 → d, batch norm, dropout, ReLU) to embeddings zA , zB ∈ Rd . The embeddings are combined as [zA , zB , |zA −zB |, zA ⊙zB ] ∈ R4d , capturing both absolute and multiplicative interactions. A classifier (FC 4d → 256, ReLU, batch norm, dropout, then FC 256 → 1) followed by a sigmoid outputs the score s that both segments belong to the same individual.
Fig. 3. Overview of the proposed Siamese ResNet1D for ECG verification. Input beats have shape (l, 400), with l leads and 400 samples (0.8 s at 500 Hz).
8
4.2
L. Thiebaud et al.
Single and Multi-lead Networks
We explored three versions based on the architecture described in Section 4.1. Two direct applications to single- and multi-lead inputs. For the multi-lead case, we propose a score fusion method, based on an ensemble of single-lead networks. Single-lead Architecture The single-lead network follows the Siamese architecture of Section 4.1 with l = 1. The embedding dimension is d = 256, and the ResNet-18 uses kernel size 3, selected via hyperparameter (HP) optimization. Multi-lead Early Fusion: ML-Early The multi-lead early fusion model uses the same architecture with input dimension l equal to the number of leads. HP optimization yields d = 512 and kernel size 7. Multi-lead Late Fusion: ML-Late ECG leads exhibit strong session-dependent correlations (e.g., electrode placement, recording conditions), so joint processing may induce overfitting, with the model capturing electrode configuration rather than subject-specific features. Following [3], we propose a score-level fusion (ML-Late), based on several independent single-lead Siamese networks si , each trained solely on its respective lead i. Let bA , bB ∈ R(l,400) be two multi-lead beats, and bA,i , bB,i ∈ R400 their i-th lead. Each single-lead Siamese expert si produces a match score si (bA,i , bB,i ) ∈ [0, 1]. The final verification score is obtained by averaging across leads: l
s(bA , bB ) =
1X si (bA,i , bB,i ). l i=1
Lead Selection Strategy for Multi-lead ECG For standard 12-lead ECG, we retain eight leads (I, II, V1–V6) and omit III, aVR, aVL, and aVF, as they are linear combinations of I and II, following Einthoven’s and Goldberger’s relationships [27]. Preliminary experiments show no performance gain from their inclusion, while increasing computational cost. The multi-lead late fusion framework extends to other acquisition setups (e.g., finger-to-finger single-lead ECG) that do not match standard leads. We propose averaging scores from a subset of standard leads selected via geometryaware criteria, based on the angular deviation between lead direction and the recording axis derived from standard electrode geometry [27].
5
Experiments and Results
All experiments were conducted under deterministic conditions (fixed seeds), with models trained on the CPET database. Intra- and cross-session evaluations on CPET highlight the superiority of late fusion. Generalizability is then assessed on public datasets without fine-tuning, and compared to state-of-the-art methods. Finally, the impact of exercise-induced stress is analyzed.
Learning Cardiac Features: ECG Biometrics Across Time and Exercise
5.1
9
Experimental Protocol
Data Splitting Strategy and Pair Construction The CPET dataset includes 1,523 single-session and 128 multi-session patients. To prevent data leakage, splitting was performed at the patient level, enabling both intra- and crosssession evaluations. Multi-session patients were evenly split (64/64), and singlesession patients were split 85%/15%, yielding 82% of patients for training, thereby limiting cross-session exposure during training while preserving it for testing. Pairs were constructed at the beat level by pairing each heartbeat with another from the same subject (genuine; different session if available, otherwise different strip, possibly within the same phase of the CPET) and with a beat from another subject (impostor; random strip). The training set comprises 720,254 pairs balanced (1:1), used identically for 1-lead and 8-lead Siamese ResNet models, differing only in input dimensionality. The same procedure yields 146,472 single-session test pairs and 70,118 cross-session test pairs for testing (from 262 single-session and 64 multi-session test patients, respectively). For public datasets, strict benchmark compliance is ensured via beat templates (Sec. 5.3). HP Optimization and Training Training was capped at 15 epochs with early stopping (patience = 7) due to rapid convergence. HP (learning rate, batch size, dropout, kernel size k, embedding dimension d) were optimized with Optuna [1] using 5-fold cross-validation on the training set. All single-lead models share the optimal HP from lead II, while ML-Early was optimized independently. For final training, the training dataset was split 80/20 (train/validation) to monitor convergence and limit overfitting. In total, eight single-lead Siamese ResNets and one 8-lead model (ML-Early) were trained. Equal Error Rate (EER) as Performance Metric The EER, a standard biometric metric, was used as the primary performance measure. It corresponds to the operating point where False Acceptance Rate equals False Rejection Rate, providing a threshold-independent assessment of discriminative performance. 5.2
Performance on the CPET Dataset: ML-Late Outperforms ML-Early
The intra-session and cross-session performance are evaluated on their respective test sets (Sec. 5.1). Results are presented on Figure 4. ML-Late fusion systematically outperforms both single-lead and early multilead fusion, achieving EER as low as 1.0% intra-session. More importantly, it is the only approach showing good performance cross-session with 5.6% EER. In contrast, ML-Early performs well intra-session (1.7% EER) but dramatically drops to 16.4% cross-session—even below the single-lead baseline (Lead II, 13.5%). For the remainder of the study, only the late fusion model is considered, as it substantially outperforms the other approaches.
10
L. Thiebaud et al.
Intra-session
Cross-session
0.8
True Positive Rate
True Positive Rate
1.0
0.6 0.4 ML-Early(AUC=99.9) ML-Late (AUC=99.9) Single-lead models
0.2 0.0 0.0
0.2
0.4
0.6
0.8
False Positive Rate
1.0
0.0
EER [%] Model
Intra Cross
ML-Early ML-Late Lead II
1.7 1.0 5.4
16.4 5.6 13.5
ML-Early(AUC=91.1) ML-Late (AUC=98.1) Single-lead models
0.2
0.4
0.6
0.8
False Positive Rate
1.0
Fig. 4. CPET test set performance. Left: ROC curves (AUC, %) in intra-session. Middle: ROC curves in cross-session. Right: EER (%) in intra- and cross-session (Lead II is shown as the best single-lead, likely due to its alignment with the heart’s mean electrical axis). ML-Late outperforms ML-Early, particularly in the cross-session setting.
5.3
Cross-Dataset Generalizability: SOTA Performance on Public Benchmarks
After evaluating CPET performance, we assess ML-Late generalizability on public datasets against state-of-the-art cross-session results, without dataset-specific fine-tuning. This preserves data in small datasets and ensures a simple, accessible methodology. The same preprocessing is applied across datasets (Sec. 3.3). To ensure full methodological consistency with the ECGXtractor benchmark [23], we replicated its protocol. Session templates were built by averaging the five beats closest to the mean in Euclidean distance and used as network inputs instead of raw beats. All models were retrained with the same 1:5 genuineto-impostor ratio. For each dataset, EER was averaged over ten comparison lists: those provided by the benchmark for PTB and CYBHi, and ten generated similarly for Heartprint (which was not included in the benchmark). Cross-session EERs are reported in Table 3. On PTB, our method achieves 2.1% EER, matching the ECGXtractor baseline [23]. This may be explained by the modest dataset size (113 subjects/templates), which likely promotes stable performance across the ten predefined comparison lists. Table 3. Cross-session EER (%) on datasets. ∗ Fine-tuned on target data. Our method matches or outperforms previous results without any fine-tuning. Work
PTB (12L) CYBHi (1L) Heartprint (1L)
EDITH [13] 5.7* [6] 10.2 [14] ECGXtractor benchmark [23] 2.1 Proposed (µ ±σ) 2.1 ±0.3
8.0 (5.4* ) 3.9 ±0.9
53.9 10.0 ±1.0
Learning Cardiac Features: ECG Biometrics Across Time and Exercise
11
The CYBHi and Heartprint datasets are more challenging due to their single finger-to-finger lead. We thus apply the proposed geometry-aware adaptation (Sec. 4.2), selecting an ML-Late model based on I, V5, V6. Lead I captures a horizontal vector in the frontal plane, matching the acquisition geometry, while V5 and V6, though defined in the horizontal plane, form relatively small angles with I and provide complementary, consistent information. Other leads are excluded due to larger angular deviations. For comparison, lead I alone yields EERs of 4.9% on CYBHi and 10.8% on Heartprint, versus 4.2% and 9.8% using all eight leads. Overall, our method matches ECGXtractor on PTB, significantly outperforms it on CYBHi, and establishes a new benchmark on the less-studied Heartprint dataset, without fine-tuning, achieving state-of-the-art performance. 5.4
Impact of Exercise
One of the main contributions of this work is to assess robustness to exerciseinduced ECG physiological variations. Phase-specific evaluation sets were derived from the CPET test subset by selecting beats from rest, warm-up, peak, and recovery phases. Two pairing schemes were considered: (i) same-phase comparisons across tests, and (ii) phase-to-rest comparisons. This was done in both intra- and cross-session scenarios. For all configurations, EER was computed along with 95% confidence intervals (CI) via 2000-bootstrap resampling. Given the dependence of bootstrap CIs’ on dataset size, evaluation was restricted to the 64-patient multi-session CPET test dataset, using only their first test session for the intra-session configuration. Results are presented in Figure 5. Intra-session The pooled intra-session EER (Fig. 5) confirms excellent overall performance. Phase-specific analysis reveals a moderate drop for Peak–Rest (EER 1.7%, CI [0.8; 2.8]), which remains low, indicating strong robustness to intensity variations. This contrasts with prior studies [33, 31, 16, 8], which report larger degradations under mixed physiological states, possibly due to smaller
Comparison All (pooled) Rest vs. Rest Warm-up vs. Warm-up Peak vs. Peak Recovery vs. Recovery Warm-up vs. Rest Peak vs. Rest Recovery vs. Rest
EER (%)
0.7
Single-session
5.6
0.2 0.5 0.6 0.1 0.7
4.5 4.9 8.1 4.7 4.9 1.7
7.8
0.4
0
Cross-session
5.7
2
4
6
8
10
12
Fig. 5. Phase-specific intra- and cross-session EER comparisons, using the 64 multisession test patients (mean and 95% CI, intra-session on first test only). For “All”, intra-session EER differs from Sec. 5.2 due to a different patient subset.
12
L. Thiebaud et al.
datasets. In addition, the Recovery–Recovery condition exhibits the lowest error and narrowest CI, an observation of interest given the specific physiological characteristics of this phase, particularly the predominance of vagal tone. Cross-session Figure 4(right) showed the overall performance decrease to 5.6%. However, Fig. 5 indicates a sharp increase in variability, with CI widths of at least 5. The largest performance drops are observed for Peak–Rest (EER = 7.8% [5.5; 10.7]) and, interestingly, also for Peak–Peak (EER = 8.1% [4.7; 12.2]). Three main observations emerge. First, the shift in EER distributions between intra- and cross-session settings—marked by degraded Peak–Peak performance and wider CI—indicates a fundamental difference beyond proportional decay, suggesting a non-linear temporal evolution of ECG information. This degradation likely reflects both intrinsic temporal variability and changes in acquisition conditions (e.g., electrode placement, batches, skin preparation, input impedance). Second, the increased difficulty of Peak–Peak in cross-session indicates a loss of discriminative information across recordings. This may be further exacerbated by differences in exercise conditions between sessions, such as variations in workload, pedaling patterns (and associated motion artifacts), as well as changes in baseline physiological state and effort response. Finally, the similar performance of Peak–Rest and Peak–Peak shows that ML-Late remains effective even when comparing markedly different physiological states across sessions, a noteworthy contribution. This suggests that, despite variability induced by both physiological changes and acquisition-related factors, the model captures features that retain a degree of invariance across conditions.
6
Conclusion
This work proposes ECG-based biometric methods relying on single-lead Siamese ResNet architectures, combined with a late-fusion strategy for multi-lead signals. Trained on an extensive in-house CPET dataset, they outperform state-of-theart results on PTB and CYBHi (3.94% EER). It provides the first extensive evaluation of ECG biometrics under exercise conditions, showing strong robustness to both exercise and temporal drift, suggesting the existence of a largely preserved electrocardiographic signature. However, it is constrained by the scarcity of public datasets combining intersession variability and physical exertion, and by the limited interpretability of the learned discriminative ECG features, which remains a prerequisite for clinical translation. Future research will focus on transferring representations learned through self-supervised learning-based ECG biometrics (as a pretext task) to prognostic applications, such as mortality and care intensity prediction, alongside improving interpretability. Acknowledgments. This study was funded by the Excellence Initiative of AixMarseille Université (A*Midex, AMX-21-IET-017, “Investissements d’Avenir”).
Learning Cardiac Features: ECG Biometrics Across Time and Exercise
13
References 1. Akiba, T., Sano, S., Yanase, T., Ohta, T., Koyama, M.: Optuna: A nextgeneration hyperparameter optimization framework. In: Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery; Data Mining. pp. 2623–2631. KDD ’19, ACM (2019). https://doi.org/10/gf7mzz 2. Al-Jibreen, A., Al-Ahmadi, S., Islam, S., Artoli, A.M.: Person identification with arrhythmic ecg signals using deep convolution neural network. Scientific Reports 14(1) (2024) 3. Aublin, P.G., Ben Ammar, M., Fix, J., Barret, M., Behar, J.A., Oster, J.: Predict alone, decide together: Cardiac abnormality detection based on single lead classifier voting. Physiological Measurement 43(5), 054001 (2022) 4. Balady, G.J., Arena, R., Sietsema, K., Myers, J., Coke, L., Fletcher, G.F., Forman, D., Franklin, B., Guazzi, M., Gulati, M., Keteyian, S.J., Lavie, C.J., Macko, R., Mancini, D., Milani, R.V.: Clinician’s guide to cardiopulmonary exercise testing in adults: A scientific statement from the american heart association. Circulation 122(2), 191–225 (2010) 5. Bousseljot, R.D., Kreiseler, D., Schnabel, A.: The ptb diagnostic ecg database (2004). https://doi.org/10/hbtxw2 6. Chee, K.J., Ramli, D.A.: Electrocardiogram biometrics using transformer’s selfattention mechanism for sequence pair feature extractor and flexible enrollment scope identification. Sensors 22(9), 3446 (2022) 7. Chen, Y.: A benchmark study of deep learning methods for multi-label pediatric electrocardiogram-based cardiovascular disease classification (2025). https://doi.org/10/hbtxwq 8. Cui, W., Wang, Z., Li, Y.: Ecg-based biometric recognition under exercise and rest situations. Biomedical Engineering Advances 2, 100008 (2021) 9. De Giovanni, E., Teijeiro, T., Meier, D., Millet, G., Atienza, D.: Ecg in high intensity exercise dataset (2021). https://doi.org/10/hbtxxk 10. Diamant, N., Reinertsen, E., Song, S., Aguirre, A.D., Stultz, C.M., Batra, P.: Patient contrastive learning: A performant, expressive, and practical approach to electrocardiogram modeling. PLOS Computational Biology 18(2), e1009862 (2022) 11. Fratini, A., Sansone, M., Bifulco, P., Cesarelli, M.: Individual identification via electrocardiogram analysis. BioMedical Engineering OnLine 14(1) (2015) 12. He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition (2015). https://doi.org/10/gpptvk 13. Ibtehaz, N., Chowdhury, M.E.H., Khandakar, A., Kiranyaz, S., Rahman, M.S., Tahir, A., Qiblawey, Y., Rahman, T.: Edith: Ecg biometrics aided by deep learning for reliable individual authentication. IEEE Transactions on Emerging Topics in Computational Intelligence 6(4), 928–940 (2022) 14. Islam, M.S., Alhichri, H., Bazi, Y., Ammour, N., Alajlan, N., Jomaa, R.M.: Heartprint: A dataset of multisession ecg signal with long interval captured from fingers for biometric recognition. Data 7(10), 141 (2022). https://doi.org/10/hbtxxh 15. Kligfield, P., Gettes, L.S., Bailey, J.J., Childers, R., Deal, B.J., Hancock, E.W., van Herpen, G., Kors, J.A., Macfarlane, P., Mirvis, D.M., Pahlm, O., Rautaharju, P., Wagner, G.S.: Recommendations for the standardization and interpretation of the electrocardiogram: Part i: The electrocardiogram and its technology. Circulation 115(10), 1306–1324 (2007). https://doi.org/10/bcmzc9 16. Komeili, M., Louis, W., Armanfard, N., Hatzinakos, D.: On evaluating human recognition using electrocardiogram signals: From rest to exercise. In: 2016 IEEE
14
L. Thiebaud et al.
Canadian Conference on Electrical and Computer Engineering (CCECE). pp. 1–4. IEEE (2016). https://doi.org/10/hbtxw8 17. Li, Y., Pang, Y., Wang, K., Li, X.: Toward improving ecg biometric identification using cascaded convolutional neural networks. Neurocomputing 391, 83–95 (2020) 18. Liu, X., Wang, H., Li, Z., Qin, L.: Deep learning in ecg diagnosis: A review. Knowledge-Based Systems 227, 107187 (2021) 19. Lugovaya, T.: The ecg-id database (2011). https://doi.org/10/hbtxw3 20. Macierzanka, K., Sau, A., Patlatzoglou, K., Pastika, L., Sieliwonczyk, E., Gurnani, M., Peters, N.S., Waks, J.W., Kramer, D.B., Ng, F.S.: Siamese neural networkenhanced electrocardiography can re-identify anonymized healthcare data. European Heart Journal - Digital Health 6(3), 417–426 (2025) 21. Makowski, D., Pham, T., Lau, Z.J., Brammer, J.C., Lespinasse, F., Pham, H., Schölzel, C., Chen, S.H.A.: Neurokit2: A python toolbox for neurophysiological signal processing. Behavior Research Methods 53(4), 1689–1696 (2021). https://doi.org/10/gm9ckz 22. Meltzer, D., Luengo, D.: Ecg-based biometric recognition: A survey of methods and databases. Sensors 25(6), 1864 (2025). https://doi.org/10/hbtxwb 23. Melzi, P., Tolosana, R., Vera-Rodriguez, R.: Ecg biometric recognition: Review, system proposal, and benchmark evaluation. IEEE Access 11, 15555–15566 (2023) 24. Mincholé, A., Camps, J., Lyon, A., Rodríguez, B.: Machine learning in the electrocardiogram. Journal of Electrocardiology 57, S61–S64 (2019) 25. Moody, G.B., Mark, R.G.: Mit-bih arrhythmia database (1992). https://doi.org/10/gkr8xd 26. Nonaka, N., Seita, J.: In-depth benchmarking of deep neural network architectures for ecg diagnosis. In: Jung, K., Yeung, S., Sendak, M., Sjoding, M., Ranganath, R. (eds.) Proceedings of the 6th Machine Learning for Healthcare Conference. Proceedings of Machine Learning Research, vol. 149, pp. 414–439. PMLR (2021), https://proceedings.mlr.press/v149/nonaka21a.html 27. Ramirez, E., Ruiperez-Campillo, S., Casado-Arroyo, R., Merino, J.L., Vogt, J.E., Castells, F., Millet, J.: The art of selecting the ecg input in neural networks to classify heart diseases: A dual focus on maximizing information and reducing redundancy. Frontiers in Physiology 15 (2024). https://doi.org/10/hbtxxb 28. da Silva, H.P., Lourenço, A., Fred, A., Raposo, N., Aires-de Sousa, M.: Check your biosignals here: A new dataset for off-the-person ecg biometrics. Computer Methods and Programs in Biomedicine 113(2), 503–514 (2014). https://doi.org/10/f5qdgj 29. Simoons, M.L., Hugenholtz, P.G.: Gradual changes of ecg waveform during and after exercise in normal subjects. Circulation 52(4), 570–577 (1975) 30. Srivastva, R., Singh, A., Singh, Y.N.: Plexnet: A fast and robust ecg biometric system for human recognition. Information Sciences 558, 208–228 (2021) 31. Sung, D., Kim, J., Koh, M., Park, K.: Ecg authentication in post-exercise situation. In: 2017 39th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). pp. 446–449. IEEE (2017). https://doi.org/10/hbtxw7 32. Zhang, K., Wen, Q., Zhang, C., Cai, R., Jin, M., Liu, Y., Zhang, J.Y., Liang, Y., Pang, G., Song, D., Pan, S.: Self-supervised learning for time series analysis: Taxonomy, progress, and prospects. IEEE Transactions on Pattern Analysis and Machine Intelligence 46(10), 6775–6794 (2024). https://doi.org/10/gt9kdw 33. Zhou, R., Wang, C., Zhang, P., Chen, X., Du, L., Wang, P., Zhao, Z., Du, M., Fang, Z.: Ecg-based biometric under different psychological stress states. Computer Methods and Programs in Biomedicine 202, 106005 (2021)