1
EERLoss: A Novel Loss Function for Training Deep Biometric Models. A Case Study in Keystroke Dynamics
arXiv:2606.24586v1 [cs.CV] 23 Jun 2026
Nahuel Gonzalez† , Marta Robledo-Moreno∗ , Ivan DeAndres-Tame∗ , Ruben Vera-Rodriguez∗ , and Ruben Tolosana∗
Abstract—Deep learning approaches to biometric verification are commonly trained by optimizing indirect objectives, creating a misalignment between the optimization process and the primary evaluation metric, typically the Equal Error Rate (EER). This paper introduces EERLoss: a subdifferentiable, arbitrarily accurate approximation to EER for training deep biometric models. Furthermore, this framework has the potential to be adapted to optimize any specific operating point on the DET curve, enhancing its generalizability. To validate this approach, EERLoss is evaluated on a particularly demanding behavioral biometric modality: keystroke dynamics verification. This task is characterized by its high intra-class and low interclass variability. Experiments are conducted on the large-scale KVC-onGoing benchmark, incorporating data from over 185,000 subjects across different scenarios. A comprehensive ablation study initially demonstrates the superiority of EERLoss in comparison to existing state-of-the-art loss functions. It also converges substantially faster compared to other losses, reducing the overall training cost. Additionally, a comparison is made between the proposed loss and the KVC-winning architecture by re-training it with EERLoss, demonstrating that the proposed approach significantly outperforms the original SoTA, achieving a relative EER reduction of up to approx. 30%. This improvement on a challenging, large-scale benchmark validates the effectiveness of EERLoss as a task-aligned training objective specifically suited for high-variance biometric traits. Index Terms—Loss Function, Behavioral Biometrics, Keystroke Dynamics, Verification, Feature Learning
I. Introduction Biometric verification systems decide whether two samples belong to the same identity while trading off security and usability. The level of risk tolerance varies depending on the application [1]. For instance, high-security applications like banking logins prioritize the absolute exclusion of impostors, accepting occasional user lockouts as a necessary operational cost. Conversely, friction-sensitive scenarios, such as smartphone unlocking or continuous desktop re-authentication, prioritize a seamless user experience, tolerating transient false rejections that allow for immediate recovery. In this context, the Equal Error Rate (EER) denotes performance at the operating point where the False Acceptance Rate (FAR) is equal to the False Rejection Rate (FRR). A lower EER is indicative of a system that effectively admits genuine users and rejects impostors. † LSIA, University of Buenos Aires, Argentina ([email protected]) ∗ BiometricsAI, Universidad Autonoma de Madrid, Spain ({marta.robledo, ivan.deandres, ruben.vera, ruben.tolosana}@uam.es) Manuscript received May 2026
Deep biometric models are commonly trained to optimize for embedding separability, creating a fundamental misalignment with the common evaluation standard, typically the EER. Current systems rely on diverse indirect objectives, including cross-entropy and various margin-based metric losses such as contrastive [2], triplet [3], and sophisticated angular losses like ArcFace, CosFace or AdaFace [4]–[6]. While these methods successfully enhance the clustering of embeddings, their primary goal is to maximize the distance between classes. This inherent mismatch between the training objective (separability) and the evaluation metric (EER) can result in sub-optimal performance, particularly when the decision boundary is critical to the security and usability trade-off. The present study aims to address this gap by introducing EERLoss, an arbitrarily accurate, subdifferentiable approximation of the EER for training deep biometric models. Unlike conventional losses, EERLoss aligns the training objective directly with the evaluation metric, optimizing the actual operational performance of a biometric system. The proposed loss function also has the potential to be adapted to optimize any specific operating point on the DET curve, enhancing its generalizability. To empirically validate EERLoss, this work focuses on the challenging domain of behavioral biometrics, specifically keystroke dynamics (KD) [7]. This modality is notorious for its high temporal variability, noise, and data sparsity. Demonstrating robust optimization of EER in this complex, large-scale scenario provides strong evidence of the method’s efficacy and its potential for generalization to other biometric traits, such as gait [8] or signature verification [9], [10]. Our main contributions are: - The formulation of a novel loss function named EERLoss: a subdifferentiable, and arbitrarily accurate approximation of the EER metric, enabling direct end-to-end optimization. - The code’s implementation of EERLoss has been made available1 . - A comprehensive comparison study that validates the design choices of EERLoss and compares its performance and training time against multiple SoTA loss functions. - A new state-of-the-art result on the large-scale KVConGoing benchmark, where EERLoss significantly outperforms the previous best method, achieving a relative EER reduction of up to 30.7%. 1 https://github.com/BiDAlab/EERLoss
2
The remainder of the paper is structured as follows. Sec. II provides an overview of loss functions for biometric models and keystroke dynamics for biometric verification. The problem statement is described in Sec. III. Sec. IV presents the necessary definitions for the proposed loss function, which is further elaborated in Sec. V. The experimental setup is described in Sec. VI, while the obtained results are presented in Sec. VII. In Sec. VIII, it is provided a theoretical justification for the performance gains of the proposed loss function. Finally, Sec. IX summarizes the findings and outlines directions for future work. II. Related Work This section provides a review of the most relevant literature to contextualize the need for the proposed loss function. While the present study is focused on the use of KD for biometric verification, the proposed formulation is general and can be applied to other biometric tasks. A. Loss Functions for Biometric Models In the context of deep learning, the loss function defines the optimization target that guides the updating of model parameters during training. Meanwhile, performance metrics quantify the ability of the trained model to generalize to unseen data. The selection of an appropriate loss function is essential, as it determines the extent to which a model aligns its internal representations with the final evaluation criteria. As stated in [11], two key properties of a well-defined loss function are differentiability (or subdifferentiability, as discussed in Sec. III) and smoothness (meaning that is has a continuous gradient without abrupt discontinuities that could destabilize training). In computer vision and biometric recognition, several wellestablished loss functions have proven effective for learning discriminative embeddings, including Contrastive loss [2], Triplet loss [3], CosFace [5], ArcFace [4], AdaFace [6] or Set2Set [12]. These formulations promote intra-class compactness and inter-class separation in the embedding space, indirectly improving verification metrics such as accuracy, precision, recall, or F1-score. However, none of these losses explicitly target the EER, which remains one of the most relevant metric in biometric verification systems. B. Keystroke Dynamics for Biometric Verification Keystroke dynamics refers to the analysis of a person’s typing behavior used to infer or verify identity. As summarized in the recent comprehensive survey by Shadman et al. [7], KD has become one of the most promising behavioral biometric modalities due to its low acquisition cost, non-intrusive nature, and seamless integration with existing human-computer interfaces. In contrast to physiological biometrics, such as fingerprint or face recognition, which require dedicated sensors, KD can operate continuously in the background during normal interaction without the need for additional sensors. The KVC-onGoing Challenge2 [13] has become the reference benchmark for keystroke-based biometric verification. 2 https://sites.google.com/view/bida-kvc/
It provides a unified experimental framework with a standard evaluation protocol, large-scale datasets, and consistent performance metrics, allowing researchers to benchmark new systems under identical conditions. The challenge is based on the Aalto University Keystroke Databases [14] [15], which are the largest public collections of typing data to date. These databases include over 185,000 subjects across desktop and mobile scenarios. These datasets capture transcript-level typing sessions in natural conditions, enabling realistic modeling of behavioral variability while maintaining demographic diversity and sufficient enrollment samples per user. In contrast to earlier KD datasets (e.g., GREYC [16], RHU [17], Clarkson II [18], HuMIdb [19]), the Aalto database introduced massive scale and cross-device realism. The KVC-onGoing challenge establishes the current state of the art. It reports the best global EER performances of 3.33% for the desktop task and 3.61% for the mobile task. An analysis of the competing methods reveals a critical trend: the vast majority of systems rely on standard margin-based metric losses, such as triplet loss. The systems that demonstrated the highest levels of performance, LSIA and VeriKVC, exhibited distinct characteristics that set them apart from their competitors. The winning team from the LSIA employed a dualbranch architecture, i.e., Convolutional Neural Network (CNN) + Recurrent Neural Network (RNN) with Attention, which was trained using a novel custom loss function called Set2Set [12] that extends the SetMargin [20] loss. The VeriKVC team (runner-up) employed a pure CNN architecture that was trained with the ArcFace loss function. In [13], Stragapede et al. concluded that this focus on loss was a key factor, indicating that the selection of an optimal training objective is crucial. In this context, the KVC-onGoing is as an appropriate and rigorous benchmark for our proposed loss function, as it is precisely in these high-variability scenarios where a loss function that directly optimizes for the EER can demonstrate a measurable advantage over conventional methods that only optimize for separability. III. Problem Statement A distance metric learning model computes embeddings for the labeled samples in a training batch. For each training batch, we define 𝐿 as the set of distances between embeddings of samples that share the same label; that is, samples originated from the same user. The set 𝐼 will consist of the distances between the embeddings of samples with different labels; that is, samples from different users. Informally, our objective is to build a loss function that outputs the EER of such training batch using 𝐿 and 𝐼 as input. However, despite the apparent simplicity of this task, it is inherently constrained by the need for loss functions to be subdifferentiable. As stated before, the EER is defined as the specific point at which the FAR and FRR intersect. Unfortunately, the operation of finding the intersection of generic curves is inherently non-differentiable. Given this consideration, we must abandon the idea of calculating the EER exactly within a loss function, leading to the following. Problem statement. Find a subdifferentiable function L (𝐿, 𝐼) that, given the sets 𝐿 and 𝐼 containing the distances between
3
samples of the same user and those of different users, respectively, approximates 𝐸 𝐸 𝑅(𝐿, 𝐼) arbitrarily well.
IV. Definitions Let 𝑑 denote a positive real number. Concretely, 𝑑 (sometimes with subscripts) will have two possible denotations: (i) the distance between the embedding vectors of a pair of samples, as calculated by a neural network, and (ii) the uniform global threshold of a distance metric learning (DML) verification system. We can assume that 𝑑 is bounded; that is, there is a constant 𝑀 such that 0 ≤ 𝑑 ≤ 𝑀. We now define the legitimate and impostor distance sets: 𝐿 = {𝑑1𝐿 , . . . , 𝑑 𝑛𝐿 }
(1)
𝐼 𝐼 = {𝑑1𝐼 , . . . , 𝑑 𝑚 }
(2)
where 𝑛 = |𝐿|, 𝑚 = |𝐼 |, and 𝑑𝑖𝐿 , 𝑑 𝐿𝑗
However, the EER is a local, stepwise objective that fails to encourage the network to separate samples when their distance is far from 𝑑 𝐸𝐸 𝑅 . To overcome this problem, we will propose an alternative formulation that offers a smooth target extending to the entire area of overlap between the FAR and FRR curves. In addition, we will explore adding a margin away from 𝑑 𝐸𝐸 𝑅 and delinearizing the contribution of distant values. A. Smooth FAR and FRR Given a positive 𝑑, together with sets 𝐿 and 𝐼 as in Section IV, the FAR and FRR of the batch at threshold 𝑑 can be approximated using Algorithms 1 and 2, which are subdifferentiable renditions of equations (3) and (4). Our points of departure are the following smoothings of the equations for FAR and FRR. Let
∈ R+ . In concreteness,
the sets 𝐿 and 𝐼 will correspond to the distances between the embedding vectors of the samples of the same user and of different users in a training batch, respectively. From 𝐿, 𝐼, and a given threshold 𝑑, we can calculate the False Rejection Rate (FRR) and False Acceptance Rate (FAR) for a batch, in the form: |{𝑑 𝐿 ∈ 𝐿 : 𝑑 𝐿 > 𝑑}| 𝐹 𝑅𝑅(𝐿, 𝑑) = 1 − |𝐿| |{𝑑 𝐼 ∈ 𝐼 : 𝑑 𝐼 < 𝑑}| 𝐹 𝐴𝑅(𝐼, 𝑑) = |𝐼 |
(7)
𝑀 𝐿 (𝑑, 𝑑 𝐼𝑗 ) = tanh [𝐾.(𝑑 − min {𝑑, 𝑑 𝐼𝑗 })]
(8)
and define the smooth counterparts of 3 and 4 as 1 ∑︁ 𝑀𝑅 (𝑑, 𝑑𝑖𝐿 ) 𝐹 𝑅𝑅𝑆 (𝐿, 𝑑) = 1 − |𝐿| 𝐿
(4)
(5)
(9)
𝑑𝑖 ∈ 𝐿
1 ∑︁ 𝐹 𝐴𝑅𝑆 (𝐼, 𝑑) = 𝑀 𝐿 (𝑑, 𝑑 𝐼𝑗 ) |𝐼 | 𝐼
(10)
𝑑 𝑗 ∈𝐼
(3)
Note that 𝐹 𝑅𝑅(𝐿, 𝑑) is a step function, monotonously decreasing with 𝑑. 𝐹 𝐴𝑅(𝐼, 𝑑) is also a step function, but monotonously increasing. Thus, there is a unique value 𝑑 𝐸𝐸 𝑅 in [0, 𝑀] such that 𝐹 𝑅𝑅(𝐿, 𝑑 𝐸𝐸 𝑅 ) ≤ 𝐹 𝐴𝑅(𝐼, 𝑑 𝐸𝐸 𝑅 )
𝑀𝑅 (𝑑, 𝑑𝑖𝐿 ) = tanh [𝐾.(max {𝑑, 𝑑𝑖𝐿 } − 𝑑)]
To understand how equation (9) provides a smooth approximation to equation (3), observe that ( 0 if 𝑑𝑖𝐿 ≤ 𝑑 𝐿 max {𝑑, 𝑑𝑖 } − 𝑑 = (11) > 0 if 𝑑𝑖𝐿 > 𝑑 Hence, for large 𝐾, we have that ( ≈0 𝑀𝑅 (𝑑, 𝑑𝑖𝐿 ) = ≈1
if 𝑑𝑖𝐿 ≤ 𝑑 if 𝑑𝑖𝐿 > 𝑑
(12)
Thus,
and 𝐹 𝑅𝑅(𝐿, 𝑑) > 𝐹 𝐴𝑅(𝐼, 𝑑)
∀𝑑 < 𝑑 𝐸𝐸 𝑅
(6)
Using the above, we define the batch EER as the average of the FAR and FRR at 𝑑 𝐸𝐸 𝑅 . In the algorithms that follow, 𝐾 is a constant that will control the quality of the approximation. The error bound will decrease as 𝐾 increases, approaching zero as 𝐾 goes to infinity. In practical cases we can have 𝐾 = 1000. The subscripts 𝐿 and 𝑅 (as in 𝑀 𝐿 , 𝑀𝑅 , or 𝐼 𝐿 , 𝐼 𝑅 , or 𝑑 𝐿 , 𝑑 𝑅 ) will denote the left and right of 𝑑 𝐸𝐸 𝑅 , respectively. V. EERLoss: Proposed Loss Function We will find L (𝐿, 𝐼) as in the Problem Statement by building the objective function from the bottom up. We will first show, in Section V-A, how to approximate the FAR and FRR of a batch at a given threshold in a subdifferentiable way. Using these approximations, in Section V-B we will estimate 𝑑 𝐸𝐸 𝑅 through a subdifferentiable binary search over a reasonably chosen interval. Once 𝑑 𝐸𝐸 𝑅 is determined, the calculation of the EER of the batch to be used as an objective function is immediate.
𝐹 𝑅𝑅𝑆 (𝐿, 𝑑) ≈ 𝐹 𝑅𝑅(𝐿, 𝑑)
(13)
and identical reasoning yields 𝐹 𝐴𝑅𝑆 (𝐼, 𝑑) ≈ 𝐹 𝐴𝑅(𝐿, 𝑑)
(14)
The FRR algorithm 1 is a direct tensorial implementation of equation (9); the only difference is that it multiplies the result by 100 to scale it to a human-readable percentage. The FAR algorithm 2 implements equation (10) similarly. Algorithm 1: Smooth False Rejection Rate (𝐹 𝑅𝑅𝑆 ) Data: 𝐿, 𝑑 ≥ 0 Result: 𝑓 𝑟𝑟 𝐿 𝑅 ← 𝑚𝑎𝑥(𝑑, 𝐿) − 𝑑 𝑀𝑅 ← 𝑡𝑎𝑛ℎ(𝐾 ∗ 𝐿 𝑅 ) return 100 ∗ 𝑎𝑣𝑔(𝑀𝑅 ) Note that using the 𝑡𝑎𝑛ℎ function both in equations (9) and (10), as well as in Algorithms 1 and 2, is not mandatory. Any continuous, monotonously increasing function 𝑓 such that 𝑓 (0) = 0 and 𝑓 (𝑥) → 1 fast enough when 𝑥 → ∞ should
4
Algorithm 2: Smooth False Acceptance Rate (𝐹 𝐴𝑅𝑆 ) Data: 𝐼, 𝑑 ≥ 0 Result: 𝑓 𝑎𝑟
sets 𝐿 and 𝐼. The third parameter 𝑛 is a positive integer that controls the number of steps.
𝐼 𝐿 ← 𝑑 − 𝑚𝑖𝑛(𝑑, 𝐼) 𝑀 𝐿 ← 𝑡𝑎𝑛ℎ(𝐾 ∗ 𝐼 𝐿 ) return 100 ∗ 𝑎𝑣𝑔(𝑀 𝐿 )
Algorithm 4: Binary Search for 𝑑 𝐸𝐸 𝑅 (DEER) Data: 𝐿, 𝐼, 𝑛 > 0 Result: 𝑑 𝐸𝐸 𝑅 𝑑 𝐿 ← 0.5 ∗ 𝑎𝑣𝑔(𝐿) 𝑑 𝑅 ← 1.5 ∗ 𝑎𝑣𝑔(𝐼) for 𝑖 ≤ 𝑛 do 𝑑 𝐿 , 𝑑 𝑅 ← 𝐵𝑆𝑆(𝑑 𝐿 , 𝑑 𝑅 ) end return (𝑑 𝐿 + 𝑑 𝑅 )/2
suffice for the purpose. In particular, 𝑡𝑎𝑛ℎ was chosen for behaving well when training neural networks, and having fast tensorial implementations in modern GPUs The proof that Algorithms 1 and 2 are subdifferentiable is almost trivial. As subtraction, multiplication, 𝑚𝑖𝑛, 𝑡𝑎𝑛ℎ, and 𝑎𝑣𝑔 are subdifferentiable, and compositions of subdifferentiable functions are subdifferentiable, we have that both algorithms compute subdifferentiable approximations to the FAR and FRR of the batch at a threshold 𝑑. The proofs are similar for the algorithms that follow. Therefore, we will not repeat them. B. A Smooth Implementation of Binary Search We now want to use the proposed FAR and FRR approximations to perform a binary search for 𝑑 𝐸𝐸 𝑅 . The subdifferentiable binary search step is shown in Alg. 3. It takes for input an interval (𝑑 𝐿 , 𝑑 𝑅 ), where 𝑑 𝐿 < 𝑑 𝐸𝐸 𝑅 < 𝑑 𝑅 , and returns a ′ ) of half the size. similar interval (𝑑 ′𝐿 , 𝑑 𝑅
Note that the initial endpoints are based on the average of the distance values in 𝐿 and 𝐼. The objective is to make the endpoints less sensitive to a choice of constants or to any variations in the embeddings. During the first training epoch, it is possible that the average distance between embeddings of samples of the same user is larger than the average distance between embeddings of samples of different users. Thus, the 0.5 and 1.5 constants are used to make sure that 𝑑 𝐿 < 𝑑 𝑅 in the initial search interval.
C. Encoding the EER as a Loss Function Algorithm 3: Binary Search Step for 𝑑 𝐸𝐸 𝑅 (BSS) Data: 𝐿, 𝐼, 𝑑 𝐿 < 𝑑 𝑅 ′ Result: 𝑑 ′𝐿 < 𝑑 𝑅 𝑑 ← (𝑑 𝐿 + 𝑑 𝑅 )/2 𝑓 𝑎𝑟 ← 𝐹 𝐴𝑅𝑆 (𝐼, 𝑑) 𝑓 𝑟𝑟 ← 𝐹 𝑅𝑅𝑆 (𝐿, 𝑑) 𝑚 ← 𝑚𝑎𝑥{ 𝑓 𝑎𝑟, 𝑓 𝑟𝑟 } 𝑐 𝐿 ← 𝑒𝑥 𝑝(𝐾 ∗ ( 𝑓 𝑟𝑟 − 𝑚)) 𝑐 𝑅 ← 𝑒𝑥 𝑝(𝐾 ∗ ( 𝑓 𝑎𝑟 − 𝑚)) 𝑑 ′𝐿 ← 𝑑 𝐿 ∗ (1 − 𝑐 𝐿 ) + 𝑑 ∗ 𝑐 𝐿 ′ ← 𝑑 ∗ (1 − 𝑐 ) + 𝑑 ∗ 𝑐 𝑑𝑅 𝑅 𝑅 𝑅 ′ return 𝑑 ′𝐿 , 𝑑 𝑅 To understand the correctness of the algorithm, note that 𝑐 𝐿 , calculated in line 5, acts as a flag signaling whether the midpoint 𝑑 is to the left of 𝑑 𝐸𝐸 𝑅 or not. If 𝑑 ≤ 𝑑 𝐸𝐸 𝑅 , then 𝑓 𝑟𝑟 ≥ 𝑓 𝑎𝑟. Hence 𝑓 𝑟𝑟 − 𝑚𝑎𝑥( 𝑓 𝑎𝑟, 𝑓 𝑟𝑟) = 0 yielding
We have all the building blocks needed to encode the EER as a loss function, which is shown in Alg. 5. We simply estimate 𝑑 𝐸𝐸 𝑅 with Alg. 4, compute the FRR and FAR at 𝑑 𝐸𝐸 𝑅 using Algorithms 1 and 2, take the EER to be its average, and use it as the loss value. This algorithm fulfills the Problem Statement of Section III. Algorithm 5: EER Loss Data: 𝐿, 𝐼, 𝑛 > 0 Result: 𝑒𝑒𝑟 𝑑 𝐸𝐸 𝑅 ← 𝐷𝐸 𝐸 𝑅(𝐿, 𝐼, 𝑛) return [𝐹 𝐴𝑅𝑆 (𝐼, 𝑑 𝐸𝐸 𝑅 ) + 𝐹 𝑅𝑅𝑆 (𝐿, 𝑑 𝐸𝐸 𝑅 )]/2
Although the above algorithm encodes the EER objective directly, it is not optimal for training a neural network. To see why, note that the EER value can only be reduced by eliminating false negatives and positives. This is an inherently discrete process that doesn’t provide a clear target for optimization when no point can be made to cross the 𝑑 𝐸𝐸 𝑅 threshold.
𝑐 𝐿 = 𝑒𝑥 𝑝(𝐾 ∗ 0) = 1 whereas if 𝑑 > 𝑑 𝐸𝐸 𝑅 , then 𝑓 𝑟𝑟 < 𝑓 𝑎𝑟 and 𝑓 𝑟𝑟 − 𝑚𝑎𝑥( 𝑓 𝑎𝑟, 𝑓 𝑟𝑟) < 0 Thus, 𝑐 𝐿 ≈ 0 because the term inside the exponential is negative and large. Similarly, 𝑐 𝑅 marks whether 𝑑 ≥ 𝑑 𝐸𝐸 𝑅 is true or not. Using this subdifferentiable binary search step, Alg. 4 implements a subdifferentiable binary search for 𝑑 𝐸𝐸 𝑅 given
D. Area-Based Formulation of the Loss Function To overcome the limitation discussed in the previous subsection, we propose a loss function that measures the area of overlap between the FAR and FRR curves. As described in Section IV, the FAR and FRR curves are stepwise, with steps of 1/|𝐼 | and 1/|𝐿| respectively. Thus, letting 𝐴 𝑅 (𝐿) be the
5
area of the FRR curve to the right of 𝑑 𝐸𝐸 𝑅 and 𝐴 𝐿 (𝐼) the area of the FAR curve to the left of 𝑑 𝐸𝐸 𝑅 , we have 𝐴 𝑅 (𝐿) =
1 |𝐿|
1 𝐴 𝐿 (𝐼) = |𝐼 |
∑︁
(𝑑𝑖𝐿 − 𝑑 𝐸𝐸 𝑅 )
(15)
2, to delinearize the contribution of points further from 𝑑 𝐸 𝑅𝑅 , in the form !𝛽 ∑︁ max {𝜖, 𝑑𝑖𝐿 − 𝑑 𝐸𝐸 𝑅 } 1 (19) 𝐴 𝑅 (𝐿) = |𝐿| 𝑑 𝐸𝐸 𝑅 𝐿 𝑑𝑖 ∈ 𝐿 𝑑𝑖𝐿 >𝑡𝐿 ( 𝛼)
𝑑𝑖𝐿 ∈ 𝐿 𝑑𝑖𝐿 >𝑑𝐸𝐸𝑅
∑︁
(𝑑 𝐸𝐸 𝑅 − 𝑑 𝐼𝑗 )
(16)
𝑑 𝐼𝑗 ∈ 𝐼 𝑑 𝐼𝑗 <𝑑𝐸𝐸 𝑅
∑︁
max {𝜖, 𝑑 𝐸𝐸 𝑅 − 𝑑 𝐼𝑗 }
𝑑 𝐼𝑗 ∈𝐼
!𝛽 (20)
𝑑 𝐸𝐸 𝑅
𝑑 𝐼𝑗 <𝑡𝑅 ( 𝛼)
Hence, the total area of overlap between the FAR and FRR curves is given by 𝐴(𝐿, 𝐼) = 𝐴 𝑅 (𝐿) + 𝐴 𝐿 (𝐼)
(17)
By itself, the area of overlap does not provide a good target for minimization, as it can be easily exploited by the neural network under training without improving the intended authentication performance. Preliminary experiments showed that, to avoid this undesirable behavior, the area of overlap needs to be scaled down by 𝑑 𝐸𝐸 𝑅 , giving L 𝐴𝑅𝐸 𝐴 (𝐿, 𝐼) =
1 𝐴 𝐿 (𝐼) = |𝐼 |
𝐴(𝐿, 𝐼) 𝑑 𝐸𝐸 𝑅
(18)
Alg. 6 shows a subdifferentiable tensorial implementation of the above equation. Unlike the previous algorithms, this one is straightforward and requires no further explanation. Algorithm 6: Area-based Loss Data: 𝐿, 𝐼, 𝑛 > 0 Result: 𝑎𝑟𝑒𝑎 𝑑 𝐸𝐸 𝑅 ← 𝐷𝐸 𝐸 𝑅(𝐿, 𝐼, 𝑛) 𝐿 𝑅 ← 𝑚𝑎𝑥(0, 𝐿 − 𝑑 𝐸𝐸 𝑅 ) 𝐴 𝑅 ← 𝑠𝑢𝑚(𝐿 𝑅 ) / |𝐿| 𝐼 𝐿 ← 𝑚𝑎𝑥(0, 𝑑 𝐸𝐸 𝑅 − 𝐼) 𝐴 𝐿 ← 𝑠𝑢𝑚(𝐼 𝐿 ) / |𝐼 | return ( 𝐴 𝑅 + 𝐴 𝐿 )/𝑑 𝐸𝐸 𝑅
where the small 𝜖 term has been introduced to avoid arithmetic errors if the differences are too small. With the above modification, the 𝛼/𝛽 area-based loss function takes the form L 𝐴𝑅𝐸 𝐴 (𝐿, 𝐼) = 𝐴 𝑅 (𝐿) 1/𝛽 + 𝐴 𝐿 (𝐼) 1/𝛽
(21)
Algorithm 7 shows the subdifferentiable, tensorial implementation of the above formula. Algorithm 7: EERLoss (𝛼, 𝛽) Data: 𝐿, 𝐼, 𝑛 > 0, 𝛼 > 0, 𝛽 > 0 Result: 𝛼/𝛽 area 𝑑 𝐸𝐸 𝑅 ← 𝐷𝐸 𝐸 𝑅(𝐿, 𝐼, 𝑛) 𝑡 𝑅 ← (1 − 𝛼) ∗ 𝑑 𝐸𝐸 𝑅 𝐿 𝑅 ← 𝑝𝑜𝑤(𝑚𝑎𝑥(𝜖, 𝐿 − 𝑑 𝑅 )/𝑑 𝐸𝐸 𝑅 , 𝛽) 𝐴 𝑅 ← 𝑝𝑜𝑤(𝑠𝑢𝑚(𝐿 𝑅 ) / |𝐿|, 1/𝛽) 𝑡 𝐿 ← (1 + 𝛼) ∗ 𝑑 𝐸𝐸 𝑅 𝐼 𝐿 ← 𝑝𝑜𝑤(𝑚𝑎𝑥(𝜖, 𝑡 𝐿 − 𝐼)/𝑑 𝐸𝐸 𝑅 , 𝛽) 𝐴 𝐿 ← 𝑝𝑜𝑤(𝑠𝑢𝑚(𝐼 𝐿 ) / |𝐼 |, 1/𝛽) return 𝐴 𝐿 + 𝐴 𝑅
VI. Experimental Setup In order to validate the proposed loss function, experiments are designed to isolate its impact and compare it directly against state-of-the-art methods. All experiments are conducted based on the KVC-onGoing benchmark [13]. This benchmark uses the Aalto Desktop [14] and Aalto Mobile [15] datasets, which together comprise data from over 185,000 subjects. A. Evaluation Protocol
E. Generalizations It is customary for distance metric learning loss functions to include a parameter 𝛼 that enforces a larger separation between positive and negative samples. For example, the well-known TripletLoss [21] and, for keystroke dynamics verification, the recent SetMarginLoss [20] and Set2Set loss [12]. Extending equations (15) and (16) to include a margin is straightforward. Note that in the aforementioned equations, those values of 𝑑𝑖𝐿 and 𝑑 𝐼𝑗 further away from the thresholds contribute more to the sums, linearly. This is not optimal; the efforts of the loss function are better spent focusing on values near the threshold. For this purpose, we introduce a constant 𝛽 such that 0 < 𝛽 <
Our empirical evaluation protocol proceeds in two distinct stages: 1. Ablation Study (Sec. VII-A): The goal of this first phase is to efficiently select the optimal hyperparameters for the proposed loss described in Algorithm 7, i.e., EERLoss (𝛼, 𝛽) and to conduct a fair comparison against other losses, in terms of performance and training time. This study is performed using a reduced-parameter version of the main architecture (as specified in Sec. VI-B), trained on a subset of 1,000 users randomly sampled from the KVC-Desktop development set. Evaluation is then conducted on a fixed 1,000-user subset of the official KVC-Desktop evaluation set. 2. SoTA Comparison (Sec. VII-B): The second stage provides a definitive, large-scale comparison of the winning configuration against the strongest published SoTA baseline,
6
the Set2Set loss [12]. For this comparison, the full-capacity model architecture (Sec. VI-B) is used. The model is trained on the entire KVC development sets for both Desktop (168k users) and Mobile (37k users) tasks. Evaluation is performed on the full, official KVC evaluation sets, strictly adhering to the official benchmark protocol. The evaluation relies on the primary metrics defined by the KVC benchmark. The main ranking metric is the Global EER, which evaluates performance using a single, fixed decision threshold for all subjects. Additionally, the Mean per-subject EER (Avg. per-user EER) is reported. This metric is critical as it calculates an optimal, user-specific threshold, reflecting a more realistic deployment scenario on personal devices. To analyze performance under varying enrollment conditions, both EER metrics are evaluated as a function of 𝐺, the number of enrollment samples, where 𝐺 ∈ {1, 2, 5, 7, 10}.
VII. Results In this section, we present the empirical results of our proposed method. We first conduct an comparative study to determine the optimal configuration of EERLoss (Sec. VII-A). We then compare our optimized configuration against the SoTA baseline on the full KVC benchmark (Sec. VII-B). A. Comparative Study and Hyperparameter Selection Our EERLoss variant (Alg. 7) is controlled by the parameters 𝛼 and 𝛽. We first perform a grid search to find the optimal combination of these parameters. Figure 1 visualizes the resulting average per-user EER surface. We observe a clear global minimum achieved near 𝛼 = 0 and 𝛽 = 0.85.
B. Model Architecture
C. Baseline Loss Functions We compare our proposed EERLoss (both the direct formulation described in Alg. 5 and the refined version of Alg. 7) against a comprehensive set of SoTA and widely-used loss functions identified in the KVC analysis: • Semi-Hard Triplet Loss [3]: A foundational metric learning objective that minimizes the distance between an anchor and a positive sample while maximizing the distance to a negative sample. This was the most common baseline adopted by multiple teams in the KVC [13]. • ArcFace [4]: A widely-adopted loss that introduces an additive angular margin to the target logit, enforcing a more discriminative embedding space on the hypersphere. It was used by the VeriKVC team [13]. • CosFace [5]: A related approach that introduces an additive cosine margin to the target logits to improve intra-class compactness and inter-class separation. • Set2Set: This is the custom loss function used by the KVC winning LSIA team [12]. It is an extension of SetMargin Loss described in [20].
4 EER [%]
To isolate the impact of the loss function, our experiments are based on two scaled versions of a single base architecture. We adopt the model from the LSIA team [12], which won the KVC-onGoing Challenge [13]. This model is a dualbranch (RNN + CNN) embedding network. As detailed in [12], the recurrent branch consists of two bidirectional GRU layers with Self-Attention. The convolutional branch uses a stack of three 1D convolutions with progressively increasing filter counts (𝐹, 2𝐹, 4𝐹), interspersed with Channel Attention modules. Both branches use Temporal Attention at their input. The outputs of both branches are concatenated and passed through a final MLP to produce a fixed 256-dimensional embedding. The capacity of this architecture is controlled by two hyperparameters: the width of the GRU and dense layers (𝑊) and the base number of filters in the convolutional branch (𝐹). We define the two configurations for our experiments: the reduced-capacity one (𝑊 = 256, 𝐹 = 128) and the fullcapacity one (𝑊 = 512, 𝐹 = 1024).
3
2 5 0.0 04 0. 03 0. 02 0. 01 0. .0 𝛼 0
1 1. 1. 1.2 1.3 .4 5 1 0 1 0 . 0. 0.6 0.7 .8 9 5 𝛽
Fig. 1. Average per-user EER (interpolated) for different values of 𝛼 and 𝛽 in Alg. 7. The minimum is achieved near 𝛼 = 0, 𝛽 = 0.85.
Table I details the configurations for all evaluated loss functions. The Parameters column denotes the optimal hyperparameters employed (even though more configurations where investigated), such as the margin (𝑚) and scale (𝑠) for the ArcFace and CosFace baselines, and the 𝛽 parameter for our EERLoss. The 𝑃 column indicates the number of distinct users (sets) per batch, a value maximized to fit GPU capacity, as higher 𝑃 is better. The preliminary findings indicate that standard losses are highly sensitive and frequently suboptimal in this task. However, Semi-Hard Triplet loss has been demonstrated to provide a stable baseline performance. Our proposed EERLoss demonstrates a significantly superior performance. While the base direct EER approximation of Alg. 5 is not competitive, the EERLoss (𝛽 = 0.85, 𝐾 = 40) variant of Alg. 7 outperforms all baselines. This enhancement is refined through the implementation of a data augmentation technique, where Gaussian noise with a 10 millisecond variance is introduced to the keystroke timing features, simulating minor temporal jitter. With this final configuration, we achieve the best overall Global EER: 7.74% in the hardest enrollment scenario (𝐺 = 1). However, we argue that the Avg. per-user EER is a more significant and practical metric for this application. Biometric verification on personal devices (desktop or mobile) naturally supports user-specific decision thresholds rather than a single
7
TABLE I Comparative study results for different loss functions and their hyperparameters. Parameter settings are described in the text, and P denotes the number of distinct users (sets) per batch. Alg. 5 refers to the direct formulation of the proposed loss, while Alg. 7 is the refined formulation. We compare the Global EER (%) and the Average per-user EER (%) varying the number of enrollment samples (G).
Loss
Parameters
Global EER (%)
P
Avg. per-user EER (%)
G=1
G=2
G=5
G=7
G=10
G=1
G=2
G=5
G=7
G=10
Semi-Hard Triplet [3]
-
-
8.28
6.75
5.67
5.60
5.35
6.92
5.83
5.27
5.29
6.04
ArcFace [4]
𝑚 = 0.2; 𝑠 = 16 𝑚 = 0.1; 𝑠 = 8
50 50
13.62 14.30
12.05 12.75
10.89 11.59
10.38 11.35
10.20 10.84
10.50 10.62
9.03 9.35
7.97 8.41
7.10 7.45
7.03 7.73
CosFace [5]
𝑚 = 0.1; 𝑠 = 8
50
13.05
11.67
10.23
9.97
9.70
10.28
8.73
8.02
6.85
7.53
Set2Set [12]
-
-
8.02
6.60
5.47
5.26
4.70
6.92
5.73
4.86
4.75
5.03
EERLoss (Alg. 5)
-
40
13.63
11.98
10.71
10.20
9.99
12.19
10.67
9.37
9.07
9.73
EERLoss (Alg. 7)
𝛽=1 𝛽 = 0.85 𝛽 = 0.85 𝛽 = 0.85; data aug.
40 10 40 40
8.20 11.83 8.58 7.74
6.81 10.10 6.86 6.27
5.84 8.77 6.00 5.41
5.70 8.49 5.71 5.29
5.58 8.42 5.54 4.80
6.83 10.04 6.89 6.51
5.67 8.22 5.39 5.07
4.94 6.84 4.68 4.51
4.91 5.97 3.86 4.28
5.37 5.94 3.64 4.61
14
EER (%)
12 10 8 6 4
5
10
20
30
Samples per Class (K)
40
Fig. 2. Effect of Batch Size. Error decreases consistently as the number of samples per class grows, yet the method remains stable even at 5 samples per batch.
global one. In this more realistic evaluation setting, the proposed system (without data augmentation) achieves an average per-user EER of 3.64% for 𝐺 = 10 enrollment samples and 3.86% for 𝐺 = 7. This is a remarkable result, especially given that this comparative study was performed on a reduced dataset with a smaller-capacity model. To further validate the stability of the proposed objective, we analyzed the impact of batch composition, in particular, the number of samples per class. As illustrated in Figure 2, while the error rate consistently decreases as the number of samples per class increases (due to a better approximation of the global distribution), EERLoss converges effectively and remains stable even with minimal sampling (e.g., 5). This robustness to batch sparsity represents a significant operational advantage over traditional margin-based losses, which often require large, carefully mined batches to approximate global class centers.
B. Comparison with State-of-the-Art Following the comparative study, we compare our optimized model using the proposed loss EERLoss with 𝛼 = 0, 𝛽 = 0.85, 𝑃 = 40, against the same model with the strongest SoTA baseline, the Set2Set loss. Both methods are trained from scratch on the full KVC-onGoing development sets (Desktop
and Mobile) using the full-capacity model architecture, as described in Sec. VI-B, to ensure a fair and direct comparison. The final results are presented in Table II. Our proposed EERLoss outperforms the Set2Set baseline across both challenging, large-scale scenarios. On the Desktop Task, the performance gains are substantial. The proposed method achieves a Global EER of 1.15% (at 𝐺 = 10), which corresponds to a 30.7% relative reduction in error compared to the 1.66% achieved by the Set2Set baseline. The improvement in the more practical Avg. per-user EER is even more pronounced. We achieve 0.92% (at 𝐺 = 10) and 0.87% (at 𝐺 = 7), demonstrating a clear advantage over the baseline’s 1.29% and 1.27% (both with a relative reduction of approx. 30%). On the Mobile Task, the efficacy of our proposed method is also evident, outperforming the baseline in four out of the five enrollment scenarios for the Avg. per-user EER. It is notable that EERLoss shows a particularly strong performance advantage in the few-shot enrollment scenarios (𝐺 = 1 and 𝐺 = 2). For example, in the Desktop task at 𝐺 = 1, we achieve a Global EER of 2.74% vs. 3.20% for Set2Set. This suggests that the proposed method produces a more robust and well-structured embedding space, which allows for a more accurate definition of a user’s identity from fewer samples. C. Computational Efficiency and Parallelism Beyond verification accuracy, computational cost is a critical factor for training large-scale models. Table III reports the total training time required for our two experimental stages. The comparison study (Exp. 1) was conducted on an NVIDIA RTX 3070. Due to the significant computational demands of training on the full-scale datasets, the SoTA comparison (Exp. 2) was performed on a NVIDIA A100. All experiments used the AdamW optimizer with an initial learning rate of 1 × 10−4 and a patience of 40 epochs. In both settings, the proposed EERLoss is dramatically more efficient. In Exp. 1, EERLoss converges in only 23 minutes (vs. 4h 46min for Set2Set). This efficiency scales directly to the large-scale task (Exp. 2), where EERLoss converges in 6h 52min, a greater than 13.5×
8
TABLE II Quantitative comparison of our proposed EERLoss against the Set2Set (SoTA) baseline on the full KVC Mobile and Desktop tasks. We report Global EER (%) and Avg. per-user EER (%) as a function of the number of enrollment samples (G).
Scenario
Global EER (%)
Loss
Avg. per-user EER (%)
G=1
G=2
G=5
G=7
G=10
G=1
G=2
G=5
G=7
G=10
Mobile
Set2Set [12] EERLoss (Alg. 7)
3.44 3.35
2.53 2.53
2.11 2.04
2.02 2.03
2.03 2.04
2.41 2.31
1.97 1.68
1.49 1.42
1.51 1.57
1.71 1.60
Desktop
Set2Set [12] EERLoss (Alg. 7)
3.20 2.74
2.36 1.79
1.80 1.37
1.69 1.29
1.66 1.15
2.36 2.05
1.65 1.43
1.34 0.92
1.27 0.87
1.29 0.92
TABLE III Training time comparison. Exp. 1 (Comparison Study) was run on an NVIDIA RTX 3070. Exp. 2 (SOTA Comparison - Desktop Task) was run on an NVIDIA A100.
Experiment
Loss Function
Training Time
Exp. 1 (VII-A)
Set2Set [12] Semi-Hard Triplet [3] CosFace [5] ArcFace [4] EERLoss (Alg. 7)
4h 46min 2h 7min 46 min 42 min 23 min
Exp. 2 (VII-B)
Set2Set [12] EERLoss (Alg. 7)
92h 47min 6h 52min
speedup over the 92h 47min required by the Set2Set baseline. This shows that EERLoss achieves its state-of-the-art accuracy (Table I) through a more efficient training objective, with a training cost even better than the simple ArcFace and CosFace baselines. This efficiency of the proposed algorithm is remarkable. Alg. 7 reveals that the loss computation is embarrassingly parallel. This design avoids the iterative or complex samplemining bottlenecks of other objectives, allowing it to scale efficiently on modern GPU hardware.
VIII. Theoretical Grounding This section provides a theoretical justification for the performance gains observed with the area formulation of EERLoss relative to state-of-the-art alternative functions. Although angular losses are extensively adopted and optimized for physiological biometrics, they demonstrate suboptimal performance in comparison to EERLoss when applied to behavioral modalities such as keystroke dynamics. The underlying reasons for this phenomenon require further clarification. A fundamental distinction between physiological and behavioral biometrics lies in their relative levels of signal stability. Physiological traits, such as faces or fingerprints, are anchored to permanent anatomical structures and thus exhibit significantly lower intra-subject variability than behavioral modalities. In the physiological domain, noise is typically extrinsic, arising from environmental interference or sensor corruption of an otherwise constant biological trait. Conversely, behavioral biometrics demonstrate inherent variability. In contrast to physiological modalities, behavioral profiles lack a fixed representation. In this case, intra-subject variability is
not just a measurement artifact but an inherent property of the signal itself. While a face remains a definitive reference for a subject, there is no such definitive keystroke dynamics profile. Consequently, while physiological biometrics are characterized by well-behaved, predictable distributions, behavioral biometrics are inherently stochastic and non-ideal. Angular loss functions impose four a priori assumptions on the structure of the embedding space: one topological and three geometrical. Topologically, they assume that embedding clusters lie on the surface of a hypersphere; that is, the manifold is closed and every geodesic eventually wraps back on itself. Geometrically, they assume that clusters have comparable radii, that they are approximately equidistant, and that inter-cluster overlap is negligible, allowing a separating margin. Like the physiological signals they encode, the resulting embedding clusters are well-behaved: compact, separable, and geometrically regular. However, when applied to behavioral biometrics, these a priori assumptions become a liability. When the underlying dimensions of the space are not closed, projecting embeddings onto a hypersphere hinders classification by imposing false topological regularity. When inter-cluster overlap is significant, angular loss functions expend capacity attempting to enforce separation that the data cannot support; intrinsically inseparable clusters function as hard triplets, which are welldocumented to destabilize training in metric learning frameworks. Similarly, constraining clusters to comparable radii distorts the embedding space when intra-subject variability profiles differ substantially across subjects. EERLoss, by contrast, imposes no such priors and is free to learn the true topology and geometry of the embedding space from data. We now substantiate the claim that keystroke dynamics profiles, across sufficiently large user populations, violate at least the three geometrical a priori assumptions imposed by angular loss functions: comparable cluster radii, approximate equidistance, and negligible inter-cluster overlap. The geometric constraints are easier to disprove than the topological one. Figure 3 (a) shows the distribution of persubject cluster radii. Rather than concentrating around a common value, the distribution is broad, with radii spanning from 3 to 13, a fourfold range, indicating that cluster radii are far from uniform. Similarly, Figure 3 (b) shows a broad distribution of distances to the nearest cluster center, disproving approximate equidistance. Finally, Figure 3 (c) shows the distribution of cluster overlap between each cluster and its ten nearest neigh-
9
(a)
(b) 1400
1400
120
1200
1200
100
1000
1000
Count
140
80
Count
1600
160
Number of Subjects
(c)
1600
800
800
60
600
600
40
400
400
20
200
200
0
0
4
6
8
10
12
Cluster Radius (mean 2)
14
2
4
6
8
10
12
Cluster Distance
14
16
18
0
0.0
0.2
0.4
0.6
Cluster Overlap
0.8
1.0
Fig. 3. Empirical violation of geometric priors. We illustrate the non-ideal nature of keystroke embeddings through the distribution of: (a) cluster radii, showing high variance; (b) inter-cluster distances, showing non-equidistance; and (c) cluster overlap, which is non-negligible for the majority of subjects.
bours. For clusters 𝑖 and 𝑗, overlap is defined as 𝑑𝑖 𝑗 𝑜𝑖 𝑗 = max 0, 1 − 𝑟𝑖 + 𝑟 𝑗
(22)
where 𝑑𝑖 𝑗 is the Euclidean distance between centroids and 𝑟 𝑖 , 𝑟 𝑗 are the respective cluster radii. Only a minority of nearest-neighbour cluster pairs are disjoint (leftmost bar, label 0.0, subfigure c); most exhibit overlaps averaging around 50% and extending well beyond, directly violating the assumption of negligible inter-cluster overlap.
Fig. 5. Distribution of subject clusters in the embedding space on the 1,000user test set (ArcFace), dimensionality reduced via t-SNE. ArcFace enforces a more uniform use of the embedding space, in alignment with its a priori assumptions, at the cost of greater inter-cluster overlap compared to EERLoss.
Fig. 4. Distribution of subject clusters in the embedding space on the 1,000-user test set (EERLoss), dimensionality reduced via t-SNE. Rather than concentrating on a closed surface, clusters distribute along open, irregular manifolds.
Figure 4 shows a random selection of 50 clusters in the embedding space, reduced via t-SNE. The figure illustrates the violation of geometric assumptions at a glance: clusters exhibit significant mutual overlap and broad variation in radii and inter-cluster distances. Notably, the figure also refutes
the topological assumption. Although dimensionality reduction via t-SNE introduces known local geometric distortions, the macroscopic open-manifold arrangement observed here remains fundamentally incompatible with a closed, hyperspherical prior. Instead, the clusters are arranged as elongated, irregular manifolds emanating from a common origin, consistent with an open rather than closed topology. Figure 5 shows 50 clusters under ArcFace training, reduced via t-SNE for direct comparison. Here, the visual projection reflects ArcFace’s a priori assumptions: the space is utilized more uniformly, suggesting a tendency towards comparable cluster radii and equidistant centroids. However, the trade-off of imposing this rigid geometry is visible in the boundaries between clusters, which are noticeably more porous than in
10
TABLE IV Preliminary evaluation on Face Verification (LFW) under different data regimes to assess generalization limits.
Setup (𝑁𝑖𝑚𝑔 )
Loss
Acc (%)
EER (%)
Gain (EER)
16 imgs/ID
ArcFace EERLoss
96.76 95.97
3.38 4.08
– -0.7%
8 imgs/ID
ArcFace EERLoss
92.08 93.28
7.55 6.80
– +0.75%
4 imgs/ID
ArcFace EERLoss
86.43 87.96
13.92 11.75
– +2.17%
Figure 4, exhibiting greater overlap between neighbouring subjects. The resulting closed, radial arrangement stands in clear topological contrast to the open manifolds naturally learned by EERLoss. The superior performance of EERLoss in behavioral biometric tasks, and in keystroke dynamics in particular, can now be explained as stemming from its ability to adapt to the underlying geometry of the embedding space, rather than expending capacity enforcing a rigid geometric structure that the data do not support. Conversely, on theoretical grounds, angular loss functions should outperform EERLoss in wellbehaved classification tasks such as those arising in physiological biometrics, where their a priori assumptions align with the intrinsic structure of the data.
A. Case Study on Face Recognition Under Data Scarcity The substantial improvements reported in our results indicate that EERLoss effectively exploits the irregular nature of behavioral biometrics. To explore the boundaries of this performance and test how our approach generalizes to more structured data, we conducted a preliminary evaluation on face recognition, one of the main physiological modalities, characterized by more stable geometric profiles. We trained an IResNet50 on subsets of CASIA-WebFace [22] and evaluated on LFW [23], deliberately limiting the training samples per identity (𝑁𝑖𝑚𝑔 ) to simulate different levels of data scarcity. Table IV presents the results of this small-scale comparison. As expected, in a well-behaved modality like face recognition with abundant data (𝑁𝑖𝑚𝑔 = 16), ArcFace [4] maintains its superiority, as its hyperspherical priors align well with the densely sampled embedding space. However, as we move towards more restrictive regimes, the gap narrows. Notably, in the most extreme case of data scarcity (𝑁𝑖𝑚𝑔 = 4), EERLoss manages to outperform ArcFace, reducing the EER from 13.92% to 11.75%. This suggests that while angular losses remain the optimal choice for physiological biometrics under ideal sampling conditions, EERLoss offers a competitive alternative when global geometric proxies (like class centers) cannot be robustly estimated. More importantly, it confirms that the strength of EERLoss lies in its ability to operate without rigid assumptions, making it particularly valuable for biometric traits where intra-subject variability is high and training samples are limited.
IX. Conclusion In this paper we introduce EERLoss, a novel loss function that enables direct end-to-end optimization of the target metric. A rigorous evaluation was conducted on the large-scale KVConGoing benchmark. The empirical results are conclusive: EERLoss significantly outperforms the SoTA, achieving a relative EER reduction of up to 30.7%. Furthermore, the proposed method converges substantially faster compared to other losses, reducing the SoTA’s overall training cost from 92 hours and 47 minutes to only 6 hours and 52 minutes in the final experiment. These results strongly validate this approach and suggest that direct metric optimization, via EERLoss, is a superior paradigm for training deep biometric verification models. The scope of this study’s empirical validation was focused on the keystroke dynamics modality. This was a deliberate methodological choice, using the KVC-onGoing benchmark as a rigorous, large-scale test known for its high intra-class and low inter-class variability. While there is already strong evidence of the method’s efficacy, future work will investigate the generalization of EERLoss to other verification tasks. Additionally, the underlying subdifferentiable approximation framework itself could be adapted to optimize other nondifferentiable evaluation metrics present in computer vision and beyond. X. Acknowledgment This project has been supported by Cátedra ENIA UAMVERIDAS en IA Responsable (NextGenerationEU PRTR TSI100927-2023-2) and TRUST-ID (PID2025-173396OB-I00 MICIU/AEI and the EU). Robledo-Moreno is supported by a FPI Fellowship (FPI-UAM-2025). References [1] A. K. Jain, A. Ross, and S. Pankanti, “Biometrics: A Tool for Information Security,” IEEE Transactions on information forensics and security, 2006. [2] S. Chopra, R. Hadsell, and Y. LeCun, “Learning a Similarity Metric Discriminatively, with Application to Face Verification,” in In Proc. of the IEEE Conference on Computer Vision and Pattern Recognition, 2005. [3] M. Schultz and T. Joachims, “Learning a Distance Metric from Relative Comparisons,” Advances in neural information processing systems, 2003. [4] J. Deng, J. Guo, N. Xue et al., “Arcface: Additive Angular Margin Loss for Deep Face Recognition,” in In Proc. of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019. [5] H. Wang, Y. Wang, Z. Zhou et al., “Cosface: Large Margin Cosine Loss for Deep Face Recognition,” in In Proc. of the IEEE Conference on Computer Vision and Pattern Recognition, 2018. [6] M. Kim, A. K. Jain, and X. Liu, “Adaface: Quality adaptive margin for face recognition,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022. [7] R. Shadman, A. A. Wahab, M. Manno et al., “Keystroke Dynamics: Concepts, Techniques, and Applications,” ACM Computing Surveys, 2025. [8] P. Delgado-Santos, R. Tolosana et al., “M-GaitFormer: Mobile Biometric Gait Verification Using Transformers,” Engineering Applications of Artificial Intelligence, 2023. [9] R. Tolosana, R. Vera-Rodriguez, J. Fierrez et al., “DeepSign: Deep On-Line Signature Verification,” IEEE Transactions on Biometrics, Behavior, and Identity Science, 2021. [10] E. Maiorana et al., “Signature Biometrics,” in Encyclopedia of Cryptography, Security and Privacy. Springer, 2025.
11
[11] J. Terven, D.-M. Cordova-Esparza, J.-A. Romero-Gonzalez et al., “A Comprehensive Survey of Loss Functions and Metrics in Deep Learning,” Artificial Intelligence Review, 2025. [12] N. Gonzalez, G. Stragapede, R. Vera-Rodriguez et al., “Type2Branch: Keystroke Biometrics Based on a Dual-Branch Architecture With Attention Mechanisms and Set2set Loss,” IEEE Transactions on Information Forensics and Security, 2025. [13] G. Stragapede, R. Vera-Rodriguez, R. Tolosana et al., “KVC-onGoing: Keystroke Verification Challenge,” Pattern Recognition, 2025. [14] V. Dhakal, A. M. Feit, P. O. Kristensson et al., “Observations on Typing from 136 Million Keystrokes,” in In Proc. of the 2018 CHI Conference on Human Factors in Computing Systems, 2018. [15] K. Palin, A. M. Feit, S. Kim et al., “How do People Type on Mobile Devices? Observations from a Study with 37,000 Volunteers,” in In Proc. of the 21st International Conference on Human-Computer Interaction with Mobile Devices and Services, 2019. [16] R. Giot, M. El-Abed, and C. Rosenberger, “GREYC Keystroke: A Benchmark for Keystroke Dynamics Biometric Systems,” in In Proc. IEEE International Conference on Biometrics: Theory, Applications, and Systems, 2009. [17] M. El-Abed, M. Dafer, and R. El Khayat, “RHU Keystroke: A MobileBased Benchmark for Keystroke Dynamics Systems,” Proceedings International Carnahan Conference on Security Technology, 2014. [18] C. Murphy, J. Huang, D. Hou et al., “Shared Dataset on Natural HumanComputer Interaction to Support Continuous Authentication Research,” in In Proc. IEEE International Joint Conference on Biometrics, 2017. [19] A. Acien, A. Morales, J. Fierrez et al., “BeCAPTCHA: Behavioral Bot Detection using Touchscreen and Mobile Sensors benchmarked on HuMIdb,” Engineering Applications of Artificial Intelligence, 2021. [20] A. Morales, J. Fierrez, A. Acien et al., “SetMargin Loss Applied to Deep Keystroke Biometrics with Circle Packing Interpretation,” Pattern Recognition, 2022. [21] G. Chechik, V. Sharma, U. Shalit et al., “Large Scale Online Learning of Image Similarity Through Ranking.” Journal of Machine Learning Research, vol. 11, no. 3, 2010. [22] D. Yi, Z. Lei, S. Liao, and S. Z. Li, “Learning Face Representation from Scratch,” arXiv preprint arXiv:1411.7923, 2014. [23] G. B. Huang, M. Mattar, T. Berg, and E. Learned-Miller, “Labeled Faces in the Wild: A Database for Studying Face Recognition in Unconstrained Environments,” in Workshop on faces in’Real-Life’Images: detection, alignment, and recognition, 2008.
XI. Biography Section Nahuel González received his Ph.D. degree in Computing Science from Universidad Nacional de La Plata (UNLP), Argentina, in 2022. His Ph.D. thesis was awarded the Raúl Gallard Prize, handed by the Network of Argentinian Universities of Computing Science (RedUNCI), the following year. He has been affiliated with the Laboratorio de Sistemas de Información Avanzados (LSIA) of the University of Buenos Aires (UBA) since 2013. His main research interests are behavioral biometrics and time series prediction/classification using deep learning. He is a member of the editorial board of Data in Brief, Elsevier, since 2024. Marta Robledo-Moreno received her B.Sc. and M.Sc. degrees in Telecommunication Engineering from Universidad Autonoma de Madrid (Spain) in 2022 and 2024, respectively. The beginning of her research career was at the Biometric System Laboratory of the University of Bologna (Italy). Subsequently, she joined the BiometricsAI Lab (UAM), where she is currently collaborating as an assistant researcher pursuing a Ph.D. degree. Her research interests are mainly focused on signal processing and deep learning for behavioral biometrics.
Ivan Deandres-Tame received the B.Sc. in Computer Science and Engineering in 2021 and the M.Sc. in Deep learning for Image, Audio and Video processing in 2022 from the Universidad Autonoma de Madrid. In May 2023, he joined the BiometricsAI Lab (UAM), where he is currently collaborating as an assistant researcher pursuing a Ph.D. degree. The research activities he is currently working are focused in Person recognition in complex scenarios and the use of synthetic data to mitigate capture challenges. Ruben Vera-Rodriguez received his PhD degree in electrical and electronic engineering from Swansea University, U.K., in 2010. Since then, he has been affiliated with the BiometricsAI Lab, Universidad Autonoma de Madrid, Spain, where he is currently an Associate Professor since 2018. His research interests include signal and image processing, pattern recognition, HCI and biometrics, with emphasis on signature, face, gait verification, mobile biometrics and forensic applications of biometrics. He is actively involved in several national and European projects focused on biometrics. He has been awarded recently with a Medal from the Spanish Royal Academy of Engineering for his research contributions. He is member of ELLIS Society. Ruben Tolosana received the M.Sc. degree in Telecommunication Engineering, and the Ph.D. degree in Computer and Telecommunication Engineering, from Universidad Autonoma de Madrid, in 2014 and 2019, respectively. In 2014, he joined the BiometricsAI Lab at the Universidad Autonoma de Madrid, where he is currently an Assistant Professor. He is a member of the ELLIS Society, the Technical Area Committee of EURASIP, and the Editorial Board of the IEEE Biometrics Council Newsletter. His research interests are mainly focused on signal and image processing, pattern recognition, and machine learning, particularly in the areas of DeepFakes, Human-Computer Interaction, Biometrics, and Health. Dr. Tolosana has also received several awards such as the European Biometrics Industry Award (2018) from the European Association for Biometrics (EAB) and the Best Ph.D. Thesis Award in 2019-2022 from the Spanish Association for Pattern Recognition and Image Analysis (AERFAI).