Conceptio › Archive › arXiv CS
arXiv CSopen access

Domain-Incremental Learning for Multi-Channel Replay Speech Detection

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

DOMAIN-INCREMENTAL LEARNING FOR MULTI-CHANNEL REPLAY SPEECH DETECTION Michael Neri

arXiv:2609.11194v1 [eess.AS] 10 Sep 2026

Faculty of Information Technology and Communication Sciences, Tampere University, Tampere, Finland [email protected] task-specific beamformer heads

ABSTRACT Replay attacks are the most accessible threat to voice-controlled systems, and the acoustic cues that expose them are strongly modulated by the environment in which the attack is mounted. A detector deployed in the field therefore has to absorb new acoustic conditions over time, ideally without revisiting past recordings, since retaining speech indefinitely is both expensive and legally constrained. We frame this as domain-incremental learning (DIL) over acoustic environments and present the first continual learning benchmark for multi-channel replay speech detection, evaluating a state-of-theart beamformer-based detector over all 24 environment orderings of the ReMASC corpus with five seeds. Sequential fine-tuning forgets severely, raising the error rate on previously learned environments by 18.8 points. Elastic weight consolidation (EWC) halves forgetting but loses plasticity, gradient projection memory (GPM) is statistically indistinguishable from naive fine-tuning, and the proposed task-specific beamformer (TSB) that keeps one spatial front-end per environment significantly improves final and incremental accuracy. We further show that the last environment of the sequence dominates final performance. Code, results, and analysis are available at https://github.com/michaelneri/replay-speech-continual. Index Terms— Continual Learning, Physical Access, Deep Learning, Deepfake Detection, Spatial Audio. 1. INTRODUCTION Voice-controlled systems are now a standard interface to smart speakers, smartphones, and voice-based banking, where automatic speaker verification (ASV) separates a legitimate user from an impostor. Among the attacks that defeat it, replay is by far the most accessible [1, 2]. In fact, an adversary only needs to record a victim’s utterance and play it back to the target device, with no technical expertise and no access to synthesis or conversion models. A replayed utterance goes through one acquisition-reproduction cycle more than a genuine one, so traces are left at the spoofing stage as well. The attacker’s microphone and the playback loudspeaker add their own frequency response and non-linear distortion, and the reverberation of the room where the utterance was captured is already imprinted in the signal reaching the target device. Detectors look for these traces, and recent ones exploit the multiple microphones already present in commercial devices to extract spatial cues that a single channel cannot provide [3, 4, 5]. The cues separating the two classes are spatial. A talker is a compact source at a distance and orientation typical of a person addressing the device, whereas a loudspeaker sits closer to the array, radiates with a different directivity, and emits an already reverberated signal, so that inter-channel time and level differences, coherence, and direct-to-reverberant ratio differ between the two. These

(1)

fBM (2)

fBM XSTFT

CN ×T ×F

. . .

shared CRNN classifier gϕ

1 K

P

softmax

ŷ genuine / replay

(k)

fBM X̂ (i) ∈ CT ×F {θ (i) }i<k frozen

z ∈ R2 θ (k) , ϕ optimised

Fig. 1. Task-specific beamformer (TSB). When the model is adapted to the k-th domain, only the head θ(k) and the shared classifier ϕ are optimised (solid path), while all earlier heads {θ(i) }i<k are frozen (dashed paths). At inference the logits of the heads trained so far are averaged before the softmax, so no domain label is required. cues are modulated by the enclosure itself, hence they manifest differently in an office, in a corridor, or inside a moving car, and detectors degrade when the recording conditions change with respect to training [3]. A deployed device keeps encountering new acoustic conditions, and retraining from scratch on the union of all data is impractical. Additionally, storing past recordings for later reuse is also undesirable and impractical, as speech is personal and it is classified as biometric data [6]. Continual learning [7, 8] addresses exactly this scenario, in which a model is updated on a stream of data increments. Updating a replay detector as new acoustic environments are encountered one after the other is a domain-incremental problem, the scenario in which the label space is fixed and only the input distribution changes. Unlike generalisation to unseen environments, the detector is expected to learn each environment it is deployed in while retaining the previous ones [9, 10]. An exemplar-free regularisation approach was introduced in [11] to adapt a detector to new attacks without revisiting past data, and subsequent work refined the weight-modification strategy to better balance stability and plasticity [12, 13], with further contributions using feature distillation and class balancing [14]. These works, however, are single-channel and attack-incremental: each increment introduces new spoofing algorithms while the acoustic conditions remain comparatively fixed. Physical-access replay poses a different problem, since the discriminative cue is modulated by the recording environment itself, which prior work has addressed with domain adaptation between environments [15], but not incrementally. The contributions of this work are as follows: (i) to the best of our knowledge, this is the first work to frame multi-channel physical-access replay detection as a domain-incremental problem over acoustic environments; (ii) we define the task-specific

beamformer (TSB), an architecture-based procedure that instantiates one learnable beamformer per environment on top of a shared classification convolutional recurrent neural network (CRNN) and requires no environment label at inference; (iii) we benchmark a stateof-the-art multi-channel detector together with TSB and two standard continual learning approaches over all environment orderings and five seeds, and show that the order of the environments affects the outcome more than the choice of algorithm.

2. METHOD This section first introduces the problem of the DIL for replay speech detection. Then, it describes the components of the study: the multichannel replay speech detector we adopt and the continual learning strategies we compare, including the proposed one.

2.3. Spatial Continual Learning Approach We propose task-specific beamformer (TSB), an architecture-based continual learning strategy that decouples M-ALRAD [4] into an environment-specific spatial front-end and an environment-agnostic classifier, shown in Figure 1. This design follows the observation that the multi-channel cues exploited by the beamformer are strongly environment-dependent (array geometry, reverberation, noise field). Reusing the beamformer fBM : CN ×T ×F → CT ×F , we instan(k) tiate K independent heads {fBM }K k=1 , one per domain/environment, (k) each with parameters θ and producing a beamformed spectro(k) gram X̂ (k) = fBM (XSTFT ). All heads share a single CRNN classifier gϕ , mapping a beamformed spectrogram to the class logits z = gϕ (X̂ (k) ) ∈ R2 . When the model is adapted to the k-th domain, only its head θ(k) and the shared CRNN classifier ϕ are optimised, while all earlier heads {θ(i) }i<k are frozen:     θ(k) , ϕ = arg min L gϕ X̂ (k) , y ,

2.1. DIL for Replay Speech Detection Replay speech detection is the binary classification of a multichannel recording as genuine, i.e., uttered by a talker in front of the array, or replayed, i.e., a previously captured utterance reproduced by a loudspeaker. In this work a detector is updated on a stream of K acoustic environments, observed one at a time and never jointly: since only the input distribution changes across increments, the setting is domain-incremental. At step k the model is initialised from the previous set of parameters and optimised on the new environment alone, without revisiting past utterances and without access to the environment identity at inference.

2.2. Replay Speech Detector In our experiments we employ the state-of-the-art replay speech detector M-ALRAD [4], a CRNN that jointly processes N complex short-time Fourier transforms (STFTs) {XSTFTn,T ,F , n = 1, . . . , N } of a multi-channel recording, with T and F time and frequency bins respectively, to produce a single-channel beamformed spectrogram. Each recording is resampled to 16 kHz, cropped to its first second, zero-padding shorter utterances, and peak-normalised before computing the STFTs with a 1024-point window and 50% overlap, which gives T = 32 and F = 513. First, a convolutional neural network (CNN) fBM : CN ×T ×F → CT ×F predicts the beamforming weights W ∈ CN ×T ×F through a Conv2D-BatchNorm-ELU-Conv2D stack, where real and imaginary parts are concatenated along the channel dimension. The beamformed spectrogram is then computed as X̂STFTT ,F = P X̂STFT is classified by a CRNN, n XSTFTn,T ,F · Wn,T,F . denoted with gϕ with ϕ the set of learnable parameters, previously used for speaker distance estimation [16, 17]. Its magnitude, together with the sine and cosine of its phase, forms a T × F × 3 tensor that is passed through three convolutional layers with 1 × 3 filters and batch normalization, each followed by parallel max and average pooling. Two bi-directional gated recurrent unit (GRU) layers with 128 neurons refine the feature maps, and the final hidden state is mapped to the binary prediction ŷ ∈ R2 by a fully connected layer with softmax activation function, where 0 denotes the genuine class and 1 the replay one. Training minimizes the labelweighted binary cross-entropy loss, together with the orthogonality and sparsity losses on W using the same hyperparameters as in [4].

θ (k) ,ϕ

(1)

where L is the binary weighted cross-entropy with the orthogonality and sparsity regularisers, as in [4]. Freezing past heads makes the spatial filters of earlier environments immutable, structurally preventing forgetting at the beamforming stage, while the shared classifier stays fully plastic so the discriminative cues keep improving. At inference no domain label is required: the input is processed by all trained heads and their logits are averaged before the softmax, ŷ = softmax

! K−1   1 X (k) gϕ X̂ , K

(2)

k=0

where at an intermediate domain only the heads trained so far take part in the average. The overhead over the baseline is one head per (k) domain, growing linearly with K; each fBM adds only ≈9.4k parameters, a small fraction of the shared classifier. 2.4. Continual Learning Baselines In this work, we focus on regularization-based, optimization-based, and architecture-based continual learning approaches. We discard replay- and representation-based [18] techniques because (i) speech is personal and biometric data under the GDPR [6], hence retaining utterances of past environments for the whole life of the model conflicts with data minimisation, storage limitation and the right to erasure; and (ii) representation-based methods are built on large-scale pre-trained or self-supervised encoders, which for multi-channel spatial audio do not exist in the literature. We adapt two continual learning strategies to the multi-channel replay speech detection task: (i) Elastic weight Consolidation (EWC) [19] (regularizationbased) away from the previous solution through P penalises ⋆movement 2 λ F (θ − θ ) , where F is the diagonal empirical Fisher i i k−1,i i 2 accumulated over one pass on the environment just learned; we keep a single anchor, so the carried state stays at 2|θ| scalars, and set λ = 5 000. (ii) Gradient Projection Memory (GPM) [20] (optimization-based) constrains the direction of the update instead: after each step the per-batch parameter gradients of every layer are stacked and a singular value decomposition (SVD) retains the leading vectors covering εth = 0.97 of the singular-value energy; the resulting bases are extended online and later gradients are projected onto their orthogonal complement, g ← g − M M ⊤ g. Both EWC and GPM hyperparameters are reported in Table 1.

Table 1. Configuration of the DIL benchmark: setting shared by all methods (top) and method-specific hyper-parameters (bottom). Shared setting Input STFT Optimiser Scheduler Batch / epochs per env. Model selection Sequences

4-channel, 1 s, 16 kHz (from 44.1 kHz) NFFT = 1024, hop = 512 Adam, η = 10−4 , weight decay 10−4 cosine annealing, Tmax = 100, ηmin = 10−5 8 / 50, gradient clipping 1.0 lowest EER on the held-out partition 4! = 24 orderings × 5 seeds = 120

BWT =

Method-specific EWC GPM TSB

BWT — Backward Transfer. BWT measures how much the EER on each previously-learned environment j changed between the step at which j was first learned and the end of the full sequence:

λ = 5 000; diagonal empirical Fisher, single anchor εth = 0.97, residual tol. 0.1; parameter-gradient bases K = 4 heads, 9 416 params. each (+11.3%); logit ensemble

K−2 X  1 E[j, j] − E[K−1, j] . K − 1 j=0

(6)

A negative BWT means the final EER on j is higher than it was when j was just learned, i.e., forgetting occurred. IM — Intransigence Measure. IM requires a jointly-trained reference model: Eref [k] is the EER, evaluated on environment k, of a model trained on environments 0, . . . , k simultaneously. IM is the gap between the sequential (continual) model and the reference model at the final step: (7)

3. CONTINUAL LEARNING BENCHMARK DESIGN

IM = E[K−1, K−1] − Eref [K−1].

We define here how the approaches of Section 2 are evaluated. Detection performance at each step is measured with the equal error rate (EER), while the metrics below summarise an entire incremental sequence in terms of accuracy, forgetting, and transfer.

A positive IM means the continually-trained model could not reach the performance a jointly-trained model achieves on the same data. FWT — Forward Transfer. FWT requires single-task reference models: Esref [j] is the EER of a model trained only on environment j, with no exposure to any other environment. For every environment except the first, FWT compares this single-task baseline against the EER achieved by the continual model at the step where j was first introduced

3.1. Continual Learning Metrics Let K denote the number of sequentially-learned environments (K = 4 in our experiments), indexed 0, . . . , K − 1, and let E[k, j] denote the EER on environment j’s test set after the model has been trained sequentially on environments 0, . . . , k (only j ≤ k is meaningful). It is worth noting that for the computation of the metrics the indices j, k denote the position of an environment within a given ordering, not a fixed environment identifier. Following [18], we define the following metrics: AE — Average Error. The mean EER across all K environments after the full training sequence has completed

FWT =

4. EXPERIMENTAL RESULTS (3)

AE summarises overall error rate at deployment time, after the model has seen every environment. AIE — Average Incremental Error. Since AE only looks at the final step, AIE rewards models that perform well throughout training, not only at the end. Let AEk be the running average at intermediate step k. Then, we compute the mean over all steps AEk =

1 E[k, j], k + 1 j=0

AIE =

1 K

K−1 X

AEk .

(4)

k=0

FM — Forgetting Measure. For each environment j learned before the final step, FM compares its best-ever EER (over all steps where it could be evaluated, up to but excluding the final step) against its EER at the final step: FM =

(8)

A positive FWT means the model, having already learned earlier environments, performs better on a new environment j than a model trained on j alone, i.e., prior environments provided useful forward transfer.

K−1

1 X E[K−1, j]. AE = K j=0

k X

K−1 X  1 Esref [j] − E[j, j] . K − 1 j=1

K−2  X 1 E[K −1, j] − min E[i, j] . j≤i≤K−2 K − 1 j=0

(5)

A positive FM means the EER on environment j rose above its best previously-achieved value by the end of training, i.e., the model forgot what it had learned about j.

We first introduce the corpus which has been used to carry out the experiments. Then, we report the outcome of the benchmark over the 24 environment orderings and five seeds, first comparing the four strategies on the continual learning metrics and then isolating the effect of the position that each environment occupies in the sequence. 4.1. Dataset Experiments are conducted on Realistic Replay Attack Microphone Array Speech (ReMASC) [21], the only publicly available corpus providing synchronised multi-channel recordings of replay attacks, collected with four microphone arrays across four acoustic environments: an outdoor scenario (Env-A), two enclosed spaces (Env-B and Env-C), and a moving vehicle (Env-D). Among the four arrays we retain only D2, a linear array of four omnidirectional microphones sampled at 44.1 kHz, and leave the remaining ones to future work. Fixing the array makes the acoustic environment the only axis along which the domain changes: the arrays differ in the number of microphones and in geometry, which would alter the input dimensionality of the beamformer and confound the spatial mismatch with the environment shift. D2 provides 10,664 utterances (2,452 genuine and 8,212 replay), distributed over Env-A to Env-D as 1,942, 3,614, 2,423, and 2,685 samples respectively, and we adopt the speaker-disjoint partition released with the dataset, with 40 speakers for training and 11 for testing.

Table 2. Continual learning benchmark on ReMASC (D2, 16 kHz), pooled over 24 orderings × 5 runs (n = 120). ↓/↑ = lower/higher is better; best per column in bold. AE and AIE in %, others in percentage points. Values are mean ± 95% CI half-width (t-distribution). Significance is assessed against the naive fine-tuning baseline with Wilcoxon signed-rank tests paired by (ordering, run) (∗ p < 0.05, ∗∗ p < 0.01, ∗∗∗ p < 0.001); shading marks a significant (pW < 0.05) improvement / degradation . Algorithm Naive fine-tuning (baseline) EWC [19] GPM [20] TSB (ours)

AE ↓

AIE ↓

FM ↓

BWT ↑

IM ↓

FWT ↑

24.09 ±1.15 23.61 ±1.13 24.15 ±1.13 23.06∗ ±1.14

19.56 ±0.76 19.92 ±1.11 19.69 ±0.77 18.75∗∗∗ ±0.82

18.79 ±1.58 10.18∗∗∗ ±1.45 18.45 ±1.58 17.77 ±1.56

-18.66 ±1.59 -9.63∗∗∗ ±1.45 -18.36 ±1.59 -17.64 ±1.54

-4.34 ±1.09 1.84∗∗∗ ±1.47 -4.17 ±1.08 -6.69∗∗ ±1.70

3.42 ±0.62 -5.12∗∗∗ ±1.08 3.42 ±0.74 4.04 ±0.65

4.2. Results

Baseline EWC

GPM TSB

30 Mean AE (%)

Each algorithm is compared against the naive fine-tuning baseline with paired tests matched by (ordering, run), exploiting the fact that all algorithms are evaluated on the identical set of 4! × 5 = 120 environment-ordering and run combinations. Pairing removes the substantial variance attributable to ordering difficulty, isolating the effect of the algorithm itself. Following [22], we adopt the Wilcoxon signed-rank test [23] as the decisive criterion, as EER distributions are bounded and right-skewed. Table 2 shows that sequential finetuning forgets severely. Specifically, the EER of the baseline on the environments it has already learned rises by 18.79 points, and the sequence closes at 24.09% average EER. TSB is the only method that significantly improves the error-oriented metrics, reducing AE by 1.02 ± 0.92 points and AIE by 0.81 ± 0.42, while FM and BWT remain statistically unchanged: per-environment beamformer heads lower the EER by adapting to the acoustics of each environment rather than by preventing forgetting. TSB also attains the best intransigence (−2.35 ± 1.30 against the baseline), i.e., sequential training over-specialises to the environment seen last and surpasses the joint reference on it. EWC exhibits the expected stability-plasticity trade-off. The Fisher-based penalty anchors the weights to the previous optimum and nearly halves forgetting (10.18 against 18.79 points of FM, pW < 0.001), but the same rigidity prevents the model from reaching the joint-training optimum (+6.18 points of IM) and produces negative forward transfer (−8.54 points of FWT), showing that the penalties of previous environments interfere with the adaptation to new ones. GPM, instead, is statistically indistinguishable from the baseline on all six metrics: projecting gradients onto the null space of the subspaces of previous environments brings no measurable benefit when all environments share the same binary objective and their gradient directions largely overlap.

40

20

10

0 Outdoor

Indoor quiet

Indoor lounge

Vehicle

Fig. 2. Effect of the last environment of the sequence on the mean 1 AE (95% CI, t-distribution), averaged over all orderings ending with that environment (6 × 5 = 30 observations per bar). The four environments are the outdoor scenario (traffic and wind), the quiet enclosed space, the indoor lounge (background music and TV), and the moving vehicle. expense of previously learned environments. Notably, EWC partially mitigates this effect, achieving a mean AE of approximately 28% on the vehicle-last condition compared to 32–33% for the other methods. This is consistent with EWC’s Fisher-based regularisation, which prevents excessive weight drift when adapting to the final environment, inadvertently preserving representations learned from earlier and easier conditions. Environment ordering is therefore a practically relevant deployment variable, and EWC offers partial robustness to a difficult final condition at the cost of plasticity. 5. CONCLUSION

4.3. Analysis on the environment order Figure 2 shows the mean AE as a function of the last environment in the DIL sequence, averaged across all orderings ending with that environment and all five runs. The results reveal that the final environment in the sequence is the dominant factor driving overall performance, while the choice of first environment has negligible influence. Ending the sequence on the vehicle environment (moving car, non-stationary engine and road noise) consistently yields the highest EER across all four methods, with mean AE between 28% and 33%, compared to 19 − 21% when the indoor lounge environment (stationary background music and TV) is last. This pattern is primarily a data-driven effect: the vehicle environment is intrinsically the most acoustically challenging condition, and sequential fine-tuning on it last causes the model to over-adapt to its distribution at the

We presented the first continual learning benchmark for multichannel replay speech detection, framing the acoustic environments of ReMASC as a domain-incremental sequence and evaluating four exemplar-free strategies over all 24 orderings with five seeds. None of them removes catastrophic forgetting. EWC halves the forgetting measure but pays for it with negative forward transfer, GPM is indistinguishable from naive fine-tuning on all six metrics, and the proposed TSB improves final and incremental accuracy without reducing forgetting, indicating that its gain comes from spatial specialisation rather than from added stability. The environment that closes the sequence drives the final EER far more than the algorithm does. Future work will address increments over microphone arrays with different geometries and numbers of channels, and combine spatial specialisation with an explicit stability mechanism.

6. REFERENCES [1] Z. Wu, N. Evans, T. Kinnunen, J. Yamagishi, F. Alegre, and H. Li, “Spoofing and countermeasures for speaker verification: A survey,” Speech Communication, vol. 66, pp. 130–153, 2015. [2] T. Kinnunen, M. Sahidullah, H. Delgado, M. Todisco, N. Evans, J. Yamagishi, and K. A. Lee, “The ASVspoof 2017 challenge: Assessing the limits of replay spoofing attack detection,” in Interspeech, 2017. [3] M. Neri and T. Virtanen, “Impact of Microphone Array Mismatches to Learning-Based Replay Speech Detection,” in 2025 33rd European Signal Processing Conference (EUSIPCO), 2025, pp. 1243–1247. [4] M. Neri and T. Virtanen, “Multi-Channel Replay Speech Detection Using an Adaptive Learnable Beamformer,” IEEE Open Journal of Signal Processing, vol. 6, pp. 530–535, 2025. [5] Y. Meng, J. Li, M. Pillari, A. Deopujari, L. Brennan, H. Shamsie, H. Zhu, and Y. Tian, “Your microphone array retains your identity: A robust voice liveness detection system for smart speakers,” in USENIX Security Symposium, 2022. [6] European Union, “Regulation (EU) 2016/679 (General Data Protection Regulation),” Official Journal of the European Union, L 119, 2016. [7] G. M. Van de Ven, N. Soures, and D. Kudithipudi, “Continual learning and catastrophic forgetting,” in Learning and Memory: A Comprehensive Reference, John Wixted, Ed., pp. 153– 168. Academic Press, Oxford, third edition edition, 2025. [8] R. M. French, “Catastrophic forgetting in connectionist networks,” Trends in Cognitive Sciences, vol. 3, no. 4, pp. 128– 135, 1999. [9] M. Todisco, X. Wang, V. Vestman, M. Sahidullah, H. Delgado, A. Nautsch, J. Yamagishi, N. Evans, T. Kinnunen, and K. A. Lee, “ASVspoof 2019: Future horizons in spoofed and fake audio detection,” in Interspeech, 2019. [10] M. Neri, A. Ferrarotti, L. D. Luisa, A. Salimbeni, and M. Carli, “ParalMGC: Multiple Audio Representations for Synthetic Human Speech Attribution,” in EUVIP, 2022. [11] H. Ma, J. Yi, J. Tao, Y. Bai, Z. Tian, and C. Wang, “Continual Learning for Fake Audio Detection,” in Interspeech, 2021. [12] X. Zhang, J. Yi, J. Tao, C. Wang, and Chu Y. Zhang, “Do You Remember? Overcoming Catastrophic Forgetting for Fake Audio Detection,” in ICML, 2023. [13] X. Zhang, J. Yi, C. Wang, C. Y. Zhang, S. Zeng, and J. Tao, “What to Remember: Self-Adaptive Continual Learning for Audio Deepfake Detection,” in AAAI, 2024. [14] T. M. Wani and I. Amerini, “Audio Deepfake Detection: A Continual Approach with Feature Distillation and Dynamic Class Rebalancing,” in Pattern Recognition (ICPR), 2025. [15] H. Wang, H. Dinkel, S. Wang, Y. Qian, and K. Yu, “Dualadversarial domain adaptation for generalized replay attack detection,” in Interspeech, 2020. [16] M. Neri, A. Politis, D. A. Krause, M. Carli, and T. Virtanen, “Speaker Distance Estimation in Enclosures From SingleChannel Audio,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 32, pp. 2242–2254, 2024.

[17] M. Neri, A. Politis, D. A. Krause, M. Carli, and T. Virtanen, “Single-Channel Speaker Distance Estimation in Reverberant Environments,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2023. [18] L. Wang, X. Zhang, H. Su, and J. Zhu, “A Comprehensive Survey of Continual Learning: Theory, Method and Application,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 8, pp. 5362–5383, 2024. [19] J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, D. Hassabis, C. Clopath, D. Kumaran, and R. Hadsell, “Overcoming Catastrophic Forgetting in Neural Networks,” Proceedings of the National Academy of Sciences, vol. 114, no. 13, pp. 3521–3526, 2017. [20] G. Saha, I. Garg, and K. Roy, “Gradient Projection Memory for Continual Learning,” in ICLR, 2021. [21] Y. Gong, J. Yang, J. Huber, M. MacKnight, and C. Poellabauer, “ReMASC: Realistic Replay Attack Corpus for Voice Controlled Systems,” Interspeech, 2019. [22] J. Demšar, “Statistical comparisons of classifiers over multiple data sets,” Journal of Machine Learning Research, vol. 7, pp. 1–30, 2006. [23] F. Wilcoxon, “Individual comparisons by ranking methods,” Biometrics Bulletin, vol. 1, no. 6, pp. 80–83, 1945.

Record · ID 673415 · SHA-256 8e412d038dfd91b4
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.