ConceptioArchivearXiv CS
arXiv CSopen access

A Large-Scale Per-Speaker Analysis of Re-identification Risk in Speech Anonymization

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptography, security, privacy, cybersecurity

A Large-Scale Per-Speaker Analysis of Re-identification Risk in Speech Anonymization Orane Dufour1 , Paul Magron1 , Mickael Rouvier2 , Emmanuel Vincent1 1

Université de Lorraine, CNRS, Inria, LORIA, F-54000 Nancy, France 2 LIA, Avignon University, F-84911 Avignon, France

{orane.dufour, paul.magron, emmanuel.vincent}@inria.fr, [email protected]

arXiv:2606.07210v1 [cs.SD] 5 Jun 2026

Abstract Speech anonymization is commonly evaluated using averagecase metrics such as the equal error rate, which can hide large disparities in re-identification risks across individuals. In this paper, we conduct a large-scale per-speaker privacy analysis using a linkability-based metric under a worst-case scenario. Nearly 5,000 speakers are evaluated across multiple anonymization systems, attacker architectures, and conversation lengths. While linkability scores are highly polarized at the speaker level, the sets of easy to re-identify and hard to re-identify speakers vary substantially across configurations. We show that no single factor explains speaker vulnerability. Instead, the reidentification risk emerges from the interaction between the attacker, the anonymizer, and the amount of available speech. These results challenge the notion of intrinsic speaker-level privacy risks and emphasize the need for evaluation protocols that are explicitly conditioned on the attacker and anonymizer. Index Terms: voice anonymization, speaker recognition, reidentification risk

1. Introduction Speech conveys a multitude of personal information about the speaker such as biometric identity, age, gender, health condition, or emotional state [1]. As such, the storage and processing of large speech datasets puts privacy protection at risk [1, 2]. Existing data protection regulations, such as those outlined in the General Data Protection Regulation (GDPR) [3] by the European Parliament, have led to the development of speech anonymization methods. These methods aim to remove speaker-specific attributes from a speech signal, while preserving a signal that remains usable for other speech processing tasks. The Voice Privacy Challenge (VPC), first introduced in 2020 [4], has played a central role in fostering progress by providing common baselines and benchmarks for anonymization systems. In this context, privacy is typically evaluated through an automatic speaker verification (ASV) [5] system, also referred to as the attacker, whose task is to re-identify speakers by comparing anonymized trial utterances against enrollment utterances with known speaker identities. However, until the 2025 Voice Privacy Attacker Challenge [6], limited attention was paid to the adversarial perspective. This latest edition highlighted how severely privacy risks had been underestimated once stateof-the-art ASV attackers were considered. Still, the performance of ASV systems is most often reported using the equal error rate (EER) [7], which only provides an average-case view of privacy risks. Such aggregate measures can obscure substantial variability across speakers, masking cases of specific individuals that are highly vulnerable to re-identification. To remain fair, privacy protection should not

hinge on population averages, but instead ensure guarantees for all users. Therefore a more appropriate evaluation must adopt a worst-case perspective that explicitly accounts for individual exposure. In this direction, several works have proposed metrics that explicitly target worst-case leakage, such as the ZEBRA framework [8] or GDPR inspired notions of linkability and singling-out [9]. These approaches highlight the limitations of the EER, showing that it fails to capture the nuanced variations in the residual privacy risk under different attack conditions. While such metrics are indeed more suitable for analyzing speaker-level vulnerabilities than the EER, in practice their outcomes are typically reported as averages over the dataset, which obscures these vulnerabilities and prevents these metrics from being fully exploited for worst-case risk evaluation. So far, the literature on the analysis of ASV scores remains relatively limited. In [10], Williams and al. investigated speaker-level score distributions by identifying subpopulations with distinct behaviors, but only under ignorant and lazy-informed attack models and within the scope of the VPC, which includes only 40 test speakers. In [11], the authors examined bias in a VPC baseline system by conducting subgroup analyses based on gender and dialect, revealing disparities across groups. The authors in [12] carried out similar experiments across multiple datasets and observed that disparities (e.g., gender) were not always present in all of them. They also report that men were more protected in some cases and women in some others. Finally, [13] conducted a speaker-level study under a semi-informed as defined in [14] (worst-case) attacker against all three baseline anonymization systems, again restricted to the VPC test set. They showed that the anonymity is repeatedly compromised for certain speakers and, for some of these, the linguistic content alone suffices for re-identification. While theses studies have shown that anonymization affects speakers unevenly, they mostly rely on subgroup analyses based on predefined labels which limits scalability and may overlook unknown factors influencing privacy risks. They also typically involve small datasets and limited diversity in ASV architectures. To address these limitations, in this paper we move beyond label-based subgrouping toward large-scale perspeaker analysis driven directly by attacker behavior. Building on [9], we conduct a systematic per-speaker evaluation under multiple anonymization strategies and attacker models, with the goal of disentangling the respective impacts of speaker identity, anonymization, and ASV architecture on the re-identification risk. Comparing speaker score distributions from diverse attack settings reveals that few speakers are consistently easy or hard to re-identify, and that attack and anonymization conditions strongly influence the individual privacy risk. The rest of this paper is structured as follows. Section 2 presents our per-speaker evaluation framework. Section 3 de-

scribes the data, attackers, and anonymizers. Section 4 reports and discusses the results. Finally, Section 4 concludes the paper.

2. Methodology 2.1. Attacker strategy Within the speech anonymization framework, the attacker is modeled as an ASV system. This deep neural network, trained on a speaker identification task, learns to extract speakerdiscriminative acoustic features from speech. The last hidden layer is used as speaker representation (speaker embedding) of the utterance. During an attack, embeddings are extracted from a set of trial utterances whose speaker identities the attacker aims to recover. Embeddings are also obtained from enrollment utterances for which the speaker identities are known. The attacker then compares pairs constituted of a trial embedding xtest and an enrollment embedding xenroll , with a similarity metric s(·, ·), in order to identify potential matches. Several evaluation metrics can be used to determine whether a match occurs and to quantify the attacker’s success. In this work, we adopt linkability1 as our evaluation metric. 2.2. Linkability As defined in [9], linkability refers to the probability that an attacker can associate a test speaker embedding with its corresponding enrollment embedding for the same speaker i, among a set of N enrollment speakers. A single attempt of linkage is said to succeed if   (i) (i) (i) (j) s xtest , xenroll > max s xtest , xenroll . (1) j̸=i

The linkability metric is then defined as the probability of successful linkage across all test samples and all speakers:     (i) (i) (i) (j) Prx(i) s xtest , xenroll > max s xtest , xenroll . (2) test

j̸=i

To compute linkability on our test set, we adopt the protocol introduced in [9]. For each test speaker in the trial set, the distance between the test embedding and a pool of N enrollment speakers is computed using cosine similarity. The number of enrollment speakers N takes 11 values, ranging from 22,024 (the maximum number of available speakers in our enrollment set) down to 21. These values are obtained by repeatedly dividing the pool size: starting from 22,024, we divide by two to obtain the next value, and continue this process until reaching 21. This ensures coverage of a wide range of values while preserving a reasonable computational cost. To reduce the influence of any particular selection of enrollment speakers, we perform five random draws of the enrollment pool for each test speaker. For every draw, enrollment embeddings are computed using all available utterances of the selected speakers. The result depends on two parameters: the number of enrollment speakers N and the conversation length L. The conversation length corresponds to the number of utterances used to compute the test speaker embedding. We consider three values: L = 1, 3, and 5. This design enables us to quantify to what extent the amount of speech data available per speaker impacts the re-identification risk and also, as stated in [15], to reduce the variance of the scores when L > 1. 1 The singling-out metric [9] was excluded from the experiments because no implementation was available at the time of writing this paper.

Table 1: Datasets used for training and evaluation. Dataset

Speakers Dur. (h) Utterances

LibriSpeech CV 11.0 A CV 11.0 B

921 22,024 4,949

360 323 1,409

104,014 234,945 996,971

Split Train Test (enrol.) Test (trials)

2.3. Easy- and hard-to-link speakers Since our experiments focus on individual speakers, we do not report linkability as a function of N over the entire dataset. Instead, for each trial speaker, we compute an average linkability score obtained over multiple linkage trials. This score is averaged over 55 tests per speaker, corresponding to 5 random draws of N enrollment speakers for each of the 11 values of N . We repeat this process for each attack scenario involving a different pair of attacker/anonymizer and a different value of conversation length L, and we generate a per-speaker score distribution for each configuration. We are interested in the speakers who obtained a high linkability in those distributions and want to know if they remain the same in each of them. To do so, for each distribution, we compute the 3rd quartile (Q3) value (i.e., the value below which 75% of the scores fall) and we generate lists with speakers who obtained a linkability above this value for all utterances. For comparison purposes, we also compute the 1st quartile value (Q1), and generate lists of speakers with a linkability equal or inferior to this value. These are denoted easy-to-link speakers and hard-to-link speakers. Then, we compute the intersection of all easy-to-link speakers lists across distributions (i.e., considering all combinations of attackers, anonymizers, and conversation length L). We also compute the union of such lists, and similarly for hard-to-link speakers. Furthermore, to have a more granular analysis and to determine the influence of each variable (the attacker, the anonymizer, and the conversation length), we also compare each pair of speaker lists using the Jaccard similarity. The Jaccard similarity between two speaker lists S1 and S2 is defined as: Jaccard(S1 , S2 ) =

|S1 ∩ S2 | , |S1 ∪ S2 |

(3)

where |S1 ∩S2 | is the number of speakers common to both lists, and |S1 ∪ S2 | is the total number of unique speakers across the two lists. This measure reflects the proportion of shared items relative to the total size of the two combined sets.

3. Experimental setup This section presents our experimental protocol. Our code is available for a reproductibility purpose2 . 3.1. Datasets Table 1 describes the datasets used in our experiments. We use LibriSpeech [16] to train the ASV models (more specifically we use the train-clean-360 split), and CommonVoice’s 11th English release (CV 11.0) [17] as the test set. CV 11.0 is split in an enrollment set A and a trial set B. Every speaker of set B is present in set A, with disjoint utterances. We use the exact same split as in [9]. 2 https://github.com/OraneD/Speaker-Linkability

3.2. Anonymization systems 3

As anonymization systems, we employ two baseline systems of VPC 2025 both based on neural voice conversion: B3 [19] and B5 [20]. B3 extracts the linguistic content using explicit phonetic transcriptions and performs speech resynthesis with pseudo-speaker embeddings generated by a generative adversarial network [21, 22]. In contrast, B5 encodes linguistic content using vector-quantized bottleneck features extracted from a wav2vec 2.0-based [23] speech recognition model, and synthesizes speech with a HiFi-GAN vocoder trained on a set of (real) target speakers. Among the baselines, B5 achieves the highest performance. Anonymization for both the training and test set is consistently applied at the utterance level which means that each utterance of a given speaker is anonymized with the a different target speaker. This ensures that the evaluation reflects speaker identity rather than memorization of a fixed target. As shown by [24], with speaker level anonymization, an attacker may learn to recognize that target. By assigning a different target speaker for each utterance we mitigate biases introduced by potentially vulnerable source-target pairs. 3.3. Attackers In order to assess robustness under a worst-case threat model we deliberately employ a diverse set of attacker architectures. Indeed, differences in model capacity, inductive biases, and input features can affect which speakers are re-identified, so architectural diversity will allow us to determine whether vulnerabilities are model-specific. We employ three different attackers: 1. ECAPA4 is the baseline attacker of the VPC 2025. It is a TDNN/x-vector style architecture [25] that integrates channel and context-adaptive modules to produce highly discriminative speaker embeddings. 2. WavLM ECAPA5 is the same as ECAPA architecture-wise, but it processes WavLM [26] input features instead of log Mel filterbank features. 3. ResNet6 uses residual connections to enable very deep convolutional networks that learn hierarchical spectral-temporal representations [27]. We consider the deeper ResNet-101 trained with the Kiwano toolkit [28], which offers increased capacity to model subtle speaker cues. Note that we ran additional experiments with two more attackers [29, 30]. The results are consistent with those obtained using the 3 attackers detailed above, but we do not include these due to space constraints. Each attacker has been trained on LibriSpeech-train-clean-360 in a semi-informed scenario, which means that they have been trained on data anonymized with the same anonymizer being attacked. Note however that the source-to-target speaker mapping remains unknown by the attacker since it is done randomly. As a preliminary experiment, we evaluated each attacker following the 2025 VPC’s evaluation protocol and computed the EER on LibriSpeech-test for both B3 and B5 anonymizers, in the semi-informed setting. Results match state-of-the-art performances for both WavLM ECAPA and ResNet, ensuring that the attackers are sufficiently strong for the subsequent analysis. Note that while the VPC 2025 win3 Baseline B4 [18] was excluded due to its encoder being trained on one of the CommonVoice releases. 4 https://github.com/Voice-Privacy-Challenge/Voice-PrivacyChallenge-2024/tree/main 5 https://github.com/deep-privacy/sidekit 6 https://github.com/kiwano-toolkit/kiwano/

Table 2: Number of speakers in the intersection and union of easy- and hard-to-link speakers across the 18 distributions. Speakers

Intersection

Union

Easy-to-link Hard-to-link

5 166

4,300 4,574

ner [31] would have been a natural choice for this study, no implementation is available thus we do not consider it here.

4. Results 4.1. Per-speaker score distributions Figure 1 shows the 18 speaker score distributions (2 anonymization systems × 3 attackers × 3 conversation length L). As shown in the figure, the WavLM ECAPA attacker consistently outperforms ResNet and ECAPA, and B3 is systematically easier to attack than B5. When L = 1 most speakers obtain a linkability close to 0, whereas the opposite occurs when L = 5 where most speakers reach a linkability of 1, meaning they were successfully linked to themselves in all 55 attempts. This confirms that increasing the conversation length drastically strengthens the attacker. Across all architectures and configurations, we observe a highly polarized distribution of speaker linkability: a substantial proportion of speakers achieve zero linkability, while another large group exhibits near-perfect linkability. This could reveal the presence of distinct speaker clusters, corresponding to inherently linkable versus inherently unlinkable speakers, highlighting a large heterogeneity in speaker vulnerability. However, when looking at the intersections and unions of easy- and hard-to-link speakers in the 18 different configurations, we observe that the identities of the speakers at both end of the distributions vary a lot. Indeed, considering the intersections as described in Table 2, only 5 speakers are easy-to-link consistently across the 18 distributions, and only 166 of them are consistently hard-to-link. However, unions, for both easy- and hard-to-link speakers are very high: 86.9% of the 4,949 test speakers are at least considered once as easy-to-link and 92.4% of them are at least once hard-to-link. These results indicate that speaker vulnerability is strongly influenced by multiple system-related factors, namely the ASV architecture of the attacker, the anonymization model being attacked, and the value of the conversation length L. Importantly, these results suggest that vulnerability cannot be explained solely by intrinsic speaker characteristics, but also depends on the attack and anonymization configurations. 4.2. Similarities between distributions To quantify the respective impact of the anonymizer, the ASV architecture, and the conversation length, on the linkability score distributions, we compute the Jaccard similarity between pairs of lists for both easy- and hard-to-link speakers. In each comparison, only one factor varies while the other two are kept fixed, isolating its effect on list composition. For each factor, we report the average Jaccard similarity across all such comparisons in Figure 2, for easy- and hard-to-link speakers separately. First, we observe that the average Jaccard similarity never exceeds 0.47 for any group of speakers. This indicates that, regardless of the variable considered, the overlap between easyor hard-to-link speaker lists is limited. In other words, changing any single variable leads to substantial changes in which speakers are considered easy or hard to link, suggesting that no single

L=1 % speakers

80

ECAPA B3

ECAPA B5

WavLM ECAPA B3

WavLM ECAPA B5

ResNet B3

ResNet B5

60 40 20

L=3 % speakers

0 80 60 40 20

L=5 % speakers

0 80 60 40 20 00

0.25 0.50 0.75 Linkability

10

0.25 0.50 0.75 Linkability

10

0.25 0.50 0.75 Linkability

10

0.25 0.50 0.75 Linkability

10

0.25 0.50 0.75 Linkability

10

0.25 0.50 0.75 Linkability

1

Figure 1: Distributions of the average linkability score obtained by all test speakers for all sets of attacker, anonymizer and conversation length L. Each line corresponds to a value of L and each column to a pair of attacker and anonymizer. The purple and blue bars indicate an attack performed on B3 and B5, respectively. Dashed bars indicate the WavLM ECAPA attacker while crossed ones indicate ResNet. The green and red vertical dashed lines indicate Q1 and Q3, respectively.

Mean Jaccard

Easy-to-link speakers 0.5 0.4 0.3 0.2 0.1 0.0

0.29

0.34

Anonymizer

0.47

Hard-to-link speakers

0.38

0.35 0.27

Attacker

Conversation length (L)

Figure 2: Mean Jaccard similarities for easy- and hard-to-link speakers. Black lines are standard deviations and the x-axis indicates the varying factor, while the other two are fixed. factor fully explains speaker vulnerability. Second, Figure 2 allows us to compare the relative impact of the different variables on the linkability score distributions. A higher Jaccard similarity indicates a smaller impact, as the identities of the speakers in the lists are more consistent across configurations. • Attacker architecture: this variable yields the highest Jaccard similarities (approximately 0.39 for easy-to-link speakers and 0.47 for hard-to-link speakers). This indicates that the composition of the easy- and hard-to-link speaker lists is more stable across different attacker architectures than across changes in the other variables, suggesting that the attacker architecture has the smallest impact on the distributions. • Anonymizer: The anonymizer has a stronger influence than the attacker, with similarities around 0.29 and 0.34 for easyand hard-to-link speakers respectively. • Conversation length: The conversation length L yields Jaccard similarities very close to the anonymizer indicating that they have a comparable impact on the distributions.

Note that the differences in mean Jaccard similarities remain small, the largest gap being approximately 0.13. In particular, while the attacker architecture exhibits higher Jaccard similarities and can therefore be identified as having the smallest impact on the linkability score distributions, the anonymizer and the conversation length L show very close mean values. Moreover, the standard deviations associated with the anonymizer and the conversation length L are of the same order of magnitude as the gap between their respective mean Jaccard similarities between easy- and hard-to-link speakers. This indicates that the observed differences between these two variables are not sufficiently pronounced to support a strict ordering. As a result, rather than establishing a definitive ranking, these results suggest that the anonymizer and the conversation length L have a comparable impact on the linkability score distributions.

5. Conclusion We conducted the first large-scale per-speaker privacy risk analysis across multiple anonymization and attacker architectures and showed that the privacy risk does not depend solely on speakers. Instead, vulnerability to re-identification emerges from the interaction of the attacker’s ASV architecture, the anonymization system, and the amount of speech for each test speaker. These results raise a fundamental challenge for privacy evaluation. The privacy risk under a given anonymization and attack configuration does not transfer across threat models, highlighting the need for context-dependent privacy guarantees. Future work will focus on defining an a priori re-identification risk [32] conditioned on a fixed anonymization and attacker pair. We plan to explore reliability measures of ASV speaker embeddings as a means to estimate the privacy risk without explicit score-based evaluation. This direction is promising for more predictive and user-oriented privacy guarantees.

6. Acknowledgments This work was supported by the French Agence Nationale de la Recherche via the SpeechPrivacy project (ANR-23-CE230022). It was provided with computer and storage resources by GENCI at IDRIS thanks to the grant 2025-AD011015838R1 on the supercomputer Jean Zay’s V100 partition.

7. Use of Generative AI Disclosure

Acoustics, Speech and Signal Processing (ICASSP), 2024, pp. 4785–4789. [12] C. Franzreb, T. Polzehl, and S. Möller, “A comprehensive evaluation framework for speaker anonymization systems,” in 3rd Symposium on Security and Privacy in Speech Communication, 2023, pp. 65–72. [13] U. E. Gaznepoglu, A. Leschanowsky, A. Aloradi, P. Singh, D. Tenbrinck, E. A. P. Habets, and N. Peters, “You are what you say: Exploiting linguistic content for VoicePrivacy attacks,” in Interspeech, 2025, pp. 4238–4242.

We acknowledge the ISCA policy regarding the use of generative AI tools. The authors declare that generative AI tools have been used solely to correct grammar. No such tools were used to write significant parts of this manuscript.

[14] B. M. L. Srivastava, “Speaker anonymization : representation, evaluation and formal guarantees,” Ph.D. dissertation, Université de Lille, 2021. [Online]. Available: https://theses.hal.science/ tel-03674540

8. References

[15] C. Franzreb, A. Das, T. Polzehl, and S. Möller, “Optimizing the dataset for the privacy evaluation of speaker anonymizers,” in 5th Symposium on Security and Privacy in Speech Communication, 2025, pp. 18–26.

[1] J. L. Kröger, O. H.-M. Lutz, and P. Raschke, “Privacy implications of voice and speech analysis – information disclosure by inference,” in Privacy and Identity Management. Data for Better Living: AI and Privacy: 14th IFIP WG 9.2, 9.6/11.7, 11.6/SIG 9.2.2 International Summer School, 2020, pp. 242–258. [2] A. Nautsch, A. Jiménez, A. Treiber, J. Kolberg, C. Jasserand, E. Kindt, H. Delgado, M. Todisco, M. A. Hmani, A. Mtibaa, M. A. Abdelraheem, A. Abad, F. Teixeira, D. Matrouf, M. GomezBarrero, D. Petrovska-Delacrétaz, G. Chollet, N. Evans, T. Schneider, J.-F. Bonastre, B. Raj, I. Trancoso, and C. Busch, “Preserving privacy in speaker and speech characterisation,” Computer Speech & Language, vol. 58, pp. 441–480, 2019. [3] “Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 april 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing directive 95/46/EC (General Data Protection Regulation),” http://data.europa.eu/eli/ reg/2016/679/2016-05-04, 2016, official Journal of the European Union, L 119, pp. 1–88. [4] N. Tomashenko, X. Wang, E. Vincent, J. Patino, B. M. L. Srivastava, P.-G. Noé, A. Nautsch, N. Evans, J. Yamagishi, B. OBrien, A. Chanclu, J.-F. Bonastre, M. Todisco, and M. Maouche, “The VoicePrivacy 2020 Challenge: Results and findings,” Computer Speech & Language, vol. 74, p. 101362, 2022. [5] Z. Bai and X.-L. Zhang, “Speaker recognition based on deep learning: An overview,” Neural Networks, vol. 140, pp. 65–99, 2021. [6] N. Tomashenko, X. Miao, E. Vincent, and J. Yamagishi, “The first VoicePrivacy Attacker Challenge,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025, pp. 1–2. [7] N. Tomashenko, X. Miao, P. Champion, S. Meyer, X. Wang, E. Vincent, M. Panariello, N. Evans, J. Yamagishi, and M. Todisco, “The VoicePrivacy 2024 Challenge evaluation plan,” 2024, arXiv preprint arxiv:2404.02677. [8] A. Nautsch, J. Patino, N. Tomashenko, J. Yamagishi, P.-G. Noé, J.-F. Bonastre, M. Todisco, and N. Evans, “The privacy ZEBRA: Zero evidence biometric recognition assessment,” in INTERSPEECH, 21st Annual Conference of the International Speech Communication Association, Shanghai, China (Virtual Conference), 2020, pp. 1698–1702. [9] N. Vauquier, B. M. L. Srivastava, S. A. Hosseini, and E. Vincent, “Legally validated evaluation framework for voice anonymization,” in Interspeech, 2025, pp. 3229–3233.

[16] V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015, pp. 5206–5210. [17] R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber, “Common Voice: A massively-multilingual speech corpus,” in 12th Language Resources and Evaluation Conference (LREC), 2020, pp. 4218–4222. [18] M. Panariello, F. Nespoli, M. Todisco, and N. Evans, “Speaker anonymization using neural audio codec language models,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024, pp. 4725–4729. [19] S. Meyer, F. Lux, J. Koch, P. Denisov, P. Tilli, and T. Vu, “Prosody is not identity: A speaker anonymization approach using prosody cloning,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5. [20] P. Champion, “Anonymizing speech: Evaluating and designing speaker anonymization techniques,” Ph.D. dissertation, Université de Lorraine, 2024. [Online]. Available: https://hal.univ-lorraine. fr/tel-04218098v1 [21] S. Meyer, F. Lux, P. Denisov, J. Koch, P. Tilli, and N. T. Vu, “Speaker anonymization with phonetic intermediate representations,” in Interspeech, 2022, pp. 4925–4929. [22] S. Meyer, P. Tilli, P. Denisov, F. Lux, J. Koch, and N. T. Vu, “Anonymizing speech with generative adversarial networks to preserve speaker privacy,” in IEEE Spoken Language Technology, 2022, pp. 912–919. [23] A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” pp. 12 449–12 460, 2020. [Online]. Available: https://proceedings.neurips.cc/paper files/paper/2020/file/ 92d1e1eb1cd6f9fba3227870bb6d7f07-Paper.pdf [24] C. Franzreb, A. Das, T. Polzehl, and S. Möller, “Improving the speaker anonymization evaluation’s robustness to target speakers with adversarial learning,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026. [25] B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPATDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,” in Interspeech, 2020, pp. 3830–3834.

[10] J. Williams, K. Pizzi, N. Tomashenko, and S. Das, “Anonymizing speaker voices: Easy to imitate, difficult to recognize?” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024, pp. 12 491–12 495.

[26] S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y. Qian, Y. Qian, J. Wu, M. Zeng, X. Yu, and F. Wei, “WavLM: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1505–1518, 2022.

[11] A. Leschanowsky, . E. Gaznepoglu, and N. Peters, “Voice anonymization for all-bias evaluation of the voice privacy challenge baseline systems,” in IEEE International Conference on

[27] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778.

[28] M. Rouvier and P.-M. Bousquet, “Kiwano: A Cutting-Edge OpenSource Toolkit for Speaker Verification,” in Odyssey 2026, 2026. [29] R. Arefeen, X. Miao, R. Tong, A. B. Ng, S. See, and T. Liu, “DAST: A dual-stream voice anonymization attacker with staged training,” 2026. [Online]. Available: https: //arxiv.org/abs/2603.12840 [30] I. Yakovlev, R. Makarov, A. Balykin, P. Malov, A. Okhotnikov, and N. Torgashov, “Reshape dimensions network for speaker recognition,” in Interspeech, 2024, p. 32353239. [31] X. Lyu, Y. Wang, T. Zhao, and H. Liu, “Fast adaptation of pretrained speaker verification system for source speaker tracking,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025, pp. 1–2. [32] P.-M. Bousquet, M. Rouvier, and J.-F. Bonastre, “Reliability criterion based on learning-phase entropy for speaker recognition with neural network,” in Interspeech, 2022, pp. 281–285.

Record · ID 266114 · SHA-256 9b843491e2169b74
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.