ConceptioArchivearXiv CS
arXiv CSopen access

VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion Aastha Sharma, Guangjing Wang University of South Florida, Tampa, FL, USA [email protected], [email protected]

arXiv:2607.11706v1 [cs.SD] 13 Jul 2026

Abstract Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks. This mismatch creates a temporal generalization gap that can overestimate detector robustness under real-world postprocessing conditions. We bridge this gap by introducing VoxENES 2026, a bilingual (English and Spanish) benchmark of 53,628 audio samples generated using 10 contemporary speech synthesis methods and evaluated under 10 standardized postprocessing conditions. Using VoxENES 2026, we benchmark eight pretrained detectors without fine-tuning and observe substantial performance degradation: the best model achieves 28.98% EER overall, while most perform near or below random chance across modern generators and perturbations. Our results highlight the reliance on brittle artifacts in current detectors and establish VoxENES 2026 as a practical testbed for developing robust audio spoofing countermeasures. Index Terms: audio deepfake detection, speech spoofing detection, benchmark dataset

1. Introduction Robust speech spoofing and deepfake detection are essential to preserve trust in speech-based authentication [1, 2, 3, 4]. This need is growing as voice becomes a biometric and a control channel for speech-driven agents and assistive technologies. If synthetic speech becomes indistinguishable from genuine speech, detection failures not only compromise security protocols but also erode trust in audio evidence. Many benchmark datasets are proposed for speech spoofing and deepfake detection evaluation. For example, the ASVspoof challenge series has been the primary driver of spoofing countermeasure development. ASVspoof 2019 [5] introduces logical access (LA) with TTS and VC, and physical access (PA) with replay tracks. ASVspoof 2021 [6] adds the deepfake task targeting compressed manipulated speech, and ASVspoof 5 [7] introduces crowdsourced data with adversarial attacks at scale. In addition to ASVspoof, WaveFake [8] provides a multilingual dataset from six neural vocoder architectures. The In-the-Wild dataset [9] includes real-world deepfakes of celebrities. The MLAAD [10] dataset expands coverage to 23 languages and 54 TTS models. The VoiceWukong [11] benchmarks 12 detectors against 34 commercial and open-source tools with postprocessing manipulations. Yet, existing benchmarks primarily rely on speech synthesis systems before 2024 and fail to capture artifact patterns produced by the modern large language model (LLM)-driven generation pipelines. For example, text-to-speech (TTS) designs include autoregressive language-model-based synthesis, such

as VoxCPM [12] and Qwen3-TTS [13]; flow-matching models, including GLM-TTS [14], CosyVoice 3 [15], and Chatterbox [16]; diffusion-based systems like FlashLabs Chroma [17]; and hybrid DiT architectures, such as VibeVoice [18]. For voice conversion (VC), zero-shot approaches such as Seed-VC [19], tone-color extraction methods such as OpenVoice v2 [20], and retrieval-based systems like RVC v2 [21] have substantially improved naturalness and speaker similarity. Evaluation on temporally stale benchmarks can overestimate real-world robustness as speech synthesis techniques improve. LLM-driven TTS and VC systems produce synthetic audio with acoustic characteristics that differ substantially from earlier spoofing corpora, creating a data drifting issue for existing detectors. Deepfake detection models that perform well on older benchmarks may fail when deployed against newer deepfake generators. This mismatch reflects a cat-and-mouse dynamic in which detectors learn cues tied to previous synthesis artifacts, while generators and post-processing steps progressively suppress or conceal those cues. As a consequence, the rapid LLM-driven TTS and VC evolution motivates updated evaluation benchmarks. To study the generalization ability of deepfake detectors under the data drifting issue in the era of LLM, we introduce a modern bilingual benchmark dataset VoxENES 2026. We generate synthetic audios from 10 LLM-driven synthesis methods (7 TTS, 3 VC), covering English and Spanish audios. In addition, real-world audios are frequently subjected to transformations such as compression, noise, and resampling, which can substantially affect detector behavior [6, 11, 10]. Therefore, VoxENES 2026 explicitly models the post-processing setting by applying a standardized set of audio augmentation methods that simulate common transmission and manipulation effects, including codec compression, additive noise, resampling, speed perturbation, and loudness normalization. With VoxENES 2026, we evaluate eight pretrained deepfake detectors without fine-tuning to measure out-ofdistribution generalization. Our results show that existing detectors suffer from substantial degradation relative to reported performance on legacy benchmarks. The best detector achieves only 28.98% EER overall, while most detectors perform near or below random chance. These findings suggest that many current detectors rely on brittle artifact cues that may not generalize across new TTS and VC generations and realistic postprocessing. By providing a modern benchmark and a controlled out-of-distribution evaluation, our work establishes a necessary testbed for measuring real-world speech spoofing detector robustness and for driving detector development that keeps pace with the evolving synthesis frontier. In summary, the main contributions of our work are: • We introduce a modern bilingual benchmark dataset Vox-

Table 1: VoxENES 2026 dataset summary. Component

Count

Notes

Real speech Original synthetic Augmented synthetic

3,028 4,600 46,000

LibriSpeech + VoxPopuli TTS + VC originals 10× post-processing

Total

53,628

EN + ES combined

Table 2: Language and source breakdown in VoxENES 2026. Category

EN

ES

Total

Real speech TTS (7 methods) VC (3 methods)

1,500 11,000 16,500

1,528 6,600 16,500

3,028 17,600 33,000

Total

29,000

24,628

53,628

ENES 2026 using LLM-driven TTS and VC systems with realistic post-processing to emulate deploymenttime distribution shifts, totaling 53,628 audio samples available at https://www.kaggle.com/datasets/ interspeech2712/voxenes-2026. • We benchmark eight pretrained detectors and reveal a substantial generalization gap, exposing detector-specific blind spots and robustness failures across synthesis methods.

2. VoxENES 2026 VoxENES 2026 consists of three components: (1) real speech from established corpora, (2) synthetic audio generated by modern TTS and VC systems, and (3) post-processed augmented variants simulating real-world transmission conditions. As shown in Table 1, VoxENES includes 3,028 real speech samples and 50,600 synthetic samples (original synthetic 4,600 and augmented synthetic 46,000). 2.1. Real Speech and Standardization As shown in Table 2, we drew 1,500 real English speech samples from 40 speakers from LibriSpeech [22]. We drew 1,528 real Spanish speech samples from 96 speakers from VoxPopuli [23]. All audio was standardized to 16 kHz mono WAV. For detector compatibility, samples were capped at 4 seconds with truncation and zero-padding. This fixed-length standardization was applied uniformly to both bonafide and synthetic samples so that any padding-induced cues are shared across classes rather than correlated with the spoofing label. We, therefore, do not expect zero-padding to serve as a discriminative shortcut, though a detailed analysis of padding artifacts is left to future work. Note that for the downstream voice conversion tasks, the people whose voices appear in the real speech samples are not used as target speakers.

matching model from Zhipu AI; (iv) FlashLabs Chroma [17], a diffusion-based system; (v) VibeVoice [18], combining DiT with flow matching; (vi) CosyVoice 3 [15], an LM plus flowmatching system from Alibaba; and (vii) Chatterbox ML [16], a flow-matching model from Resemble AI. The three VC systems represent distinct paradigms: (i) Seed-VC [19] uses diffusionbased zero-shot conversion with Whisper and WavLM features; (ii) OpenVoice v2 [20] performs tone-color extraction and transfer; and (iii) RVC v2 [21] uses HuBERT-based retrieval with HiFi-GAN vocoding. All TTS systems were released or updated in 2025, ensuring they represent modern threats that older detection models have never encountered. VC samples are generated using disjoint source audio partitions and multiple target speaker references to maintain diversity. 2.3. Post-Processing Augmentations Each original synthetic sample was subjected to one postprocessing operation in Table 4 that reflect common transmission and editing effects: lossy codec compression (MP3 at 64 kbps and AAC at 128 kbps), additive noise perturbations (white noise at 10/20 dB SNR and babble noise at 15 dB SNR), bandwidth and sampling-rate distortions (downsampling to 8 kHz with restoration to 16 kHz, plus a 16 kHz resampling control), playback-rate changes (1.1× and 0.9× speed), and amplitude normalization (peak normalization to −3 dBFS). This pipeline expands the synthetic subset and enables systematic robustness analysis across realistic post-processing conditions.

3. Evaluation Setup 3.1. Detection Baselines We evaluated eight pretrained audio deepfake detection systems spanning graph neural networks, raw-waveform CNNs, self-supervised learning (SSL) models, transformer-based detectors, and speaker-embedding anomaly detection as shown in Table 5. None were retrained or fine-tuned on VoxENES 2026, enabling an honest temporal generalization evaluation. The detectors differ in their training corpora (e.g., ASVspoof 2019 LA, ASVspoof 2021 DF, ASVspoof 5, and VoxCeleb), so absolute EER values should be read as out-of-distribution generalization indicators rather than as a strictly controlled head-to-head comparison; the most directly comparable cases are detectors sharing the same training source. All detectors were evaluated in inference-only mode using published pretrained weights. We report Equal Error Rate (EER) and accuracy in the results. EER was computed from the ROC operating point where the false acceptance rate (FAR) and the false rejection rate (FRR) are equal. In particular, ECAPATDNN [27] is a speaker verification model rather than a dedicated spoof detector. Thus, we use an anomaly-based scoring scheme, where a reference centroid is computed from real speech audio, and the cosine distance from this centroid is used as the spoofing score. The system’s performance is then evaluated using the EER, determined by the threshold at which the FAR and FRR are equivalent.

2.2. Synthesis Systems We selected seven TTS and three VC systems representing state-of-the-art LLM-driven synthesis techniques across diverse architectures as shown in Table 3. The TTS systems include: (i) VoxCPM 1.5 [12], an autoregressive language model from HKUST; (ii) Qwen3-TTS [13], a streaming LM from Alibaba with native multilingual support; (iii) GLM-TTS [14], a flow-

4. Evaluation Results 4.1. Qualitative Spectrogram Analysis Figures 1 and 2 visualize the spectral variability across bonafide (real) speech, synthetic speech, and post-processed synthetic samples in VoxENES 2026. The examples highlight that re-

Architecture

Developer

EN

ES

Total

Year

ES Mode

VoxCPM 1.5 [12] Qwen3-TTS [13] GLM-TTS [14] FlashLabs Chroma [17] VibeVoice [18] CosyVoice 3 [15] Chatterbox ML [16]

TTS TTS TTS TTS TTS TTS TTS

Autoregressive LM Streaming LM Flow matching Diffusion DiT + flow matching LM + flow matching Flow matching

HKUST Alibaba Zhipu AI FlashLabs Vibe AI Alibaba Resemble AI

200 200 200 200 200 – –

– 200 – – – 200 200

200 400 200 200 200 200 200

2025 2025 2025 2025 2025 2025 2025

– Native – – – Native Native

Seed-VC [19] OpenVoice v2 [20] RVC v2 [21]

VC VC VC

Diffusion + zero-shot Tone cloning Retrieval + HiFi-GAN

ByteDance MyShell AI RVC-Project

500 500 500

500 500 500

1,000 1,000 1,000

2025 2024 2023

– – –

1

1.5

2

2.5

Time (s)

3

3.5

4

VibeVoice + Resample 16 kHz

(e)

0

0.5

1

1.5

2

2.5

Time (s)

3

3.5

4

8 7 6 5 4 3 2 1 0.0

0

0.5

(f)

1

1.5

2

2.5

Time (s)

3

3.5

4

Seed-VC + Speed 1.1×

0

0.5

1

1.5

2

2.5

Time (s)

3

3.5

4

8 7 6 5 4 3 2 1 0.0

8 7 6 5 4 3 2 1 0.0

Chroma + White Noise (20 dB)

(c)

Frequency (kHz)

(b)

0

0.5

1

1.5

2

2.5

Time (s)

3

3.5

4

OpenVoice v2 + Speed 0.9×

(g)

0

0.5

1

1.5

2

2.5

Time (s)

3

3.5

4

8 7 6 5 4 3 2 1 0.0

8 7 6 5 4 3 2 1 0.0

CosyVoice 3 Spanish (Original)

(d)

0 10 20 30

0

0.5

1

1.5

2

2.5

Time (s)

3

3.5

4

Power (dB)

0.5

GLM-TTS + White Noise (10 dB)

Frequency (kHz)

0

8 7 6 5 4 3 2 1 0.0

Frequency (kHz)

VoxCPM 1.5 + AAC 128k

(a)

Frequency (kHz)

8 7 6 5 4 3 2 1 0.0

Type

Frequency (kHz)

8 7 6 5 4 3 2 1 0.0

System

Frequency (kHz)

Frequency (kHz)

Frequency (kHz)

Table 3: TTS and VC systems used in VoxENES 2026.

40

RVC v2 + Volume Norm

50

(h)

60 70 80 0

0.5

1

1.5

2

2.5

Time (s)

3

3.5

4

Figure 1: Representative spectrograms across multiple TTS and VC systems and augmentation conditions in VoxENES 2026. Table 4: Post-processing augmentation techniques.

Table 5: Detection baselines evaluated on VoxENES 2026.

Augmentation

Description

Detector

Family

Training Data

mp3 64k aac 128k noise white 10db noise white 20db noise babble 15db resample 8k resample 16k speed fast speed slow volume norm

MP3 encoding at 64 kbps AAC encoding at 128 kbps White Gaussian noise at 10 dB SNR White Gaussian noise at 20 dB SNR Multi-speaker babble noise at 15 dB SNR Downsample to 8 kHz, upsample to 16 kHz Resample to 16 kHz (control) 1.1× playback speed 0.9× playback speed Peak normalization to −3 dBFS

AASIST2 [24] RawNet2 [25] Wav2Vec2-AASIST Wav2Vec2-DF Wav2Vec2-Large AST-ASVspoof5 [26] Wav2Vec2-ASVspoof5 ECAPA-TDNN [27]

Graph NN Raw waveform CNN SSL + classifier SSL + classifier SSL + classifier Spectrogram Transformer SSL + classifier Speaker embeddings

ASVspoof 2019 LA ASVspoof 2021 DF ASVspoof 2019 Mixed deepfake Mixed deepfake ASVspoof 5 ASVspoof 5 VoxCeleb

alistic perturbations, such as additive noise and lossy codec compression, can reshape time-frequency structure and attenuate synthesis artifacts that many detectors implicitly rely on, providing a qualitative rationale for the observed performance changes under post-processing.

to 27.9%). This is plausibly because noise perturbs synthetic and bonafide speech differently and attenuates generatorspecific artifacts, thereby inducing alternative discriminative cues. In contrast, MP3 compression degrades AST-ASVspoof5 (from 26.7% to 48.4% EER), consistent with a reliance on finegrained spectral structure, particularly in higher frequencies, that is suppressed by lossy codec encoding.

4.2. Impact of Post-Processing

4.3. Overall Detection Performance

As shown in Table 6, post-processing impacts detector performance in a highly non-uniform and occasionally counterintuitive manner. Adding white noise reduces EER for ASTASVspoof5 (from 26.7% to 17.4%) and RawNet2 (from 51.3%

As shown in Table 7, among the evaluated models, only ASTASVspoof5 achieves an EER below 30% (i.e., 28.98%), a result that remains insufficient for reliable field deployment. In addition, five of the eight detectors perform at or below the stochas-

Table 6: Per-augmentation EER (%) across all detectors. AASIST2

RawNet2

W2V-AASIST

W2V-DF

W2V-Large

AST-ASV5

W2V-ASV5

ECAPA

54.4 53.1 53.7 67.5 59.0 65.6 59.1 54.1 56.0 52.7 56.7

51.3 50.2 50.1 27.9 40.0 33.0 55.6 49.7 53.4 47.8 47.1

41.0 42.5 42.5 32.3 44.5 22.9 37.5 42.3 38.6 45.0 42.0

53.9 53.6 54.1 69.5 68.9 50.8 42.0 54.1 52.2 52.5 53.9

39.9 37.9 39.8 61.2 63.3 45.1 34.3 39.7 37.5 42.4 40.1

26.7 31.8 48.4 17.4 18.3 25.7 29.5 27.1 26.1 30.2 28.5

52.6 51.5 52.0 52.5 49.8 46.3 52.1 52.3 59.0 49.8 52.2

46.5 46.9 47.2 33.3 36.2 44.5 46.4 46.1 38.5 42.0 46.6

original aac 128k mp3 64k noise white 10db noise white 20db noise babble 15db resample 8k resample 16k speed fast speed slow volume norm

(a)

Bona fide Speech (English)

4

Frequency (kHz)

Frequency (kHz)

4 3 2 1 0.0

0

0.5

1

1.5

2

Time (s)

2.5

3

3.5

Qwen3-TTS Spanish + Babble Noise (15 dB SNR)

(c)

3 2 1 0.0

0

0.5

1

1.5

2

Time (s)

2.5

3

3.5

4

0 10

2 20

1

30 0

0.5

1

1.5

2

Time (s)

2.5

3

3.5

4

40

Qwen3-TTS Spanish + MP3 64k Codec

4

Frequency (kHz)

Frequency (kHz)

4

(b)

3

0.0

4

Original Qwen3-TTS Spanish (Synthetic)

Power (dB)

Augmentation

50

(d)

3

60

2

70

More broadly, no single detector is consistently reliable across all synthesis methods. For example, ECAPA-TDNN performs well on several TTS systems (e.g., GLM-TTS at 10.2% and CosyVoice 3 at 11.5%) but degrades substantially on VC (e.g., OpenVoice v2 at 62.7%). The results suggest that speakerembedding approaches may capture anomalies introduced by certain TTS pipelines yet struggle when VC more faithfully preserves speaker identity and natural speech structure. A promising future direction is to explore alternative modalities, such as acoustic sensing data [28], for deepfake audio detection.

1 80 0.0

0

0.5

1

1.5

2

Time (s)

2.5

3

3.5

Detector AASIST2 RawNet2 W2V-AASIST W2V-DF W2V-Large AST-ASV5 W2V-ASV5 ECAPA

70

4

60

Figure 2: Qualitative spectrogram comparison.

50

EER (%)

Table 7: Overall detection results on VoxENES 2026.

40 30

Detector

EER (%)

Acc (%)

TTS EER (%)

VC EER (%)

20

AASIST2 RawNet2 Wav2Vec2-AASIST Wav2Vec2-DF Wav2Vec2-Large AST-ASVspoof5 Wav2Vec2-ASVspoof5 ECAPA-TDNN

57.86 47.03 39.16 55.51 44.38 28.98 51.53 43.22

42.13 52.97 60.85 44.49 55.62 75.94 48.49 56.78

61.2 53.6 43.7 59.3 40.0 20.9 50.6 28.5

49.9 49.8 38.4 49.2 39.7 29.8 53.5 52.5

10

4.4. Per-Method Detection Performance We evaluate different detectors on each speech synthesis method. As shown in Figure 3, Seed-VC is the most challenging synthesis method overall: no evaluated detector achieves an EER below 41%. We hypothesize that its diffusion-based conversion pipeline, augmented with strong speech representations (e.g., Whisper- and WavLM-derived features), yields highly natural outputs that preserve many bonafide time–frequency characteristics, thereby reducing detector-accessible artifacts.

VoxCPM 1.5

Qwen3-TTS

GLM-TTS

Method

Chroma

VibeVoice Detector AASIST2 RawNet2 W2V-AASIST W2V-DF W2V-Large AST-ASV5 W2V-ASV5 ECAPA

80

EER (%)

tic baseline (EER ≥ 47%), most notably AASIST2 (57.86%), which exhibits inverted prediction behavior on this benchmark and suggests a severe domain mismatch between the training distribution and the VoxENES 2026 benchmark. This phenomenon occurs when a classifier learns dataset-specific artifacts, such as silence patterns or channel characteristics, that are reversed or absent in the target domain. Consequently, the model assigns higher bonafide scores to synthetic samples, misinterpreting spoofing cues as genuine speech features. Across models, TTS-generated samples are marginally easier to detect than VC-generated samples for some detectors, but the relative difficulty varies by detector and does not hold uniformly.

0

60 40 20 0

CosyVoice 3 Chatterbox ML

Seed-VC

Method

OpenVoice v2

RVC v2

Figure 3: Per-method EER on Original Samples

5. Conclusion We presented VoxENES 2026, a modern bilingual benchmark for speech spoofing detection that reflects LLM-era TTS and VC generation and realistic post-processing. Using VoxENES 2026, we evaluated eight pretrained detectors without finetuning and observed a substantial generalization gap: the best model achieves only 28.98% EER overall, while many detectors perform near or below chance. These results suggest that many current countermeasures rely on brittle, benchmark-specific artifacts and remain sensitive to modern generators and routine post-processing. We encourage future work on deploymentoriented spoofing detection that improves generalization under data drift and advances evaluation protocols that continuously track the fast-evolving TTS and VC frontier.

6. Use of Generative AI Disclosure Generative AI tools were used for language editing and manuscript polishing. All authors reviewed, verified, and take full responsibility for the content, experiments, and conclusions presented in this paper.

7. References [1] Y. Wang, B. Chen, H. Guo, G. Wang, W. Ding, and Q. Yan, “Clearmask: Noise-free and naturalness-preserving protection against voice deepfake attacks,” in Proceedings of the 20th ACM Asia Conference on Computer and Communications Security, 2025, pp. 696–709. [2] H. Guo, G. Wang, B. Chen, Y. Wang, X. Zhang, X. Chen, Q. Yan, and L. Xiao, “Wavepurifier: Purifying audio adversarial examples via hierarchical diffusion models,” in Proceedings of the 30th Annual International Conference on Mobile Computing and Networking, 2024, pp. 1268–1282. [3] H. Guo, G. Wang, Y. Wang, B. Chen, Q. Yan, and L. Xiao, “Phantomsound: Black-box, query-efficient audio adversarial attack via split-second phoneme injection,” in Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses, 2023, pp. 366–380. [4] Y. Wang, H. Guo, G. Wang, B. Chen, and Q. Yan, “Vsmask: Defending against voice synthesis attack via real-time predictive perturbation,” in Proceedings of the 16th ACM Conference on Security and Privacy in Wireless and Mobile Networks, 2023, pp. 239–250. [5] M. Todisco, X. Wang, V. Vestman, M. Sahidullah, H. Delgado, A. Nautsch, J. Yamagishi, N. Evans, T. Kinnunen, and K. A. Lee, “ASVspoof 2019: Future horizons in spoofed and fake audio detection,” in Proc. Interspeech, 2019, pp. 1008–1012.

[15] Z. Du, C. Gao, Y. Wang, F. Yu, T. Zhao, H. Wang, X. Lv, H. Wang, X. Shi, K. An et al., “CosyVoice 3: Towards in-thewild speech generation via scaling-up and post-training,” arXiv preprint arXiv:2505.17589, 2025. [16] Resemble AI, “Chatterbox: Open-source flow-matching TTS with emotion exaggeration control,” https://huggingface.co/ resemble-ai/chatterbox, 2025. [17] T. Chen, T. Chen, K. Shen, Z. Bao, Z. Zhang, M. Yuan, and Y. Shi, “FlashLabs Chroma 1.0: A real-time end-to-end spoken dialogue model with personalized voice cloning,” arXiv preprint arXiv:2601.11141, 2026. [18] Z. Peng, J. Yu, W. Wang, Y. Chang, Y. Sun, L. Dong, Y. Zhu, W. Xu, H. Bao, Z. Wang et al., “VibeVoice: Expressive podcast generation with next-token diffusion,” in Proc. ICLR, 2025. [19] S. Liu et al., “Zero-shot voice conversion with diffusion transformers,” arXiv preprint arXiv:2411.09943, 2024. [20] MyShell AI, “OpenVoice V2,” https://huggingface.co/myshell-ai/ OpenVoiceV2, 2024. [21] RVC-Project, “Retrieval-based-voice-conversionwebui (RVC),” https://github.com/RVC-Project/ Retrieval-based-Voice-Conversion-WebUI, 2024. [22] V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: An ASR corpus based on public domain audio books,” in Proc. ICASSP, 2015, pp. 5206–5210. [23] C. Wang, M. Riviere, A. Lee, A. Wu, C. Talber, J. Bhosale et al., “VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,” in Proc. ACL, 2021, pp. 993–1003. [24] J.-w. Jung, H.-S. Heo, H. Tak, H.-j. Shim, J. S. Chung, B.-J. Lee, H.-J. Yu, and N. Evans, “AASIST: A new end-to-end antispoofing system using integrated spectro-temporal graph attention networks,” in Proc. ICASSP, 2022, pp. 6367–6371.

[6] J. Yamagishi, X. Wang, M. Todisco, M. Sahidullah, J. Patino, A. Nautsch, X. Liu, K. A. Lee, T. Kinnunen, N. Evans, and H. Delgado, “ASVspoof 2021: accelerating progress in spoofed and deepfake speech detection,” IEEE/ACM Trans. Audio, Speech, and Language Processing, 2024.

[25] H. Tak, J.-w. Jung, J. Patino, M. Kamble, M. Todisco, and N. Evans, “End-to-end spectro-temporal graph attention networks for speaker verification anti-spoofing and speech deepfake detection,” in Proc. 2021 Edition of the Automatic Speaker Verification and Spoofing Countermeasures Challenge, 2021, pp. 1–8.

[7] X. Wang, H. Delgado, H. Tak, J.-w. Jung, H.-j. Shim, M. Todisco, I. Kukanov, X. Liu, M. Sahidullah, T. Kinnunen, N. Evans, K. A. Lee, and J. Yamagishi, “ASVspoof 5: Design, collection and validation of resources for spoofing, deepfake, and adversarial attack detection using crowdsourced speech,” Computer Speech & Language, vol. 95, 2026.

[26] Y. Gong, Y.-A. Chung, and J. Glass, “AST: Audio spectrogram transformer,” in Proc. Interspeech, 2021, pp. 571–575.

[8] J. Frank and L. Schönherr, “WaveFake: A data set to facilitate audio deepfake detection,” in Proc. NeurIPS Datasets and Benchmarks Track, 2021. [9] N. M. Müller, P. Czempin, F. Dieckmann, A. Frober, and K. Böttinger, “Does audio deepfake detection generalize?” arXiv preprint arXiv:2203.16263, 2022. [10] N. M. Müller, P. Kawa, W. H. Choong, E. Casanova, E. Gölge, T. Müller, P. Syga, P. Sperl, and K. Böttinger, “MLAAD: The multi-language audio anti-spoofing dataset,” in Proc. IJCNN, 2024. [11] Z. Yan, Y. Zhao, and H. Wang, “VoiceWukong: Benchmarking deepfake voice detection,” arXiv preprint arXiv:2409.06348, 2024. [12] Y. Zhou, G. Zeng, X. Liu, X. Li, R. Yu, Z. Wang, R. Ye, W. Sun, J. Gui, K. Li et al., “VoxCPM: Tokenizer-free TTS for contextaware speech generation and true-to-life voice cloning,” arXiv preprint arXiv:2509.24650, 2025. [13] H. Hu, X. Zhu, T. He, D. Guo, B. Zhang, X. Wang, Z. Guo, Z. Jiang, H. Hao, Z. Guo et al., “Qwen3-TTS technical report,” arXiv preprint arXiv:2601.15621, 2026. [14] J. Cui, Z. Yang, N. Li, J. Tian, X. Ma, Y. Zhang, G. Chen, R. Yang, Y. Cheng, Y. Zhou et al., “GLM-TTS technical report,” arXiv preprint arXiv:2512.14291, 2025.

[27] B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPATDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,” in Proc. Interspeech, 2020, pp. 3830–3834. [28] G. Wang, Q. Yan, S. Patrarungrong, J. Wang, and H. Zeng, “Facer: Contrastive attention based expression recognition via smartphone earpiece speaker,” in IEEE INFOCOM 2023-IEEE conference on computer communications. IEEE, 2023, pp. 1– 10.

Record · ID 363292 · SHA-256 0188eb20a15a40a4
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.