ConceptioArchivearXiv CS
arXiv CSopen access

From Self-Supervised Speech Models to Mixture-of-Experts for Robust Anti-Spoofing

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

From Self-Supervised Speech Models to Mixture-of-Experts for Robust Anti-Spoofing Hugo Daumain1,2 , Driss Matrouf1 , Khaled Khelif2 , Mickael Rouvier1 1

LIA, Université d’Avignon, France 2 Airbus Defence & Space, France

[email protected]

arXiv:2606.14639v1 [cs.SD] 12 Jun 2026

Abstract Recent advances in speech generation have significantly improved the naturalness of synthetic speech, making spoofing detection increasingly challenging. A key limitation of current anti-spoofing systems is their limited robustness to unseen synthesis methods. In this work, we transform a self-supervised speech representation model into a Mixture-of-Experts (MoE) architecture to improve generalization. Feed-forward blocks in selected encoder layers are replaced by multiple expert networks controlled by a layer-wise gating mechanism, allowing experts to capture complementary acoustic patterns while preserving the representations learned during self-supervised pretraining. We further analyze the architectural choices affecting the performance of this MoE conversion and investigate the activation behavior of the experts. The proposed approach is evaluated on 14 spoofing datasets and reduces the macro EER from 5.46% to 4.81%, corresponding to 11.9% relative improvement over the baseline. Index Terms: speech anti-spoofing, mixture of experts, selfsupervised learning, robustness

1. Introduction In recent years, speech spoofing technologies have made significant progress. Modern voice conversion systems and text-tospeech synthesizers now rely on neural audio codec [1], flowmatching approaches [2], and diffusion-based architectures [3], producing speech of high perceptual quality. As a result, distinguishing bonafide from spoofed speech (synthetic or manipulated speech) has become increasingly challenging. This technological progress also raises critical societal concerns as highly realistic speech synthesis can be exploited for impersonation, large-scale misinformation, and public opinion manipulation [4, 5]. Developing reliable speech anti-spoofing systems has therefore become a research priority. Classical countermeasure approaches rely on supervised deep neural networks (DNNs) trained to detect artifacts left by speech generation systems. These artifacts manifest as residual traces in the temporal and spectral structure of the signal, often imperceptible to human listeners. Early end-to-end approaches focused on convolutional architectures operating directly on raw waveforms, such as SincNet [6] and RawNet2 [7], which learn discriminative temporal representations for spoofing detection. More recent supervised architectures introduce explicit modeling of spectro-temporal dependencies to better capture these artifacts. In particular, graph-based approaches such as RAWGAT-ST [8] and AASIST [9] leverage graph attention mechanisms to capture structured relationships across time and frequency components of the acoustic signal. However, a fundamental limitation persists: spoofing arti-

facts are often synthesizer-dependent, and models trained on a given set of generation techniques may fail to generalize to unseen attacks. In practical scenarios, where attack methods continuously evolve, robustness to distribution shifts becomes the central challenge. Recently, Self-Supervised Learning (SSL) models such as Wav2vec2 [10], WavLM [11] or HuBERT [12] have emerged as powerful backbone architectures for speech processing tasks. Pre-trained on large-scale unlabeled corpora, these models capture rich acoustic and linguistic representations, yielding highly transferable representations. Their strong generalization capacity has made them particularly attractive for anti-spoofing systems [13]. Beyond representation learning, recent work has explored architectural strategies to further enhance model capacity, among which Mixture-of-Experts (MoE) has gained increasing attention. MoE architectures scale model capacity through conditional expert activation [14], making them well suited to heterogeneous data such as spoofing artifacts, which vary across synthesizers, languages, and acoustic conditions [15]. Furthermore, recent studies have shown that dense pretrained models can be converted into MoE architectures through weight reuse, preserving pretrained knowledge while expanding model capacity [16, 17]. In speech anti-spoofing, MoE approaches have recently been explored at various architectural levels. Some introduce MoE at the model level, treating experts as complete detectors, with either identical [18] or diverse [19] lightweight architectures. Within SSL-based systems, MoE has been applied in two main ways: as a mechanism to aggregate multi-layer representations extracted from the backbone [20, 21], or as a set of low-rank (LoRA) modules [22] integrated directly into the SSL architecture [23, 24, 25]. In the latter setting, each expert is implemented as a low-rank weight correction applied to selected linear layers of a frozen pretrained model. A routing mechanism dynamically selects or combines these corrections depending on the input, enabling input-dependent modulation of the SSL representations. Since these adaptations are parameterized in low rank and the backbone remains frozen, only a small number of additional parameters need to be learned, resulting in efficient and stable task adaptation. However, by constraining each expert to a low-rank correction, these strategies can limit the degree to which expert specialization can reshape internal representations. In this work, we explore the conversion of a self-supervised speech model into a full Mixture-of-Experts (MoE) architecture to improve generalization in speech anti-spoofing. Feedforward blocks in selected encoder layers are replaced by multiple expert networks controlled by a layer-wise gating mechanism, encouraging experts to capture complementary spoofing-

related patterns while preserving the representations learned during self-supervised pretraining. The main contributions of this paper are: • A full MoE conversion paradigm for speech antispoofing: to the best of our knowledge, this is the first work to investigate the conversion of a pretrained SSL speech model into a full Mixture-of-Experts architecture for this task, rather than relying on LoRA-based expert adaptation. • An extensive architectural study: we analyze the impact of several key design choices, including expert placement, number of experts, and pooling strategy for the gating network. • An analysis of expert activation behavior: We examine whether the best-performing MoE configuration shows signs of expert specialization, particularly with regard to its intersynthesizer behavior. Evaluated on 14 spoofing datasets, the best MoE approach based on WavLM-Large reduces the macro EER from 5.46% to 4.81%, corresponding to 11.9% relative improvement over the baseline. This article is organized as follows. We present a stateof-the-art SSL-based anti-spoofing architecture in Section 2 and the proposed Mixture-of-Experts (MoE) conversion in Section 3. After describing the experimental setup in Section 4, we present the results of various experiments on architectural choices in Section 5 and analyze potential expert specialization in Section 6. Finally, Section 7 concludes this work.

2. State-of-the-Art Approach The current state-of-the-art approaches to anti-spoofing are based on a neural architecture composed of a self-supervised learning model, which extracts rich feature representations from the input speech signal (Section 2.1), and a classifier, which discriminates between bonafide and spoofed speech signals (Section 2.2). 2.1. Self-Supervised Learning model Self-Supervised Learning (SSL) models are considered as powerful backbone architectures for speech anti-spoofing. Unlike conventional supervised approaches that learn task-specific representations from limited labeled spoofing datasets, SSL models are pre-trained on large-scale unlabeled speech corpora. This pre-training enables the extraction of rich, hierarchical representations that encode fine-grained acoustic structure as well as higher-level phonetic and prosodic information. Most of these SSL models follow a similar two-stage design (Wav2vec2, HuBERT, WavLM,...) composed of a convolutional feature extractor followed by a transformer encoder. Let x ∈ RTraw be a raw waveform, the feature extractor implemented as a stack of strided 1D convolutions produces a latent sequence H0 ∈ RT ×Fcnn , T < Traw . This sequence is then processed by L transformer encoder layers. Each layer l ∈ {1, . . . , L} is based on a multi-head self-attention (MHA) mechanism and a feed-forward network coupled with residual connections and layer normalization (Figure 2). The output representation Hl ∈ RT ×F of each layer l captures signal-level information in the lower layers, while higher layers encode progressively richer and more abstract representations. 2.2. Classification Following the SSL backbone, a Multi-Head Factorized Attention (MHFA) module [26] is used as the classification head, as

Multi-Head Factorized Attention Q1

Q2

···

QH

QH

SSL model K wL

Transformer L

+

K

Linear

V wL

.. .

AH

+

V

Linear

CNN Encoder

Attentive pooling Concat

w1V

Key flow

w0K

Value flow

w0V

Learned queries

Linear Prediction spoof/bonafide

Figure 1: Representations extracted from the SSL feature extractor and selected Transformer layers are routed to the MHFA module. MHFA attentively aggregates multi-layer representations to produce an utterance-level embedding used for the final prediction.

illustrated in Figure 1. Unlike approaches relying solely on the final Transformer layer, MHFA leverages representations from multiple SSL layers. MHFA applies an attention mechanism with learnable query vectors. Let Hl ∈ RT ×F denote the output of layer l. The aggregated keys and values are computed as K=

L X

wlk Hl Sk ,

H

H

w1K Transformer 1

Head-wise attention

V =

l=0

L X

wlv Hl Sv ,

(1)

l=0

where Sk , Sv ∈ RF ×D are projection matrices, and wlk , wlv are layer-wise scalar weights. The attention weights and head-wise outputs are computed as Ah = softmax(KQh ),

Zh = A ⊤ h V,

(2)

where Qh denotes a learnable query vector associated with head h. The head-wise outputs Zh are then pooled and concatenated to produce the final utterance-level embedding. The utterancelevel embedding is finally passed through a linear layer with a sigmoid activation to produce the bonafide/spoof probability.

3. Mixture-of-Experts Based on the classical approach described above, this section presents the proposed Mixture-of-Experts (MoE) architecture applied to selected layers of the SSL backbone. We first describe the MoE transformation of the Transformer layers (Section 3.1), then detail the gating network used for expert routing (Section 3.2), and finally introduce the auxiliary loss used to prevent expert collapse (Section 3.3). 3.1. SSL Mixture-of-Experts layers We convert a subset {Li | i ∈ S ⊆ {1, . . . , L}} of SSL transformer layers into Mixture-of-Experts (MoE) layers, where each selected layer Li is replaced by its MoE counterpart. As depicted in Figure 2, a MoE layer is obtained by replacing the original dense feed-forward module by E new dense feedforward modules considered as experts. Each one of these E experts is initialized with the weights of the original dense feedforward module, to prevent forgetting knowledge from the pretraining stage. Expert selection is performed by a gating net-

work (described below), which computes a routing distribution over the E experts of the post-MHA representations and activates only the highest top-k scoring experts for the entire utterance. This gating network is layer-dependent; each MoE layer possesses its own gating network.

Let E denote the number of experts. We define pi as the mean routing probability assigned to expert i, and fi as the fraction of utterances routed to expert i. The auxiliary loss is defined as: Laux = E

E X

pi f i .

(6)

i=1 Multi-Head Self-Attention

Multi-Head Self-Attention

Let LBCE denote the binary cross-entropy classification loss. The final training objective is defined as

Add & Norm

Add & Norm Gating Expert 1

Expert 2

×

×

...

Ltotal = LBCE + λLaux .

Expert N

(7)

FFN

Add & Norm

(a) Standard layer

×

4. Experimental setup

Add & Norm

This section describes the experimental protocol used to train and evaluate the proposed approach.

Transformer (b) Mixture-of-Experts Transformer layer

Figure 2: Comparison between (a) a standard SSL model Transformer layer and (b) a Mixture-of-Experts Transformer layer. In (b), the feed-forward network (FFN) is replaced by N parallel FFNs, each considered as an expert; a gating module computes a routing distribution and selects the Top-K experts, whose outputs are combined through a weighted sum. 3.2. Gating network The gating network assigns routing probabilities to the E experts and selects the k most relevant ones according to a top-k routing strategy. To compute these routing probabilities, the frame-level representations produced by the transformer layer must first be aggregated into an utterance-level embedding. Let [h1 , . . . , hT ] ∈ RT ×F denote the sequence of framelevel representations after the MHA module at layer l. An utterance-level representation g is obtained by applying a pooling operation along the temporal dimension. We investigate several pooling strategies such as mean, max, statistical and attentive pooling. Given the pooled representation g, the gating network produces routing probabilities for each expert: p = Softmax(Wg g),

where p ∈ R denotes the probability of selecting each expert. Following the top-k routing strategy, only the k experts with the highest probabilities are activated: (4)

The corresponding probabilities are then re-normalized as p̃i = P

pi

j∈K pj

.

Training datasets: We train our model using the 6 spoofing corpora ASVspoof5 [28], ASVspoof2019 LA [29], ADD22 [30], FakeOrReal [31], Codecfake [32], and MLADD [33]. These datasets cover a wide range of spoofing conditions, synthesis methods, languages, and recording conditions, enabling the model to learn robust and generalizable representations. The distribution of spoof and bonafide samples in these corpora is reported below (Table 1). Table 1: Distribution of spoof and bonafide samples across the training corpora.

Dataset

Spoof

Bonafide

Total

ASVspoof5 ASVspoof2019 LA ADD2022 FakeOrReal Codecfake MLAAD

163 560 22 800 24 072 26 927 634 926 243 000

18 797 2 580 3 012 26 939 105 821 139 085

182 357 25 380 27 084 53 866 740 747 382 085

Total

1 115 285

296 234

1 411 519

(3)

E

K = TopK(p, k).

4.1. Datasets and metrics

(5)

and used as the combination weights of the selected experts. The non-differentiable top-k selection does not prevent training as gradients still flow through the renormalized softmax weights back to the gating module. 3.3. Auxiliary loss To encourage a balanced utilization of experts and prevent expert collapse, we employ an auxiliary load-balancing loss following [27].

Development datasets: Model checkpoint selection is performed using the development splits of several training corpora. Specifically, we use the official development sets of Codecfake [32], ASVspoof2019 LA [29], ASVspoof5 [28], and ADD2022 [30]. Evaluation datasets: For evaluation, we consider 14 corpora to assess generalization under a wide variability of voice generation and manipulation methods. ASVspoof2019 LA [29], ASVspoof2021 LA, ASVspoof2021 DF [34], and ASVspoof5 [28] include a large variety of TTS and VC techniques. Sonar [35] and FakeOrReal [31] focus on recent TTS systems, while DFADD [36] evaluates diffusion and flow-matching-based generation. Codecfake [32] and LibriSeVoc [37] target codec and vocoder-related artifacts respectively. ADD2022 [30] (Track 1 and Track 3/Round 2) and ADD2023 [38] (Track 1/Round 1 and Track 1/Round 2) provide Chinese-language spoofing conditions, and InTheWild [39] assesses performance in unconstrained real-world scenarios. Our evaluation follows the protocol introduced in Speech DF

Arena [40]. Accordingly, for the Sonar corpus we adopt the same preprocessing strategy and exclude samples generated with the SeedTTS synthesizer from the evaluation. Evaluation metrics: We evaluate performance using the Equal Error Rate (EER). The EER corresponds to the operating point where the false acceptance rate is equal to the false rejection rate. To assess cross-dataset generalization, we report both macro and micro EER across the test corpora. The micro (pooled) EER is computed by pooling the scores from all datasets and estimating a single global EER. The macro EER is computed as the average of the EER values obtained independently on each dataset, giving equal weight to all corpora, and is used as our primary metric. 4.2. Implementation Details The toolkit used for the speaker embedding extractor is based on the Kiwano [41] toolkit 1 . We train the model for 80k optimization steps using the AdamW optimizer. The initial learning rate is set to 2 × 10−5 and decayed with a cosine annealing schedule to a final learning rate of 1 × 10−6 . A linear warmup is applied over the first 10% of the total training steps to stabilize optimization in the early phase. Weight decay is set to 1 × 10−2 . Training is performed with a batch size of 128. Training audio samples are standardized to 2-second audio segments, obtained by randomly cropping a 2 s window from each training utterance at every iteration (when longer than 2 s). The model is trained with a binary cross-entropy loss. For the MoE configuration, the previously introduced auxiliary loss is weighted by 1 × 10−2 . At the beginning of training, the parameters of the self-supervised learning (SSL) backbone are frozen, and only the gating network, expert layers, and the MHFA module are optimized. The remaining SSL parameters are then progressively unfrozen following a linear schedule over the first 15% steps. To improve robustness, we apply on-the-fly data augmentation during training, including codec-based augmentation, additive noise augmentation using MUSAN, and reverberation augmentation via convolution with room impulse responses.

acoustic information. Table 2: Baseline performance comparison using 13 or 24 layers (EER ↓ in %). SSL

Layer used

Macro EER (%) ↓

Micro EER (%) ↓

WavLM Large

24 13

5.61 5.46

15.61 14.95

Wav2vec2 XLSR

24 13

6.01 6.01

14.89 13.71

HuBERT Large

24 13

6.21 6.24

11.97 12.91

As shown in Table 2, the 13-layer configuration achieves better performance for WavLM-Large, supporting this choice. 5.2. Impact of MoE Layer Position In the first experiment, we investigate the effect of inserting MoE layers at different positions within the 13 used Transformer layers of WavLM-Large. An attentive statistical pooling method is employed to produce the gating representation, and the routing strategy is fixed to top-k = 1 (hard routing). We evaluate four insertion strategies: early insertion (first six layers), late insertion (last six layers), full insertion (all 13 layers), and alternating insertion (one layer out of two). Table 4: Impact of MoE insertion position. All variants use E = 4 experts, top-k = 1 routing, and attentive pooling on the 13 used WavLM-Large layers. The Insertion column indicates which Transformer layers are converted to MoE layers. Type

SSL

Insertion

Macro EER (%) ↓

Baseline

WavLM-L (13)

5.46

14.95

WavLM-L (13)

first 6 last 6 all 13 alternating

5.60 5.21 5.77 5.42

15.35 13.80 14.13 14.27

MoE

Micro EER (%) ↓

5. Results This section presents the experimental evaluation of the proposed MoE conversion. The used SSL backbone is first identified (Section 5.1), followed by a study of each architectural design choice (Sections 5.2 to 5.4), and a comparison with LoRAbased alternatives (Section 5.5). Detailed per-corpus results are provided in Table 3.

The best performance, as shown in Table 4, is consistently obtained when MoE layers are inserted in the last six layers. This suggests that such layers are more effective when applied to higher-level representations rather than low-level features. 5.3. Impact of the Pooling Strategy for Gating

5.1. SSL Backbone Selection We evaluate three SSL backbone encoders: WavLM-Large [11], Wav2vec2-XLSR [10] pre-trained on 128 languages, and HuBERT-Large [12]. As shown in Table 2, WavLM-Large consistently achieves the lowest macro EER and is therefore selected as the backbone for all subsequent experiments. Although WavLM-Large consists of 24 Transformer layers, we restrict MHFA to only use the first 13 layers, following the MHFA reference implementation [26]. Since speech antispoofing relies predominantly on detecting low-level acoustic artifacts rather than linguistic information, restricting the model to the first 13 layers allows us to preserve the most relevant 1 https://github.com/kiwano-toolkit/kiwano

Using the configuration identified in the previous experiment where six MoE layers are inserted in the last Transformer layers, we now investigate how the pooling strategy used to compute the gating representation influences performance. The gating network requires an utterance-level embedding to determine which expert should be activated. The quality of this representation may therefore directly affect the routing decisions and, consequently, the overall system performance. To study this aspect, we evaluate several pooling strategies, presented in Section 3.2, that aggregate frame-level representations into a single utterance-level vector. In particular, we evaluate attentive statistical pooling, statistical pooling, max pooling and mean pooling.

Table 3: Detailed anti-spoofing results (EER ↓, in %) with Codec = Codecfake and LSV = LibriSeVoc. The MoE Config columns specify the insertion position (first 6, last 6, all 13, or alternating), the number of experts E, the top-k value, and the pooling strategy (att = attentive, stat = statistical). Macro EER averages EERs across corpora; Micro EER pools all trials. Best in bold, second-best underlined. Type

Baseline

MoE

SSL

MoE Config

#Params

WavLM-L (13) WavLM-L (24) XLSR (13) XLSR (24) HuBERT-L (13) HuBERT-L (24)

WavLM-L (13)

Overall

ASVspoof

MoE

LSV

ITW

Sonar

16.34 18.61 15.37 13.74 15.38 17.57

0.00 0.00 0.13 0.00 0.00 0.13

2.05 2.09 5.48 5.35 1.57 1.92

0.00 0.00 0.35 0.43 0.00 0.09

0.21 0.39 0.52 0.98 3.10 2.30

1.47 1.54 1.84 2.36 3.21 2.45

0.14 0.21 3.63 1.53 5.85 4.66

9.43 8.48 9.86 9.33

17.01 15.46 18.83 16.92

0.13 0.00 0.13 0.10

2.26 2.44 2.53 1.80

0.00 0.00 0.04 0.13

0.88 0.38 0.87 0.50

1.94 1.60 2.27 1.86

0.56 0.08 0.76 0.29

4.07 4.26 5.23

7.89 8.18 9.42

14.46 14.23 17.63

0.12 0.13 0.00

1.28 1.60 1.82

0.04 0.04 0.04

0.18 0.14 0.42

1.49 1.47 1.62

0.08 0.14 0.08

4.65 4.94 5.32 5.66 5.52 4.72 5.02 5.67

7.99 7.90 8.20 9.07 8.90 8.43 7.96 10.05

14.96 14.12 16.54 17.66 16.02 15.21 14.43 18.22

0.00 0.12 0.25 0.02 0.00 0.02 0.00 0.13

1.30 1.90 2.50 1.73 2.75 1.97 1.35 1.95

0.04 0.09 0.00 0.04 0.00 0.00 0.04 0.04

0.14 0.17 0.16 0.40 0.45 0.32 0.17 0.36

1.56 1.48 1.39 1.52 1.53 1.60 1.36 1.61

0.14 0.19 0.27 0.27 0.06 0.08 0.35 0.49

Micro

19LA

21DF

21LA

24

T1

T3-R2

T1-R1

T1-R2

178M 317M 178M 317M 178M 317M

— — — — — —

-

-

— — — — — —

5.46 5.61 6.01 6.01 6.24 6.21

14.95 15.61 13.71 14.89 12.91 11.97

0.04 0.04 0.33 0.68 0.15 0.19

0.30 0.43 0.89 1.31 0.72 1.02

2.38 2.32 3.60 4.83 4.36 4.55

17.12 14.70 15.80 17.89 16.62 16.23

21.46 21.89 22.70 21.07 23.08 21.89

5.50 6.04 4.81 5.26 4.85 4.45

9.45 10.33 8.67 8.75 8.46 9.47

330M 330M 507M 355M

first 6 last 6 all 13 alternating

4 4 4 4

1 1 1 1

att att att att

5.60 5.21 5.77 5.42

15.35 13.80 14.13 14.27

0.05 0.03 0.12 0.04

0.43 0.30 0.45 0.43

3.14 2.67 3.09 2.62

15.29 17.34 15.83 15.85

21.16 19.32 20.85 20.42

6.13 4.88 5.10 5.59

329M 329M 329M

last 6 last 6 last 6

4 4 4

1 1 1

stat max mean

4.81 4.91 5.35

12.34 12.73 13.75

0.04 0.05 0.04

0.29 0.42 0.43

3.14 3.20 2.62

16.28 16.13 15.81

18.01 18.70 19.72

227M 278M 278M 329M 329M 378M 378M 378M

last 6 last 6 last 6 last 6 last 6 last 6 last 6 last 6

2 3 3 4 4 5 5 5

1 1 2 2 3 1 2 3

stat stat stat stat stat stat stat stat

4.98 5.13 5.33 5.41 5.40 5.17 5.00 5.60

12.81 13.81 13.07 13.99 14.46 13.60 13.00 14.23

0.04 0.05 0.05 0.03 0.03 0.03 0.04 0.03

0.30 0.43 0.43 0.43 0.29 0.31 0.32 0.29

2.85 3.13 2.97 2.28 2.56 2.85 2.78 2.56

16.09 16.84 16.56 17.01 17.23 17.60 16.68 16.67

19.60 20.40 20.03 19.62 20.23 19.30 19.56 20.35

Pooling

Macro EER (%) ↓

Micro EER (%) ↓

WavLM-L (13)

attentive statistical max mean

4.99 4.81 4.91 5.35

13.39 12.34 12.73 13.75

The last experiment was conducted by combining the best configurations from previous experiments. We fixed the 6 MoE layers as the last WavLM-Large layers and used statistical pooling as the gating pooling method. We analyze the effect of varying the number of experts E and the top routing parameter k. These two parameters affect different aspects of the model. The routing parameter k controls the inference cost by determining how many experts are activated per input, whereas the number of experts E mainly impacts the model size. Table 6: Impact of the number of experts E and top-k routing. All variants use statistical pooling, with MoE layers inserted in the last 6 of the 13 used WavLM-Large layers.

WavLM-L (13)

FoR

Macro

5.4. Impact of the Number of Experts and Top-k

MoE

Codec

Pool

SSL

SSL

DFADD

k

Table 5 shows that simpler pooling strategies lead to improved performance compared to the attentive statistical pooling used in the previous configuration. In particular, both statistical pooling and max pooling yield lower EER values, with statistical pooling providing the best overall results. Based on these observations, we adopt statistical pooling as the default pooling strategy in the following experiments.

Type

ADD23

E

Table 5: Impact of the pooling strategy. All variants use E = 4 experts and top-k = 1 routing, with MoE layers inserted in the last 6 of the 13 used WavLM-Large layers. Type

ADD22

Insertion

#Params

E

k

Macro EER (%) ↓

227M

2

1

4.98

12.81

278M 278M

3 3

1 2

5.13 5.33

13.81 13.07

329M 329M 329M

4 4 4

1 2 3

4.81 5.41 5.40

12.34 13.99 14.46

378M 378M 378M

5 5 5

1 2 3

5.17 5.00 5.60

13.60 13.00 14.23

Table 6 reports the impact of varying the number of experts E and the routing parameter k. The best performance is obtained with E = 4 experts and k = 1, achieving a macro EER of 4.81% and a micro EER of 12.34%. Configurations with larger routing values (k ≥ 2) tend to degrade performance, suggesting that activating multiple experts simultaneously may reduce the specialization effect expected from the MoE mechanism. Interestingly, the configuration with E = 2 experts and k = 1 also achieves competitive performance while requiring a significantly smaller model (227M parameters). Although this setting does not reduce inference cost compared to E = 4 (since k = 1 in both cases), it provides a more memory-efficient alternative. 5.5. Comparison with LoRA-based MoE Approaches To validate our design choice using dense experts over LoRAbased experts adopted in prior works, we compare both strategies under identical conditions: MoE layers in the last six layers, statistical pooling, E = 4 experts, and top-k = 1 routing. In the LoRA-based setting, each layer expert is designed as a pair of LoRA modules applied to the two dense layers of the feed forward network block. Only these adapters and the gating network are trained, while the rest of the model remains frozen. Table 7: Comparison between full fine-tuning and LoRA-based adaptation for WavLM-Large MoE. Our MoE approach finetunes all 329M parameters, while LoRA-based approaches update only a small subset of a smaller model. Type MoE (ours)

SSL

Rank

EER (%) ↓

Trainable

WavLM-L (13)

329M

329M

4.81

12.34

WavLM-L (13)

178M 182M 186M 194M

2M 4M 8M 16M

8 16 32 64

6.80 6.84 6.73 6.66

16.23 15.92 15.79 16.06

Micro EER (%) ↓

MoE-LoRA

#Params Total

Macro

Micro

As shown in Table 7, our MoE conversion, at the cost of more parameters, outperforms LoRA-based approaches across all ranks, with LoRA performance improving as rank increases. This gap can be attributed to both the removal of the low-

rank constraint on experts, and the fact that attention layers are jointly fine-tuned, allowing the model to co-adapt with the experts. While this comparison does not fully isolate the effect of the MoE structure from that of full fine-tuning, it suggests that both factors are important for expert expressiveness in this task.

6. Analysis

across synthesis methods, we compute the Jensen–Shannon (JS) divergence between the expert activation distributions p(e|s, l) defined above. The JS divergence, which is based on the Kullback–Leibler divergence, quantifies the dissimilarity between two probability distributions: JS(P ∥ Q) =

To assess whether MoE routing exhibits specialization with respect to particular synthesizers or, more broadly, attack types, we first analyze the distribution of expert activations for each synthesizer in the Sonar corpus. For each MoE layer l, the activation counts of expert e for synthesizer s are normalized to obtain a probability distribution p(e|s, l). Figure 3 shows this distribution across the 6 MoE layers used in our best-performing model.

1 1 KL(P ∥ M ) + KL(Q ∥ M ) 2 2

(8)

where M = 12 (P + Q) and KL(·) denotes the Kullback– Leibler divergence. For each layer, pairwise JS divergence is computed for all synthesizer pairs, and the mean divergence is reported per dataset. This analysis is conducted on three test corpora: Sonar, ASVspoof2024 and Codecfake. Table 8: Mean pairwise Jensen–Shannon divergence between synthesizer-specific expert activation distributions.

Layer 8 p(e|s, l=8)

0.75

MoE layer

Codecfake

ASV24

Sonar

0.5

8

0.086

0.189

0.273

0.25

9

0.113

0.171

0.286

10

0.121

0.273

0.210

11

0.156

0.210

0.201

12

0.159

0.253

0.258

13

0.162

0.299

0.291

0 n Ge dio Au

h ec

e Sp sh

Fla

h3 ec

pe

alS

tur

Na

I nA

e

Op

2 TS

ptT

om

Pr

LE

L VA

TS

xT

Layer 9 p(e|s, l=9)

0.75 0.5 0.25 0 ch ee

ee Sp sh

lSp

Fla

I

3

ch

en

G dio Au

Na

a tur

2

A en

Op

TS

ptT

om

Pr

LE

L VA

TS

xT

p(e|s, l=10)

Layer 10 0.75 0.5 0.25 0 ch

n

Ge dio Au

ee Sp

sh

Fla

h3 ec pe

alS

tur Na

AI

2

TS

en

Op

ptT

om

Pr

LE

L VA

TS

xT

p(e|s, l=11)

Layer 11 0.75 0.5 0.25 0 ch

en

ioG

d Au

3

ch ee

ee Sp

lSp

sh

Fla

a tur

Na

en

Op

AI

2

TS

ptT

om

Pr

LE

L VA

TS

xT

p(e|s, l=12)

Layer 12 0.75 0.5

7. Conclusion

0.25 0 en ioG ud

A

h3 ec

h

c ee Sp

sh

Fla

pe

alS

tur Na

I nA

2 TS

e

Op

ptT

om

Pr

E LL

VA

S TT

x

Layer 13 p(e|s, l=13)

Table 8 summarizes the resulting JS mean divergences. For Codecfake and ASVspoof2024, we observe that the divergence slightly increases in deeper MoE layers. In contrast, Sonar shows a different trend, with divergence decreasing in the first layers before increasing again in the deepest ones. These trends suggest that routing patterns become slightly more distinct across synthesis methods in the deeper layers. However, the magnitude of the mean divergence remains relatively low. This indicates that expert activation distributions remain largely similar across synthesizers and do not provide evidence of specialization to specific spoofing artifacts. More generally, interpreting the role of individual experts remains challenging. Even when variations in routing patterns are observed, it is difficult to associate them with specific attacks (spoofing artifacts), as experts may capture complex acoustic characteristics.

0.75 0.5 0.25 0 ch

n Ge

dio Au

ee Sp

sh

Fla

Expert 1

alS

tur

Na

h3

ec

pe

en

Op

Expert 2

AI

2

TS

ptT

om

Pr

Expert 3

L VA

LE

TS

xT

Expert 4

Figure 3: Expert activation distribution p(e | s, l) per layer l and Sonar synthesizer s. The histograms reveal that expert activations remain relatively balanced across synthesizers. Although some synthesizers, such as AudioGen or OpenAI, show slightly different activation patterns, these differences do not indicate a clear routing specialization. To quantitatively measure differences in expert routing

In this work, we proposed a method to convert a self-supervised speech model into a Mixture-of-Experts architecture for speech anti-spoofing. Unlike recent LoRA-based MoE approaches that keep the backbone frozen, our method replaces the dense feed-forward modules with full expert networks initialized from the pretrained weights, and jointly fine-tunes the entire model, prioritizing performance and robustness over parameter efficiency. Validated on WavLM-Large across 14 evaluation corpora, the best configuration reduces the macro EER from 5.46% to 4.81%, an 11.9% relative improvement over the baseline, and outperforms LoRA-based alternatives across all tested ranks. Our analysis of expert activation patterns showed no clear specialization with respect to specific attacks. However, experts may capture complex acoustic patterns which are not straightforward to interpret. Future work will conduct more in-depth analyses to better characterize potential expert specialization, and explore strategies to explicitly guide experts toward distinct spoofing methods.

8. Acknowledgements This work was performed using HPC resources from GENCI–IDRIS (Grant 2025-AD011017195).

9. References [1] C. Wang, S. Chen, Y. Wu, Z. Zhang, L. Zhou, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li et al., “Neural codec language models are zero-shot text to speech synthesizers,” arXiv preprint arXiv:2301.02111, 2023. [2] S. Mehta, R. Tu, J. Beskow, É. Székely, and G. E. Henter, “Matcha-tts: A fast tts architecture with conditional flow matching,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp. 11 341–11 345. [3] V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, and M. Kudinov, “Grad-tts: A diffusion probabilistic model for text-to-speech,” in International conference on machine learning. PMLR, 2021, pp. 8599–8608. [4] C. Vaccari and A. Chadwick, “Deepfakes and disinformation: Exploring the impact of synthetic political video on deception, uncertainty, and trust in news,” Social media+ society, vol. 6, no. 1, p. 2056305120903408, 2020. [5] J. Yi, C. Wang, J. Tao, X. Zhang, C. Y. Zhang, and Y. Zhao, “Audio deepfake detection: A survey,” arXiv preprint arXiv:2308.14970, 2023. [6] M. Ravanelli and Y. Bengio, “Interpretable convolutional filters with sincnet,” arXiv preprint arXiv:1811.09725, 2018. [7] H. Tak, J. Patino, M. Todisco, A. Nautsch, N. Evans, and A. Larcher, “End-to-end anti-spoofing with rawnet2,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021, pp. 6369–6373. [8] H. Tak, J.-w. Jung, J. Patino, M. Kamble, M. Todisco, and N. Evans, “End-to-end spectro-temporal graph attention networks for speaker verification anti-spoofing and speech deepfake detection,” arXiv preprint arXiv:2107.12710, 2021. [9] J.-w. Jung, H.-S. Heo, H. Tak, H.-j. Shim, J. S. Chung, B.-J. Lee, H.-J. Yu, and N. Evans, “Aasist: Audio anti-spoofing using integrated spectro-temporal graph attention networks,” in ICASSP 2022-2022 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2022, pp. 6367–6371. [10] A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in neural information processing systems, vol. 33, pp. 12 449–12 460, 2020. [11] S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al., “Wavlm: Large-scale selfsupervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1505–1518, 2022. [12] W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM transactions on audio, speech, and language processing, vol. 29, pp. 3451–3460, 2021. [13] X. Li, P.-Y. Chen, and W. Wei, “Measuring the robustness of audio deepfake detectors,” arXiv preprint arXiv:2503.17577, 2025. [14] H. Huang, N. Ardalani, A. Sun, L. Ke, H.-H. S. Lee, S. Bhosale, C.-J. Wu, and B. Lee, “Toward efficient inference for mixture of experts,” Advances in Neural Information Processing Systems, vol. 37, pp. 84 033–84 059, 2024. [15] S. Mu and S. Lin, “A comprehensive survey of mixture-ofexperts: Algorithms, theory, and applications,” arXiv preprint arXiv:2503.07137, 2025. [16] A. Komatsuzaki, J. Puigcerver, J. Lee-Thorp, C. R. Ruiz, B. Mustafa, J. Ainslie, Y. Tay, M. Dehghani, and N. Houlsby, “Sparse upcycling: Training mixture-of-experts from dense checkpoints,” arXiv preprint arXiv:2212.05055, 2022.

[17] L. Fu, S. Yu, S. Li, L. Fan, Y. Wu, and X. He, “Ume: Upcycling mixture-of-experts for scalable and efficient automatic speech recognition,” in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, pp. 1–5. [18] V. Negroni, D. Salvi, A. I. Mezza, P. Bestagini, and S. Tubaro, “Leveraging mixture of experts for improved speech deepfake detection,” in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, pp. 1–5. [19] ——, “Attention-based mixture of experts for robust speech deepfake detection,” arXiv preprint arXiv:2509.17585, 2025. [20] Z. Wang, R. Fu, Z. Wen, J. Tao, X. Wang, Y. Xie, X. Qi, S. Shi, Y. Lu, Y. Liu et al., “Mixture of experts fusion for fake audio detection using frozen wav2vec 2.0,” in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, pp. 1–5. [21] Y. Hao, Y. Chen, M. Xu, J. Zhan, L. He, L. Fang, S. Fang, and L. Liu, “Wav2df-tsl: Two-stage learning with efficient pretraining and hierarchical experts fusion for robust audio deepfake detection,” in 2025 International Joint Conference on Neural Networks (IJCNN). IEEE, 2025, pp. 1–8. [22] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen et al., “Lora: Low-rank adaptation of large language models.” Iclr, vol. 1, no. 2, p. 3, 2022. [23] J. Laakkonen, I. Kukanov, and V. Hautamäki, “Mixture of lowrank adapter experts in generalizable audio deepfake detection,” in 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC). IEEE, 2025, pp. 2211–2216. [24] Z. Pan, S. H. Bhupendra, and J. Wu, “Molex: Mixture of lora experts in speech self-supervised models for audio deepfake detection,” arXiv preprint arXiv:2509.09175, 2025. [25] Q. Chen, Y. Xu, S. Mandelli, S. Li, and B. Li, “Adaptive mixture of low-rank experts for robust audio spoofing detection,” IEEE Signal Processing Letters, 2025. [26] J. Peng, O. Plchot, T. Stafylakis, L. Mošner, L. Burget, and J. Černockỳ, “An attention-based backend allowing efficient finetuning of transformer models for speaker verification,” in 2022 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2023, pp. 555–562. [27] W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” Journal of Machine Learning Research, vol. 23, no. 120, pp. 1–39, 2022. [28] X. Wang, H. Delgado, H. Tak, J.-w. Jung, H.-j. Shim, M. Todisco, I. Kukanov, X. Liu, M. Sahidullah, T. Kinnunen et al., “Asvspoof 5: Crowdsourced speech data, deepfakes, and adversarial attacks at scale,” arXiv preprint arXiv:2408.08739, 2024. [29] A. Nautsch, X. Wang, N. Evans, T. H. Kinnunen, V. Vestman, M. Todisco, H. Delgado, M. Sahidullah, J. Yamagishi, and K. A. Lee, “Asvspoof 2019: Spoofing countermeasures for the detection of synthesized, converted and replayed speech,” IEEE Transactions on Biometrics, Behavior, and Identity Science, vol. 3, no. 2, pp. 252–265, 2021. [30] J. Yi, R. Fu, J. Tao, S. Nie, H. Ma, C. Wang, T. Wang, Z. Tian, Y. Bai, C. Fan et al., “Add 2022: the first audio deep synthesis detection challenge,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, pp. 9216–9220. [31] R. Reimao and V. Tzerpos, “For: A dataset for synthetic speech detection.” in SpeD, 2019, pp. 1–10. [32] Y. Xie, Y. Lu, R. Fu, Z. Wen, Z. Wang, J. Tao, X. Qi, X. Wang, Y. Liu, H. Cheng et al., “The codecfake dataset and countermeasures for the universally detection of deepfake audio,” IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 386–400, 2025.

[33] N. M. Müller, P. Kawa, W. H. Choong, E. Casanova, E. Gölge, T. Müller, P. Syga, P. Sperl, and K. Böttinger, “Mlaad: The multilanguage audio anti-spoofing dataset,” in 2024 International Joint Conference on Neural Networks (IJCNN). IEEE, 2024, pp. 1–7. [34] J. Yamagishi, X. Wang, M. Todisco, M. Sahidullah, J. Patino, A. Nautsch, X. Liu, K. A. Lee, T. Kinnunen, N. Evans et al., “Asvspoof 2021: accelerating progress in spoofed and deepfake speech detection,” arXiv preprint arXiv:2109.00537, 2021. [35] X. Li, P.-Y. Chen, and W. Wei, “Sonar: A synthetic ai-audio detection framework and benchmark,” 2024. [36] J. Du, I.-M. Lin, I.-H. Chiu, X. Chen, H. Wu, W. Ren, Y. Tsao, H.-Y. Lee, and J.-S. R. Jang, “Dfadd: The diffusion and flowmatching based audio deepfake dataset,” in 2024 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2024, pp. 921– 928. [37] C. Sun, S. Jia, S. Hou, and S. Lyu, “Ai-synthesized voice detection using neural vocoder artifacts,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 904–912. [38] J. Yi, J. Tao, R. Fu, X. Yan, C. Wang, T. Wang, C. Y. Zhang, X. Zhang, Y. Zhao, Y. Ren et al., “Add 2023: the second audio deepfake detection challenge,” arXiv preprint arXiv:2305.13774, 2023. [39] N. M. Müller, P. Czempin, F. Dieckmann, A. Froghyar, and K. Böttinger, “Does audio deepfake detection generalize?” arXiv preprint arXiv:2203.16263, 2022. [40] S. Dowerah, A. Kulkarni, A. Kulkarni, H. M. Tran, J. Kalda, A. Fedorchenko, B. Fauve, D. Lolive, T. Alumäe, and M. M. Doss, “Speech df arena: A leaderboard for speech deepfake detection models,” IEEE Open Journal of Signal Processing, 2026. [41] M. Rouvier and P.-M. Bousquet, “Kiwano: A Cutting-Edge OpenSource Toolkit for Speaker Verification,” in Odyssey 2026, 2026.

Record · ID 271873 · SHA-256 72642d3dba9e0191
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.