Conceptio › Archive › arXiv CS
arXiv CSopen access

XSQ-AST: An Explainable Audio Spectrogram Transformer Framework for Localising Synthetic Speech Artifacts

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

XSQ-AST: AN EXPLAINABLE AUDIO SPECTROGRAM TRANSFORMER FRAMEWORK FOR LOCALISING SYNTHETIC SPEECH ARTIFACTS Ben Heritage1,∗ , Luca Resti1,2,∗ , Mónica Villanueva Aylagas3 , Timothy Mehlenbacher4 , Konrad Tollmar3 , James Alfred Walker2 AudioLab, School of Physics, Engineering and Technology, University of York, United Kingdom 2 Department of Computer Science, University of York, United Kingdom 3 SEED – Electronic Arts (EA), Sweden 4 Electronic Arts (EA), United States ABSTRACT

∗ Authors contributed equally. This work was supported by EPSRC Impact Accelerator award EP/X525856/1 and by CoSTARLive Lab, funded by AHRC grant reference AH/Y001079/1.

Loudness Saliency

The advancement of deep generative architectures has enabled Text-to-Speech (TTS) systems to synthesise increasingly natural speech. Traditionally, the evaluation of synthetic speech has relied on subjective Mean Opinion Scores (MOSs) or objective estimators of such global scores, e.g. [1]. While aggregate metrics provide a useful baseline for assessing overall system performance, they offer little to no interpretability. Furthermore, as recent zero-shot neural codec language models and diffusion-based architectures approach parity with human speech [2, 3], global naturalness scores are becoming increasingly saturated, while highly localised signal degradations remain an open challenge. These local artifacts are comprised of transient temporal discontinuities,

Coloration Saliency

1. INTRODUCTION

Input Config

Discontinuity Saliency

Index Terms— Speech synthesis, speech processing, saliency detection, feature extraction, quality of experience

Input Dataset

SQ-AST

Noisiness Saliency

Localising artifacts in synthetic speech remains challenging, as most evaluation methods yield only global quality scores. This paper presents XSQ-AST, a framework that combines the SQ-AST speech quality model with WhisperX phoneme alignment and multiple saliency methods to produce temporally localised artifact diagnostics without model retraining. Saliency maps are projected onto continuous distributions via kernel density estimation and onto phoneme boundaries via phoneme-discretised saliency maps. A 40-participant listening test validated the framework across five perceptual dimensions. Attention Rollout, Attention Flow and an adapted GradCAM produced temporal distributions that correlated with listener highlights, with different methods best suited to different artifact types. An AUC-ROC analysis confirmed discrimination above chance.

Single Utterance Processing

Extract Utterances

Overall MOS Saliency

arXiv:2609.24770v1 [eess.AS] 21 Sep 2026

1

tsv data Score threshold

WhisperX

PDSM, KDE, and ASR (for each dimension under threshold)

SQ-AST Scores

SystemLevel Plots

UtteranceLevel Plots

Fig. 1. Signal flow for XSQ-AST, containing SQ-AST [8] and WhisperX [9] models.

localised distortions or momentary prosodic drift due to cumulative error propagation in sequence generation [4, 5]. Localised distortions have also been more broadly investigated in the context of synthetic speech in videogames [6]. To overcome the limited interpretability of global MOSs, recent research has pivoted towards more fine-grained approaches. Kuhlmann et al. [7] introduced an approach to extract frame-level scores that correlate with human perception of localised artifacts. While effective, this solution relies on modifying the training objective with segment-based consistency constraints to compensate for the lack of frame-level quality annotations in standard datasets. In this paper, eXplainable SQ-AST (XSQ-AST) is presented as a framework for localising synthesis artifacts directly from pre-trained architectures, avoiding the need for specialised training objectives. By integrating SQ-AST [8] with WhisperX [9], the proposed approach extracts feature importance via Attention Rollout [10] and an adapted GradCAM [11]. Because SQ-AST evaluates perceptual dimensions like noisiness and colouration alongside MOSs, degradations are directly associated with specific artifact

classes. Time-frequency saliency maps are then projected onto phoneme boundaries to generate Phoneme Discretized Saliency Maps (PDSMs) [12]. Finally, this localised information is combined with Kernel Density Estimation (KDE), translating model outputs into explainable diagnostics without requiring model retraining. 2. METHODOLOGY The proposed XSQ-AST framework acts as a wrapper for SOTA models in speech sound quality estimation [8] and Automatic Speech Recognition (ASR) [9] to gather temporal and frequency-dependent metrics that are interpretable1 . This section will outline the model input, signal flow as shown in Fig. 1, and output reports.

Flow [10], and an altered Gradient Class Attribution Maps (GradCAM) [11] are implemented. Raw Attention, Attention Rollout, and Attention Flow are implemented as described in [10], whereas GradCAM is an adapted version of the method outlined in [11]. While this method is defined for classification tasks in convolutional neural networks, our approach averages the gradients across the last transformer layer, and since it is a regression model, the scoring of 1 acts as a proxy for the confidence in class prediction. All of these methods return a saliency map regarding to time-frequency patches in utterances, indicating patches that degrade the score the most. This relies on the assumption that the SQ-AST model works by assuming a perfect score, and negating from this with the presence of audio artifacts for a given dimension in the linear layers. The resulting saliency maps are scaled and trimmed to the size of the input Mel-filter bank features, and interpolated.

2.1. Input and Preprocessing The method takes a dataset of multiple-speaker synthetic speech (in one language) of float or PCM audio files up to the sample rate of 48 kHz, and a configuration file. The configuration file allows the user to specify the settings regarding the score threshold, saliency and PDSMs [12], model settings, KDE, and output report types. The score threshold denotes the threshold for utterances where a full analysis is to be conducted. For large datasets, this both controls the number of reports generated and reduces analysis time. The scores for each of the dimensions (MOS, Noisiness, Discontinuity, Colouration, and Loudness) range from 1 − 5, where 5 denotes a perfect sample with no artifacts [8]. Preprocessing is conducted to the input dataset by asserting the audio file is above 2 seconds, handling multiple channel audio files as separate audio files, resampling audio to target of 48 kHz, and segmenting audio files longer than 10 seconds due to constraints in SQ-AST. The segmentation works by finding the minimum number of segments with overlap of 1.5 seconds for the length of audio, allowing the boundaries of the segmented utterances to be accounted for in both neighbouring utterances. Following steps outlined in [8], the utterances are processed by a pre-trained Mel-filter bank feature extractor, and are then normalised to a fixed mean and standard deviation outlined in the SQ-AST code. 2.2. Saliency Extraction When the utterances are parsed through to SQ-AST, based upon the Audio Spectrogram Transformer (AST) model [13], the method checks if any of the dimension scores are below the threshold outlined in the configuration file, and extracts the saliency via the method requested. The saliency extraction methods of Raw Attention, Attention Rollout, Attention 1 Code is available at https://github.com/luca-resti/ synth-speech-eval.

2.3. Automatic Speech Recognition The utterances that have at least one dimension score under the threshold are downsampled to 16 kHz and concatenated into a large array with padding of one second of silence between utterances to increase efficiency in inference. The WhisperX model [9] parses the data to determine the language, transcribe the utterances alongside confidence scores and then align the phonemes from the transcription to the utterances. From the aligned phonemes, we are able to extract the Phoneme Posteriorgrams (PPGs) [14] used in the PDSM method [12]. 2.4. Utterance-Level Reports PPGs and the extracted saliency maps for each dimension are combined by pooling for each phoneme as shown in [12], to extract the PDSM. As well as the pooling methods of sum and mean as shown in [12], the additional methods of median, max, ℓ1 norm, and ℓ2 norm are made available to the user. The PDSM values are ranked, and the top 10% of troublesome phonemes by default, in accordance with [12], are highlighted in a report alongside the phoneme transcription and Mel-filter feature representation of the speech. Additionally the saliency maps are aggregated across the time and frequency domains using KDE determining a smoothed distribution of the troublesome areas. These KDE plots are shown against a waveform and transcription of the speech for explainability, as shown in Fig. 2. Using utterance level ASR confidence, this value is plotted against the waveform and transcription, to show any words or phrases the system has struggled to transcribe. Measures on KDE flatness across the time dimension, inspired by spectral [16] and histogram [17] flatness, are also exported and added to the utterance metadata. This allows users to find the utterances where there are specific troublesome areas or general quality artifacts.

Perceived Influence vs SQ-AST Scores It

a that is booklet records

all information the you prescribed. the onmedication. were

MOS

Extremely

Dimension-Specific Score

=-0.638 p=0.0001

=-0.373 p=0.0424

100

Mean Perceived Influence

Mel Frequency Bins

120

80 60 40 20 0 0.0

0.5

1.0

1.5

2.0

2.5

Time in Seconds

3.0

3.5

4.0

Very

Somewhat

MOS Noisiness Discontinuity Colouration Loudness

Slightly

4.5

Not at all 2.5

Fig. 2. Example saliency output with KDE (in blue) and transcription temporally localising a discontinuity artifact on the last consonant at 4.2s (sample from AudioMOS 2025 [15]). 2.5. System-Level Reports The output reports include two tsv files with a dump of the whole system SQ-AST scores for each dimension requested, and the raw analysis values for the samples that lie under the threshold. Sorted tsv files are also provided for each dimension showing the lowest scoring samples first (along with other analysis details) allowing users to investigate utterance level reports by order of severity. Reports that give a visual aid in showing the distributions of each SQ-AST dimension are given as violin plot and bar chart indicating the amount of samples that fall below the threshold. The system-wide frequency aggregated KDE is shown indicating the troublesome frequency bands for a given system. Additionally, correlation heatmaps between dimensions are shown for the whole dataset and the below-threshold dataset, to give an overview of inter-dimension trends. ASR confidence intervals are shown as a distribution of utterance means and medians, indicating ease of transcription. Histograms of system-wide troublesome phoneme pairs are shown as described in [12], capturing artifacts occurring at specific phoneme boundaries. 3. LISTENING TEST DESIGN A listening test was designed in order to validate the efficacy of our pipeline in regards to the temporal extraction of troublesome features. Participants were asked to highlight a transcribed waveform for areas of the speech that they deem to be troublesome [18, 19] for a given dimension. For each sample in the test, the results from all the participants were aggregated and normalised per sample. The discretised data is used to conduct PDSM analysis, while applying KDE, using the same bandwidth as the model, is used to detail agreement in the KDE output of the system. The synthetic speech used in the test comprised of 30 handpicked samples from the VoiceMOS 2022 datasets [20, 21]. The samples used have localized temporal artifacts, as-

3.0

3.5

SQ-AST MOS

4.0

2.2

2.4

2.6

2.8

3.0

3.2

SQ-AST Dimension Score

3.4

Fig. 3. Perceived influence versus SQ-AST scores. Each marker corresponds to one of the listening test items. sociated with low KDE flatness scores, along with variance in the synthesis methodology and scoring for each dimension. The test was designed using Psychopy [22] and run online with a headphone screening method [23] to ensure adequate responses. 4. RESULTS The listening test was completed by 40 participants, comprised of 20 female, 19 male and 1 other participant. The mean age was 32.8 years (SD = 8.1, range 24-55). No diagnosed hearing impairments were reported by participants. 4.1. SQ-AST Score Validation Before assessing temporal localisation, the relationship between SQ-AST scores and perceived artifact severity was examined. As expected, a significant negative correlation was found between mean perceived influence ratings and SQ-AST MOS values (Spearman’s ρ = −0.638, p = 0.0001). Lower SQ-AST MOS values thus corresponded to higher listenerreported annoyance from artifacts. SQ-AST dimensionspecific scores were also correlated with perceived influence, though more weakly (ρ = −0.373, p = 0.042). The relationship is illustrated in Fig. 3. 4.2. Temporal Artifact Localisation Temporal localisation was evaluated by comparing the KDEs derived from each saliency method with the KDE derived from aggregated listener highlight annotations. Spearman’s ρ was used to measure rank agreement between model and listener attention over time. The methods considered were Raw Attention, Attention Rollout, Attention Flow, GradCAM+ (positive) and GradCAM− (negative). Attention Rollout and Attention Flow propagate attention through the AST layers of the SQ-AST model. Table 1 reports the median sample-wise correlation for each method and quality dimension.

Table 2. Best method combination chosen by the maximum median Spearman’s ρ between model and listener PDSM. Dimension Best Method Combination Overall MOS Nois. Discont. Colour. Loud.

GradCAM+ (ℓ1 norm, sum) Rollout (ℓ2 norm) Rollout (median) Flow (ℓ2 norm) Flow (ℓ2 norm) Rollout (median)

Spearman’s ρ 0.634 0.469 0.806 0.560 0.474 0.806

The highest overall median correlation was observed for Attention Rollout (ρ = 0.444), followed by GradCAM+ (ρ = 0.403) and Attention Flow (ρ = 0.378). Wilcoxon signed-rank tests on the sample-wise correlations indicated that Attention Rollout and Attention Flow each displayed higher median correlations than Raw attention and GradCAM− (p < 0.001). GradCAM+ showed a similar pattern (p ≤ 0.001). GradCAM+ yielded the highest values for MOS (ρ = 0.649) and loudness (ρ = 0.625), Attention Flow for colouration (ρ = 0.732) and noisiness (ρ = 0.389), and Attention Rollout for discontinuity (ρ = 0.358). This indicates that the different attention mechanisms are sensitive to different artifact classes. At the per-sample level, positive correlations were observed in 83.3% of samples for Attention Rollout, 86.7% for Attention Flow and 83.3% for GradCAM+. With a threshold of ρ > 0.3, the corresponding values were 66.7%, 60% and 63.3%. Peak alignment was also examined. The time bin with the highest model saliency fell within the top 50% of listener-highlighted bins on 70% of samples for Attention Rollout, 73.3% for Attention Flow and 60% for GradCAM+. Table 2 presents the best saliency and pooling method combinations across each dimension for PDSM median Spearman’s ρ between the model and listener outputs. The PDSM pooling method was applied to the full model and listener discretised values, giving ranked phonemes. A threshold of 50% was used, resulting in rank analysis between the most troublesome phonemes while avoiding low listener consensus areas. In the dimensions of MOS, discontinuity

0.662 0.686 0.665 0.628 0.501 0.861

− ra G

0.626 0.640 0.659 0.467 0.737 0.505

dC A M

+

0.666 0.691 0.570 0.542 0.741 0.690

dC A M

0.496 0.552 0.465 0.522 0.503 0.431

G ra

Overall MOS Nois. Discont. Colour. Loud.

ow Fl

ut

n en sio

Ro llo

0.034 −0.521 0.155 0.017 0.089 0.041

Ra w

0.403 0.649 0.380 0.257 0.088 0.625

Table 3. Participant-level median AUC-ROC for discriminating highlighted and non-highlighted time bins. The highest value per dimension is shown in bold.

D im

M − dC A G ra

0.378 0.473 0.389 0.234 0.732 0.264

CA

Fl ow

0.444 0.542 0.070 0.358 0.620 0.398

ra d

Ro llo ut

−0.020 0.111 −0.226 0.076 −0.019 −0.230

G

Ra w

Overall MOS Nois. Discont. Colour. Loud.

en im D

sio

M

n

+

Table 1. Median Spearman’s ρ between model and listener KDEs. The highest value per dimension is shown in bold.

0.468 0.375 0.610 0.449 0.576 0.340

and colouration, no other method combinations achieved a median ρ in the margin of 0.05 from the best combination, whereas in the other dimensions, there were multiple methods that achieve this margin. Additionally, all saliency extraction methods achieved overall median ρ > 0.5 using the ℓ1 norm and sum pooling methods. 4.3. Per-Participant Validation As aggregated KDEs can smooth over disagreement between individual participants, a further analysis was performed on a per-participant basis. Each participant’s binary highlight vector was compared against corresponding model-derived KDEs using AUC-ROC. Of the 1200 participant-sample pairs, 1061 contained both highlighted and non-highlighted bins. Attention Rollout, Attention Flow and GradCAM+ yielded median overall AUCs of 0.666, 0.626 and 0.662, while Raw attention and GradCAM− were close to chance. Median AUCs for each dimension for the three best methods are shown in Table 3. 5. CONCLUSION XSQ-AST provides a post-hoc framework for localising artifacts in synthetic speech without retraining. A listening-test conducted on 40 participants confirmed that SQ-AST scores reflect perceived artifact severity, and that model-derived temporal KDEs align with listener highlight annotations at both the aggregate and participant level. Notably, different saliency mechanisms appear to be best suited to different perceptual dimensions. PDSM analysis further demonstrated that phoneme-level localisation is achievable with appropriate pooling strategies. Future work will scale the evaluation to larger datasets, incorporate prosodic analysis via appropriate models (e.g., WavLM [24]), and refine phoneme-level localisation accuracy. The results demonstrate that post-hoc saliency from pretrained speech quality models can yield interpretable, temporally localised diagnostics for synthetic speech artifacts.

6. REFERENCES [1] Takaaki Saeki et al., “UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022,” in Interspeech 2022, 2022, pp. 4521–4525. [2] Sanyuan Chen et al., “Vall-e 2: Neural codec language models are human parity zero-shot text to speech synthesizers,” arXiv preprint arXiv:2406.05370, 2024. [3] Zeqian Ju et al., “Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,” in Forty-first International Conference on Machine Learning, 2024. [4] Paarth Neekhara et al., “Improving Robustness of LLMbased Speech Synthesis by Learning Monotonic Alignment,” in Interspeech 2024, 2024, pp. 3425–3429. [5] Jixun Yao et al., “Fpo: Fine-grained preference optimization improves zero-shot text-to-speech,” IEEE Transactions on Audio, Speech and Language Processing, vol. 34, pp. 557–566, 2025. [6] Luca Resti et al., “Acoustic indicators of join quality in concatenated video game commentary,” in UK and Ireland Speech Workshop 2025. York, 2025. [7] Michael Kuhlmann et al., “Speech quality-based localization of low-quality speech and text-to-speech synthesis artefacts,” in ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026, pp. 15002–15006. [8] Wafaa Wardah et al., “Sq-ast: A transformer-based model for speech quality prediction,” in Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH. 2025, pp. 2335–2339, International Speech Communication Association. [9] Max Bain et al., “WhisperX: Time-Accurate Speech Transcription of Long-Form Audio,” in Interspeech 2023, 2023, pp. 4489–4493. [10] Samira Abnar et al., “Quantifying attention flow in transformers,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online, July 2020, pp. 4190–4197, Association for Computational Linguistics.

[13] Yuan Gong et al., “AST: Audio Spectrogram Transformer,” in Interspeech 2021, 2021, pp. 571–575. [14] Timothy Hazen et al., “Query-by-example spoken term detection using phonetic posteriorgram templates,” 01 2010, pp. 421 – 426. [15] Wen-Chin Huang et al., 2025,” 9 2025.

“The audiomos challenge

[16] N. Madhu, “Note on measures for spectral flatness,” Electronics Letters, vol. 45, pp. 1195–1196, 2009. [17] Abhishek Kumar Tripathi et al., “Performance metrics for image contrast,” in 2011 International Conference on Image Information Processing. IEEE, 2011, pp. 1–4. [18] Fritz Seebauer et al., “Re-examining the quality dimensions of synthetic speech,” in 12th Speech Synthesis Workshop (SSW) 2023, 2023. [19] Michael Kuhlmann et al., “Towards Frame-level Quality Predictions of Synthetic Speech ,” in Interspeech 2025, 2025, pp. 2300–2304. [20] Wen Chin Huang et al., “The VoiceMOS Challenge 2022,” in Interspeech 2022, 2022, pp. 4536–4540. [21] Erica Cooper et al., “How do Voices from Past Speech Synthesis Challenges Compare Today?,” in 11th ISCA Speech Synthesis Workshop (SSW 11), 2021, pp. 183– 188. [22] Jonathan Peirce et al., “Psychopy2: Experiments in behavior made easy,” Behavior research methods, vol. 51, no. 1, pp. 195–203, 2019. [23] Kevin JP Woods et al., “Headphone screening to facilitate web-based auditory experiments,” Attention, Perception, & Psychophysics, vol. 79, no. 7, pp. 2064– 2072, 2017. [24] Sanyuan Chen et al., “Wavlm: Large-scale selfsupervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1505–1518, 2022.

[11] Ramprasaath R Selvaraju et al., “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in Proceedings of the IEEE international conference on computer vision, 2017, pp. 618–626. [12] Shubham Gupta et al., “Phoneme Discretized Saliency Maps for Explainable Detection of AI-Generated Voice,” in Interspeech 2024, 2024, pp. 3295–3299.

The authors/contributors from Electronic Arts (EA) collaborated on this academic research project through supervision of the work. The views, conclusions, methods, and results presented do not necessarily represent the official position, technical claims, product plans, or future direction of EA.

Record · ID 1028694 · SHA-256 77794bc9e8e2698e
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.