ConceptioArchivearXiv CS
arXiv CSopen access

How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech Nityanand Mathur1 , Hamees Sayed1 , Wasim Madha1 , Apoorv Singh1 , Sameer Khurana1 , Akshat Mandloi1 , Sudarshan Kamath1 1

Smallest.ai

[email protected]

arXiv:2606.20532v1 [cs.AI] 18 Jun 2026

Abstract Style-captioned text-to-speech systems use natural language to control voice characteristics, but how individual words influence acoustic output remains unclear. Understanding this is critical for diagnosing failure modes and improving controllability in expressive TTS. We propose cross-attention attribution for speech diffusion models, adapting the DAAM framework to the speech domain for the first time, and apply it to CapSpeechTTS. Our method extracts per-token heatmaps across 25 layers and 24 ODE steps. We analyze 3,600 (style caption, text transcript) combinations comprising 120 style captions conditioning the generation of 30 text transcripts each, revealing how caption tokens shape waveforms. Results show: (1) style tokens have lower temporal variance than content/function tokens, confirming global conditioning; (2) style attention correlates with F0 and energy; (3) style conditioning peaks in early steps and deep layers; (4) attention entropy reaches its minimum at layer 17, co-occurring with the style importance peak, indicating maximal network selectivity at the most style-critical stage. This is the first study of how natural language influences cross-attention in speech diffusion models. Index Terms: text-to-speech, cross-attention, interpretability, diffusion models, style conditioning

1. Introduction Modern text-to-speech (TTS) systems have moved beyond fixed speaker embeddings [1] toward natural language conditioning, where free-form captions such as “a calm, deep voice speaking slowly” control the style of generated speech [2, 3, 4]. This paradigm, building on advances in generative spoken language modeling [5], offers powerful expressivity but introduces a fundamental interpretability question: how do individual words in a style caption influence the synthesized audio? In text-to-image generation, Diffusion Attentive Attribution Maps (DAAM) [6] answered an analogous question by extracting cross-attention maps from Stable Diffusion [7] and attributing spatial regions to prompt tokens. Prompt-to-Prompt [8] and generic attention rollout [9] further established cross-attention as a faithful proxy for conditioning influence. No equivalent framework exists for speech. Unlike images, speech is inherently temporal: attention must map caption tokens to spectrogram time frames rather than spatial pixels. Furthermore, the distinction between global properties (style, emotion) and local properties (phoneme identity, duration) creates an interpretability challenge unique to speech. We address this gap by adapting DAAM to CapSpeech [2], a non-autoregressive TTS model that uses flow matching [10] conditioned on T5-encoded [11] style captions. Our contributions are:

❶ The first cross-attention attribution analysis for TTS, extracting per-token temporal heatmaps across all 25 transformer layers and 24 ODE steps over 3,600 (style caption, text transcript) combinations. ❷ A global-vs-local analysis showing that style tokens exhibit significantly lower temporal variance (p < 10−43 , d = −1.16) than content and function tokens. ❸ An acoustic grounding analysis demonstrating that styletoken attention correlates with F0 and energy with semantic coherence (e.g., “loud” ↔ energy, r = +0.64). ❹ A layer-step dynamics study revealing hierarchical conditioning: style importance peaks in early ODE steps (5.2× decay) and deepens through transformer layers.

2. Related Work Style-conditioned TTS. Recent TTS systems accept natural language descriptions to control voice style. CapSpeech [2] conditions a flow-matching transformer on T5-encoded captions, while VoiceBox [3] and NaturalSpeech 3 [4] explore similar paradigms. Earlier autoregressive approaches like Tacotron [12], Tacotron 2 [13], and FastSpeech [14] established attention-based TTS but relied on fixed speaker embeddings [1] rather than flexible text conditioning. Zero-shot multi-speaker systems like YourTTS [15] and NaturalSpeech 2 [16] further demonstrated the potential of latent diffusion for voice control. Despite impressive generation quality, the internal mechanisms by which these systems process style instructions remain unexplored. Interpretability in generative models. DAAM [6] showed that cross-attention maps in Stable Diffusion [7] faithfully attribute image regions to individual prompt tokens. Prompt-toPrompt [8] leveraged these maps for controlled image editing, while Chefer et al. [9] proposed generic attention rollout for encoder-decoder transformers. In speech, attention analysis has been limited to alignment visualization in autoregressive models [17]. Our work brings DAAM-style attribution to the speech domain for the first time. Flow matching for speech. Flow matching [10] offers a simulation-free alternative to diffusion probabilistic models [18, 19] by directly learning the velocity field of the probability flow ODE. Unlike diffusion models that require iterative denoising through a learned score function, flow matching provides a deterministic mapping from noise to data via continuous normalizing flows. This formulation has gained traction in speech synthesis, with models like Grad-TTS [20] pioneering diffusion-based mel-spectrogram generation, GlowTTS [21] using normalizing flows, and F5-TTS [22] demonstrating flow matching for high-quality synthesis. CapSpeech adopts flow matching with a Diffusion Transformer (DiT) [23]

backbone, processing mel-spectrogram latents through a transformer [24] with explicit cross-attention at every layer and ODE step. This design makes it particularly amenable to DAAMstyle attribution: unlike autoregressive models where attention primarily serves alignment, the DiT’s cross-attention directly conditions the velocity field on caption embeddings at each generation stage, providing a clear interpretable pathway from text to acoustics.

3. Method 3.1. CapSpeech architecture overview Figure 1 illustrates the CapSpeech pipeline and our DAAM attribution mechanism. The model comprises four components working in sequence to transform a style caption and text transcript into a waveform. The pipeline begins with a T5 caption encoder [11], which maps the style caption c = (c1 , . . . , cTc ) to contextual embeddings Ec ∈ RTc ×d ; T5’s pre-trained encoder provides rich semantic representations that capture linguistic nuances in style descriptions. These text embeddings are complemented by a CLAP encoder [25], which produces ′ a global style tag embedding eclap ∈ Rd from a short style label, providing acoustic grounding alongside the T5 representations. The central component is a flow-matching DiT [10, 22], which serves as the core generative model: a transformer with L=25 layers that iteratively refines a mel-spectrogram latent xs ∈ RTa ×dmel from Gaussian noise x0 ∼ N (0, I) over S=24 ODE steps. Each layer of the DiT contains self-attention, crossattention, and feed-forward sublayers, where the cross-attention mechanism is the locus of style conditioning: queries from the current latent xs attend to keys and values derived from the caption embeddings Ec . Finally, a HiFi-GAN vocoder [26] converts the refined mel-spectrogram into the output waveform w ∈ RN . In flow matching, the model learns a velocity field vθ (xs , s, Ec ) satisfying the ODE:  dx = vθ xs , s, Ec , ds

s ∈ [0, 1]

(1)

transporting noise x0 ∼ N (0, I) to the data distribution. At each ODE step s and each layer l, the transformer computes multi-head cross-attention: ! (l,s) (l,s) ⊤ Qh Kh (l,s) √ Ah = softmax (2) ∈ RTa ×Tc dk where h ∈ {1, . . . , H} indexes attention heads, and Q(l,s) ∈ RTa ×dk is the audio query, K(l,s) ∈ RTc ×dk is the caption key. This cross-attention is the mechanism through which the caption directly influences audio generation at every layer and every step. 3.2. DAAM adaptation for speech We register forward hooks on each cross-attention module to in(l,s) tercept Ah (Eq. 2) at every layer l ∈ {0, . . . , L−1} and ODE step s ∈ {0, . . . , S−1}. For each tensor A(l,s) ∈ RH×Ta ×Tc , we first average over heads: H

Ā(l,s) =

1 X (l,s) Ah ∈ RTa ×Tc H h=1

(3)

Style Caption

T5 Encoder

Self-Attn

Cross-Attn

FFN

Vocoder

N (0, I)

Waveform hook

DAAM Hook

Heatmaps Mj

Figure 1: CapSpeech pipeline with our DAAM attribution hook. Style captions condition the Flow-Matching DiT. Crossattention maps are intercepted at every layer and ODE step (dashed arrows), then aggregated into per-token temporal heatmaps Mj . We then aggregate across all layers and steps to obtain a per-token heatmap: L−1 S−1

Mj =

1 X X (l,s) Ā:, j ∈ RTa L·S s=0

(4)

l=0

where Mj is the temporal attribution map for caption token j. This 1-D heatmap is analogous to the 2-D spatial maps in image DAAM [6], but over the audio time axis. In total, we capture L × S = 600 attention matrices per generation. 3.3. Token categorisation We classify each caption token into three functional categories: • Style (Csty ): 30 adjectives describing voice quality, emotion, or pace (e.g., “calm”, “bright”, “harsh”). 7,968 instances across all generations. • Content (Ccon ): 20 nouns describing the speaker or voice (e.g., “voice”, “speaker”, “male”). 8,480 instances across all generations. • Function (Cfn ): articles, prepositions, punctuation. 38,432 instances across all generations. The style adjectives and content nouns were curated from TTS literature [2, 3], covering emotional, prosodic, and quality attributes. Since T5’s SentencePiece tokenizer splits words into subword units, we merge contiguous subwords by detecting word-boundary prefixes before classification. 3.4. Analysis metrics We define five complementary metrics on each token’s aggregated heatmap Mj . P Temporal variance σj2 = T1a t (Mj,t − M̄j )2 measures how concentrated attention is over time. Low variance indicates temporally diffuse (global) influence; high variance signals time-localized attention. Peak-to-mean ratio PMRj = maxt Mj,t / (M̄j + 10−8 ) captures the sharpness of the attention peak relative to the background level. A ratio near 1 indicates uniform attention; values ≫ 1 indicate a dominant spikePat a particular time region. Temporal P entropy Hj = − t pj,t log2 pj,t (where pj,t = Mj,t / t′ Mj,t′ ) quantifies distributional uniformity in the information-theoretic sense [27]. Maximum entropy Hmax = log2 Ta corresponds to perfectly uniform attention across all audio frames. Acoustic correlation. We extract frame-level F0 (pYIN algorithm [28], 50–600 Hz range to cover typical male and female speech) and RMS energy from each generated waveform using librosa [29]. Both features are linearly interpolated to Ta frames to match the attention time axis, and we compute Pearson correlation r between each token’s heatmap Mj and each acoustic feature f ∈ {F0, energy}. A positive correlation indicates that

more salient acoustic correlates “loud” (6.3 × 10−5 ), “nasal” (4.2 × 10−5 ) exhibit higher variance, suggesting they partially modulate specific temporal regions where their acoustic effect is strongest. Even the least global style word (“loud”) still has 3.1× lower variance than the function-token average. Figure 1: Global vs Local Conditioning Analysis

0.0014

(A) Temporal Variance

3.5

***

(5)

Peak-to-Mean Ratio

Ta X S−1 XX 1 (l,s) Ā S · Ta · |C| s=0 t=1 j∈C t,j

Temporal Variance

(l)

0.0008 0.0006 0.0004

We construct 120 style captions by systematically combining 30 style adjectives (see token vocabulary after references) with 20 content nouns across 6 caption templates (e.g., “a {adjective} {noun} speaking”). Each style caption serves as the crossattention conditioning input to the model. We pair these with 30 diverse text transcripts (varying in length, phonetic complexity, and content) for which audio is generated. This yields 120 × 30 = 3,600 (style caption, text transcript) combinations. All combinations are processed through CapSpeech, yielding 3,520 successful generations (80 excluded due to duration estimation failures). Across 3,520 generations, we analyze 54,880 token instances and capture 3,520 × 600 = 2,112,000 attention matrices.

4. Analysis and Results 4.1. Experiment 1: Global vs. local conditioning We compute temporal variance, PMR, and entropy for each token’s aggregated heatmap and group by category. To quantify effect magnitude beyond p-values, we report Cohen’s d [30] (standardized mean difference) alongside Mann–Whitney U pvalues [31]: the former provides interpretable effect sizes, while the latter avoids parametric assumptions about the underlying distributions. Results are shown in Figure 2 and Table 1. Style tokens exhibit the lowest temporal variance (σ̄ 2 = 2.1 × 10−5 ), significantly lower than content tokens (7.0 × 10−5 ; p < 10−43 , d = −1.16) and function tokens (19.2 × 10−5 ; p < 10−44 , d = −0.72). The 9.2× variance ratio between function and style tokens confirms that style adjectives distribute attention uniformly across the utterance, acting as global modulators rather than aligning to specific temporal regions. Conversely, style tokens show the highest PMR (1.74 vs. 1.48 for content, 1.36 for function; all p < 10−10 ). This seemingly paradoxical combination, low variance yet high PMR, reveals that style tokens produce compact, characteristic attention signatures: a consistent shape with a distinctive peak, spread uniformly across time. Function tokens show the opposite pattern (high variance, low PMR). Temporal entropy shows no significant inter-category differences (p > 0.05). Per-word variance analysis. Table 2 reports temporal variance for individual style words, revealing a clear hierarchy. Words describing global prosodic properties “cheerful” (σ 2 = 1.0 × 10−5 ), “deep” (1.1 × 10−5 ), “harsh” (1.1 × 10−5 ) show the most temporally diffuse attention, while words with

10.0

2.5 2.0

n.s.

9.5 9.0 8.5 8.0

0.0000

3.5. Dataset

*

10.5

1.5

0.0002

(s)

Step importance IC is defined analogously by averaging across layers. We further summarise these as a late-to-early ratio RC , comparing mean importance in layers 13–24 to layers 0–12, and (0) (S−1) a step decay ratio DC = IC /IC .

(C) Temporal Entropy

***

***

3.0

0.0010

IC =

(B) Peak-to-Mean Ratio

***

0.0012

Temporal Entropy

the model allocates more attention to a given token in temporal regions where that acoustic feature is elevated. Layer/step importance. To understand where in the generation process each token category is most influential, we com(l) pute per-category importance IC at each layer l by averaging, for all tokens j in category C, the mean attention weight received across all ODE steps and audio frames:

Style

Content

1.0

Function

Style

Content

Function

Style

Content

Function

Figure 2: Global vs. local conditioning. Boxplots of (A) temporal variance, (B) PMR, and (C) temporal entropy by category: Style, Content, Function. Brackets: Mann–Whitney U ; *** p < 0.001. Table 1: Token-level metrics (mean ± std). All significance tests: Mann–Whitney U . d: Cohen’s effect size vs. style. Cat. Style Content Function

n

σ̄ 2 (×10−5 )

PMR

H̄ (bits)

dvar

7,968 8,480 38,432

2.1 ± 2.2 7.0 ± 5.6 19.2 ± 33.5

1.74 ± 0.48 1.48 ± 0.30 1.36 ± 0.43

8.72 ± 0.36 8.74 ± 0.36 8.76 ± 0.36

— −1.16 −0.72

Table 2: Per-word temporal variance (×10−5 ) for selected style words, sorted by σ 2 . Lower values indicate more global attention. Word

n

σ̄ 2

Word

n

σ̄ 2

cheerful deep harsh soft cold smooth excited

640 320 320 320 448 416 384

1.0 1.1 1.1 1.3 1.3 1.4 1.4

nervous calm robotic clear dramatic nasal loud

352 224 384 416 544 288 256

2.2 2.4 2.7 3.7 3.7 4.2 6.3

4.2. Experiment 2: Acoustic feature correlations We compute Pearson correlations between each token’s attention heatmap Mj and the F0 and energy contours of the generated waveform to test whether attention reflects measurable acoustic influence. Figure 3 and Table 3 show style tokens have moderate positive correlations with F0 (r̄ = +0.21) and energy (r̄ = +0.28), substantially stronger than function tokens (F0: +0.11, energy: +0.09). The energy difference is highly significant (p < 10−8 ); the F0 difference is also significant (p = 0.02). Content tokens show the strongest correlations overall (F0: +0.50, energy: +0.54), because speaker-identity words (“male”, “female”) impose broad, categorical constraints on the pitch and energy envelope that dominate the entire utterance. Semantic coherence. Table 3 (bottom) reveals acoustically meaningful patterns: “loud” correlates most strongly with energy (r = +0.64), confirming the model’s attention peaks where audio is loudest. “Nasal” shows highest energy correlation (r = +0.67), consistent with increased spectral energy. Words describing emotional arousal (“nervous”: r = +0.47; “dramatic”: r = +0.46) correlate more with energy than F0, while “confident” shows the opposite (rF0 = +0.40), consistent with pitch-raising in confident speech. These patterns demonstrate cross-attention functionally grounds style semantics in acoustic reality.

We decompose attention importance by transformer layer and ODE step separately for each token category, and also track attention entropy to measure selectivity. Figure 2: Correlation with Acoustic Features (B) Corr(Attention, Energy)

(C) Case Study 1.0

nervous nasal loud confident dramatic calm deep robotic warm harsh soothing cold soft slow gentle cheerful energetic smooth bright excited fast clear

0.2

0.1 0.0

0.1

0.2 0.3 Pearson r

0.4

0.5

0.6

Attn: "soft" Energy (norm) F0 (norm)

Normalised magnitude

0.8 0.6 0.4 0.2 0.0 0.2

0.0

0.2 0.4 Pearson r

0.6

0

1

2 Time (s)

3

Figure 3: Layer-Wise & Step-Wise Attention Evolution

4

Figure 3: Acoustic grounding. (A) Mean r(attention, F0) per style word. (B) Same for energy. (C) Case study: normalised attention, F0, and energy overlaid.

(C) Attention Entropy vs Depth (A) Attention Importance by Layer

(B) Attention Importance by ODE Step 0.10

Function

0.06

r̄Energy

n

By category Style +0.21 Content +0.50 Function +0.11

+0.28 +0.54 +0.09

7,968 8,480 38,432

Selected style words loud +0.49 nasal +0.41 confident +0.40 nervous +0.37 robotic +0.32 dramatic +0.30 calm +0.27

+0.64 +0.67 +0.30 +0.47 +0.56 +0.46 +0.40

256 288 256 352 384 544 224

Layer dynamics (Figure 4A, Table 4): Style-token importance increases monotonically from early to late layers, with (17) peak importance at layer 17 (Isty = 0.034). The late-to-early ratio is Rsty = 1.28, meaning the average style importance in layers 13–24 is 28% higher than in layers 0–12. Content tokens (22) peak even later at layer 22 (Icon = 0.061, Rcon = 1.07), suggesting the network resolves style modulation first and then refines speaker identity in the deepest layers. Function tokens, by contrast, show flat or slightly declining importance across depth (Rfn = 0.98), meaning deeper layers selectively amplify semantically meaningful tokens while suppressing grammatical scaffolding. This layer-wise hierarchy is reminiscent of findings in vision transformers [9, 32], where deeper layers produce more semantically refined attention patterns, and aligns with hierarchical feature learning observed in diffusion models [18, 23]. Step dynamics (Figure 4B): Style importance peaks at (0) ODE step s = 0 (Isty = 0.053) and declines to 0.010 by s = 23, a 5.2× decay, the largest among all categories. Content tokens decay more gradually (1.7×: 0.048 → 0.029), suggesting speaker-identity features remain relevant throughout denoising. Most strikingly, function tokens show the opposite trend: their importance increases from 0.086 at step 0 to 0.103 at step 23 (Dfn = 0.84×, i.e., rising). This crossover reveals that in early steps, the network primarily attends to style and content tokens to establish the global acoustic scaffold, while in later steps it shifts attention toward function tokens for finegrained sequential structure such as phrasing and timing. This mirrors the coarse-to-fine dynamics observed in image diffusion [8].

0.06

Content

0.04 0.02

Style

0.04 0.02

Style

0

5

ODE step 10 15

20

0

5

10 15 Layer index

20

8.750 8.725 8.700 8.675 8.650 8.625 8.600 Layer entropy Step entropy

8.575 0

r̄F0

0.08

0.08 Content

Table 3: Pearson r between attention and acoustic features. Top: by category. Bottom: selected style words (n ≥ 224).

0.10

Function

Mean importance

(A) Corr(Attention, F0) loud confident deep nervous calm harsh dramatic warm cold slow nasal robotic gentle soft cheerful soothing energetic smooth bright excited fast clear

Entropy dynamics (Figure 4C): Layer entropy ranges from 8.54 to 8.76 bits. The minimum occurs at layer 18 (H (18) = 8.54 bits), directly adjacent to the style importance peak at layer 17. This co-occurrence is not coincidental: it indicates the network becomes maximally selective, concentrating attention on fewer but more relevant tokens, at precisely the layers most critical for style conditioning. In contrast, step entropy is nearly constant (∆H < 0.03 bits across all 24 steps), indicating that the breadth of attention remains stable throughout denoising even as its distribution across token categories shifts dramatically.

Mean importance Mean entropy (bits)

4.3. Experiment 3: Layer and step dynamics

5

10 15 Layer index

20

0

5

10 15 ODE step

20

(l)

25

(s)

Figure 4: Layer and step dynamics. (A) IC by layer. (B) IC by ODE step. (C) Entropy vs. layer (solid) and step (dashed).

Table 4: Layer and step dynamics summary. R: late-to-early importance ratio (layers 13–24 vs. 0–12). D: first-to-last step decay ratio. Peak: layer or step of maximum importance. Cat. Style Content Function

Layer dynamics (l) R Peak l Ipeak 17 22 18

0.034 0.061 0.108

1.28 1.07 0.98

Step dynamics Peak s D Trend 0 0 23

5.2× 1.7× 0.84×

↘ ↘ ↗

5. Discussion and Conclusion Our three experiments converge on a unified picture: stylecaptioned TTS implements cross-attention as a hierarchical global conditioning channel. Style tokens distribute attention uniformly across time (9.2× lower variance than function tokens, d = −1.16), correlate with acoustic features in semantically coherent patterns (“loud” ↔ energy, r = +0.64), and follow a coarse-to-fine schedule where early ODE steps establish global structure and deep transformer layers progressively refine acoustic details. This constitutes the first quantitative evidence that cross-attention in flow-matching TTS functions as a global modulation mechanism, fundamentally distinct from the temporal alignment role it plays in autoregressive TTS [17]. The per-word analysis reveals a clear division of labor: acoustically grounded words (“loud”, “nasal”) partially localize to regions where their effect is strongest, while abstract descriptors (“cheerful”, “deep”) remain uniformly distributed. Early ODE steps establish coarse global structure, deeper layers refine acoustic details, and attention narrows at critical layers (17–18) to focus on informative tokens—enabling both efficient global style conditioning and fine-grained sequential control within a single mechanism. Limitations. Our analysis is limited to one model (CapSpeech) and synthetic prompts over 30 style words. Future work should extend to: (1) other flow-matching and diffusionbased TTS architectures, (2) naturally occurring prompts from user studies, (3) causal intervention via attention editing [8], (4) per-head analysis to identify specialized style-encoding heads, and (5) comparison with baseline attention patterns to quantify learned structure.

6. Use of Generative AI Disclosure In preparing this manuscript, the authors used generative AI tools for language refinement (rephrasing and improving the clarity of author-written text) and as a coding assistant (helping write and debug software for experiments and analysis). All research contributions, including the methodology, experimental design, results, and scientific claims, are the authors’ own. The authors reviewed and verified all AI-assisted text and code, and take full responsibility for the content of this paper.

7. References [1] N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,” IEEE Trans. Audio Speech Lang. Process., vol. 19, no. 4, pp. 788–798, 2011. [2] H. Wang, J. Hai, D. Chong, K. Song, and D. Yang, “CapSpeech: Enabling downstream applications in style-captioned text-to-speech,” arXiv preprint arXiv:2406.02391, 2024. [3] M. Le, A. Vyas, B. Shi, B. Karrer, L. Sari, R. Moritz, M. Williamson, V. Manohar, Y. Adi, J. Mahadeokar, and W.-N. Hsu, “Voicebox: Text-guided multilingual universal speech generation at scale,” Advances in Neural Information Processing Systems, vol. 36, 2024. [4] X. Tan, J. Chen, H. Liu, J. Cong, C. Zhang, Y. Liu, X. Wang, Y. Leng, Y. Yi, L. He, F. Soong, T. Qin, S. Zhao, and T.-Y. Liu, “NaturalSpeech 3: Zero-shot speech synthesis with a factorized codec and diffusion models,” Proc. ICML, 2024. [5] K. Lakhotia, E. Kharitonov, W.-N. Hsu, Y. Adi, A. Polyak, B. Bolte, T.-A. Nguyen, J. Copet, A. Baevski, A. Mohamed, and E. Dupoux, “On generative spoken language modeling from raw audio,” in Trans. Assoc. Comput. Linguistics, vol. 9, 2021, pp. 1336–1354.

[15] E. Casanova, J. Weber, C. Shulby, A. C. Junior, E. Gölge, and M. A. Ponti, “YourTTS: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone,” in Proc. ICML, 2022, pp. 2709–2720. [16] K. Yuan, X. Tan, D. Yang, Y. Yi, C. Zhang, J. Cong, and S. Zhao, “NaturalSpeech 2: Latent diffusion models are natural and zeroshot speech and singing synthesizers,” in Proc. ICLR, 2024. [17] Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “FastSpeech 2: Fast and high-quality end-to-end text to speech,” in Proc. ICLR, 2021. [18] J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in Neural Information Processing Systems, vol. 33, pp. 6840–6851, 2020. [19] Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” Proc. ICLR, 2021. [20] V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, and M. Kudinov, “Grad-TTS: A diffusion probabilistic model for text-to-speech,” in Proc. ICML, 2021, pp. 8599–8608. [21] J. Kim, S. Kim, J. Kong, and S. Yoon, “Glow-TTS: A generative flow for text-to-speech via monotonic alignment search,” in Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 8067–8077. [22] Y. Chen, Z. Niu, Z. Ma, K. Deng, C. Wang, J. Zhao, K. Yu, and X. Chen, “F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching,” in Proc. ICLR, 2025. [23] W. Peebles and S. Xie, “Scalable diffusion models with transformers,” Proc. ICCV, pp. 4195–4205, 2023. [24] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems, vol. 30, 2017.

[6] R. Tang, L. Liu, A. Pandey, Z. Jiang, G. Yang, K. Kumar, P. Stenetorp, J. Lin, and F. Ture, “What the DAAM: Interpreting stable diffusion using cross attention,” in Proc. ACL, 2023, pp. 5644– 5659.

[25] Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in Proc. ICASSP, 2023, pp. 1–5.

[7] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proc. CVPR, 2022, pp. 10 684–10 695.

[26] J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” Advances in Neural Information Processing Systems, vol. 33, pp. 17 022–17 033, 2020.

[8] A. Hertz, R. Mokady, J. Tenenbaum, K. Aberman, Y. Pritch, and D. Cohen-Or, “Prompt-to-prompt image editing with crossattention control,” in Proc. ICLR, 2023.

[27] C. E. Shannon, “Communication in the presence of noise,” Proc. IRE, vol. 37, no. 1, pp. 10–21, 1949.

[9] H. Chefer, S. Gur, and L. Wolf, “Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers,” in Proc. ICCV, 2021, pp. 397–406.

[28] M. Mauch and S. Dixon, “pYIN: A fundamental frequency estimator using probabilistic threshold distributions,” in Proc. ICASSP, 2014, pp. 659–663.

[10] Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le, “Flow matching for generative modeling,” Proc. ICLR, 2023.

[29] B. McFee, C. Raffel, D. Liang, D. P. W. Ellis, M. McVicar, E. Battenberg, and O. Nieto, “librosa: Audio and music signal analysis in Python,” Proc. SciPy, pp. 18–24, 2015.

[11] C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” J. Mach. Learn. Res., vol. 21, no. 140, pp. 1–67, 2020. [12] Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards end-to-end speech synthesis,” in Proc. Interspeech, 2017, pp. 4006–4010. [13] J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, “Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,” in Proc. ICASSP, 2018, pp. 4779–4783. [14] Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “FastSpeech: Fast, robust and controllable text to speech,” in Advances in Neural Information Processing Systems, vol. 32, 2019.

[30] J. Cohen, “Statistical power analysis for the behavioral sciences,” Lawrence Erlbaum Associates, 1988. [31] H. B. Mann and D. R. Whitney, “On a test of whether one of two random variables is stochastically larger than the other,” Ann. Math. Stat., vol. 18, no. 1, pp. 50–60, 1947. [32] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” Proc. ICLR, 2021.

Record · ID 290606 · SHA-256 cc5efd6829693390
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.