TASU2: Controllable CTC Simulation for Alignment and Low-Resource Adaptation of Speech LLMs Jing Peng1,∗ , Chenghao Wang2,∗ , Yi Yang1 , Lirong Qian1 , Junjie Li1 , Yu Xi1 , Shuai Wang3 , Kai Yu1,∗∗ 1
X-LANCE Lab, Department of Computer Science and Engineering, Shanghai Jiao Tong University, China 1 MoE Key Lab of Artificial Intelligence, 1 Jiangsu Key Lab of Language Computing 2 AISpeech Ltd, Suzhou, China 3 Nanjing University, China
arXiv:2604.08384v1 [eess.AS] 9 Apr 2026
{jing.peng, kai.yu}@sjtu.edu.cn, [email protected], [email protected]
Abstract Speech LLM post-training increasingly relies on efficient crossmodal alignment and robust low-resource adaptation, yet collecting large-scale audio-text pairs remains costly. Text-only alignment methods such as TASU reduce this burden by simulating CTC posteriors from transcripts, but they provide limited control over uncertainty and error rate, making curriculum design largely heuristic. We propose TASU2, a controllable CTC simulation framework that simulates CTC posterior distributions under a specified WER range, producing textderived supervision that better matches the acoustic decoding interface. This enables principled post-training curricula that smoothly vary supervision difficulty without TTS. Across multiple source-to-target adaptation settings, TASU2 improves indomain and out-of-domain recognition over TASU, and consistently outperforms strong baselines including text-only finetuning and TTS-based augmentation, while mitigating sourcedomain performance degradation. Index Terms: speech large language models, speech recognition, domain adaptation
1. Introduction The rapid progress of large language models has accelerated Speech LLM research [1, 2]. However, strong Speech LLM performance often comes with heavy reliance on large-scale audio–text pairs and compute-intensive pipelines [3], making post-training, adaptation, and reproduction costly. Recent studies therefore revisit lightweight alignment between speech and text representations, including TASU [4], LegoSLM [5], and AlignFormer [6]. Among them, TASU (Text-only Alignment for Speech Understanding) is particularly appealing because it enables textonly post-training: it stochastically simulates CTC posteriors from transcripts, allowing training without paired audio while retaining real-audio inference. In practice, TASU can serve as an effective curriculum and improves recognition both on the source domain and under domain shift. Yet its simulation provides limited contro l over posterior uncertainty and the resulting error regime, making difficulty scheduling largely heuristic. In parallel, low-resource speech understanding remains a persistent challenge [7]. Many target domains lack sufficient paired audio, and straightforward audio-based fine-tuning can yield unstable gains and noticeable source-domain degradation [8, 9]. Text-centric adaptation has been explored to reduce * These authors contributed equally. ** indicates the corresponding author.
the need for paired audio, including parameter-efficient tuning [10] and text-only updates of the language component [11]. However, using plain text as the training signal still suffers from a mismatch to the acoustic decoding interface, and its improvements can be limited compared to stronger audio-based augmentation baselines such as TTS [12]. To address both limitations, we propose TASU2, a controllable text-to-CTC simulation framework for Speech LLM posttraining. Instead of relying on unconstrained stochastic simulation, TASU2 generates pseudo CTC posterior distributions under a specified target WER range, producing text-derived supervision that better matches the acoustic decoding interface. This enables a principled curriculum that smoothly varies supervision difficulty and error profiles without TTS or paired audio. Experiments show that TASU2 strengthens text-only alignment beyond TASU, improving recognition on the source domain and under cross-domain generalization without any audiotext training. More importantly, under low-resource transfer settings, TASU2 consistently outperforms strong baselines including text-only adaptation [11] and TTS-based augmentation [12], while better preserving source-domain performance. In summary, our contributions are: • We introduce a WER-conditioned text-derived post-training signal by simulating calibrated CTC posterior distributions from transcripts, bridging the gap between plain text supervision and the acoustic decoding interface. • We develop a controllable text-to-CTC simulator that generates posterior sequences under a specified WER range, enabling explicit control over supervision difficulty and error profiles for curriculum design. • We demonstrate consistent gains over TASU on both sourcedomain and generalization evaluations without audio training, and competitive improvements in low-resource transfer over text-only and TTS-based augmentation baselines.
2. Text-only Alignment: From TASU to TASU2 CTC introduces a blank symbol and marginalizes over alignments, collapsing frame-level posteriors into compact label sequences via blank removal and repetition merging [13, 14]. This compact representation has inspired several speech-LLM alignment methods, like AlignFormer and LegoSLM. However, these methods still rely on paired audio–text data and exhibit performance drops under limited supervision. TASU eliminates the need for paired data by simulating CTC posteriors from text alone. During inference, Label-
Low-WER Subset (WER < 5%)
LibriSpeech Corpus
Simulator
Hello, this is Daniel speaking. LP
One-hot Vector
One-hot Vector Pseudo CTC
20% wer
LP
Encoder Concat
Q
Real CTC
Simulator
5% wer
wer
K, V 3-Level Data Augmentation
Decoder LP
50% wer CE Loss
Evenly divided into 3 WER Intervals
Pseudo CTC
Augmented + Original Dataset
a. The function of Simulator
Interval 1 wer: 0–6% Interval 2 wer: 10–40% Interval 3 wer: 50%+
c. Simulator Training Data Construction
b. The Structure of Simulator and Training details
Figure 1: An Overview of TASU2.
Synchronous Decoding (LSD) [15, 16] compresses real audioderived CTC posteriors P ∈ RT ×V by removing frames where the blank probability exceeds a threshold τ : P′t =
( ∅, Pt ,
if Pt (< blank >) > τ, otherwise,
(1)
1 X ′ Pt , |Sj | t∈S
j = 1, . . . , J.
As described in detail in Algorithm 1, for each utterance, a teacher ASR system provides a real CTC posterior sequence P = (p1 , . . . , pT ⋆ ). Since T ⋆ varies, we pad to a fixed horizon Ttrain with a validity mask m ∈ {0, 1}Ttrain . We train the simulator with posterior-level cross-entropy: "
and then merges consecutive identical frames via averaging: P′′t =
3.2. Training Signal: Distribution-level Supervision
# Ttrain V X X 1 mt pt,v log p̂t,v . t mt t=1 v=1
Lsim (θ) = E − P (2)
j
For text-only training, CTC Posterior Simulation (CPS) generates pseudo-posteriors S̃ from token sequences by applying random label smoothing (p̃ = αδy + (1 − α) V1 1), deletions, and insertions (blanks or duplicates). A projector trained on these simulated posteriors maps them into a frozen LLM, enabling zero-shot speech recognition and, when used as a curriculum pre-training stage, improved domain generalization. Despite its effectiveness, TASU’s simulation remains uncontrolled and may not fully capture real acoustic-phonetic confusions. This raises two questions: controllability and f idelity, which we focus on in TASU2.
3. TASU2: WER-Controllable Text-to-CTC Posterior Simulation As motivated in related work, a key question in text-only alignment is whether a simulator can generate CTC-like posteriors that are both (i) close to real acoustic posteriors and (ii) controllable to support principled curricula and low-resource adaptation. We address this with TASU2, which learns a controllable text-to-CTC posterior simulator that outputs pseudo CTC posterior sequences conditioned on a transcript and a discrete WER control code. Fig. 1 overviews the pipeline.
(3)
This distribution-matching objective encourages CTC-like structure like blank dominance and token confusability, which is more faithful to acoustic decoding than text-only perturbations. 3.3. Architecture Design of Simulator We instantiate the simulator as a lightweight Transformer encoder–decoder (Fig. 1(b)). The transcript is embedded and combined with a learned embedding for the WER code c, which conditions generation. The decoder autoregressively outputs one posterior frame per step. During training we use teacher forcing; at inference time, the simulator runs autoregressively conditioned only on (y, c). 3.4. WER-conditioned Data Construction To obtain controllable supervision, we need the same transcript paired with teacher posteriors spanning different error regimes. Following Fig. 1(c), we sample only a one-seventh portion of LibriSpeech and generate multiple augmented variants (noise/reverb). For each variant, we run a teacher ASR model to obtain a greedy hypothesis ỹ and CTC posteriors P. We compute WER(ỹ, y) and map it to one of K WER intervals to form the control code c. Using discrete intervals reduces sensitivity to measurement noise and alleviates skewed WER distributions, yielding a more stable control signal. And we train the simulator on this data.
3.1. Task and Notation Let the transcript be a token sequence y = (y1 , . . . , yU ) from a tokenizer with vocabulary size V (including the CTC blank). TASU2 learns a simulator that maps (y, c) to a posteriorframe sequence P̂ = (p̂1 , . . . , p̂T ), where p̂t ∈ [0, 1]V and PV v=1 p̂t,v = 1. The control code c ∈ {1, . . . , K} indicates a target WER interval (e.g., 1,2,3 or low/medium/high).
3.5. Simulator Fidelity Analysis We further assess whether the simulator indeed reproduces acoustic-like CTC posteriors and whether the WER control behaves as intended. We consider two complementary aspects: (i) Posterior similarity to teacher CTC (unconditioned). We compare TASU2 against the TASU baseline simulation (without
Algorithm 1: TASU2: WER-Conditioned Text-toCTC Simulation Input: Transcript y, vocab size V (incl. blank), WER bins {Ik }K k=1 , teacher ASR T , augmentor A, horizon Ttrain Output: Simulator parameters θ
WER Distribution
0.14
bin 1 bin 2 bin 3
0.1
Probability Density
0.06
(1) Build multi-WER supervision 2 foreach utterance (x, y) do 3 Generate augmented waveforms {x(j) } ← A(x) 4 foreach x(j) do 5 Obtain teacher posteriors P(j) ← TCTC (x(j) ) 6 Obtain greedy hypothesis ỹ(j) ← Tgreedy (x(j) ) 7 Compute WER(j) = WER(ỹ(j) , y) and assign c(j) ← bin(WER(j) ; {Ik }) 8 Store tuple (y, c(j) , P(j) ) 1
0.02 0.01 0.005
bin 1 (0-6%)
bin 2 (10-40%)
bin 3 (50-150%)
Figure 2: WER controllability under WER-bin conditioning.
(2) Train conditional simulator while not converged do 11 Sample a tuple (y, c, P) 12 Pad P to length Ttrain and build mask m 13 Predict P̂ ← pθ (· | y, c) with teacher forcing 14 Update θ by minimizing Lsim in Eq. (3) 9
generalization, and (ii) low-resource domain adaptation.
10
4.1. Model Architecture
Table 1: Posterior similarity metrics (lower is better except Acc). Method
CE ↓
KL ↓
Acc ↑
ProbDiff ↓
TASU-style baseline TASU2 (AR simulator)
2.5149 1.2296
2.3372 1.0497
0.8220 0.8782
0.0813 0.0615
4.2. Datasets
WER conditioning) using distribution-level metrics between teacher posteriors P and simulated posteriors P̂:
CE =
N V 1 X X − pt,v log p̂t,v , N t=1 v=1 N
N 1 X Acc = I arg max pt,v = arg max p̂t,v , v v N t=1
(4)
N
ProbDiff =
We pre-train the CTC simulator on one-seventh of LibriSpeech (960h) [19] with augmentation, to learn a general mapping from linguistic units to acoustic-like posterior distributions. For Stage 2 post-training, we synthesize pseudo CTC supervision from text-only corpora in LibriSpeech and Medical [20]. We evaluate on LibriSpeech (test-clean/other), Medical (8h), TEDLIUM 3 [21], SlideSpeech [22], and CoVoST2 En→Zh [23]. 4.3. Training and Evaluation Setup
V
1 XX pt,v KL = pt,v log , N t=1 v=1 p̂t,v
Our simulator is a Transformer encoder–decoder with 6 encoder and 6 decoder layers and hidden size 512. It takes a transcript and a WER-bin ID as input and autoregressively generates pseudo CTC posterior frames. For the Speech LLM, we follow TASU [4] and use SenseVoice-Small [17] as the speech encoder and Qwen2.5-1.5B [18] as the LLM, connected by a Linear–SiLU–Linear projector.
1 X max pt,v − max p̂t,v , v N t=1 v
where N is the number of valid frames. As shown in Table 1, TASU2 yields substantially lower CE/KL and improved argmax agreement, indicating closer matching to real CTC posteriors. (ii) WER controllability (conditioned). To verify controllability, we inject WER-bin codes and decode the simulated posteriors with a fixed decoder, then measure the realized WER for each bin. Fig. 2 reports the empirical WER distributions across bins, showing that TASU2 reliably separates error regimes and tracks the target intervals.
4. Experiments To validate the effectiveness and controllability of TASU2, we evaluate (i) text-to-CTC alignment quality and cross-domain
WER-conditioned simulation. We discretize the realized teacher WER into three bins: 0–6% (low), 10–40% (medium), and 50–150% (high). Unless otherwise stated, we use bin 1 for conditioning in the main experiments. Two-stage post-training. Stage 1 trains on simulated supervision with learning rate 5 × 10−5 for 5 epochs. Stage 2 adapts the model using simulated text from the target domain (Medical or SlideSpeech), following the same optimization protocol. Optimization. We use AdamW with DeepSpeed ZeRO-2 on 8×Ascend 910B NPUs. LoRA (rank=16, α=32) is used only for the two-stage adaptation setting in Table 4.
5. Evaluation And Analysis We evaluate TASU2 from three perspectives. Table 2 probes whether simulated posteriors improve CTC-style alignment and cross-domain generalization without audio training. Table 3 provides a lightweight multi-task sanity check beyond ASR. Our key result is Table 4, where TASU2 achieves strong low-resource target gains while largely preserving sourcedomain performance, outperforming both text-only adaptation and TTS-based augmentation.
Table 2: Alignment and generalization of simulated CTC training. All models are trained on LibriSpeech and evaluated on LibriSpeech, SlideSpeech, and TED-LIUM; WER (%)↓ is reported. TASU2 conditions: ● None (no WER conditioning); ❍ WER-binned (coarse 3-level WER bins; we use bin 1 here).
System
Train data Stage 1 Stage 2
Audio
Train Data Libri Slide TED Text (Audio, Text) clean/other
System
TASU Libri TASU Libri+Slide SLAM-CTC – TASU2● TASU2❍ TASU2❍
Table 4: Two-stage domain adaptation (source → target). All systems are first trained on the source domain (LibriSpeech) and then adapted to the target domain (Medical). We report WER%↓. LoRA is used for all systems for better performance.
– – Libri
4.57 / 9.90 24.07 19.36 4.21 / 10.31 18.70 13.23 3.13 / 8.59 18.59 14.61
SLAM-CTC
– – –
4.63 / 9.82 16.41 14.02 3.41 / 8.15 17.67 14.49 3.94 / 8.25 17.31 13.93
TASU2❍
Libri Libri Libri+Slide
5.1. Alignment and Generalization We evaluate simulated CTC training on both in-domain and outof-domain test sets to probe alignment and generalization of TASU2. We start from ASR and further extend the evaluation to a multi-task speech understanding setting. ASR performance. As illustrated in Table 2, relative to the original TASU random simulation, TASU2 consistently improves in-domain recognition while yielding much stronger cross-domain transfer. In particular, TASU2 reduces WER markedly on SlideSpeech and TED-LIUM3, and in several cases becomes competitive with, or even surpasses, the strong SLAM-CTC audio baseline, despite using text-only training inputs. Moreover, when we expand the text side by adding SlideSpeech transcripts (Libri+Slide), TASU2 further improves on the target and TED sets, indicating that text expansion and posterior simulation are complementary: additional target-style text benefits transfer, while TASU2 provides the CTC-like supervision signal needed to make such text effective. To understand the role of controllability, we compare two simulator variants for generating pseudo posteriors: unconditioned simulation (None) and WER-binned simulation (coarse 3-level bins; we use bin 1 here). The ablation reveals a clear trade-off. The unconditioned simulator tends to improve outof-domain recognition (better Slide/TED WER), but it incurs noticeable regression on the source domain (worse Libri). In contrast, WER-binned conditioning yields a better balance between alignment and transfer: it preserves strong generalization while substantially reducing source-domain degradation. These results support our design choice that lightweight, discrete error control is useful not only for curriculum design, but also for stabilizing domain shift behavior in simulated-CTC training. Table 3: TASU2 performance under multi-task evaluation. TASU2 uses WER-conditioned simulation with discretized WER bins; here we use only bin 1 for conditioning.
LibriSpeech Medical clean/other test
– Raw text [11] TTS audio Raw audio
2.43 / 6.07 2.56 / 6.62 2.72 / 6.80 2.70 / 6.76
17.55 13.62 12.79 12.35
Text Sim-CTC – Text Sim-CTC Text Sim-CTC
2.94 / 7.16 2.96 / 7.23
15.34 12.12
task sanity check. Since TASU2 simulates CTC posteriors with WER-conditioned control, its supervision is most directly aligned with token recognition and decoding, so gains beyond ASR are not necessarily expected. Still, TASU2 substantially improves over TASU on LibriSpeech without using any training audio, and remains competitive on CoVoST2, suggesting that CTC-style simulated supervision can transfer modestly beyond pure ASR under a zero-audio regime. 5.2. Domain Adaptation Section 5.1 indicates that TASU2 achieves stronger alignment and generalization. In this section, we turn to its domain adaptation capability, with a particular focus on low-resource settings. We study two-stage domain adaptation from a source domain to a low-resource target domain. LibriSpeech serves as the sourcedomain pre-training data (Stage 1), while the medical set simulates a low-resource target domain for adaptation (Stage 2). We report WER on both LibriSpeech and the medical test set to jointly measure target improvement and source retention. Table 4 highlights a clear trade-off in conventional adaptation: target-domain gains often come with reduced source retention. For SLAM-CTC, adapting with target-domain audio (raw or TTS) improves the Medical WER to 12.35/12.79, but also shifts LibriSpeech from 2.43/6.07 to around 2.70–2.72/6.76– 6.80, indicating a non-trivial source-domain drop. In contrast, TASU2 yields a more favorable low-resource transfer profile. Using only text-derived simulated CTC supervision for Stage 2, TASU2 achieves the best Medical WER of 12.12, outperforming text-only adaptation (13.62) and TTS augmentation (12.79), and even slightly surpassing rawaudio adaptation (12.35). Meanwhile, source-domain WER remains nearly unchanged from 2.94/7.16 to 2.96/7.23 (only +0.02/+0.07), demonstrating strong source retention. These results suggest that WER-controlled posterior simulation provides an effective and stable alternative to target-audio-heavy finetuning when paired data is scarce.
6. Conclusion Model TASU TASU2 SLAM Step-Audio Qwen2.5-Omni
Train audio duration(h)
LibriSpeech ↓
CoVoST2 ↑
clean/other (WER↓)
En2Zh (BLEU↑)
0 0 1.8k >1000k >1000k
6.47 / 10.35 3.68 / 8.25 3.30 / 7.24 2.36 / 6.32 2.37 / 4.21
33.35 33.08 37.34 – 41.40
Multi-task performance. Table 3 provides a lightweight multi-
We presented TASU2, a WER-controllable text-to-CTC simulator trained with distribution-level supervision. By generating pseudo posteriors that better match acoustic CTC behavior, TASU2 provides stronger alignment signals for speech foundation models and Speech LLMs. Across evaluations, it improves robustness and cross-domain generalization, and enables effective source-to-target adaptation. In low-resource targets, TASU2 can outperform TTS-based augmentation, offering a practical alternative when paired audio is scarce.
7. Generative AI Use Disclosure During the preparation of this work, we used generative AI tools for assistance. The AI tools were only employed for improving the presentation, readability, and formatting of the manuscript, as well as for auxiliary support in code development and verification. They were not used to generate any substantial content, core ideas, experimental design, analysis, or conclusions of the paper.
8. References [1] J. Peng, Y. Wang, B. Li, Y. Guo, H. Wang, Y. Fang, Y. Xi, H. Li, X. Li, K. Zhang, S. Wang, and K. Yu, “A survey on speech large language models for understanding,” IEEE Journal of Selected Topics in Signal Processing, p. 1–32, 2025. [Online]. Available: http://dx.doi.org/10.1109/JSTSP.2025.3640535
[15] Z. Chen, W. Deng, T. Xu, and K. Yu, “Phone synchronous decoding with ctc lattice.” in Interspeech, 2016, pp. 1923–1927. [16] K. Deng and P. C. Woodland, “Label-synchronous neural transducer for adaptable online e2e speech recognition,” 2023. [Online]. Available: https://arxiv.org/abs/2311.11353 [17] K. An, Q. Chen, C. Deng, Z. Du, C. Gao, Z. Gao, Y. Gu, T. He, H. Hu, K. Hu, S. Ji, Y. Li, Z. Li, H. Lu, H. Luo, X. Lv, B. Ma, Z. Ma, C. Ni, C. Song, J. Shi, X. Shi, H. Wang, W. Wang, Y. Wang, Z. Xiao, Z. Yan, Y. Yang, B. Zhang, Q. Zhang, S. Zhang, N. Zhao, and S. Zheng, “Funaudiollm: Voice understanding and generation foundation models for natural interaction between humans and llms,” 2024. [Online]. Available: https://arxiv.org/abs/2407.04051
[2] S. Arora, K.-W. Chang, C.-M. Chien, Y. Peng, H. Wu, Y. Adi, E. Dupoux, H.-Y. Lee, K. Livescu, and S. Watanabe, “On the landscape of spoken language models: A comprehensive survey,” 2025. [Online]. Available: https://arxiv.org/abs/2504.08528
[18] Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu, “Qwen2.5 technical report,” 2025. [Online]. Available: https://arxiv.org/abs/2412.15115
[3] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via largescale weak supervision,” 2022. [Online]. Available: https: //arxiv.org/abs/2212.04356
[19] V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An asr corpus based on public domain audio books,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015, pp. 5206–5210.
[4] J. Peng, Y. Yang, X. Li, Y. Xi, Q. Tang, Y. Fang, J. Li, and K. Yu, “Tasu: Text-only alignment for speech understanding,” 2026. [Online]. Available: https://arxiv.org/abs/2511.03310
[20] Figure Eight Inc., “Medical speech, transcription, and intent,” Kaggle Dataset, 2019. [Online]. Available: https://www.kaggle.com/datasets/paultimothymooney/ medical-speech-transcription-and-intent
[5] R. Ma, T. Chen, K. Audhkhasi, and B. Ramabhadran, “Legoslm: Connecting llm with speech encoder using ctc posteriors,” arXiv preprint arXiv:2505.11352, 2025. [6] R. Fan, B. Ren, Y. Hu, R. Zhao, S. Liu, and J. Li, “Alignformer: Modality matching can achieve better zero-shot instructionfollowing speech-llm,” IEEE Journal of Selected Topics in Signal Processing, pp. 1–10, 2025. [7] J. Zhao and W.-Q. Zhang, “Improving automatic speech recognition performance for low-resource languages with self-supervised models,” IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1227–1241, 2022. [8] Y. Takashima, S. Horiguchi, S. Watanabe, P. Garcı́a, and Y. Kawaguchi, “Updating only encoders prevents catastrophic forgetting of end-to-end asr models,” 2022. [Online]. Available: https://arxiv.org/abs/2207.00216 [9] S. Burdisso, E. Villatoro-Tello, A. Carofilis, S. Kumar, K. Hacioglu, S. Madikeri, P. Rangappa, M. K. E, P. Motlicek, S. Venkatesan, and A. Stolcke, “Text-only adaptation in llmbased asr through text denoising,” 2026. [Online]. Available: https://arxiv.org/abs/2601.20900 [10] F.-T. Liao, Y.-C. Chan, Y.-C. Chen, C.-J. Hsu, and D.-s. Shiu, “Zero-shot domain-sensitive speech recognition with promptconditioning fine-tuning,” in 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2023, pp. 1–8. [11] Y. Fang, J. Peng, X. Li, Y. Xi, C. Zhang, G. Zhong, and K. Yu, “Low-resource domain adaptation for speech llms via text-only fine-tuning,” arXiv preprint arXiv:2506.05671, 2025. [12] E. Casanova, C. Shulby, A. Korolev, A. C. Junior, A. da Silva Soares, S. Aluı́sio, and M. A. Ponti, “Asr data augmentation in low-resource settings using cross-lingual multi-speaker tts and cross-lingual voice conversion,” 2023. [Online]. Available: https://arxiv.org/abs/2204.00618 [13] M. Jung, O. Kwon, S. Seo, and S. Seo, “Blank collapse: Compressing ctc emission for the faster decoding,” 2023. [Online]. Available: https://arxiv.org/abs/2210.17017 [14] K. Deng, S. Cao, Y. Zhang, and L. Ma, “Improving hybrid ctc/attention end-to-end speech recognition with pretrained acoustic and language model,” 2021. [Online]. Available: https://arxiv.org/abs/2112.07254
[21] F. Hernandez, V. Nguyen, S. Ghannay, N. Tomashenko, and Y. Estève, TED-LIUM 3: Twice as Much Data and Corpus Repartition for Experiments on Speaker Adaptation. Springer International Publishing, 2018, p. 198–208. [Online]. Available: http://dx.doi.org/10.1007/978-3-319-99579-3 21 [22] H. Wang, F. Yu, X. Shi, Y. Wang, S. Zhang, and M. Li, “Slidespeech: A large-scale slide-enriched audio-visual corpus,” 2023. [Online]. Available: https://arxiv.org/abs/2309.05396 [23] C. Wang et al., “CoVoST 2 and massively multilingual speech-totext translation,” in Proc. Interspeech, 2021, pp. 2247–2251.